YouTube2Text

AI And Machine Learning Full Course [FREE] | Learn AI And Machine Learning In 24 Hours | Simplilearn — Transcript

by Simplilearn · 423,981 words · 63,517 segments · language en · Watch on YouTube

Full transcript

  1. 0:05Welcome to simply learns YouTube
  2. 0:07channel. Artificial intelligence and
  3. 0:09machine learning are transforming the
  4. 0:11way business operate, make decisions and
  5. 0:13innovate. From personalized
  6. 0:14recommendation on streaming platforms
  7. 0:17and intelligent chat bots to
  8. 0:19self-driving vehicles and advanced
  9. 0:20healthcare systems, AI and machine
  10. 0:22learning are powering some of the most
  11. 0:25impactful technologies of our time. And
  12. 0:27AI and machine learning engineering
  13. 0:29combines programming, mathematics,
  14. 0:31statistics, data science, machine
  15. 0:33learning expertise to create intelligent
  16. 0:35applications that deliver business
  17. 0:37value. In this complete AI and machine
  18. 0:39learning engineering course, you will
  19. 0:41learn everything from the fundamentals
  20. 0:42of AI and machine learning to advanced
  21. 0:45concept used in the modern intelligent
  22. 0:47systems. We will start with the
  23. 0:48programming and mathematical foundation
  24. 0:50then gradually move into data analysis,
  25. 0:52machine learning algorithms, deep
  26. 0:54learning and generative AI. Throughout
  27. 0:56the course, you will gain hands-on
  28. 0:57experience with industry standard tools
  29. 0:59and frameworks such as Python, NumPy,
  30. 1:02Pandas, Kikit Learn, TensorFlow,
  31. 1:04PyTorch, and others. You'll also learn
  32. 1:06how to collect and prepare data, build
  33. 1:08predictive models, train neural
  34. 1:10networks, evaluate model performance,
  35. 1:12and deploy machine learning solution in
  36. 1:14real world environments. By the end of
  37. 1:16this course, you'll have a strong
  38. 1:17understanding of complete AI and machine
  39. 1:20learning life cycle and practical skills
  40. 1:22required to pursue a career as a AI and
  41. 1:25machine learning engineer. Having said
  42. 1:27that, let's take a look at today's
  43. 1:28agenda. We'll start off with module one,
  44. 1:30which is introduction to artificial
  45. 1:32intelligence and machine learning.
  46. 1:33Module two is Python programming for AI
  47. 1:36and ML. Module three is mathematics,
  48. 1:39statistics, and probability for machine
  49. 1:41learning. Module four is data
  50. 1:43collection, cleaning, and
  51. 1:44pre-processing. Module five is
  52. 1:46exploratory data analysis and data
  53. 1:48visualization. Module six is machine
  54. 1:51learning fundamentals. Module seven is
  55. 1:53supervised learning algorithms. Module
  56. 1:55eight is unsupervised learning
  57. 1:56algorithms. Module 9 is model evaluation
  58. 1:59and feature engineering. Module 10 is
  59. 2:02deep learning and neural networks.
  60. 2:04Module 11 is natural language
  61. 2:05processing. Module 12 is computer vision
  62. 2:08fundamentals. Module 13 is generative
  63. 2:11AI, LLMS and AI agents. Module 14 is
  64. 2:15envelopes, model deployment and AI
  65. 2:17engineering workflows. Module 15 is real
  66. 2:20world AI projects. Module 16 is
  67. 2:23interview question and answers. Hope I
  68. 2:25made myself clear with that agenda.
  69. 2:27That's it. If these are the type of
  70. 2:28videos you would like to watch, then hit
  71. 2:29that subscribe button with the bell icon
  72. 2:31to get notified whenever we host. Also,
  73. 2:33just so that you know, if you want to
  74. 2:34upskill yourself, master generative AI
  75. 2:36and land your dream job or even grow in
  76. 2:38your career, then you must explore
  77. 2:40Simply Learn's cohort of various
  78. 2:42generative AI training and professional
  79. 2:44certification programs. Simply learn
  80. 2:46offers a variety of masters
  81. 2:47certification and post-graduate programs
  82. 2:48in collaboration with some of the
  83. 2:50world's leading universities. Through
  84. 2:51our courses, you will gain knowledge
  85. 2:53along with work ready expertise in
  86. 2:55skills like Python, Aentic AI, AI
  87. 2:57automation systems, LLMs, and over a
  88. 3:00dozen others. And that's not all. You
  89. 3:02will also get the opportunity to work on
  90. 3:03multiple projects led by industry
  91. 3:05experts working on top tier
  92. 3:07service-based and product companies.
  93. 3:09After completing these courses,
  94. 3:10thousands of learners have transition
  95. 3:12into an AI and machine learning role as
  96. 3:14a fresher or moved on to your higher
  97. 3:16paying job and profile. If you're
  98. 3:17passionate about making your career in
  99. 3:19this field, then make sure to check out
  100. 3:20the link in the pinned comments and in
  101. 3:22the description box to find an AI and
  102. 3:25machine learning program that fits your
  103. 3:27experience and areas of interest. So
  104. 3:29let's get started with our AI and
  105. 3:31machine learning engineer full course
  106. 3:33with a small quiz. What is NLP? Is it
  107. 3:36network layer processing, natural
  108. 3:38language processing, neural learning
  109. 3:40platform, or is it native language
  110. 3:42programming? Please let us know your
  111. 3:44answers in the comment section below.
  112. 3:45Now over to our training experts.
  113. 3:48>> Once upon a time in the quiet town of
  114. 3:50Newite, there lived a curious teenager
  115. 3:53named Arya. She wasn't like most kids in
  116. 3:56her school. While others were busy with
  117. 3:58sports or music, Arya was fascinated by
  118. 4:01machines, especially the idea of making
  119. 4:03machines think like humans. Her
  120. 4:05curiosity began one evening when she
  121. 4:07asked her grandfather, who used to be a
  122. 4:09computer engineer, "Can machines ever
  123. 4:11think?" Her grandfather smiled and said,
  124. 4:14"That's what artificial intelligence is
  125. 4:16all about." Arya's eyes lit up.
  126. 4:18Artificial intelligence? What's that? So
  127. 4:21he began to tell her a story, not a
  128. 4:24fairy tale, but a real story about the
  129. 4:25science and ideas behind machines that
  130. 4:28learn, decide, and sometimes even
  131. 4:31surprise their creators. Artificial
  132. 4:34intelligence, or AI, is the science of
  133. 4:37making machines that can do things that
  134. 4:39normally require human intelligence.
  135. 4:42This includes tasks like recognizing
  136. 4:44faces, understanding speech, making
  137. 4:47decisions, and even playing games. But
  138. 4:50AI isn't magic. It's built through
  139. 4:53programming, mathematics, and data. Arya
  140. 4:56imagined a robot that could talk like a
  141. 4:58human and help with homework. Her
  142. 5:00grandfather nodded. That's one kind of
  143. 5:02AI, but there are many types. He
  144. 5:05explained that AI isn't just about
  145. 5:07robots. In fact, most AI systems are
  146. 5:10just computer programs running inside
  147. 5:12machines we already use, like phones,
  148. 5:15laptops, or even refrigerators. Her
  149. 5:17grandfather told her that AI comes in
  150. 5:19two main types, narrow AI and general
  151. 5:22AI. Narrow AI is the kind we see today.
  152. 5:25It's designed to do one specific task.
  153. 5:28For example, the AI in a smartphone that
  154. 5:30unlocks the screen by recognizing your
  155. 5:32face is only good at that one job. It
  156. 5:35can't cook or write a story. General AI,
  157. 5:38on the other hand, would be as smart as
  158. 5:41a human, able to learn anything and do
  159. 5:44many tasks. But this type of AI doesn't
  160. 5:47exist yet. It's more of a dream for now.
  161. 5:50Arya asked, "How do these machines
  162. 5:52learn? That's where machine learning
  163. 5:54comes in," her grandfather replied.
  164. 5:57Machine learning is a type of AI that
  165. 5:59learns from data instead of being told
  166. 6:01what to do step by step. "Imagine
  167. 6:04teaching a dog to sit. You show it how,
  168. 6:07give it treats, and repeat. Over time,
  169. 6:10the dog learns. Machine learning works
  170. 6:13the same way. You feed it data and it
  171. 6:15finds patterns. For example, if you want
  172. 6:18a computer to recognize pictures of
  173. 6:19cats, you show it thousands of cat
  174. 6:22pictures. It starts to see what cats
  175. 6:24usually look like. Furry whiskers,
  176. 6:27pointy ears. Over time, it learns to
  177. 6:31tell a cat apart from a dog or a chair.
  178. 6:34The program that does this learning is
  179. 6:36called a model. A model is like a brain
  180. 6:39built by the computer using the data it
  181. 6:41was given. The more data it gets, the
  182. 6:43better it learns. But how does the
  183. 6:45computer know what a cat is? Arya asked.
  184. 6:48Her grandfather said, "That's thanks to
  185. 6:51something called a neural network. It's
  186. 6:53a method used in machine learning that's
  187. 6:55inspired by how our brains work. A
  188. 6:58neural network is made up of layers of
  189. 7:00tiny parts called neurons. These are not
  190. 7:03real brain cells, but math functions.
  191. 7:06Each neuron takes in numbers, does some
  192. 7:09math, and passes the result to the next
  193. 7:11layer of neurons. Imagine passing a note
  194. 7:14through a group of friends, and each one
  195. 7:17adds or changes a word before giving it
  196. 7:19to the next. By the end, the note may
  197. 7:21have transformed in a useful way. That's
  198. 7:24what a neural network does to data. It
  199. 7:28turns it into something meaningful, like
  200. 7:30recognizing a cat in a picture. The more
  201. 7:33layers a network has, the more complex
  202. 7:36patterns it can understand. When a
  203. 7:39network has many layers, it's called
  204. 7:41deep learning. To get a neural network
  205. 7:43to work, it needs to be trained.
  206. 7:46Training is the process where the model
  207. 7:48is shown lots of examples so it can
  208. 7:50learn. Training involves giving the
  209. 7:52model data and letting it guess
  210. 7:54something like whether a picture has a
  211. 7:56cat. At first, it guessed badly, but
  212. 7:59then it compares its guess to the
  213. 8:01correct answer. If it's wrong, it
  214. 8:03adjusts itself using a method called
  215. 8:06back propagation. Back propagation is
  216. 8:08like checking your math homework. If the
  217. 8:11answer is wrong, you go back, find where
  218. 8:14you messed up, and fix it. In AI, this
  219. 8:17helps the model improve step by step.
  220. 8:20This cycle of guessing, checking, and
  221. 8:22adjusting is repeated many times. The
  222. 8:24model slowly gets better at the task.
  223. 8:27Can AI make mistakes? Arya asked. Oh
  224. 8:31yes, her grandfather said AI is smart in
  225. 8:34some ways but not perfect. AI only
  226. 8:36learns from the data we give it. If the
  227. 8:39data is bad, the AI will be bad. This is
  228. 8:42called bias. For example, if a face
  229. 8:45recognition system is trained mostly on
  230. 8:47photos of light-kinned people, it might
  231. 8:49not work well on darkerkinned people.
  232. 8:52Also, AI doesn't really understand the
  233. 8:54world. It only sees patterns in numbers.
  234. 8:58It doesn't know what a cat feels like or
  235. 9:00why we love them. That's why AI can
  236. 9:02sometimes be fooled by simple tricks
  237. 9:05like weird images that a human would
  238. 9:07never mistake for a cat. AI is
  239. 9:09everywhere, her grandfather explained.
  240. 9:12It helps recommend videos on YouTube,
  241. 9:14powers voice assistants like Siri or
  242. 9:16Alexa, drives some cars, and even helps
  243. 9:20doctors find diseases and scans. But not
  244. 9:23all AI is harmless. It can be used for
  245. 9:26spying, spreading fake news, or making
  246. 9:29decisions that affect people's lives,
  247. 9:31like who gets a loan or a job? That's
  248. 9:34why it's important for people to
  249. 9:35understand how AI works so they can ask
  250. 9:38good questions and build it responsibly.
  251. 9:41Arya asked, "Will AI take over the
  252. 9:44world?" Her grandfather laughed. Not
  253. 9:46like in the movies, but it will change
  254. 9:48the world. The future of AI depends on
  255. 9:51how people choose to use it. It can help
  256. 9:53solve big problems like climate change
  257. 9:55or disease. But it also needs rules and
  258. 9:58careful thinking. Just like fire or
  259. 10:00electricity, AI is a tool, a powerful
  260. 10:02one. If used wisely, it can do great
  261. 10:05good. Arya sat back, her mind buzzing.
  262. 10:08She had started the day wondering if
  263. 10:10machines could think. Now she knows that
  264. 10:12while they don't think like humans, they
  265. 10:15can do amazing things through learning
  266. 10:17data, and clever programming. She smiled
  267. 10:20and said, "Maybe I'll build an AI
  268. 10:22someday." Her grandfather smiled, too.
  269. 10:25Just remember, it's not about making a
  270. 10:26machine smart. It's about making it
  271. 10:28useful and fair for everyone. And from
  272. 10:32that day on, Arya started her journey
  273. 10:35not just to understand AI, but to shape
  274. 10:38it with care, creativity, and curiosity.
  275. 10:41>> We know humans learn from their past
  276. 10:43experiences, and machines follow
  277. 10:46instructions given by humans.
  278. 10:48But what if humans can train the
  279. 10:51machines to learn from their past data
  280. 10:52and do what humans can do and much
  281. 10:54faster? Well, that's called machine
  282. 10:56learning. But it's a lot more than just
  283. 10:58learning. It's also about understanding
  284. 11:00and reasoning. So today we will learn
  285. 11:02about the basics of machine learning. So
  286. 11:05that's Paul. He loves listening to new
  287. 11:08songs.
  288. 11:10He either likes them or dislikes them.
  289. 11:12Paul decides this on the basis of the
  290. 11:14song's tempo, genre, intensity, and the
  291. 11:19gender of voice. For simplicity, let's
  292. 11:21just use tempo and intensity for now.
  293. 11:24So, here tempo is on the x-axis, ranging
  294. 11:27from relaxed to fast, whereas intensity
  295. 11:30is on the y-axis, ranging from light to
  296. 11:33soaring. We see that Paul likes the song
  297. 11:36with fast tempo and soaring intensity
  298. 11:40while he dislikes the song with relaxed
  299. 11:43tempo and light intensity. So now we
  300. 11:45know Paul's choices. Let's say Paul
  301. 11:47listens to a new song. Let's name it as
  302. 11:49song A. Song A has fast tempo and a
  303. 11:53soaring intensity. So it lies somewhere
  304. 11:55here. Looking at the data, can you guess
  305. 11:58whether Paul will like the song or not?
  306. 12:00Correct. So Paul likes this song. By
  307. 12:02looking at Paul's past choices, we were
  308. 12:05able to classify the unknown song very
  309. 12:07easily, right? Let's say now Paul
  310. 12:10listens to a new song. Let's label it as
  311. 12:12song B. So song B lies somewhere here
  312. 12:16with medium tempo and medium intensity.
  313. 12:19Neither relaxed nor fast, neither light
  314. 12:22nor soaring. Now, can you guess whether
  315. 12:24Paul likes it or not? Not able to guess
  316. 12:26whether Paul will like it or dislike it.
  317. 12:29Are the choices unclear? Correct. We
  318. 12:31could easily classify song A. But when
  319. 12:34the choice became complicated as in the
  320. 12:36case of song B. Yes. And that's where
  321. 12:39machine learning comes in. Let's see
  322. 12:40how. In the same example for song B, if
  323. 12:43we draw a circle around the song B, we
  324. 12:45see that there are four votes for like
  325. 12:48whereas one vote for dislike. If we go
  326. 12:50for the majority votes, we can say that
  327. 12:53Paul will definitely like the song.
  328. 12:54That's all. This was a basic machine
  329. 12:56learning algorithm also. It's called K
  330. 12:58nearest neighbors. So this is just a
  331. 13:00small example in one of the many machine
  332. 13:03learning algorithms quite easy right
  333. 13:05believe me it is but what happens when
  334. 13:08the choices become complicated as in the
  335. 13:11case of song B that's when machine
  336. 13:13learning comes in it learns the data
  337. 13:15builds the prediction model and when the
  338. 13:17new data point comes in it can easily
  339. 13:19predict for it more the data better the
  340. 13:22model higher will be the accuracy there
  341. 13:24are many ways in which the machine
  342. 13:26learns it could be either supervised
  343. 13:29learning unsupervised learning or
  344. 13:31reinforcement learning. Let's first
  345. 13:32quickly understand supervised learning.
  346. 13:35Suppose your friend gives you 1 million
  347. 13:37coins of three different currencies. Say
  348. 13:391 rupee, 1 and 1 dirham. Each coin has
  349. 13:43different weights. For example, a coin
  350. 13:45of 1 rupee weighs 3 g. 1 euro weighs 7 g
  351. 13:49and 1 dirham weighs 4 g. Your model will
  352. 13:51predict the currency of the coin. Here
  353. 13:54your weight becomes the feature of coins
  354. 13:56while currency becomes their label. When
  355. 13:58you feed this data to the machine
  356. 14:00learning model, it learns which feature
  357. 14:03is associated with which label. For
  358. 14:05example, it will learn that if a coin is
  359. 14:07of 3 g, it will be a 1 rupee coin. Let's
  360. 14:10give a new coin to the machine. On the
  361. 14:12basis of the weight of the new coin,
  362. 14:14your model will predict the currency.
  363. 14:16Hence, supervised learning uses labeled
  364. 14:19data to train the model. Here, the
  365. 14:21machine knew the features of the object
  366. 14:23and also the labels associated with
  367. 14:25those features. On this note, let's move
  368. 14:27to unsupervised learning and see the
  369. 14:29difference. Suppose you have cricket
  370. 14:31data set of various players with their
  371. 14:33respective scores and the wickets taken.
  372. 14:35When we feed this data set to the
  373. 14:37machine, the machine identifies the
  374. 14:39pattern of player performance. So, it
  375. 14:41plots this data with the respective
  376. 14:43wickets on the x-axis while runs on the
  377. 14:45y-axis. While looking at the data,
  378. 14:47you'll clearly see that there are two
  379. 14:49clusters. The one cluster are the
  380. 14:51players who scored high runs and took
  381. 14:53less wickets while the other cluster is
  382. 14:56of the players who scored less runs but
  383. 14:58took many wickets. So here we interpret
  384. 15:00these two clusters as batsmen and
  385. 15:03bowlers. The important point to note
  386. 15:05here is that there were no labels of
  387. 15:07batsmen and bowlers. Hence the learning
  388. 15:09with unlabeled data is unsupervised
  389. 15:12learning. So we saw supervised learning
  390. 15:13where the data was labeled and the
  391. 15:15unsupervised learning where the data was
  392. 15:17unlabeled. And then there is
  393. 15:19reinforcement learning which is a
  394. 15:21reward-based learning or we can say that
  395. 15:22it works on the principle of feedback.
  396. 15:24Here let's say you provide the system
  397. 15:26with an image of a dog and ask it to
  398. 15:28identify it. The system identifies it as
  399. 15:31a cat. So you give a negative feedback
  400. 15:33to the machine saying that it's a dog's
  401. 15:35image. The machine will learn from the
  402. 15:37feedback and finally if it comes across
  403. 15:39any other image of a dog, it'll be able
  404. 15:41to classify it correctly. That is
  405. 15:43reinforcement learning. To generalize
  406. 15:45machine learning model, let's see a
  407. 15:47flowchart. Input is given to a machine
  408. 15:49learning model which then gives the
  409. 15:50output according to the algorithm
  410. 15:52applied. If it's right, we take the
  411. 15:54output as our final result. Else we
  412. 15:57provide feedback to the training model
  413. 15:58and ask it to predict until it learns. I
  414. 16:02hope you've understood supervised and
  415. 16:03unsupervised learning. So let's have a
  416. 16:05quick quiz. You have to determine
  417. 16:07whether the given scenarios uses
  418. 16:09supervised or unsupervised learning.
  419. 16:11Simple, right? Scenario one. Facebook
  420. 16:13recognizes your friend in a picture from
  421. 16:15an album of tagged photographs.
  422. 16:19Scenario two, Netflix recommends new
  423. 16:21movies based on someone's past movie
  424. 16:23choices.
  425. 16:25Scenario three, analyzing bank data for
  426. 16:28suspicious transactions and flagging the
  427. 16:30fraud transactions. Think wisely and
  428. 16:32comment below your answers. Moving on,
  429. 16:34don't you sometimes wonder how is
  430. 16:37machine learning possible in today's
  431. 16:38era? Well, that's because today we have
  432. 16:40humongous data available. Everybody's
  433. 16:43online either making a transaction or
  434. 16:46just surfing the internet and that's
  435. 16:47generating a huge amount of data every
  436. 16:50minute and that data my friend is the
  437. 16:52key to analysis. Also, the memory
  438. 16:54handling capabilities of computers have
  439. 16:56largely increased which helps them to
  440. 16:58process such huge amount of data at hand
  441. 17:01without any delay. And yes, computers
  442. 17:04now have great computational powers. So
  443. 17:06there are a lot of applications of
  444. 17:08machine learning out there. To name a
  445. 17:10few, machine learning is used in
  446. 17:12healthcare where diagnostics are
  447. 17:13predicted for doctor's review. The
  448. 17:15sentiment analysis that the tech giants
  449. 17:17are doing on social media is another
  450. 17:19interesting application of machine
  451. 17:21learning. Fraud detection in the finance
  452. 17:23sector and also to predict customer
  453. 17:25churn in the e-commerce sector. While
  454. 17:27booking a cab, you must have encountered
  455. 17:29search pricing often where it says the
  456. 17:31fair of your trip has been updated.
  457. 17:33Continue booking. Yes, please. I'm
  458. 17:35getting late for office. Well, that's an
  459. 17:38interesting machine learning model which
  460. 17:40is used by global taxi giant Uber and
  461. 17:43others where they have differential
  462. 17:44pricing in real time based on demand,
  463. 17:47the number of cars available, bad
  464. 17:49weather, rush hour, etc. So they use the
  465. 17:51search pricing model to ensure that
  466. 17:54those who need a cab can get one. Also,
  467. 17:56it uses predictive modeling to predict
  468. 17:59where the demand will be high with a
  469. 18:01goal that drivers can take care of the
  470. 18:03demand and search pricing can be
  471. 18:05minimized. Great. Hey Siri, can you
  472. 18:07remind me to book a cab at 6 p.m. today?
  473. 18:10>> Okay, I'll remind you.
  474. 18:11>> Thanks.
  475. 18:12>> No problem.
  476. 18:13>> Artificial intelligence, machine
  477. 18:15learning, and deep learning represent
  478. 18:16the evolution of computer science
  479. 18:18towards creating intelligent systems. AI
  480. 18:21is the broader concept striving to build
  481. 18:23machines capable of humanlike
  482. 18:25intelligence. ML is a subset of AI
  483. 18:27emphasizing algorithms that learn from
  484. 18:29data to make predictions or decisions.
  485. 18:32DL in turn is a specialized branch of ML
  486. 18:35that employs deep neural networks to
  487. 18:37model complex patterns. Imagine an AI
  488. 18:39powered voice assistant like Apple Siri.
  489. 18:42It utilizes ML to understand and respond
  490. 18:45to user queries, learning from
  491. 18:47interactions over time. Deep learning
  492. 18:49comes into play when Siri recognizes
  493. 18:51speech patterns or interprets natural
  494. 18:53language using neural networks to
  495. 18:55process intricate features. The better
  496. 18:57it becomes at understanding diverse
  497. 18:59accents or refining responses
  498. 19:01exemplifying the continuous learning
  499. 19:03inherent in these technologies. AI seeks
  500. 19:06to emulate human intelligence. ML
  501. 19:08harness data for learning and DL employs
  502. 19:10deep neural networks for intricate task.
  503. 19:13The integration of these technologies
  504. 19:15manifest in everyday applications,
  505. 19:17transforming how we interact with and
  506. 19:20benefit from intelligent systems. This
  507. 19:22technology enables voice interaction,
  508. 19:24allowing the device to play music, set
  509. 19:26alarms, present audio books, and provide
  510. 19:28up-to-date information on topics like
  511. 19:30news, weather, sports, and traffic
  512. 19:32reports, etc. Let's move forward and see
  513. 19:35what is machine learning. Machine
  514. 19:37learning is a subset of artificial
  515. 19:38intelligence that focuses on developing
  516. 19:40algorithms and models capable of
  517. 19:42learning and making predictions or
  518. 19:45decisions without being explicitly
  519. 19:47programmed. ML systems leverage data to
  520. 19:49recognize patterns, adapt and improve
  521. 19:51their performance over time. There are
  522. 19:53several types of machine learning.
  523. 19:55Number one, supervised learning. The
  524. 19:57algorithm is trained on a label data set
  525. 19:59where each input is associated with a
  526. 20:01corresponding output. Number two comes
  527. 20:03as unsupervised learning. Unsupervised
  528. 20:05learning deals with unlabelled data to
  529. 20:07find inherent patterns or structures
  530. 20:09within the information. And then comes
  531. 20:11the reinforcement learning. This type
  532. 20:13involves training agents to make
  533. 20:15sequences of decisions by interacting
  534. 20:17with an environment. And then comes
  535. 20:19semi-supervised learning.
  536. 20:20Semi-supervised learning combines
  537. 20:22supervised and unsupervised learning
  538. 20:24elements typically using a small amount
  539. 20:26of labelled data and a larger pool of
  540. 20:28unlabelled data. Let us move forward and
  541. 20:30see what deep learning is. Deep
  542. 20:32learning, a branch of machine learning,
  543. 20:34focuses on algorithms inspired by the
  544. 20:36human brain structure and functionality.
  545. 20:38It excels in processing vast amounts of
  546. 20:40both structured and unstructured data.
  547. 20:42At the heart of deep learning are
  548. 20:43artificial neural networks, empowering
  549. 20:45machines to make decisions. The key
  550. 20:47distinction between deep learning and
  551. 20:49machine learning lies in data
  552. 20:50presentation. Machine learning
  553. 20:52algorithms typically demand structured
  554. 20:53data while deep learning networks
  555. 20:55operate through multiple layers of
  556. 20:57artificial neural networks allowing them
  557. 20:59to handle diverse data formats. So let's
  558. 21:02start with the difference between
  559. 21:03artificial intelligence, machine
  560. 21:04learning and deep learning. And this
  561. 21:07we'll show in a table form. So starting
  562. 21:09with the definition.
  563. 21:11So definition of artificial
  564. 21:13intelligence. So broad field of machine
  565. 21:16learning or creating machines with
  566. 21:18intelligent behavior is artificial
  567. 21:19intelligence. And when we talk about
  568. 21:21machine learning, it's the subset of AI
  569. 21:23focusing on algorithms learning from
  570. 21:25data. And then comes the deep learning
  571. 21:27that is specialized subset of ML using
  572. 21:29deep neural networks. And now we'll see
  573. 21:32the difference with the learning
  574. 21:33approach between all these three. So in
  575. 21:35learning approach artificial
  576. 21:37intelligence can include rulebased
  577. 21:39systems, expert system and more. And in
  578. 21:42machine learning, it learns from data
  579. 21:43patterns without explicit programming.
  580. 21:46And then comes the deep learning where
  581. 21:48it learns hierarchical representation
  582. 21:50using neural networks. And if we talk
  583. 21:52about scope, it encompasses various
  584. 21:54techniques beyond learning from data.
  585. 21:57And in machine learning, it primarily
  586. 21:58focus on learning patterns from data.
  587. 22:01And then the deep learning, it
  588. 22:02specifically utilizes deep neural
  589. 22:04networks for complex task. And now we'll
  590. 22:07move to the next difference. And we'll
  591. 22:09start with an example. So in artificial
  592. 22:12intelligence, the example is autonomous
  593. 22:14vehicles, chatboards or expert systems.
  594. 22:17And for machine learning, it's spam
  595. 22:18filters, recommendation systems, image
  596. 22:20recognition. And in deep learning it is
  597. 22:22image and speech recognition natural
  598. 22:25language processing. And now we'll see
  599. 22:27the difference for the data
  600. 22:29requirements. So it depends on the
  601. 22:31specific application and problem solving
  602. 22:32approach. And in machine learning it
  603. 22:34requires labeled or unlabelled data or
  604. 22:36training. And for the deep learning it
  605. 22:38relies on large amounts of labelled data
  606. 22:40for training deep networks. And now for
  607. 22:44the complexity artificial intelligence
  608. 22:46addresses a wide range of task including
  609. 22:48those beyond ML. And in machine
  610. 22:50learning, it deals with moderate to
  611. 22:52complex task depending on algorithms.
  612. 22:54And for the deep learning, it is well
  613. 22:56suited for intricate task often
  614. 22:58requiring substantial computational
  615. 23:00resources. And now see the flexibility.
  616. 23:03So for the artificial intelligence, it
  617. 23:06can be rule- based, evolving and
  618. 23:08adaptive. And for the machine learning,
  619. 23:10the flexibility adapts to patterns and
  620. 23:12the changes in data. And for the deep
  621. 23:14learning, it adapts to hierarchical
  622. 23:17representations and diverse data types.
  623. 23:20And then comes the training process. So
  624. 23:23in artificial intelligence, training
  625. 23:25process varies based on specific AI
  626. 23:27techniques used. And in machine
  627. 23:28learning, training involves feeding data
  628. 23:31and adjusting model parameters. And in
  629. 23:33deep learning, training involves
  630. 23:35optimizing neural weights and
  631. 23:37structures. And now we'll talk about the
  632. 23:39applications between all these three
  633. 23:40terms that is a IML and deep learning.
  634. 23:43So for artificial intelligence the
  635. 23:45applications are robotics, natural
  636. 23:47language processing, game playing and
  637. 23:49for machine learning it's predictive
  638. 23:51analytics, fraud detection and
  639. 23:53healthcare diagnostic and for the deep
  640. 23:55learning that is image recognition,
  641. 23:57speech synthesis and language
  642. 23:59translation.
  643. 24:00>> Now you guys must be thinking why should
  644. 24:02I consider a career in AI? Well AI is
  645. 24:05not just a passing trend. It's a seismic
  646. 24:08shift that is reshaping our world and
  647. 24:11creating new venues for innovation and
  648. 24:14discovery. Now by embracing a career in
  649. 24:16AI, you become a part of dynamic field
  650. 24:19that thrives on solving complex problem,
  651. 24:22pushing boundaries and making a profound
  652. 24:25impact on society. The demand for AI
  653. 24:28professionals is skyrocketing across the
  654. 24:30industries from healthcare, finance,
  655. 24:32entertainment, transportation.
  656. 24:34Organizations are actively seeking
  657. 24:37talented individuals who can harness the
  658. 24:39power of AI and drive their business
  659. 24:42forward. But what skills does it take to
  660. 24:44become an AI engineer? How can you
  661. 24:46embark on this thrilling journey? We
  662. 24:49have the answer to all your questions.
  663. 24:51Some steps are crucial to master the
  664. 24:53field of AI and become an AI engineer.
  665. 24:56Let's go through them real quick. So the
  666. 24:59first step is to establish a strong
  667. 25:01foundation in mathematics and
  668. 25:03programming. Start by gaining a solid
  669. 25:06understanding of critical mathematical
  670. 25:08concept such as linear algebra, calculus
  671. 25:12and probability theory. Additionally, it
  672. 25:14is crucial to become proficient in
  673. 25:16programming languages like Python which
  674. 25:19is commonly used in AI and develop
  675. 25:22coding skills. Next, you need to pursue
  676. 25:25a degree in relevant field. Earn
  677. 25:27bachelor's or master's degree in
  678. 25:29computer science, data science, AI or a
  679. 25:32related discipline to acquire a
  680. 25:34comprehensive understanding of AI
  681. 25:37principle and techniques and after that
  682. 25:39you need to acquire knowledge in machine
  683. 25:42learning and deep learning. Familiarize
  684. 25:44yourself with ML algorithms, neural
  685. 25:47network and deep learning frameworks
  686. 25:50like for example TensorFlow, PyTorch to
  687. 25:53train and optimize models using real
  688. 25:56world data sets and afterward engage in
  689. 25:59practical projects. Gain hands-on
  690. 26:02experience and demonstrate your skills
  691. 26:04by working on AI projects. Building a
  692. 26:08portfolio of projects that showcase your
  693. 26:10ability to solve AI problems can make a
  694. 26:13strong impression on potential
  695. 26:15employers. After that, collaborate and
  696. 26:18network. This is really important.
  697. 26:20Engage with AR communities, attend
  698. 26:23conferences, and participate in online
  699. 26:25forums to connect with professionals in
  700. 26:28this field. Collaborating with others
  701. 26:31can enhance your learning experience and
  702. 26:33open up new opportunities.
  703. 26:36Seek internships or entrylevel positions
  704. 26:39where you can gain practical experience
  705. 26:42through AI internships or entry-level
  706. 26:44roles in industry or research
  707. 26:47institution. Now this will provide
  708. 26:49valuable exposure and help you further
  709. 26:51develop your skills. After that
  710. 26:54continuously learn and adapt. In the
  711. 26:56fast-paced world of AR, it is very
  712. 26:59important to stay updated on new
  713. 27:01developments, explore specialized areas,
  714. 27:04and embrace emerging technologies and
  715. 27:06tools. Continual learning and
  716. 27:08adaptability are essential for pursuing
  717. 27:10a successful career as an AI engineer.
  718. 27:14Now that you're familiar with the steps
  719. 27:16involved in the journey of an AI
  720. 27:17engineer, let's discuss the essential
  721. 27:19skills you need to know to become an AI
  722. 27:22engineer. So, here's a breakdown of the
  723. 27:24skills needed. First one is having
  724. 27:26strong programming abilities. This
  725. 27:29typically refers to expertise in one or
  726. 27:32more programming languages commonly used
  727. 27:34in data science and machine learning
  728. 27:37such as Python or R language. Now,
  729. 27:40proficiency in programming allows you to
  730. 27:42write efficient and scalable code for
  731. 27:45data analysis, modeling and algorithm
  732. 27:48implementation.
  733. 27:49Next, you need knowledge of machine
  734. 27:51learning algorithms. This involves
  735. 27:53understanding and familiarity with wide
  736. 27:56range of machine learning algorithms
  737. 27:58including both supervised and
  738. 28:00unsupervised techniques. You should be
  739. 28:03able to select and apply appropriate
  740. 28:05algorithms for specific problems as well
  741. 28:08as evaluate and optimize their
  742. 28:10performance. Next skill is proficiency
  743. 28:13in statistics and mathematics. Sound
  744. 28:16knowledge of statistics and mathematics
  745. 28:18is fundamental for data analysis and
  746. 28:21machine learning. You should be
  747. 28:23comfortable with statistical concepts,
  748. 28:25hypothesis testing, regression analysis,
  749. 28:28probability theory, linear algebra and
  750. 28:31calculus.
  751. 28:32Now after that you have acquired a good
  752. 28:35amount of knowledge of these skill set,
  753. 28:37we'll move on to our next skill which is
  754. 28:39having familiarity with deep learning
  755. 28:42frameworks. Now deep learning has gained
  756. 28:44significant popularity in recent years
  757. 28:47and familiarity with deep learning
  758. 28:49frameworks like TensorFlow, PyTorch or
  759. 28:52Keras is valuable. Now these frameworks
  760. 28:55provide tools and libraries for
  761. 28:57building, training and deploying deep
  762. 28:59neural networks for tasks such as image
  763. 29:02recognition, natural language processing
  764. 29:04and time series analysis. Next, you need
  765. 29:08experience with big data technologies.
  766. 29:11Dealing with large scale data sets
  767. 29:13requires knowledge of big data
  768. 29:14technologies such as Apache, Hadoop,
  769. 29:17Spark or distributed computing
  770. 29:19frameworks. Understanding how to
  771. 29:22process, store and analyze data
  772. 29:24efficiently in distributed environments
  773. 29:26is very essential. Now after you have
  774. 29:28gotten experience with big data
  775. 29:30technologies, now it's the time to move
  776. 29:32on to our next skill which is having
  777. 29:34excellent problem solving and analytical
  778. 29:36skills. Now these skills will enable you
  779. 29:39to break down complex problems, identify
  780. 29:42key factors and develop efficient
  781. 29:44solution.
  782. 29:46Now you should be able to adapt at
  783. 29:48critical thinking, troubleshooting and
  784. 29:50debugging to handle real world
  785. 29:52challenges in data science and machine
  786. 29:54learning. So guys, remember to stay
  787. 29:56updated with the latest advancements in
  788. 29:59the field and continue learning to stay
  789. 30:01at the forefront of data science and
  790. 30:03machine learning. So that's all we had
  791. 30:06for you in this AI engineer road map. Do
  792. 30:08you know how AI has become so fast? It's
  793. 30:11now replacing entire teams in some
  794. 30:13industries. Yes, it's true. Over 50% of
  795. 30:17companies are already using AI to
  796. 30:19automate jobs. AI tools are writing
  797. 30:21emails, creating content, and even
  798. 30:24giving job interviews. And while some
  799. 30:26people are worried AI will take the job,
  800. 30:28I let you in on a secret. AI is also
  801. 30:31creating tons of highpaying roles. The
  802. 30:34catch, you need the right skills to get
  803. 30:37it. And that starts with learning with
  804. 30:39the right programming language. Now,
  805. 30:41I've tested a whole bunch of them.
  806. 30:44Python, C++, R, Java, you name it. And
  807. 30:48in this video, I'm breaking down the top
  808. 30:50five programming languages for AI that
  809. 30:54you need to know if you want to build a
  810. 30:56career, land real jobs, and actually
  811. 30:58stay relevant in the age of AI. We will
  812. 31:01cover what each language is best at, how
  813. 31:04to start learning, what kinds of AI jobs
  814. 31:06they lead to, and yes, how much you can
  815. 31:09earn with each one. All right, first up,
  816. 31:12we've got Python. And honestly, this one
  817. 31:14is the most valuable player of the AI
  818. 31:17development. Just like the star player
  819. 31:19in a sports team, Python is the go-to
  820. 31:22language that everyone relies on when it
  821. 31:26comes to building AI system. So, why is
  822. 31:28Python the AI king? Let me break it
  823. 31:30down. Simplicity and readability. Now,
  824. 31:34Python is super easy to learn. It's
  825. 31:36almost like writing in plain English.
  826. 31:38You don't have to worry about
  827. 31:40complicated code. If you're just
  828. 31:42starting out in programming, that is
  829. 31:44definitely the language you are going to
  830. 31:47feel most comfortable with. It's got
  831. 31:49this userfriendly vibe that makes it
  832. 31:52simple even for people new to coding.
  833. 31:54Second of all, it has got endless
  834. 31:56libraries. Now, Python is packed with
  835. 31:58tools. We call it libraries like
  836. 32:00TensorFlow, PyTorch and Scikitlearn.
  837. 32:03Think of these library as pre-made
  838. 32:06toolkits that make AI development way
  839. 32:09easier. They save you a lot of time
  840. 32:11because instead of building everything
  841. 32:13from scratch, you can use these
  842. 32:15libraries to quickly train your models
  843. 32:18and run algorithms. It's like having a
  844. 32:20shortcut to building AI system. It has
  845. 32:23also got rapid prototyping. If you need
  846. 32:26to test your ideas quickly, Python is
  847. 32:28perfect for that. You can build a model,
  848. 32:30a simple version of your AI system in no
  849. 32:33time. So whether you're working on
  850. 32:35machine learning models or neural
  851. 32:37networks, fancy word for AI system that
  852. 32:39learn like the brain, Python let you
  853. 32:42prototype or build a quick model fast.
  854. 32:45So I know you must be wondering now what
  855. 32:47kind of AI jobs can Python land me? It's
  856. 32:49a great question. With Python, you could
  857. 32:52land jobs like data scientist, machine
  858. 32:55learning engineer, or an AI researcher.
  859. 32:58Now these jobs typically pay between
  860. 33:00around six lakh to 15 lakh peranom.
  861. 33:03That's the salary range. But the best
  862. 33:05part is as you gain more experience and
  863. 33:07expertise that number will go way
  864. 33:10higher. So how do you start learning
  865. 33:12Python? You don't have to break the bank
  866. 33:15to learn Python. You can get started
  867. 33:17with free platforms and YouTube
  868. 33:19channels. Simply learn even offers a
  869. 33:22free comprehensive course in Python and
  870. 33:24I'll leave the link for you to check it
  871. 33:26out. And the best part is Python has got
  872. 33:29huge community. So if you ever feel
  873. 33:31stuck, there's always someone out there
  874. 33:34who's ready to help you. Next, we'll
  875. 33:36talk about C++. Now C++ isn't as
  876. 33:39beginner friendly as Python, but it's a
  877. 33:42beast when it comes to performance heavy
  878. 33:44applications. If you're working on
  879. 33:46realtime AI like self-driving cars or
  880. 33:49high frequency trading algorithms, then
  881. 33:52C++ is where you want to be. But why did
  882. 33:55we choose C++ for AI? First of all,
  883. 33:58because of its speed and efficiency.
  884. 34:00Now, C++ is all about its speed. It's
  885. 34:03the language you want when you're
  886. 34:05working with large data sets or AI
  887. 34:08applications that need to be super fast.
  888. 34:11Second of all, it has got lowlevel
  889. 34:13memory management. Now, C++ gives you
  890. 34:16full control over memory, which is
  891. 34:18essential when you're building AI system
  892. 34:20that require extensive computation and
  893. 34:23realtime performance. But isn't C++ more
  894. 34:26complex than Python? Definitely, yes.
  895. 34:29But if you're diving into AI
  896. 34:30applications that require high
  897. 34:32performance, think computer vision or
  898. 34:34robotics, C++ is unmatched. It's a bit
  899. 34:38trickier to learn, but if you want to
  900. 34:40build realtime AI systems, it's worth
  901. 34:43the effort. Roles like AI software
  902. 34:45developer or computer vision engineer
  903. 34:48are your goto with C++. The salary range
  904. 34:51for these roles is around 8 lakh to 20
  905. 34:54lakh peranom depending on the project's
  906. 34:56complexity and your experience. Third on
  907. 34:59a list is Java. This one's a workhorse
  908. 35:02in the world of AI. And if you're aiming
  909. 35:05to work on enterprise level AI projects,
  910. 35:09then Java is definitely a language you
  911. 35:11want to know. Now, it's not the first
  912. 35:13choice for small scale AI projects. But
  913. 35:16when it comes to big scalable systems,
  914. 35:19Java is untouchable. So why Java for AI?
  915. 35:22Because of its scalability. Now, you can
  916. 35:24think Java as a beast when it comes to
  917. 35:27handling large scale applications. If
  918. 35:29you're working on AI system that need to
  919. 35:32process huge data sets or manage complex
  920. 35:34computations, then Java can handle it
  921. 35:37all without breaking a sweat. It's
  922. 35:39designed to scale which makes it perfect
  923. 35:41for enterprise level AI projects where
  924. 35:44big data is involved. Platform
  925. 35:46independence. One of the best things
  926. 35:48about Java is its right ones run
  927. 35:51anywhere feature. It doesn't matter
  928. 35:53which platform you're using, whether
  929. 35:54it's Windows, Mac, Linux, Java can run
  930. 35:58all of it without issue. This is a huge
  931. 36:01win when you're building AI systems that
  932. 36:03need to operate across multiple
  933. 36:05platforms. Mature libraries. Java has
  934. 36:08been around for decades and because of
  935. 36:10that, it's packed with reliable
  936. 36:12libraries for AI. Libraries like Qua,
  937. 36:15H2O make implementing machine learning
  938. 36:17models or building AI system a lot
  939. 36:20smoother. These libraries we tried and
  940. 36:22tested so you know you're working with
  941. 36:24solid tools. Let's talk about what jobs
  942. 36:27can you actually land with Java. Now
  943. 36:29with Java you're looking at some big
  944. 36:32roles in the AI world. Think of AI
  945. 36:34solution architect or AI backend
  946. 36:37developer. These positions are not just
  947. 36:39highly respected but it also comes with
  948. 36:42a solid salary range typically between 7
  949. 36:45lakh to 18 lakh peranom. And with
  950. 36:47experience, well, let's just say that
  951. 36:50number can easily climb higher. Now, you
  952. 36:52must be thinking, how do I get started
  953. 36:54with Java? Now, if you're already
  954. 36:56familiar with object- oriented
  955. 36:58programming, learning Java will be a
  956. 37:00breeze. And if you're new to it, don't
  957. 37:02worry. You can start with some great
  958. 37:04resources like a YouTube channel or
  959. 37:07LinkedIn Learning. There are plenty of
  960. 37:09courses that will teach you how to use
  961. 37:11Java for AI from the ground up. Now,
  962. 37:14let's talk about R. This one's for all
  963. 37:17data science enthusiasts out there. If
  964. 37:20you're diving into statistical AI and
  965. 37:22the language built specifically for
  966. 37:24handling massive data and performing
  967. 37:26complex statistical analysis, R is your
  968. 37:29goto. So why R for AI? Because of its
  969. 37:32statistical power. R is packed with
  970. 37:35tools for statistical modeling. So if
  971. 37:37you're working on AI projects that need
  972. 37:39data analysis before you even start
  973. 37:41applying machine learning, R makes it a
  974. 37:44breeze. It's got everything you need for
  975. 37:47analyzing trends, finding patterns and
  976. 37:50building strong predictive models. It
  977. 37:53has also got a feature of its data
  978. 37:54exploration and visualization. One of
  979. 37:57the R's strength is its data exploration
  980. 38:00and visualization capabilities. You can
  981. 38:03easily plot, chart and analyze your data
  982. 38:05to uncover insights. This makes art
  983. 38:07perfect for the datadriven side of AI
  984. 38:10development where understanding your
  985. 38:12data is just as important as building
  986. 38:14the models. But wait, can I still work
  987. 38:16in AI if I learn R or is it just for
  988. 38:19data analysis? Absolutely. R is
  989. 38:22fantastic for AI projects that rely on
  990. 38:24statistical methods and data analysis.
  991. 38:27It's actually the language of choice for
  992. 38:29roles like AI data analyst or
  993. 38:31quantitative analyst where you'll be
  994. 38:33building predictive models or analyzing
  995. 38:35data trends to make decisions. Now these
  996. 38:38roles are in high demand and the salary
  997. 38:41range typically falls between 6 lakh to
  998. 38:4312 lakh peranom but with experience you
  999. 38:46can definitely push those numbers
  1000. 38:48higher. Now to get started with art, you
  1001. 38:50can find tons of free resources on a
  1002. 38:52plat new kid on the block that's growing
  1003. 38:55fast in the AI space. It's relatively
  1004. 38:58young compared to Python or C++. But
  1005. 39:01trust me, it's making a huge impact. And
  1006. 39:04here's why. Now, Julia was created back
  1007. 39:06in 2012 by a group of researchers who
  1008. 39:09wanted a programming language that could
  1009. 39:11handle the complex calculations required
  1010. 39:14for scientific computing and they nailed
  1011. 39:17it. But why did Julia grew so fast?
  1012. 39:20Well, it's been picking up speed because
  1013. 39:22it combines the performance of C++ with
  1014. 39:25the readability of Python. You get the
  1015. 39:28speed and efficiency that C++ is known
  1016. 39:30for, but with Python's clean and easy to
  1017. 39:33write code, it's like the best of both
  1018. 39:35worlds. Let's talk about why did we
  1019. 39:38choose Julia for AI? Because of its
  1020. 39:40speed and simplicity. Now, Julia's speed
  1021. 39:42is one of the biggest advantages. It's
  1022. 39:44designed for high performance computing.
  1023. 39:46So if you need to run complex AI models
  1024. 39:49or process tons of data, Julia will do
  1025. 39:52it in a fraction of the time it would
  1026. 39:54take in other languages. And the syntax,
  1027. 39:56it's also super easy to read and write.
  1028. 39:58So you're not sacrificing convenience
  1029. 40:00for performance. It has also got the
  1030. 40:02feature of scientific computing. Now
  1031. 40:04Julia is optimized for AI task like deep
  1032. 40:07learning and numerical analysis. You can
  1033. 40:09think AI applications in robotics, data
  1034. 40:12science, and scientific research. Now,
  1035. 40:14if you're working on projects that
  1036. 40:15require heavy computations or advanced
  1037. 40:18AI models, Julia's is your go-to. So,
  1038. 40:21why isn't everyone using Julia yet? It's
  1039. 40:24still growing, but Julia community
  1040. 40:26expanding rapidly, and more libraries
  1041. 40:29and frameworks are being developed every
  1042. 40:31day. It's quickly becoming a top choice
  1043. 40:33for high performance AI, and it
  1044. 40:35continues to evolve. And of course, I
  1045. 40:38expect to be even more popular. So, is
  1046. 40:40Julia better than Python or C++? Now the
  1047. 40:43answer is it depends. Now if you're
  1048. 40:44building scientific AI applications that
  1049. 40:47require high performance, Julia is a
  1050. 40:49fantastic option. It's still growing but
  1051. 40:51the community expands. Julia will
  1052. 40:53quickly become more powerful in the AI
  1053. 40:55space. Julia is perfect for roles like
  1054. 40:57AI developer in the scientific or
  1055. 40:59numerical computing space. Salaries can
  1056. 41:01range from 7 lakh to 15 lakh peranom
  1057. 41:04especially if you're working with
  1058. 41:06advanced AI. So guys there you have it
  1059. 41:08the five best programming languages for
  1060. 41:11AI. So whether you're interested in
  1061. 41:12machine learning, realtime AI or data
  1062. 41:15science, there's language for you. Each
  1063. 41:17of these will help you land AI job you
  1064. 41:20want and give you the tools you need to
  1065. 41:21build powerful AI system. Which one are
  1066. 41:24you going to start with? Drop your
  1067. 41:26thoughts in the comment section below
  1068. 41:27and let's talk about it. And if you
  1069. 41:29found this video helpful, hit that like,
  1070. 41:32share, and subscribe button to get more
  1071. 41:34AI tips and career advice by simply
  1072. 41:36learn. get started with the onboarding
  1073. 41:38and interface including the subscription
  1074. 41:40plan. As you can see here, it is
  1075. 41:42offering us three major plans. Now, now
  1076. 41:47there is a free version of manuals you
  1077. 41:49can use on a day-to-day basis. But make
  1078. 41:51sure to know that everyday credits are
  1079. 41:54six rupees.
  1080. 41:55>> But make sure everyday credits are only
  1081. 41:57300 to limited. But only 300 credits
  1082. 42:00will be assigned to you on a everyday
  1083. 42:02basis. Now, the first plan is $20 per
  1084. 42:05month, which gives you 300 fresh credits
  1085. 42:08every day, 4,000 credits per month,
  1086. 42:11in-depth research for everyday task,
  1087. 42:13professional website for standard
  1088. 42:14outboard, insightful slides for regular
  1089. 42:17content, task scaring, and wide
  1090. 42:19research, early access beta features,
  1091. 42:21and 20 concurrent task, 20 schedule
  1092. 42:24tasks. Now again if you are working in
  1093. 42:26an organization which where you can auto
  1094. 42:28you have to auto too many stuffs you can
  1095. 42:31upgrade to a $40 or $200 plan. Now $20
  1096. 42:35is for a person single usage because
  1097. 42:37it's only 300 fresh credits per day.
  1098. 42:39It's total of 4,000 per month as well.
  1099. 42:43So when it so when it comes to $40 plan
  1100. 42:46you can consider sharing it with two to
  1101. 42:48three people. Again it's 300 credits but
  1102. 42:518,000 credits per month. all the other
  1103. 42:54things plus plus you'll get an addition
  1104. 42:57of 4,000 more credits to work on. Now
  1105. 43:00when it comes to 200 you'll get a
  1106. 43:03firstly you'll get free cloud computing
  1107. 43:05where you don't have to worry about the
  1108. 43:06storage and stuff and here it is 40,000
  1109. 43:09credits per month an organization which
  1110. 43:12uses automation tools a lot more can use
  1111. 43:15this now we are going to start by
  1112. 43:18understanding the manusi interface and
  1113. 43:20the first thing we need to lock out at
  1114. 43:22the hub which is basically your main
  1115. 43:24dashboard. Now before we start giving
  1116. 43:26task to manus AI it is very important to
  1117. 43:29understand how credits works because for
  1118. 43:32many users credits can be a little
  1119. 43:34confusing in the beginning. When you
  1120. 43:36open the dashboard you will notice that
  1121. 43:38man's AI shows two different credit
  1122. 43:40counters. The first one is the daily
  1123. 43:43refresh credits. These are the credits
  1124. 43:45that refresh every day. For example you
  1125. 43:47may see around 300 credits per day. The
  1126. 43:50important thing to remember is that
  1127. 43:51these are use them or lose them credits.
  1128. 43:55That means they reset every 24 hours and
  1129. 43:58if you don't use them, they do not carry
  1130. 44:00forward in the next day. So these are
  1131. 44:02the daily credits which are good for
  1132. 44:04regular task, quick experiments, small
  1133. 44:07research work, testing prompt or even
  1134. 44:09trying out different features inside
  1135. 44:11Manusa. The second credit counter is
  1136. 44:14your monthly pool. This is your main
  1137. 44:17credit balance for the month. For
  1138. 44:18example, if you're on a standard plan,
  1139. 44:21you may need something like 4,000
  1140. 44:23monthly credits. These credits are more
  1141. 44:26useful for larger and more complex task.
  1142. 44:29So, if you ask manus AI to do something
  1143. 44:31longunning like researching a topic
  1144. 44:34deeply, creating a report, browsing
  1145. 44:36multiple sources, analyzing information,
  1146. 44:38or even completing a multi-step
  1147. 44:40workflow, then this monthly pool gives
  1148. 44:42you the main runway to complete those
  1149. 44:45bigger tasks. So just remember this
  1150. 44:47simple difference. Daily credits are for
  1151. 44:50everyday use and reset every 24 hours.
  1152. 44:53Monthly credits are your larger credit
  1153. 44:55pool for bigger tasks throughout this
  1154. 44:57month. Now the next important thing in
  1155. 44:59the dashboard is the active task window.
  1156. 45:02This is where manusci shows the tasks
  1157. 45:04that are currently running and this is
  1158. 45:06one of the most powerful parts of the
  1159. 45:08platform.
  1160. 45:09Unlike a normal chatbot where you can
  1161. 45:11ask one question wait for one answer,
  1162. 45:14Manos AI can work on multiple task at
  1163. 45:16the same time. For example, on this
  1164. 45:18subscription you can run up to 20
  1165. 45:21concurrent task at once. That means
  1166. 45:23manos can work on multiple request in
  1167. 45:26parallel. Maybe one task is researching
  1168. 45:29a topic, another is preparing a
  1169. 45:31document, another is analyzing a website
  1170. 45:33and another is organizing the
  1171. 45:35information. For the free users, the
  1172. 45:37limit is usually lower around five
  1173. 45:40concurrent tasks. But the important
  1174. 45:42thing is not just the number of tasks.
  1175. 45:45The important thing is that these tasks
  1176. 45:47are asynchronous and cloud-based. This
  1177. 45:49means once you start a task, Manus AI
  1178. 45:52continuously working in cloud. You do
  1179. 45:54not have to keep watching the screen the
  1180. 45:56entire time. You can start a task, close
  1181. 45:59the browser, disconnect from the
  1182. 46:01internet, and even come back later. and
  1183. 46:03Manus AI can still continue to work on
  1184. 46:07that task in the background. This is
  1185. 46:09what makes it feel less like a normal AI
  1186. 46:12chat port and more like an AI worker.
  1187. 46:14You're not just asking a question and
  1188. 46:16waiting for the reply. You're assigning
  1189. 46:18work, letting the agent process it, and
  1190. 46:21then checking the results once the task
  1191. 46:23is completed. So before using Minus AI
  1192. 46:26for real workflows, always understand
  1193. 46:28these three things. Your daily credits
  1194. 46:30reset every day. Your monthly credits
  1195. 46:33support bigger and longer tasks and your
  1196. 46:36concurrent task window shows how many
  1197. 46:38jobs Manus AI is currently handling for
  1198. 46:41you. Once you understand this dashboard,
  1199. 46:43it becomes much easier to manage your
  1200. 46:45credits, plan your task properly, and
  1201. 46:47use Manus AI more efficiently. Now that
  1202. 46:50we have understood the dashboard and
  1203. 46:52credits, let's move on to the next
  1204. 46:53important part of Manus AI interface,
  1205. 46:56which is the goal, input, and task
  1206. 46:58planning area. Now this is where you
  1207. 47:00actually start working with manus. In a
  1208. 47:03normal chatbot we usually give a small
  1209. 47:05instructions one by one. But in Manus AI
  1210. 47:07the idea is slightly different. Here you
  1211. 47:10give a highle goal and manus plans the
  1212. 47:13steps needed to complete that goal. So
  1213. 47:15in the main input box let's type a
  1214. 47:17simple goal such as research the top AI
  1215. 47:21tools for content creation and create a
  1216. 47:24comparison report. So let's start.
  1217. 47:26research the top AI tools for content
  1218. 47:30creation and create comparison report.
  1219. 47:35Now here we have assigned a proper goal
  1220. 47:37to Manus AI. Now notice what happens
  1221. 47:40after we enter this prompt. Manos does
  1222. 47:43not directly jump into the final answer.
  1223. 47:45First it create a task plan. This is
  1224. 47:48where you will see a to-do list or
  1225. 47:50step-by-step structure showing how Manus
  1226. 47:52is planning to complete the task. For
  1227. 47:55example, manus may break the goal into
  1228. 47:57steps like understanding the topic,
  1229. 47:59searching the AI content, creation
  1230. 48:01tools, collecting useful information and
  1231. 48:03comparing those tools and finally
  1232. 48:05preparing the report. So here you can
  1233. 48:07see the steps. In simple words, manus is
  1234. 48:09taking one big goal and breaking it into
  1235. 48:12smaller actions. This view is very
  1236. 48:14important because it gives us a chance
  1237. 48:15to review the plan before the agent
  1238. 48:18starts doing heavy work. Before manos
  1239. 48:20begins browsing, opening pages,
  1240. 48:22analyzing sources and consuming more
  1241. 48:25credits, we can quickly check whether
  1242. 48:27the plan looks correct. For example, in
  1243. 48:29this case, we should check is manners
  1244. 48:31searching for the right type of tooth.
  1245. 48:33Is it planning to compare them properly?
  1246. 48:36Is it going to create a final report as
  1247. 48:38we asked? If the plan looks correct, we
  1248. 48:41can continue. But if the plan looks
  1249. 48:42incomplete or slightly wrong, we can
  1250. 48:44stop and adjust the prompt before moving
  1251. 48:46forward. This helps us avoid wasting
  1252. 48:49time and credits. So the key point here
  1253. 48:51is simple. In manus AI, we don't need to
  1254. 48:54write every step manually. We can give
  1255. 48:57one clear goal and manus will create a
  1256. 48:59plan for completing it. But before
  1257. 49:02allowing the task to continue, always
  1258. 49:04review the documentation decomposition.
  1259. 49:07But before allowing the task to
  1260. 49:08continue, always review the
  1261. 49:10decomposition view. This helps you
  1262. 49:13understand how the agent is thinking and
  1263. 49:14whether it is moving in the right
  1264. 49:17direction. So in this example, our goal
  1265. 49:19was to research the top AI tools for
  1266. 49:22content creation and create a comparison
  1267. 49:24report. And Manus turns that single bowl
  1268. 49:27into the structured task plan that can
  1269. 49:29review before execution. This is what
  1270. 49:32makes Manus air different from a regular
  1271. 49:34chatbot. It does not just answer
  1272. 49:36immediately. It plans the work first,
  1273. 49:38shows the direction and then start
  1274. 49:40completing the task. Now that Manus has
  1275. 49:42understood our goal, the created task
  1276. 49:45plan, the next step is execution. It's
  1277. 49:48already executing. This is where Manus
  1278. 49:50AI actually starts working on a task.
  1279. 49:52You can think of this part as a hand in
  1280. 49:54the platform. The goal input is where
  1281. 49:56Manus understands what we want. The
  1282. 49:59planning view is where it decide how to
  1283. 50:01do it and the execution view is where it
  1284. 50:04actually performs the work. Once we
  1285. 50:06approve or continue with the task,
  1286. 50:08manage begins completing the steps one
  1287. 50:11by one. The interesting part is that we
  1288. 50:13can watch this happen in real time. On
  1289. 50:16one side, you will usually see the
  1290. 50:17progress list or task steps. This shows
  1291. 50:20that manus has completed what is
  1292. 50:23currently doing and what is still
  1293. 50:25remaining. The next is that you can see
  1294. 50:28manus actually taking action. For
  1295. 50:30example, if the task is repeat, for
  1296. 50:33example, if the task requires research,
  1297. 50:35you can see the agent opening websites
  1298. 50:36and browsing pages. If the task requires
  1299. 50:39collecting information, it may take
  1300. 50:41screenshots, extract details, or even
  1301. 50:44organize the data. If the task needs a
  1302. 50:46structured output, manuals may update a
  1303. 50:49spreadsheet, write content, run the
  1304. 50:51code, or even build an interactive
  1305. 50:53artifact. So instead of only showing the
  1306. 50:56final result, manos shows the workflow
  1307. 50:58while it's happening. This is useful
  1308. 51:00because we can understand how the agent
  1309. 51:02is working not just what answer it gives
  1310. 51:05to the end. Now another important thing
  1311. 51:07is to understand here is the sandbox
  1312. 51:10environment. Manos does not directly
  1313. 51:12operate your local computer. It works
  1314. 51:14inside a cloud and is created for the
  1315. 51:17task. Inside this sandbox, miners can
  1316. 51:20browse websites, collect information,
  1317. 51:21test the ideas, run code, fill forms and
  1318. 51:24build outputs without affecting your
  1319. 51:26personal system. For example, if we ask
  1320. 51:28miners to research AI tools and prepare
  1321. 51:30a vision report, it can browse different
  1322. 51:32website, collect the required details,
  1323. 51:34organize them and then create the final
  1324. 51:36report inside this workspace. And for
  1325. 51:39more advanced task, the sandbox can also
  1326. 51:42help maners create things like websites,
  1327. 51:44slide decks, spreadsheets, dashboards,
  1328. 51:46and other interactive files. This is one
  1329. 51:49of the major reasons maners feels
  1330. 51:51different from the normal chatbot. A
  1331. 51:53regular chatbot mostly gives a text
  1332. 51:55responses. But maners can actually
  1333. 51:58perform actions inside a controlled
  1334. 52:00environment. So while the task is
  1335. 52:02running, we should keep an eye on two
  1336. 52:04things. First the progress list to
  1337. 52:06understand which step manus is working
  1338. 52:09on. Second is realtime action view to
  1339. 52:12see what the agent is actually doing.
  1340. 52:14This makes the whole process more
  1341. 52:16transparent. You're not blindly waiting
  1342. 52:18for the final output. You can see the
  1343. 52:20agent browsing, checking information,
  1344. 52:22organizing the data and building results
  1345. 52:25step by step. So in simple terms, the
  1346. 52:28execution view shows manus in action.
  1347. 52:31The sidebyside workflow helps us track
  1348. 52:33the task in real time and the sandbox
  1349. 52:36environment gives Manus a safe cloud
  1350. 52:38workspace where it can browse, run code,
  1351. 52:41collect data and create useful outputs.
  1352. 52:43This is the part where Manus moves from
  1353. 52:45planning the work to actually doing the
  1354. 52:47work. So we'll get back to this task
  1355. 52:50once it is completed. Let's start with a
  1356. 52:52new task. Now that we have seen how
  1357. 52:54Manus works inside the browser, let's
  1358. 52:56look at how can Manus be on normal web
  1359. 53:00interface. Now here you can even connect
  1360. 53:02a different apps such as Gmail, browser,
  1361. 53:05meta and you can add other connectors as
  1362. 53:08well if you're planning to automate any
  1363. 53:10kind of workflow. Now here when you come
  1364. 53:12to the desktop side you will have a
  1365. 53:14mobile app as well as the desktop app as
  1366. 53:16well. Now if you come to settings you
  1367. 53:18may find an option called integration.
  1368. 53:20This is where you can actually connect
  1369. 53:23manos with platforms like slack,
  1370. 53:25telegram or even line. So as you can see
  1371. 53:28here there are connectors. This is
  1372. 53:30useful because it allows you to interact
  1373. 53:32with manus through a messaging apps you
  1374. 53:34already use. For example, instead of
  1375. 53:37opening the browser every single time,
  1376. 53:39you can just delegate a task, check the
  1377. 53:42progress or monitor updates from a
  1378. 53:45messaging app. So if you're working with
  1379. 53:46a team, Slack can be useful. If you want
  1380. 53:49quick mobile access, WhatsApp, Telegram
  1381. 53:52or Lion can make it easier to stay
  1382. 53:54connected with the agent. Part two manus
  1383. 53:57AI. Let's continue. The main benefit is
  1384. 54:00remote control. You can start monitoring
  1385. 54:02task even when there is no sitting in
  1386. 54:05front of the main system. Now the next
  1387. 54:07advanced feature is the desktop my
  1388. 54:09computer feature. You can download the
  1389. 54:12computer version here in the desktop
  1390. 54:14app. This is available when you have
  1391. 54:16Manus desktop app installed in your Mac
  1392. 54:19or a PC. Here Manus can request access
  1393. 54:22for your local machine for specific
  1394. 54:23action. For example, it may need you to
  1395. 54:26read a local file, open a folder or run
  1396. 54:28a terminal command. But the important
  1397. 54:31thing is to notice that manus does not
  1398. 54:33get a fully access automatically. There
  1399. 54:36are permissions grades. There is a
  1400. 54:38permission gate when manus wants to
  1401. 54:40perform an action on your computer. You
  1402. 54:43will see prompts like allow once or
  1403. 54:45allow always. From a safety point of
  1404. 54:48view, allow once means you are giving
  1405. 54:50permission only for that specific
  1406. 54:52action. Always allow means you are
  1407. 54:55allowing that type of action more
  1408. 54:56regularly depending on the setup. So
  1409. 54:59while showing this, this is especially
  1410. 55:01useful when you want manos to work with
  1411. 55:04files on systems, run scripts or even
  1412. 55:06help with local development tasks. Now
  1413. 55:08the third advanced area is the web app
  1414. 55:11builder. This is where manage becomes
  1415. 55:13even more powerful. In the web app
  1416. 55:15builder interface, you can see manage
  1417. 55:17generating a live interactive web
  1418. 55:19application. This is not just writing a
  1419. 55:21text or giving code snippets. It can
  1420. 55:23actually build pages, connect the
  1421. 55:25databases, structure the app and prepare
  1422. 55:27it while working with the project. For
  1423. 55:29example, if we ask manus to create a
  1424. 55:31simple landing page or a small web app,
  1425. 55:34it can generate a layout and add
  1426. 55:35interactive sections, connect the
  1427. 55:37required backend logic, and even support
  1428. 55:40things like database setup and SEO
  1429. 55:42optimization. The best part here is that
  1430. 55:45you can watch the agent work step by
  1431. 55:47step. You can see it creating files,
  1432. 55:50updating the design, testing the pages
  1433. 55:52and even improving the final output. So
  1434. 55:54this part is useful for users who want
  1435. 55:57to build something practical like
  1436. 55:58websites, dashboard, internal tool,
  1437. 56:00product page or even prototype without
  1438. 56:02manually writing everything line of code
  1439. 56:05from scratch. To summarize this section,
  1440. 56:08so now let's test the same logic. Now
  1441. 56:10let's ask minus AI to create a web
  1442. 56:13landing page for a skincare brand. So
  1443. 56:16create a brand. Now to summarize this
  1444. 56:19section, the browser is the main place
  1445. 56:21where you use manus AI which is this.
  1446. 56:24The messaging integration help you
  1447. 56:26delegate and monitor task remotely. The
  1448. 56:28desktop app gives you manus control
  1449. 56:30access to your local machine with
  1450. 56:32permission prompts. And the web app
  1451. 56:34builder helps manus create live
  1452. 56:36interactive web project. So this is what
  1453. 56:39takes manus from being just a web- based
  1454. 56:42AI agent to something that can connect
  1455. 56:44with your communication tool, your
  1456. 56:46computer and your real project works.
  1457. 56:48Now as you can see there are approaches
  1458. 56:50here. This is the code for the entire
  1459. 56:53web page. Let it generate. I'll show you
  1460. 56:55the output since this is just running in
  1461. 56:57the first step. There are more three
  1462. 56:59steps involved in this. So we'll get
  1463. 57:01back to this once this is done. Now we
  1464. 57:03are going to see where Manus AI becomes
  1465. 57:06really powerful which is deep research
  1466. 57:08and data. The main idea here is very
  1467. 57:11simple. Manus AI is not just a chatboard
  1468. 57:13that gives one quick answer. It can work
  1469. 57:16more like an autonomous research worker.
  1470. 57:18That means you can give it a goal and it
  1471. 57:21can plan the task, browse multiple
  1472. 57:22sources, collect the information, cross
  1473. 57:24the check details and organize the
  1474. 57:26findings and finally create a proper
  1475. 57:28output. So instead of manually opening
  1476. 57:3020 tabs and copying the nodes, checking
  1477. 57:32the resources and building the report
  1478. 57:34yourself can handle a larger part of
  1479. 57:36that workflow for you. Let's start with
  1480. 57:39a we'll just use a practical prompting
  1481. 57:42as of now. Now for this demo, you can
  1482. 57:44just type in research the top CRM tools
  1483. 57:47for small business and create a
  1484. 57:49comparison report with pricing, key
  1485. 57:51features, pros, cons, best use cases and
  1486. 57:54source link. So I've given the exact
  1487. 57:56same prompting. Now once we enter this
  1488. 57:59goal, manus first creates a plan. This
  1489. 58:01is important because the task is not
  1490. 58:03just asking for a simple answer. We are
  1491. 58:05asking manage to research multiple CRM
  1492. 58:08tools, compare them and prepare a
  1493. 58:09structured report. Once the task starts,
  1494. 58:12notice how manners does not depend only
  1495. 58:14on one search result. It begins visiting
  1496. 58:17different websites and checking product
  1497. 58:19pages, pricing pages, review platform,
  1498. 58:22blogs, and others available sources.
  1499. 58:24This is what we call multi-source
  1500. 58:27research. For example, if MinusAI is
  1501. 58:29researching CRM tools, it may check
  1502. 58:31official websites for pricing, review
  1503. 58:33platforms for user feedback and
  1504. 58:35comparison articles for feature level
  1505. 58:38difference. The important thing here is
  1506. 58:40that Manos is not just collecting random
  1507. 58:42information. It is trying to cross
  1508. 58:44interface the details. So if one website
  1509. 58:46mentions a price, Manos can compare it
  1510. 58:49with the official pricing page. If one
  1511. 58:52source mentions a feature, it can check
  1512. 58:53whether the same feature is also listed
  1513. 58:55on the product website. This helps
  1514. 58:57improve the quality of the research.
  1515. 59:00Now, while the agent is working, keep
  1516. 59:02your attention on realtime interaction
  1517. 59:04view. On one side, you can see the task
  1518. 59:07progress. On the other side, you see
  1519. 59:08minus browsing websites, opening pages,
  1520. 59:11taking screenshots, reading the
  1521. 59:12information, and updating its findings.
  1522. 59:14This makes the process more transparent.
  1523. 59:17You're not blind. You're not blindly
  1524. 59:19waiting for final answer. So you are
  1525. 59:21actually seeing how the agent is
  1526. 59:23collecting and organizing the
  1527. 59:24information. Another important thing is
  1528. 59:26to notice how manage handles small
  1529. 59:28problems during the search. Sometimes a
  1530. 59:30page may not open. Sometimes a link may
  1531. 59:33be broken. Sometimes a website may be
  1532. 59:35JavaScript heavy and difficult to read.
  1533. 59:38In manual workflow we would have stopped
  1534. 59:40and find another source assets. But
  1535. 59:43maners has planning layer that can
  1536. 59:45create recovery steps. So if one source
  1537. 59:48does not work, it can try another
  1538. 59:50source. search again or adjust the path
  1539. 59:52without needing constant human help.
  1540. 59:54This is why manners is useful for
  1541. 59:56research heavy tasks. At the end, the
  1542. 59:58output should not just be a paragraph
  1543. 1:00:00summary. A good result should be a
  1544. 1:00:03structured artifact like a comparison
  1545. 1:00:05table or a full research report. For the
  1546. 1:00:08CRM example, the final output can
  1547. 1:00:10include tools, names, pricing, key
  1548. 1:00:12features, pros, cons, best use cases,
  1549. 1:00:14and source links, which we'll check back
  1550. 1:00:16in a few minutes. If you can move beyond
  1551. 1:00:19one short answers and prefer a full
  1552. 1:00:22research workflow across multiple
  1553. 1:00:24sources, manusi is the tool. Now let's
  1554. 1:00:27move on to which is wide search. This is
  1555. 1:00:30more advanced credit intensive feature.
  1556. 1:00:33So we'll get back to all the three in a
  1557. 1:00:35minute. We'll get back to all the three
  1558. 1:00:38outputs and I explain what was the exact
  1559. 1:00:40steps required. So coming back to wide
  1560. 1:00:43research. In normal research, the agent
  1561. 1:00:45may explore sources step by step. But in
  1562. 1:00:47wide research, the idea is very
  1563. 1:00:49different. Wide research is designed for
  1564. 1:00:52scaling. Instead of checking a few
  1565. 1:00:54sources one after the other, it can
  1566. 1:00:56explore many sources in parallel. Manus
  1567. 1:00:59described wide research as using
  1568. 1:01:01parallel multi-agent orchestration where
  1569. 1:01:03many agents can work across large
  1570. 1:01:06research space at the same time. So this
  1571. 1:01:08is not meant for basic questions like
  1572. 1:01:10what is CRM or even give me five tools.
  1573. 1:01:14This feature is better for high impact
  1574. 1:01:16research tasks like market analysis,
  1575. 1:01:18competitive research, industry reports,
  1576. 1:01:20investment research, product research,
  1577. 1:01:23or even strategy planning. For example,
  1578. 1:01:25we can use a large version of the same
  1579. 1:01:28CRM topic. Run wide research on the CRM
  1580. 1:01:31software market for smaller businesses.
  1581. 1:01:33Compare major players, pricing, trends,
  1582. 1:01:36AI features, and even customer
  1583. 1:01:37sentiment, market positions, or even a
  1584. 1:01:39growth opportunities. This kind of
  1585. 1:01:41prompt is much broader. Here we're not
  1586. 1:01:44only asking for a tool comparison. We
  1587. 1:01:46are asking miners to understand the
  1588. 1:01:48market from a different angles. It may
  1589. 1:01:50explore companies, websites, review
  1590. 1:01:53sites, market reports, competitors,
  1591. 1:01:55pages, product documentation, user
  1592. 1:01:57discussions, and other public sources.
  1593. 1:02:00Now, before starting wide research,
  1594. 1:02:02always explain the credit part clearly.
  1595. 1:02:05This type of task can consume a lot more
  1596. 1:02:07credits than a normal research would.
  1597. 1:02:09Since wide research explores a large
  1598. 1:02:12number of sources and runs a much
  1599. 1:02:13heavier workflow, it can cost
  1600. 1:02:15significantly more credits. So we should
  1601. 1:02:17use it for important research work, not
  1602. 1:02:20for a simple Q&A. This is important for
  1603. 1:02:22learners. Think of it like hiring a full
  1604. 1:02:25research team for one task. You would
  1605. 1:02:27not only use them for a small
  1606. 1:02:29definition. You would use it when the
  1607. 1:02:31output has real business value. So the
  1608. 1:02:33main takeaway here is use normal
  1609. 1:02:36research for focused task. Use why
  1610. 1:02:38research when you need a large scale
  1611. 1:02:40high depth analysis across many sources.
  1612. 1:02:43Now next we'll move on to it is useful
  1613. 1:02:45because it shows how manuals can move
  1614. 1:02:47from raw data to a finished business
  1615. 1:02:49report. Here for example let's say let's
  1616. 1:02:52upload a CSV file. So here I've taken a
  1617. 1:02:55random data set from Kaggle and I've
  1618. 1:02:57uploaded it. It says loan data set. Now
  1619. 1:03:00let's give it a prompt saying analyze
  1620. 1:03:02this loan data and create a report
  1621. 1:03:04showing revenue trends, top performing
  1622. 1:03:06products etc. So let's just say analyze
  1623. 1:03:09this loan data and create a report
  1624. 1:03:14showing the trends. So mind you I have
  1625. 1:03:17already cleaned this data and executed
  1626. 1:03:20using AI which is in collab but still it
  1627. 1:03:24took me like proper an hour to create
  1628. 1:03:26it. So let's just leave it. Now as you
  1629. 1:03:29can see this is where maners becomes
  1630. 1:03:31different from normal AI tools. It does
  1631. 1:03:33not only look into the file and guess
  1632. 1:03:35the answer. It can work inside a cloud
  1633. 1:03:37sandbox. Inside this sandbox, miners can
  1634. 1:03:40write and execute code such as Python to
  1635. 1:03:43process the data. So if the file
  1636. 1:03:44contains thousands of rows, miners can
  1637. 1:03:46calculate totals, averages, trends,
  1638. 1:03:48category performance, product
  1639. 1:03:50performance, regional performance, and
  1640. 1:03:51other useful metrices. While this is
  1641. 1:03:54happening, show the executional view.
  1642. 1:03:56You may see man is reading the file,
  1643. 1:03:58writing the code, running analysis,
  1644. 1:04:00checking the output and generating
  1645. 1:04:01charts. This is manus is not producing
  1646. 1:04:04text. It is actually performing mini
  1647. 1:04:06data and this is workflow. After
  1648. 1:04:07processing the data, Manus can also
  1649. 1:04:09create visualization. For example, it
  1650. 1:04:12can generate charts showing loan
  1651. 1:04:14prediction data, which category will
  1652. 1:04:17take more loan, etc. Then the final
  1653. 1:04:19step, it can synthesize everything into
  1654. 1:04:21a business report. Now, what does a good
  1655. 1:04:24report include? what the data shows,
  1656. 1:04:26which products are performing well,
  1657. 1:04:28which areas need attention, which trends
  1658. 1:04:30are visible and what actions the
  1659. 1:04:32business should take next. So from one
  1660. 1:04:34uploading of file and one prompt, Manus
  1661. 1:04:37can complete an end to end workflow. It
  1662. 1:04:40can pass the data, run the code, create
  1663. 1:04:42charts, interpret the results and write
  1664. 1:04:44a final report. This is why Minus is
  1665. 1:04:47very useful for business users,
  1666. 1:04:49analytics, marketers, sales teams,
  1667. 1:04:51founders, and students learning data
  1668. 1:04:53analysis. Now let's see how Manis AI can
  1669. 1:04:56work on autonomous research worker. It
  1670. 1:04:59can browse multiple sources, analyze the
  1671. 1:05:00data, run code and prepare structured
  1672. 1:05:02reports. Now in this module we will see
  1673. 1:05:05manus AI as a creator. This is where
  1674. 1:05:07manus move from just giving answers to
  1675. 1:05:09creating finished functional artifacts.
  1676. 1:05:12So instead of only asking manus to
  1677. 1:05:14explain something, we can ask to build
  1678. 1:05:16something. It can create web apps, slide
  1679. 1:05:18text, posters, infographic, visual
  1680. 1:05:20content and also complete project assets
  1681. 1:05:23from a single natural language prompt.
  1682. 1:05:25So let's get started. So here let's give
  1683. 1:05:27manners a single prompt. Now let's ask
  1684. 1:05:30it to create a landing page for AI
  1685. 1:05:32productivity tools for students with
  1686. 1:05:34sections for features, pricing,
  1687. 1:05:35testimonials, FAQs, and call in action.
  1688. 1:05:38Now can you notice what happens here? We
  1689. 1:05:41are not giving miners a full design
  1690. 1:05:43document. We are not writing code. We
  1691. 1:05:45are not explaining very section step by
  1692. 1:05:48step. We're only giving it an idea.
  1693. 1:05:50Manus takes this idea, understands the
  1694. 1:05:52goal, creates a plan, decides the page
  1695. 1:05:55structure, writes the content, designs
  1696. 1:05:56the layout, and starts building the
  1697. 1:05:58page. This is important. Manus is not
  1698. 1:06:01just giving us text response. It is
  1699. 1:06:03creating a clickable portfolio. So, as
  1700. 1:06:06you can see, it already started creating
  1701. 1:06:08the This can be very useful for
  1702. 1:06:10developers, product managers, startup
  1703. 1:06:12founders, marketers, and business teams.
  1704. 1:06:14If someone has an idea and want to
  1705. 1:06:16quickly see how it might look at a
  1706. 1:06:19website, manners can help create the
  1707. 1:06:21first version very quickly. Instead of
  1708. 1:06:23spending hours preparing a wireframe or
  1709. 1:06:25explaining the idea to a designer or a
  1710. 1:06:28developer, we can just use maners to
  1711. 1:06:30create a rough working version. Then we
  1712. 1:06:33can share it with the team, client or
  1713. 1:06:35stakeholder for feedback. Depending on
  1714. 1:06:37the tunnels can also help with more
  1715. 1:06:39advanced parts like databases, payment
  1716. 1:06:42flow, SEO friendly structure and
  1717. 1:06:44deployment related steps. But for
  1718. 1:06:46beginners, the main thing is to
  1719. 1:06:48understand this manual can move from an
  1720. 1:06:51idea to a functional prototype. Also,
  1721. 1:06:54this type of task is more resource
  1722. 1:06:56inensive than a simple chat response.
  1723. 1:06:59Building a web page may consume hundreds
  1724. 1:07:01of credits. Sometimes around 500 to,000
  1725. 1:07:04or even more depending on the
  1726. 1:07:06complexity. So before running a web app
  1727. 1:07:09task, always check the estimated credit
  1728. 1:07:11usage. This is because minus is not only
  1729. 1:07:13writing text, it is planning, coding,
  1730. 1:07:16testing, building and sometimes handling
  1731. 1:07:19deployment steps as well. Now let's move
  1732. 1:07:21on to the next part. Now, now let's ask
  1733. 1:07:24manus to create a slide deck. Now let's
  1734. 1:07:28ask the manus AI to create a text on the
  1735. 1:07:30future of AI agents for business teams.
  1736. 1:07:33Now once we give this prompt, Manus
  1737. 1:07:35starts planning the slide tech. It does
  1738. 1:07:37not randomly create slide. It first
  1739. 1:07:39creates a proper structure. For example,
  1740. 1:07:41it may begin with an introduction, then
  1741. 1:07:43explain what AI agents are, why
  1742. 1:07:45businesses are using them, their
  1743. 1:07:47benefits, use cases, challenges, and
  1744. 1:07:49finally a conclusion. This is what makes
  1745. 1:07:51the output useful. It's not just a set
  1746. 1:07:53of separate slides. It's a structured
  1747. 1:07:56visual story for research heavy topics.
  1748. 1:07:58Miners can also browse the web, collect
  1749. 1:08:01useful information and include cited
  1750. 1:08:03points. This makes it useful for
  1751. 1:08:05business presentation, research decks,
  1752. 1:08:07pitch decks, training models, and
  1753. 1:08:09internal reports. While the task is
  1754. 1:08:11running, look at the interaction view.
  1755. 1:08:14You can see manners creating the
  1756. 1:08:15outline, preparing slide content,
  1757. 1:08:17improving the design, building the final
  1758. 1:08:19deck and once the deck is ready, you can
  1759. 1:08:22usually download in a businessfriendly
  1760. 1:08:24format like Pex. So we can still open it
  1761. 1:08:27in PowerPoint and make final manual
  1762. 1:08:29changes. Now this is very important
  1763. 1:08:31because manus gives us a strong first
  1764. 1:08:33version but we can still fine-tune it
  1765. 1:08:35the slides based on our brand audience
  1766. 1:08:37or even presentation style. Now let's
  1767. 1:08:39move on. Let's come back to this later.
  1768. 1:08:42Let's see what are the outputs for all
  1769. 1:08:44the prompting that we have given. So
  1770. 1:08:46firstly I have asked it to create a
  1771. 1:08:48landing page for a skincare brand. Now
  1772. 1:08:51as you can see there is a skincare brand
  1773. 1:08:53where you can also edit these. So the
  1774. 1:08:56name is given the benefits products
  1775. 1:08:59purifying tensor what is the cost in
  1776. 1:09:02dollars. You can edit the landing page.
  1777. 1:09:05This usually used to take days for an
  1778. 1:09:07UIUX designer to design the entire page.
  1779. 1:09:10is just done with a small prompt. So you
  1780. 1:09:12have products, reviews, shop now and if
  1781. 1:09:15you come here ready to transform your
  1782. 1:09:17skin, the shops, new arrivals, colle
  1783. 1:09:21collection, support, etc. This doesn't
  1784. 1:09:24look like it's just done from a
  1785. 1:09:27prompting. Now if you come to the second
  1786. 1:09:29one, let's we had asked to compare the
  1787. 1:09:33top CRM tools. Let's see what's the
  1788. 1:09:36answer for that. So here the prompt was
  1789. 1:09:39to research the top CRM tools for small
  1790. 1:09:42businesses and create a comparison
  1791. 1:09:44report with pricing, key features, pros,
  1792. 1:09:46cons, best use cases and course link. So
  1793. 1:09:48as you can see let's open this report.
  1794. 1:09:52So here we have a summary where small
  1795. 1:09:54business CRM section is the best
  1796. 1:09:56approach as a trade-off among these easy
  1797. 1:09:59tools. Now is it comparing all the
  1798. 1:10:01things that we have given? The first one
  1799. 1:10:03it's HubSpot sales hub starting price
  1800. 1:10:06key features pros cons best use cases
  1801. 1:10:09and source link all the things are
  1802. 1:10:11present usually if you use a person they
  1803. 1:10:15used to browse through every single
  1804. 1:10:16website they could find and create such
  1805. 1:10:19kind of report now it's done in just a
  1806. 1:10:22small prompt that I've given now let's
  1807. 1:10:24move on to the next one which is loan
  1808. 1:10:26data analysis this is the most useful
  1809. 1:10:28tool for data analyst because we spend
  1810. 1:10:31hours. They spend hours cleaning the
  1811. 1:10:34data, visualizing trends, what graphs is
  1812. 1:10:36suitable for what kind of data,
  1813. 1:10:38normalizing the data and so many other
  1814. 1:10:41steps. Now, if you can just upload a
  1815. 1:10:43file and ask it to create all the
  1816. 1:10:45reports and all the things necessary to
  1817. 1:10:47take a business decision, this will be
  1818. 1:10:50the most useful tool for data analysts.
  1819. 1:10:53Now, let's see the answer for this. As
  1820. 1:10:55you can see the graphs are there. Let me
  1821. 1:10:57just open. You can give a prompt where
  1822. 1:10:59which kind of graph you want, what
  1823. 1:11:01against what graph you want etc. You can
  1824. 1:11:04see the credit history, marital status,
  1825. 1:11:06property area, education,
  1826. 1:11:08self-employment and dependency all
  1827. 1:11:10against approval rate. Now if you come
  1828. 1:11:13here there is a summary as well which is
  1829. 1:11:15a report. Now the summary is that the
  1830. 1:11:18report analyzes 614 do applications
  1831. 1:11:20using the uploaded loan data set
  1832. 1:11:22covering applicants demographic income
  1833. 1:11:24co-licant income etc. The data set shows
  1834. 1:11:27422 applications were approved
  1835. 1:11:30presenting an overall 68% while 192
  1836. 1:11:33applications were rejected representing
  1837. 1:11:35a rejection rate of 31%. Now as you can
  1838. 1:11:38see we have approval and rejected rate
  1839. 1:11:42and the data overview. What are the
  1840. 1:11:45data?
  1841. 1:11:47Now here this is a very small data set
  1842. 1:11:49and I took almost an day to work with
  1843. 1:11:51this data set and create modeling etc.
  1844. 1:11:54This is done within a few minutes and
  1845. 1:11:56this is amazing because it takes a lot
  1846. 1:11:59of time cleaning the data set knowing
  1847. 1:12:01the data how to understand the data.
  1848. 1:12:05This sorts out all the problem. Now
  1849. 1:12:07coming to the next one content creation.
  1850. 1:12:10So here I had asked manusi to research
  1851. 1:12:13the top AI tools for content creation
  1852. 1:12:15and create a company report. So here as
  1853. 1:12:18you can see there is a report that is
  1854. 1:12:20given. Let's preview it. So here the
  1855. 1:12:23heading is there explore tools download
  1856. 1:12:25the report again this is treating as
  1857. 1:12:27like a website that has all the
  1858. 1:12:29information. So here you can see tool
  1859. 1:12:31distribution by category text generation
  1860. 1:12:34tool is like one etc. Pricing tier
  1861. 1:12:37distribution 63.2 to AI tools directly.
  1862. 1:12:41The first thing is chat GPT which is
  1863. 1:12:44probably mostly consumed I think. Next
  1864. 1:12:46is Jasper AI. Then we have Canva AI and
  1865. 1:12:49next Grammarly Ptory Morph AI Descript
  1866. 1:12:54Midjourney
  1867. 1:12:55Surfer SEO Gemini Claude Copy.ai etc. So
  1868. 1:12:59here you can see the price also what is
  1869. 1:13:02best suited for there is a free version
  1870. 1:13:04also. So it's given free version as well
  1871. 1:13:07rating. This is amazing for content
  1872. 1:13:10creation because usually we don't get
  1873. 1:13:12pictures which give the exact direction
  1874. 1:13:14or exact ratio of the exact numbers that
  1875. 1:13:17we found online. It's either we have to
  1876. 1:13:19create from scratch. So this is amazing
  1877. 1:13:22for content like you have ratings, you
  1878. 1:13:25have pricings, you have to compare them,
  1879. 1:13:28select the top tools to compare. Let's
  1880. 1:13:30compare chat chibity and Jasper sorry
  1881. 1:13:33Jasper and chat chibity and also Canva
  1882. 1:13:35AI all three are equally used. Now let's
  1883. 1:13:38deselect them and copy.ai. Now as you
  1884. 1:13:41can see copy.ai is a little bit less on
  1885. 1:13:43ratings. Oh my god this is too good to
  1886. 1:13:47be a tool. This is literally AI to work.
  1887. 1:13:50Let's check out the next one which is
  1888. 1:13:52landing page for AI productivity. Now
  1889. 1:13:55again this also will be a landing page.
  1890. 1:13:58So it's basically like a website. Now as
  1891. 1:14:01you can see we have the heading college
  1892. 1:14:02study flow features pricing testimonials
  1893. 1:14:04FAQs study flow AI is your personal AI
  1894. 1:14:07tutor study planner productive companion
  1895. 1:14:10get instead explanation organize your
  1896. 1:14:12listings there's a free trial watch a
  1897. 1:14:14demo powerful features for the success
  1898. 1:14:16everything you need to excel in your
  1899. 1:14:18studies all in one place etc. And you
  1900. 1:14:21have the pricing as well. This looks
  1901. 1:14:24like a legit platform website that has
  1902. 1:14:29no flaws. There is literally a review
  1903. 1:14:32rating also frequently asked questions
  1904. 1:14:35which is common in most of the websites.
  1905. 1:14:38And then you have the down at 2024 study
  1906. 1:14:41flow AI all rights reserved. Next let's
  1907. 1:14:44see if the slide deck is ready. Now
  1908. 1:14:48let's play the PPT. It's about the
  1909. 1:14:51future of AI agents for business. So as
  1910. 1:14:54you can see first is the heading
  1911. 1:14:57footages. The next one is core ship with
  1912. 1:15:00this AI agents change the unit of work.
  1913. 1:15:02And then you have what and all things
  1914. 1:15:04are changing. Why now the agent stack is
  1915. 1:15:09maturing. AI agents are not just smarter
  1916. 1:15:12chat bots. What are the difference
  1917. 1:15:13between chatbot co-pilot AI agent agent
  1918. 1:15:16portfolio? The new team model in human
  1919. 1:15:18agent collaboration business teams will
  1920. 1:15:21adopt agents by functions scale agents
  1921. 1:15:24required enterprise architecture
  1922. 1:15:26governance adoption. It's a legit PPT to
  1923. 1:15:29explain each and every single step of AI
  1924. 1:15:32agents future. Manus AI is literally
  1925. 1:15:35describing how AI is put to work not
  1926. 1:15:38just give a text response.
  1927. 1:15:40>> All right, so we're ready to start. uh
  1928. 1:15:42we are going to start with this you know
  1929. 1:15:45first course which is going to study the
  1930. 1:15:46basics of Python. Python will be our
  1931. 1:15:49primary focus for the entire program. Um
  1932. 1:15:52we will use co-pilot. So there's there
  1933. 1:15:56will be co-pilot material um later on in
  1934. 1:15:59the program but like in this first
  1935. 1:16:01course we're going to be focused on
  1936. 1:16:03Python and and for most um things we
  1937. 1:16:06will be using Python. Um even when we
  1938. 1:16:08use co-pilot it will produce Python code
  1939. 1:16:11everything we do will be in Python. I
  1940. 1:16:12think one of the things is by you know
  1941. 1:16:14by the end of the program if anything
  1942. 1:16:17else you guys will be in a much better
  1943. 1:16:19position with Python. You'll be better
  1944. 1:16:21Python coders by the end by the end of
  1945. 1:16:23the program. If you don't learn anything
  1946. 1:16:24else you'll get better at Python. I
  1947. 1:16:27promise. Uh because that's you know all
  1948. 1:16:29of our examples all of our demos
  1949. 1:16:32everything we do will be in Python. So
  1950. 1:16:34you'll you'll get better at it. uh for
  1951. 1:16:37sure and we'll have a lot of practice to
  1952. 1:16:39do that. Okay. So this first lesson is
  1953. 1:16:42all about an introduction to what Python
  1954. 1:16:45is. So if you're completely unfamiliar
  1955. 1:16:47with it, totally fine. We will uh get
  1956. 1:16:51you up to speed and talk about the
  1957. 1:16:53fundamentals and how to set everything
  1958. 1:16:55up on your own computer and talk about
  1959. 1:16:57the various ways to um utilize Python.
  1960. 1:17:01that some of it will involve a setup you
  1961. 1:17:04can do on your own computer. Some of it
  1962. 1:17:05will involve some cloud resources um so
  1963. 1:17:09that you don't need to set anything up
  1964. 1:17:11on your computer if you don't want to.
  1965. 1:17:13Um we'll have options there which will
  1966. 1:17:15be nice. So I will show us those and
  1967. 1:17:17walk us through those. But this first
  1968. 1:17:19lesson all about the basics uh and
  1969. 1:17:23getting set up. So, um what's
  1970. 1:17:27interesting is like at the beginning of
  1971. 1:17:28every lesson, we usually have this uh
  1972. 1:17:31kind of um engagement or discussion. Uh
  1973. 1:17:35but you know, we've I kind of already
  1974. 1:17:37asked you guys about this of uh uh if
  1975. 1:17:40you're familiar with programming, if
  1976. 1:17:41you're familiar with Python. Um but one
  1977. 1:17:44thing I want you to think about a little
  1978. 1:17:45bit is that um especially as we go along
  1979. 1:17:48and learn about what Python is is why is
  1980. 1:17:51Python the
  1981. 1:17:53chosen language for AI? So why is it the
  1982. 1:17:57one that everyone uses uh to do AI? And
  1983. 1:18:00I think what you're going to learn is
  1984. 1:18:03that it has a really amazing ecosystem
  1985. 1:18:08that has been around for a long time
  1986. 1:18:10that um supports AI in particular. So,
  1987. 1:18:15Python is the go-to for anything AI,
  1988. 1:18:19data science, machine learning, anything
  1989. 1:18:21in that sort. Uh, because it's been used
  1990. 1:18:25for so long for that and it has such a
  1991. 1:18:28uh community and ecosystem around it.
  1992. 1:18:31That's something we're going to learn.
  1993. 1:18:32It's also really easy to learn and use,
  1994. 1:18:36which makes it nice to to be uh kind of
  1995. 1:18:39an introduction to the field. It doesn't
  1996. 1:18:42take a lot to get started in it.
  1997. 1:18:44because it's so easy to work with. Um, I
  1998. 1:18:47can tell you as someone who's gone
  1999. 1:18:49through that experience, like I studied
  2000. 1:18:51mathematics in college and in graduate
  2001. 1:18:54school and studied like probability and
  2002. 1:18:56statistics, but I was able to teach
  2003. 1:18:59myself Python primarily and use that to
  2004. 1:19:02get into kind of data science and
  2005. 1:19:04machine learning in the industry.
  2006. 1:19:06So, and I think that's a common story is
  2007. 1:19:08people and I've seen that from many
  2008. 1:19:10learners coming from uh different
  2009. 1:19:12backgrounds. Uh they've been able to
  2010. 1:19:14pick up Python pretty easily because
  2011. 1:19:17it's a very easy language to understand
  2012. 1:19:19and and syntax of it and there's so many
  2013. 1:19:22tools within it that make it really easy
  2014. 1:19:24to work with.
  2015. 1:19:26So, um I promise it won't be as uh
  2016. 1:19:31daunting as it may seem even if you're
  2017. 1:19:33coming at it from zero experience. Uh, I
  2018. 1:19:36think you'll find this is the perfect
  2019. 1:19:39way to get into programming and get into
  2020. 1:19:41data science and and AI and machine
  2021. 1:19:44learning because it's so easy to pick up
  2022. 1:19:46and learn and it has such a nice rich
  2023. 1:19:48community ecosystem.
  2024. 1:19:51So, just wanted to mention that.
  2025. 1:19:55Okay. So, some of our objectives for
  2026. 1:19:57this first lesson will be to talk about
  2027. 1:20:00programming languages in general and um
  2028. 1:20:02programming in general. So maybe you
  2029. 1:20:05know more generic than Python just you
  2030. 1:20:08know what are what do general programs
  2031. 1:20:10look like? What are some of the building
  2032. 1:20:12blocks of programs that are important?
  2033. 1:20:15What are some of those uh key principles
  2034. 1:20:17of programming that we will want to
  2035. 1:20:20follow as well? Even if we're doing
  2036. 1:20:21Python for AI purposes.
  2037. 1:20:25Um so just talk about programming in
  2038. 1:20:27general and then kind of zoom in on
  2039. 1:20:29Python as we go along. One of the things
  2040. 1:20:31we'll be interested in doing is just
  2041. 1:20:33getting you guys set up. So talk about
  2042. 1:20:34how we can configure Python for you to
  2043. 1:20:37use on your own machine. Um but also
  2044. 1:20:40have some options that don't require
  2045. 1:20:41installing anything on your own machine.
  2046. 1:20:43Uh which is nice. Um and then as I said,
  2047. 1:20:46we'll kind of zoom in on Python, talk
  2048. 1:20:47about its benefits, uh some of the nice
  2049. 1:20:50features. I've kind of already mentioned
  2050. 1:20:51it. Really big community around it, easy
  2051. 1:20:53to learn. We'll just talk about those
  2052. 1:20:55more in detail. Talk about um why it's
  2053. 1:20:58so popular in the AI world. Um,
  2054. 1:21:02and then we'll get into some very
  2055. 1:21:03fundamental things specific to Python.
  2056. 1:21:05So once we talk about the background,
  2057. 1:21:07get you guys set up, we'll go into uh
  2058. 1:21:12some of the syntax basics, things like
  2059. 1:21:15identifiers, things like indentation,
  2060. 1:21:17comments, um, some of the basics of the
  2061. 1:21:20code that are going to be important for
  2062. 1:21:21you to kind of get started with. Um and
  2063. 1:21:24then talk about some of the basic data
  2064. 1:21:26types that Python offers to manipulate
  2065. 1:21:28and work with data which of course is
  2066. 1:21:30important um when you know as we go
  2067. 1:21:33forward and and do anything with data
  2068. 1:21:35which of course with AI we will be
  2069. 1:21:37interested in doing um but that's these
  2070. 1:21:41are the objectives of just the this
  2071. 1:21:42first lesson. As we go forward we're
  2072. 1:21:45going to learn about many other basic
  2073. 1:21:48topics within Python. So things like how
  2074. 1:21:51to write functions, how to build
  2075. 1:21:54objects, how to manipulate our flow of
  2076. 1:21:57the program with like things like if
  2077. 1:21:59else statements, things like loops.
  2078. 1:22:02We'll learn all about that in kind of
  2079. 1:22:04the next lessons after this one. But
  2080. 1:22:07this is all the content for this lesson.
  2081. 1:22:09I anticipate today
  2082. 1:22:12we will get through all of this today
  2083. 1:22:13and then get into the second lesson
  2084. 1:22:15which will um get into those kind of if
  2085. 1:22:19else and loops. So we'll we'll get we'll
  2086. 1:22:21I'm sure by today we'll get into those.
  2087. 1:22:24All right. Any questions on kind of what
  2088. 1:22:26we're going to learn in this first
  2089. 1:22:27lesson? So mainly trying to get you guys
  2090. 1:22:29set up, give you some background on
  2091. 1:22:30Python and then towards the end of the
  2092. 1:22:32lesson um get into some basics of the
  2093. 1:22:35syntax is kind of the goals I would say.
  2094. 1:22:38Okay. Okay. So when we talk about
  2095. 1:22:39programming um what do we mean by
  2096. 1:22:42programming in general? It's really uh
  2097. 1:22:44synonymous with instruction. So
  2098. 1:22:47programming really means giving or
  2099. 1:22:49writing instructions for a computer to
  2100. 1:22:52perform tasks. Um so these instructions
  2101. 1:22:56we write down in what we call code. But
  2102. 1:23:00those those are just telling the
  2103. 1:23:01computer what to do. And of course the
  2104. 1:23:03computer's not going to do anything
  2105. 1:23:05unless we write down these instructions.
  2106. 1:23:08So these instructions can do really
  2107. 1:23:11powerful things. They can power, you
  2108. 1:23:13know, whole applications, things that we
  2109. 1:23:15use every day like Microsoft Word,
  2110. 1:23:17PowerPoint, Excel, those kind of things.
  2111. 1:23:19Um they can automate tasks. They can um
  2112. 1:23:22power websites. Um they can do AI,
  2113. 1:23:26right? So we can have um things like
  2114. 1:23:28chat GBT and Alexa and Siri, etc., etc.
  2115. 1:23:32Um these are all powered by instructions
  2116. 1:23:35telling the computer what to do.
  2117. 1:23:37One of the things that we will get
  2118. 1:23:38better at as we go along is figuring out
  2119. 1:23:41how to write these instructions in
  2120. 1:23:42Python. Python is going to be the
  2121. 1:23:45language we write those instructions in
  2122. 1:23:48um and and they will be executed by a
  2123. 1:23:52Python um program. But we should think
  2124. 1:23:56of programming in general as just
  2125. 1:23:58instructing the computer what to do just
  2126. 1:24:01at a high level. Right?
  2127. 1:24:03So when we talk about these
  2128. 1:24:07instructions, they have two ways of
  2129. 1:24:10being executed by the the computer. Um
  2130. 1:24:15and roughly these break down into what
  2131. 1:24:17we call interpreted languages and
  2132. 1:24:20compiled languages. So that the code
  2133. 1:24:22that we write which is um representing
  2134. 1:24:25the instructions that we write can be
  2135. 1:24:28executed um in one of these two ways.
  2136. 1:24:32Let me start with the left. So the
  2137. 1:24:33interpreted languages.
  2138. 1:24:35This means that the computer is
  2139. 1:24:38literally executing the the instructions
  2140. 1:24:41line by line by line when we run the
  2141. 1:24:45program. So there is no
  2142. 1:24:49translation of anything. It's just
  2143. 1:24:51literally taking our instructions and
  2144. 1:24:53running it line by line, instruction by
  2145. 1:24:55instruction essentially. Um, now the
  2146. 1:24:59advantage to doing this is that it's uh
  2147. 1:25:03easier to debug because the instructions
  2148. 1:25:06are going to be executed one by one. So
  2149. 1:25:07it can hit an error pretty quick. If
  2150. 1:25:09there's a mistake in one instruction,
  2151. 1:25:12nothing else will run. Um, however, it's
  2152. 1:25:15also slower because we're going to take
  2153. 1:25:18it one instruction at a time. Um, and so
  2154. 1:25:22the the uh this way of running programs
  2155. 1:25:26tends to be slower, but it's also easier
  2156. 1:25:29to work with, which is why we're so
  2157. 1:25:32interested in Python. It's in this
  2158. 1:25:34bucket of what we call interpreted
  2159. 1:25:36languages. So a lot of scripting
  2160. 1:25:38languages find themselves in this bucket
  2161. 1:25:40of being executed one line at a time. No
  2162. 1:25:43translation needed by the machine. It
  2163. 1:25:45just reads our instructions and executes
  2164. 1:25:47it. The thing that does the execution is
  2165. 1:25:50called an interpreter.
  2166. 1:25:52Um, and Python has an interpreter that
  2167. 1:25:56we will get you guys set up with on your
  2168. 1:25:59own machine that can execute Python
  2169. 1:26:01code. So you need an interpreter. The
  2170. 1:26:04interpreter just executes your
  2171. 1:26:05instructions line by line by line. Um,
  2172. 1:26:08so some examples would be like Python.
  2173. 1:26:10That's what we're going to study in this
  2174. 1:26:12um entire program. But there's other
  2175. 1:26:14languages like JavaScript, Ruby,
  2176. 1:26:17um Pearl, many others that are uh
  2177. 1:26:21interpreted. They require an
  2178. 1:26:23interpreter, but they execute line by
  2179. 1:26:24line by line and there's no intermediate
  2180. 1:26:26translation of anything. Um it's kind of
  2181. 1:26:29executed as is. Now, contrast this with
  2182. 1:26:34compiled languages, which are uh kind of
  2183. 1:26:37a different piece. they these these
  2184. 1:26:40instructions have to be translated into
  2185. 1:26:42something the machine can understand in
  2186. 1:26:45order to execute. So there is an
  2187. 1:26:47intermediate step of what we call
  2188. 1:26:50compiling the code um into uh basically
  2189. 1:26:55a translated version of your
  2190. 1:26:57instructions so that the machine can
  2191. 1:26:59execute it. Now there's a trade-off
  2192. 1:27:01there. Doing that can make it more
  2193. 1:27:03difficult to develop and it can take
  2194. 1:27:05longer to debug because you have to go
  2195. 1:27:07through this translation step every
  2196. 1:27:09single time through the compiler.
  2197. 1:27:12But when you run the code because it's
  2198. 1:27:15already been translated into this
  2199. 1:27:16machine format, it's a lot faster. Um,
  2200. 1:27:20so some examples of languages like this
  2201. 1:27:22are C, C++, Java,
  2202. 1:27:25um, Go,
  2203. 1:27:27but uh, we won't really be working with
  2204. 1:27:29those. We'll just be sticking with
  2205. 1:27:31Python. But if you have experience with
  2206. 1:27:32those languages, those you're probably
  2207. 1:27:34familiar with this, you have to compile
  2208. 1:27:36the program first before you can execute
  2209. 1:27:38it. But we are going to be in this
  2210. 1:27:41interpreted world. If you know and it's
  2211. 1:27:43okay like if none of this makes sense,
  2212. 1:27:45that's okay. Just understand that um
  2213. 1:27:48generally interpreted languages are
  2214. 1:27:50going to be more user friendly because
  2215. 1:27:52they're they're easier to execute. They
  2216. 1:27:55don't require as many moving parts as
  2217. 1:27:58what a compiled language would require.
  2218. 1:28:01which is nice for us, right? Nice for
  2219. 1:28:02Python. That's what we're going to be
  2220. 1:28:04interested in working with. Uh kind of
  2221. 1:28:08um yeah, they're kind of rel So, so the
  2222. 1:28:10question is are JavaScript and Java
  2223. 1:28:12related? Kind of. Um, JavaScript is kind
  2224. 1:28:16of like the um the the
  2225. 1:28:19scripting version of um some of the same
  2226. 1:28:22concepts we see in Java, but Java is the
  2227. 1:28:25compiled um it it requires a a special
  2228. 1:28:29kind of what's called a Java runtime,
  2229. 1:28:31which is a a compiler to translate the
  2230. 1:28:35Java code into um machine code that the
  2231. 1:28:39Java runtime will execute. JavaScript is
  2232. 1:28:42not like that at all. It can actually be
  2233. 1:28:44ran in a web browser which is um
  2234. 1:28:47JavaScript usually powers a lot of like
  2235. 1:28:49front-end websites are usually powered
  2236. 1:28:52by JavaScript and Java usually powers
  2237. 1:28:54more like backend
  2238. 1:28:57um applications like actual software
  2239. 1:28:59programs are usually would be coded in
  2240. 1:29:01Java. JavaScript is going to be used
  2241. 1:29:04more for like building a website. But,
  2242. 1:29:06you know, I'm not an expert on that
  2243. 1:29:08really, but that's kind of my
  2244. 1:29:11understanding of it. And if anyone is an
  2245. 1:29:13expert on those differences, feel free
  2246. 1:29:15to let us know in the chat. But, uh,
  2247. 1:29:18that's my that's my basic summary of
  2248. 1:29:20that. Okay. So, we have interpreted
  2249. 1:29:24languages. That's where Python falls
  2250. 1:29:25under. So, it just um summarizing that,
  2251. 1:29:29it's going to be easier to work with
  2252. 1:29:30those, which is great for us. That's
  2253. 1:29:32another reason why Python's so easy.
  2254. 1:29:34It's interpreted, meaning that
  2255. 1:29:36everything executes. We don't need to
  2256. 1:29:37worry about compiling things, which is
  2257. 1:29:40nice. Um, but also in terms of
  2258. 1:29:45programming, there's also uh categories
  2259. 1:29:48of how the instructions are written that
  2260. 1:29:51you can bucket different languages into.
  2261. 1:29:53So for example um some language are are
  2262. 1:29:56more um procedural in nature meaning
  2263. 1:29:59that you write out all the instructions
  2264. 1:30:01exactly kind of line by line by line.
  2265. 1:30:04You don't really organize things at all
  2266. 1:30:06in your instructions.
  2267. 1:30:08Um so some examples would be like C and
  2268. 1:30:10Pascal
  2269. 1:30:11are more like that. Um then on the
  2270. 1:30:15opposite end of the spectrum is kind of
  2271. 1:30:16object-oriented
  2272. 1:30:18in which case you uh build your code and
  2273. 1:30:22organize it around the idea of
  2274. 1:30:24everything being an object. And so some
  2275. 1:30:26uh Python actually falls into this
  2276. 1:30:28category where um uh most things in
  2277. 1:30:31Python are objects and you manipulate
  2278. 1:30:34objects and objects have data to them.
  2279. 1:30:37They have things they can do and
  2280. 1:30:39interact with other objects. Um, so
  2281. 1:30:42think of it just as a way we will
  2282. 1:30:44organize our instructions.
  2283. 1:30:46Python allows us to organize it around
  2284. 1:30:48the concept of an object. We'll learn
  2285. 1:30:50about what that means as we go along,
  2286. 1:30:52but just realizing that some programming
  2287. 1:30:55languages break down along these um kind
  2288. 1:31:00of buckets here. Um, Python is also a
  2289. 1:31:03scripted language, meaning you can write
  2290. 1:31:06out your code in a individual script and
  2291. 1:31:09you can e that you can have an
  2292. 1:31:11interpreter that executes that script.
  2293. 1:31:13Um, so you don't need to organize all
  2294. 1:31:15your code inside of an object. So for
  2295. 1:31:18that reason, Python super flexible.
  2296. 1:31:22That's another reason why it's so nice
  2297. 1:31:24to use. It actually falls into both of
  2298. 1:31:26these buckets on the right, which is
  2299. 1:31:28very convenient. We can have basically
  2300. 1:31:30this means we can have a lot of
  2301. 1:31:32organization or very little organization
  2302. 1:31:34depending on how we want to set it up.
  2303. 1:31:37Yeah, Roberto. So even though there are
  2304. 1:31:39different types so Java is compiled and
  2305. 1:31:42Python is interpreted
  2306. 1:31:45um they are both object-oriented meaning
  2307. 1:31:48so think of the this slide as telling
  2308. 1:31:51you how the instructions are organized.
  2309. 1:31:55So how they are executed is different.
  2310. 1:31:57So, Java requires a a compiler to
  2311. 1:32:00execute things. Python requires an
  2312. 1:32:02interpreter.
  2313. 1:32:05This is more about how the instructions
  2314. 1:32:06are organized. So, Java and Python both
  2315. 1:32:10allow you to organize your code into
  2316. 1:32:11objects.
  2317. 1:32:13Um, but what's nice about Python is it
  2318. 1:32:16also falls under the bucket of
  2319. 1:32:17scripting, meaning that it allows you to
  2320. 1:32:20organize things into scripts, which is
  2321. 1:32:23less organization than it would be in
  2322. 1:32:25into objects. We're actually going to
  2323. 1:32:26learn about objects later on in a future
  2324. 1:32:29lesson, like how to build objects and
  2325. 1:32:31what they mean.
  2326. 1:32:34So yeah, even though they're different,
  2327. 1:32:35they're both object-oriented, which just
  2328. 1:32:37means that you can organize your code
  2329. 1:32:40into objects. Python allows that. So
  2330. 1:32:42does Java. So does C++.
  2331. 1:32:45Uh many many languages allow for um
  2332. 1:32:48organizing your your code into objects.
  2333. 1:32:52So we're going to learn about that.
  2334. 1:32:54It's It's not that one's better. They're
  2335. 1:32:57just um I I would put them at different
  2336. 1:33:00So, let me draw this. I would put them
  2337. 1:33:03at different spectrum, different ends of
  2338. 1:33:05the spectrum on organization.
  2339. 1:33:07So, scripting
  2340. 1:33:10is very loose. Basically, you it's more
  2341. 1:33:14like a an individual um uh set of
  2342. 1:33:18instructions to do one task. you can
  2343. 1:33:21just have and you can have many
  2344. 1:33:22individual scripts to do many small
  2345. 1:33:24tasks. Um, and then on the other end of
  2346. 1:33:27the spectrum, think about it as like
  2347. 1:33:29you've organized your cabinet into many
  2348. 1:33:32folders and many like uh you know many
  2349. 1:33:37pieces of organization that are we would
  2350. 1:33:39call objects. Um so objectoriented
  2351. 1:33:43programming OOP is kind of on the other
  2352. 1:33:46end of the spectrum when it comes to
  2353. 1:33:48like level level
  2354. 1:33:51of organization.
  2355. 1:33:56Does that make sense? So scripting very
  2356. 1:33:58loose. It usually scripting is is um
  2357. 1:34:01reserved for like one task and it's um
  2358. 1:34:03you're just writing out your
  2359. 1:34:04instructions to accomplish that one
  2360. 1:34:06task.
  2361. 1:34:07um which is helpful for like automation
  2362. 1:34:09of things because you're going you're
  2363. 1:34:11usually automating like a single task.
  2364. 1:34:14Um so it's very loose. It's not very
  2365. 1:34:16organized into nothing is organized
  2366. 1:34:17necessarily into objects. Um very loose
  2367. 1:34:20organization. Object-oriented is much
  2368. 1:34:24more structure to it and things being
  2369. 1:34:26put into objects um in order to
  2370. 1:34:28manipulate and work with objects
  2371. 1:34:30throughout the program. Yeah, it's not
  2372. 1:34:33that one's better. I think it's more
  2373. 1:34:35just use case dependent. Um there are
  2374. 1:34:38times where it actually will benefit us
  2375. 1:34:41from using objects. Um and I think the
  2376. 1:34:45thing to pay attention to on this slide
  2377. 1:34:46is that look at where Python falls into.
  2378. 1:34:49It actually falls into both. Meaning
  2379. 1:34:51that we can have things very loose and
  2380. 1:34:54easy to work with because scripting
  2381. 1:34:56usually will be faster and easier to
  2382. 1:34:58just write something to to accomplish
  2383. 1:35:00one task. But we have the flexibility to
  2384. 1:35:03organize our code into objects if we
  2385. 1:35:05want to. which will be better for
  2386. 1:35:08bigger tasks that require more
  2387. 1:35:11organization
  2388. 1:35:13like training a neural network or
  2389. 1:35:16building an LLM.
  2390. 1:35:18Those bigger tasks would benefit from
  2391. 1:35:20organization.
  2392. 1:35:23And then uh finally on this slide um
  2393. 1:35:26there are languages that are built on
  2394. 1:35:27the concept of um their their entire way
  2395. 1:35:31of writing instructions is more in a
  2396. 1:35:33functional way meaning everything is
  2397. 1:35:35based on operating uh functions and
  2398. 1:35:38variables. Um and so there are some
  2399. 1:35:41languages like that has and scholar are
  2400. 1:35:43very popular ones. Um but that is can be
  2401. 1:35:49very difficult to learn. It's it can be
  2402. 1:35:51difficult but very nice in some ways
  2403. 1:35:53because uh it can be very natural to
  2404. 1:35:56think of um you manipulate like giving
  2405. 1:35:59instructions to computer in a functional
  2406. 1:36:01way. Think about it as like applying a
  2407. 1:36:03function to a variable.
  2408. 1:36:05Um that makes sense but writing your all
  2409. 1:36:09of your instructions in that way can be
  2410. 1:36:10kind of difficult to learn. So for that
  2411. 1:36:13reason I think these languages are more
  2412. 1:36:15difficult to learn but they can be very
  2413. 1:36:17powerful. Um and they find themselves
  2414. 1:36:20very useful in like operating on big
  2415. 1:36:23data.
  2416. 1:36:24Um so if you ever heard of like Spark um
  2417. 1:36:27Spark operates with uh Scola for
  2418. 1:36:30instance um but uh we won't really focus
  2419. 1:36:34on functional. It's kind of its own
  2420. 1:36:36paradigm.
  2421. 1:36:38Um but uh again like Python is where our
  2422. 1:36:42focus will be. It allows us to be really
  2423. 1:36:45organized, loosely organized. Nice
  2424. 1:36:48flexibility there.
  2425. 1:36:51So, so far
  2426. 1:36:53based on these two slides, I'm showing
  2427. 1:36:54you that Python is interpreted, which is
  2428. 1:36:57easier and faster to work with. Um, not
  2429. 1:37:00faster to run, but faster to get up and
  2430. 1:37:02running because you don't need to
  2431. 1:37:03compile things. That's nice from our
  2432. 1:37:06perspective.
  2433. 1:37:08And it's also has very good flexibility
  2434. 1:37:11when it comes to organizing our
  2435. 1:37:12instructions, organizing our code can be
  2436. 1:37:14very loose in scripts, could be very
  2437. 1:37:17structured in in objects.
  2438. 1:37:21Okay. Okay. So generally no matter how
  2439. 1:37:23uh no matter what language it is um when
  2440. 1:37:26you process those instructions generally
  2441. 1:37:29things are going to be organized
  2442. 1:37:32even if it's in a script or if it's
  2443. 1:37:33object-oriented
  2444. 1:37:35um you're generally going to have the
  2445. 1:37:38very beginning of the program um kind of
  2446. 1:37:40setting up the input then the middle of
  2447. 1:37:42it really processing that and doing
  2448. 1:37:45something with that. So that's usually
  2449. 1:37:46like the bulk of the logic is in the
  2450. 1:37:49processing phase and then generally
  2451. 1:37:51you're producing some output. So that
  2452. 1:37:53could be like a model prediction, that
  2453. 1:37:56could be um a a graph that you've built
  2454. 1:37:59from your code um whatever that output
  2455. 1:38:02is. But generally it flows this way.
  2456. 1:38:04This is this is makes sense, right? Of
  2457. 1:38:06course there's input, you're
  2458. 1:38:08manipulating that input in some way and
  2459. 1:38:10then you're producing some output. I
  2460. 1:38:11think that all makes sense. That's a
  2461. 1:38:13very logical way to flow.
  2462. 1:38:15Um
  2463. 1:38:17now that's not to say that within this
  2464. 1:38:19processing step there may not be
  2465. 1:38:22um iteration like of course there may
  2466. 1:38:25may be times where we need to as part of
  2467. 1:38:28the processing kind of iterate and do
  2468. 1:38:30multiple passes of processing. Um so the
  2469. 1:38:33processing could be a lot. We could be
  2470. 1:38:35doing a lot. We could be doing a little.
  2471. 1:38:37Just depends on what we're actually
  2472. 1:38:38doing. So, if we're reading in some data
  2473. 1:38:41as the input um and then we're just
  2474. 1:38:44doing some simple um slicing and dicing
  2475. 1:38:47of it, that's some easy processing and
  2476. 1:38:49maybe producing a graph or producing a
  2477. 1:38:51metric, something of that sort, that's
  2478. 1:38:53pretty easy to do. But if we're training
  2479. 1:38:56a neural network or training a model,
  2480. 1:38:59the processing step can take a while and
  2481. 1:39:01it may, you know, be very iterative in
  2482. 1:39:03nature. So it just depends on what we're
  2483. 1:39:06doing and those instructions.
  2484. 1:39:08But no matter what, most of our programs
  2485. 1:39:11will flow in this way kind of input
  2486. 1:39:14processing output. It makes sense. It's
  2487. 1:39:16very logical.
  2488. 1:39:19So what are some principles that we
  2489. 1:39:22should abide by when we're writing our
  2490. 1:39:24code? So this this would really be for
  2491. 1:39:26any language, but of course for Python
  2492. 1:39:28that we are interested in. Um so
  2493. 1:39:31something we're going to be interested
  2494. 1:39:32in doing is um basically avoiding
  2495. 1:39:36repetition where we can. So instead of
  2496. 1:39:39having copy paste everywhere, we will
  2497. 1:39:42generally favor organizing our code to
  2498. 1:39:45some degree. Meaning we will utilize
  2499. 1:39:47functions where it makes sense and
  2500. 1:39:49objects where it makes sense to organize
  2501. 1:39:50things. And also instead of um having
  2502. 1:39:54very repetitive code, we will favor
  2503. 1:39:57using uh loop structures that can
  2504. 1:40:00iterate over um things many times
  2505. 1:40:04instead of us us having to write all
  2506. 1:40:06those out one by one by one. So we're
  2507. 1:40:08going to learn about these tools that we
  2508. 1:40:10have at our disposal, but they will help
  2509. 1:40:12us organize our code, avoid repetition
  2510. 1:40:15all over the place. One of the things we
  2511. 1:40:17want to avoid is having the same code
  2512. 1:40:21repeated all over the place. If if we
  2513. 1:40:24find ourselves doing that, we should
  2514. 1:40:25really put that code into a function or
  2515. 1:40:28maybe into an object so that we can
  2516. 1:40:29reuse it. So, we're really going to
  2517. 1:40:32favor like reusability of things,
  2518. 1:40:36re recycle, reuse, you know. So, we're
  2519. 1:40:39going to learn how to do that, how to
  2520. 1:40:42build functions, how to build objects.
  2521. 1:40:44But that's something we're going to
  2522. 1:40:44favor uh when we're when we're
  2523. 1:40:46programming. It's something you should
  2524. 1:40:47be on the lookout for. If you find
  2525. 1:40:50yourself writing the same code over and
  2526. 1:40:52over just in different spots, um that's
  2527. 1:40:55probably a clue you should organize that
  2528. 1:40:57into a function so you can just call
  2529. 1:40:58that function wherever you need to
  2530. 1:41:00rather than copying all that code. Okay,
  2531. 1:41:03so we're going to avoid repetition.
  2532. 1:41:05Now, the the reason we're going to do
  2533. 1:41:06that is to uh you know keep everything
  2534. 1:41:11simple. We want to make sure things are
  2535. 1:41:13clean, simple, understandable.
  2536. 1:41:16Um, we don't want to h we don't want to
  2537. 1:41:18have overly complex things that are very
  2538. 1:41:21difficult to follow. So, one of the
  2539. 1:41:23things that is going to be really nice
  2540. 1:41:25about Python is it lends itself very
  2541. 1:41:27well to being simple because it's going
  2542. 1:41:31to be so easy to actually read and
  2543. 1:41:33understand um, you know, understand
  2544. 1:41:36what's going on. But one of the things
  2545. 1:41:38that falls in line with this is like um
  2546. 1:41:41for instance naming things
  2547. 1:41:43appropriately. So instead of just
  2548. 1:41:45calling everything in our code like X Y
  2549. 1:41:47and Z if somebody comes along and reads
  2550. 1:41:50oh I see your code has an X Y and Z that
  2551. 1:41:53may not make sense. You know we would
  2552. 1:41:56want to be more thoughtful with the
  2553. 1:41:58names of our variables and names of our
  2554. 1:42:00function. So instead of XYZ, maybe we
  2555. 1:42:02would use something like name or place
  2556. 1:42:06or you know something appropriate to
  2557. 1:42:08identify this is what this is. So think
  2558. 1:42:12about that when you're writing your code
  2559. 1:42:14is try to make it understandable. Name
  2560. 1:42:17things that somebody else reading it
  2561. 1:42:20would understand what it is if they see
  2562. 1:42:21that name. So that's that's a mistake I
  2563. 1:42:24see a lot of people make when they first
  2564. 1:42:25start. It's okay like when you're first
  2565. 1:42:27getting started and practicing to name
  2566. 1:42:28things like X, Y, and Z. I think that's
  2567. 1:42:30fine. Or like ABC.
  2568. 1:42:32Um,
  2569. 1:42:34but does that make sense? Like if
  2570. 1:42:35somebody else was reading it, they see
  2571. 1:42:38XYZ in the program, that may not make
  2572. 1:42:40sense, you know? So, but if it has a
  2573. 1:42:43good name to it, you could say, oh, like
  2574. 1:42:45I see this is somebody's name that this
  2575. 1:42:47variable is referring to or this is um a
  2576. 1:42:50particular object that this is referring
  2577. 1:42:51to. Um, it's not just kind of an
  2578. 1:42:54abstract X or Y or Z.
  2579. 1:42:58Yeah, no spaghetti. Yeah, that's that's
  2580. 1:43:02what uh that's what a lot of people
  2581. 1:43:04refer to that as. Uh just sloppy,
  2582. 1:43:07unorganized, um hard to understand code.
  2583. 1:43:12One of the things that's great about
  2584. 1:43:13Python is it's naturally very
  2585. 1:43:15understandable. So like I don't think we
  2586. 1:43:17will have that issue as much as if we
  2587. 1:43:19had other languages, but it's still
  2588. 1:43:22possible.
  2589. 1:43:23So these are things we'll learn as we go
  2590. 1:43:25along. I'm just trying to get it into
  2591. 1:43:27your mind a little early here. Name
  2592. 1:43:29things appropriately is main one of the
  2593. 1:43:31main pieces of advice I can give here.
  2594. 1:43:35Um
  2595. 1:43:36so the next tip is to organize things.
  2596. 1:43:39This goes along with avoiding
  2597. 1:43:40repetition. So organize
  2598. 1:43:43um let's put things into functions.
  2599. 1:43:45Let's put things into objects where it
  2600. 1:43:46makes sense. If we know we're going to
  2601. 1:43:48reuse that um let's put it into a
  2602. 1:43:51function. And so we're going to learn
  2603. 1:43:52about how to do that. But generally this
  2604. 1:43:55is good practice if you find yourself
  2605. 1:43:58writing um uh code to do something and
  2606. 1:44:02it turns out to be um
  2607. 1:44:06it turns out to be uh something you know
  2608. 1:44:08you're going to reuse or it turns out to
  2609. 1:44:10be more than a handful of lines of code.
  2610. 1:44:14Generally you want to organize that into
  2611. 1:44:15a function so that uh it's clear
  2612. 1:44:19this is what this code is doing. This is
  2613. 1:44:21what it's responsible for. it's obvious
  2614. 1:44:24um you know that it's organized into
  2615. 1:44:27into uh that unit of work essentially.
  2616. 1:44:32So we are going to practice this. This
  2617. 1:44:34is something we're going to get good at
  2618. 1:44:35I think as we go along because we're
  2619. 1:44:37going to favor organization where it
  2620. 1:44:40makes sense.
  2621. 1:44:42Okay.
  2622. 1:44:44So readability. One of the things is
  2623. 1:44:46using meaningful names. I kind of
  2624. 1:44:48already mentioned that. The other thing
  2625. 1:44:50is using good comments. So, we're going
  2626. 1:44:52to learn probably today how to make
  2627. 1:44:54comments in our Python code, which is
  2628. 1:44:56going to be helpful to orient yourself
  2629. 1:44:58or another reader of it to, hey, this is
  2630. 1:45:01what this function does. This is what
  2631. 1:45:03this line of code is doing. Um, I can't
  2632. 1:45:06tell you how many times, you know,
  2633. 1:45:07people write code and then it they
  2634. 1:45:09themselves come back to it a week later
  2635. 1:45:12and have no idea what it's doing. That
  2636. 1:45:15happens all the time. It's even happened
  2637. 1:45:16to me. So, uh, comments are your friend
  2638. 1:45:20in that regard. and that um they don't
  2639. 1:45:22really cost you anything to put comments
  2640. 1:45:23in there um to say to to kind of
  2641. 1:45:27highlight this is what this piece of
  2642. 1:45:30code is doing and you can make a note to
  2643. 1:45:32yourself right within the code. That's
  2644. 1:45:34what comments are. They're basically
  2645. 1:45:36notes to yourself. Um
  2646. 1:45:40so we're going to learn about that
  2647. 1:45:41today. How to write comments and and
  2648. 1:45:43what that looks like in the code. The
  2649. 1:45:46other thing is indentation. you know,
  2650. 1:45:48Python supports uh indent like you have
  2651. 1:45:50to indent. So, that's not really going
  2652. 1:45:52to be an issue. Some languages don't
  2653. 1:45:54really support that, especially the
  2654. 1:45:56compiled ones. They don't enforce
  2655. 1:45:58strictly indentation. They enforce other
  2656. 1:46:00things like braces and and semicolons
  2657. 1:46:03and such, but um our our Python code
  2658. 1:46:07will be properly indented uh by
  2659. 1:46:10necessity because otherwise it won't
  2660. 1:46:12work. So, um, that's something we're
  2661. 1:46:13going to learn about too today is how we
  2662. 1:46:16indent things and why that matters.
  2663. 1:46:19We'll talk about that.
  2664. 1:46:23Um,
  2665. 1:46:24I see a question from Sherry. Is Python
  2666. 1:46:26a program that can be programmed with
  2667. 1:46:28simple language? Yes, it's very easy to
  2668. 1:46:32uh it it's Python is a very natural
  2669. 1:46:35language to program in because um yeah
  2670. 1:46:38it's very simple uh simple languages
  2671. 1:46:41used all over the place. I think it's
  2672. 1:46:44going to be really easy to learn. I
  2673. 1:46:46think it'll be really easy to pick up.
  2674. 1:46:47At least that's my hope and I think it
  2675. 1:46:49from my experience it is. As I said I
  2676. 1:46:52was someone who did that and I've worked
  2677. 1:46:54with many learners who've done the same.
  2678. 1:46:57So yes, I think it'll be pretty easy to
  2679. 1:47:00pick up, very simple.
  2680. 1:47:04Um, and then the other thing is we can
  2681. 1:47:08do uh we can find our errors very
  2682. 1:47:10quickly. Now, because this is an
  2683. 1:47:11interpreted language, we can run things
  2684. 1:47:13one line at a time and we we will
  2685. 1:47:15quickly hit errors
  2686. 1:47:18uh early on in our code if if we have
  2687. 1:47:20them. So this will be nice and Python
  2688. 1:47:22provides really good um error messages
  2689. 1:47:25um to say hey like this is what's wrong
  2690. 1:47:28with your code you should fix it this
  2691. 1:47:31way um essentially like giving you a
  2692. 1:47:34clue into what needs to be fixed. Um so
  2693. 1:47:37so this is something uh that we will
  2694. 1:47:40practice with as we go along is kind of
  2695. 1:47:42um finding errors and what to do with
  2696. 1:47:45them. Um, but because it's interpreted,
  2697. 1:47:48we will run across those very quickly.
  2698. 1:47:50Unlike with compiled language, which is
  2699. 1:47:52harder to debug because you basically
  2700. 1:47:53have to compile everything, hope that it
  2701. 1:47:56compiles. If it does, then you have to
  2702. 1:47:58run things. Um, it just takes longer to
  2703. 1:48:01get through that debugging phase. But
  2704. 1:48:03with the with Python, it's very quick.
  2705. 1:48:05You get a very quick feedback loop on if
  2706. 1:48:08your code's working or not, which is
  2707. 1:48:10nice. A lot of votes for C. I agree. AC
  2708. 1:48:14is the correct answer here. So the
  2709. 1:48:17interpreter is the thing that will
  2710. 1:48:20execute the code line by line. So it
  2711. 1:48:23doesn't do everything at once. It
  2712. 1:48:25actually goes line by line, which is why
  2713. 1:48:29you can stumble onto your errors quickly
  2714. 1:48:32because if you're going line by line
  2715. 1:48:35um and you have an error on this first
  2716. 1:48:37line, you're never going to reach these
  2717. 1:48:39other lines, right? You're it's just
  2718. 1:48:40going to show you this is where your
  2719. 1:48:42error is. it's on line 101 or whatever
  2720. 1:48:44it is and you know it's going to show
  2721. 1:48:47you where the error is. So it's going to
  2722. 1:48:49go one at a time and execute those. Um
  2723. 1:48:53it's not going to convert the code into
  2724. 1:48:56machine language. That's what a compiled
  2725. 1:48:58language would do, not an interpreted
  2726. 1:49:00one. Um and uh they do require an
  2727. 1:49:05interpreter. So D is just completely
  2728. 1:49:07wrong. It's the opposite of that. It
  2729. 1:49:08does require it. So the interpreter is
  2730. 1:49:10the thing that is executing the uh code
  2731. 1:49:13line by line. So what is Python in
  2732. 1:49:16particular? So it is a as we've already
  2733. 1:49:20seen an interpreted language meaning
  2734. 1:49:23that it requires an interpreter to
  2735. 1:49:25execute it. It's going to be executed
  2736. 1:49:26line by line by that interpreter. Um it
  2737. 1:49:29has capability to be object-oriented. It
  2738. 1:49:32also has capability to be scripted.
  2739. 1:49:34um which is just in relation to how it's
  2740. 1:49:37organized. One of the really nice things
  2741. 1:49:41is it is what we call dynamically typed
  2742. 1:49:45or what you would say dynamic semantics.
  2743. 1:49:48We will see what this means but
  2744. 1:49:51basically it means that we don't have to
  2745. 1:49:53declare what every piece of uh what
  2746. 1:49:56every variable or every piece of data is
  2747. 1:49:59inside of Python. We can let the
  2748. 1:50:00interpreter interpret that which is
  2749. 1:50:03nice. It makes things really easy to
  2750. 1:50:05work with. We don't need to say okay
  2751. 1:50:07this is an integer. This is a
  2752. 1:50:08floatingoint number. This is an array.
  2753. 1:50:10This is you know with a lot of program
  2754. 1:50:14especially compiled languages
  2755. 1:50:15programming languages you have to do
  2756. 1:50:18that because you have to tell the
  2757. 1:50:20compiler this is what this piece of data
  2758. 1:50:23is. But with an interpreter the
  2759. 1:50:25interpreter can as the name suggests
  2760. 1:50:28interpret that. It doesn't need to know
  2761. 1:50:30what everything is in terms of its data
  2762. 1:50:33type, which is which makes it really
  2763. 1:50:35easy to code. On the cons of that, it
  2764. 1:50:39can make it more prone to error because
  2765. 1:50:41you're not really enforcing types. So,
  2766. 1:50:44there is somewhat of a trade-off there.
  2767. 1:50:46But um for our purposes the dynamic
  2768. 1:50:49semantics make make it so that um the
  2769. 1:50:52interpreter can dynamically understand
  2770. 1:50:55what data is um based on how it's being
  2771. 1:50:59used which is great um for us like it
  2772. 1:51:02makes it just quicker to get up and
  2773. 1:51:04running and started and and working with
  2774. 1:51:06data. We don't need to declare what its
  2775. 1:51:08type is which is um static semantics.
  2776. 1:51:12Um now Python itself amazing programming
  2777. 1:51:16language that's used across many
  2778. 1:51:18different applications um such as data
  2779. 1:51:20science, automation, machine learning,
  2780. 1:51:23AI. It's also used in to build software
  2781. 1:51:27even um not sure if you guys know this
  2782. 1:51:29but there's um some really famous
  2783. 1:51:32software that's written in Python. Um,
  2784. 1:51:35one of the most famous is Instagram at
  2785. 1:51:38Meta is completely coded in Python,
  2786. 1:51:41which is it's over like 20,000 lines of
  2787. 1:51:43Python code, which is pretty amazing.
  2788. 1:51:46But um so of course it's been really um
  2789. 1:51:51heavily used in AI and machine learning
  2790. 1:51:53and such but it's also as a programming
  2791. 1:51:56language been used for other things like
  2792. 1:51:58more pure software applications which is
  2793. 1:52:00what makes Python really nice is it's so
  2794. 1:52:02simple so easy to learn. Um so for that
  2795. 1:52:06reason uh it is going to be great for us
  2796. 1:52:10to get started with especially if you're
  2797. 1:52:11coming in with basically no programming
  2798. 1:52:13experience. The other thing about Python
  2799. 1:52:16is it has uh as I said earlier like a
  2800. 1:52:20really big ecosystem uh meaning that
  2801. 1:52:22there's many different packages and
  2802. 1:52:25modules within those package packages
  2803. 1:52:28that do things already. So we don't
  2804. 1:52:31what's great about Python is we won't
  2805. 1:52:33need to reinvent the wheel on so many
  2806. 1:52:36different things like if we need to
  2807. 1:52:38build a plot if we need to train a model
  2808. 1:52:41and and use a specific type of model
  2809. 1:52:44that likely already exists in a package
  2810. 1:52:47somewhere. And what's great is they're
  2811. 1:52:50almost always open source meaning we
  2812. 1:52:52don't have to pay for anything. You can
  2813. 1:52:54just use it out of the box which is
  2814. 1:52:56fantastic. So there's within Python
  2815. 1:52:59there's so many ways to do things
  2816. 1:53:01especially in the AI and machine
  2817. 1:53:02learning world that we'll just borrow
  2818. 1:53:05those and use them in our own code um
  2819. 1:53:08which helps uh you know with um getting
  2820. 1:53:12up and running very quickly. We don't
  2821. 1:53:14need to reinvent things. We can just use
  2822. 1:53:16things that already exist um which is
  2823. 1:53:19fantastic. So that ecosystem really
  2824. 1:53:22benefits machine learning AI. Um because
  2825. 1:53:26they they already exist. We don't need
  2826. 1:53:28to spend our time rewriting all those
  2827. 1:53:30things. Um and so that's something we're
  2828. 1:53:34going to learn as we go along is like
  2829. 1:53:35how to install those, how to import
  2830. 1:53:38those, how to use those in our own code,
  2831. 1:53:40those those packages that already do
  2832. 1:53:44something for us. So we don't need to
  2833. 1:53:46come up with it on our own. we just need
  2834. 1:53:48to use it properly. Okay, so there's a
  2835. 1:53:51little bit of history. Python was first
  2836. 1:53:54invented in the late 1980s by a guy
  2837. 1:53:57named Guido Van Rossom in Amsterdam. Um,
  2838. 1:54:01where it gets its name is after the old
  2839. 1:54:04comedy series, you guys might be
  2840. 1:54:05familiar with it, the Monty Python
  2841. 1:54:07Flying Circus Show. Um, and so that's
  2842. 1:54:11where it's got its name. um you know it
  2843. 1:54:14was first created then but has since
  2844. 1:54:16taken on a really big role in the
  2845. 1:54:20especially you know I keep saying in the
  2846. 1:54:22AI community so much so that it has its
  2847. 1:54:25own software foundation that kind of is
  2848. 1:54:27responsible for maintaining it they meet
  2849. 1:54:29regularly they come up with improvements
  2850. 1:54:33um they come up with new versions of
  2851. 1:54:35Python
  2852. 1:54:36uh for example Python 3.14 just released
  2853. 1:54:40in October which is a major release. Uh
  2854. 1:54:44they hadn't had one in a while and that
  2855. 1:54:47one is uh 3.14. So it's kind of known as
  2856. 1:54:50Python.
  2857. 1:54:52Um which was a big milestone. Um but you
  2858. 1:54:56know they have uh they've had many
  2859. 1:54:58different versions over the years. It's
  2860. 1:55:00been maintained and developed by this
  2861. 1:55:01software foundation. Um and people are
  2862. 1:55:06actively working on it at many large
  2863. 1:55:09companies. So for instance, Meta has a
  2864. 1:55:11big group that is working on um Python
  2865. 1:55:14improvements. Microsoft as well, um
  2866. 1:55:17Google, all of those guys have groups
  2867. 1:55:19kind of working to improve Python
  2868. 1:55:20because they all use it. And so what
  2869. 1:55:22they typically do is work on it, open
  2870. 1:55:25source it, and then the community gets
  2871. 1:55:27to use those tools, those packages,
  2872. 1:55:29those tools, those improvements. Um so
  2873. 1:55:32it's it's actively um utilized across
  2874. 1:55:35many big companies actively uh
  2875. 1:55:37maintained by them or contributed to by
  2876. 1:55:40them. So that's that's really great. Um
  2877. 1:55:43you know Python was originally derived
  2878. 1:55:46from other language um other languages
  2879. 1:55:51uh as kind of a trying to find like a
  2880. 1:55:54mixture of some of the best of all
  2881. 1:55:56worlds. But its main like driving force
  2882. 1:56:00in why Python came to existence from
  2883. 1:56:02these other languages is it just its
  2884. 1:56:05ease of use. People really wanted
  2885. 1:56:07something like super easy to get up and
  2886. 1:56:08running and something really natural.
  2887. 1:56:11Um and so we will as we start learning
  2888. 1:56:14the syntax of it I think you guys will
  2889. 1:56:16understand why it's so easy. But um
  2890. 1:56:18that's that's what led to the
  2891. 1:56:19inspiration is just people wanted
  2892. 1:56:21something easier to work with not as not
  2893. 1:56:23as uh strenuous to kind of get up and
  2894. 1:56:26running.
  2895. 1:56:29What open source license is it? Um,
  2896. 1:56:32that's a good question. I think it's the
  2897. 1:56:34MIT license, but I could be wrong on
  2898. 1:56:36that.
  2899. 1:56:37You could look it up. If you go to
  2900. 1:56:39python.org.
  2901. 1:56:41Yeah, if you go to python.org, I think
  2902. 1:56:43it might talk more about what the uh
  2903. 1:56:46license structure is there. I want to
  2904. 1:56:48say it's MIT open license, but
  2905. 1:56:52I've I'm really not 100% sure on that.
  2906. 1:56:57Okay. So, what are some of the benefits
  2907. 1:56:59of working with Python? And these are
  2908. 1:57:00things you will experience as we go
  2909. 1:57:02along, but just wanted to call them out.
  2910. 1:57:04Um, the flexibility of it. As I said, it
  2911. 1:57:07can be really organized into
  2912. 1:57:09object-oriented or it can be loosely
  2913. 1:57:11organized into scripts. So, that
  2914. 1:57:14flexibility alone is really awesome. um
  2915. 1:57:17which has allowed it to power many
  2916. 1:57:20different things like um APIs, web
  2917. 1:57:22pages, full-blown applications like
  2918. 1:57:24Instagram, um chat, GPTs, like actual uh
  2919. 1:57:29AI, LLMs.
  2920. 1:57:32Um you know, it has so much flexibility
  2921. 1:57:35there to power so many different
  2922. 1:57:36applications.
  2923. 1:57:38Um probably the biggest benefit,
  2924. 1:57:40especially to us, is its ease of use.
  2925. 1:57:43Um,
  2926. 1:57:45uh,
  2927. 1:57:48oh, thank you. Some Tim just posted it.
  2928. 1:57:50It's the the GNU,
  2929. 1:57:53uh, public license. Yes.
  2930. 1:57:58Oh, never mind. It's a Python software.
  2931. 1:58:00It has its own. Okay, perfect. Thanks
  2932. 1:58:02for sharing that. Thanks for sharing
  2933. 1:58:04that. Yeah, I wasn't completely sure
  2934. 1:58:06which which license it was, but it is
  2935. 1:58:09open source. Um, and people do make
  2936. 1:58:12their own kind of derivations of Python.
  2937. 1:58:16But as I was saying, one of the benefits
  2938. 1:58:17of Python is how easy it is to learn. I
  2939. 1:58:20keep emphasizing that because it's true.
  2940. 1:58:22Once we get into it, you will see this.
  2941. 1:58:24I promise it'll be easy to learn, easy
  2942. 1:58:26to pick up. Um, and it's designed in
  2943. 1:58:30that way. Designed to be very minimalist
  2944. 1:58:32as a language, which is great.
  2945. 1:58:35um it has a lot of things that come with
  2946. 1:58:38it and it's kind of built into Python, a
  2947. 1:58:41lot of capability. So we call that the
  2948. 1:58:43standard library. It's just the things
  2949. 1:58:45built into Python. It has a lot of
  2950. 1:58:46capability out of the box. Um you know,
  2951. 1:58:50not only that, but it has a large
  2952. 1:58:51community that's developed so many
  2953. 1:58:52different packages that do things for
  2954. 1:58:54us, especially in the AI world. So
  2955. 1:58:57that's another great thing, kind of a
  2956. 1:58:59robust community developing these
  2957. 1:59:01packages that help us get things done.
  2958. 1:59:05Um, readability. So because the code is
  2959. 1:59:08so simple, it's also easy to read. So
  2960. 1:59:11you can usually read other Python code
  2961. 1:59:13and quickly understand what it's doing
  2962. 1:59:15which you know makes for easy um easy
  2963. 1:59:20understanding of other people's code
  2964. 1:59:21easy understanding of code in the
  2965. 1:59:23community and kind of almost like it's
  2966. 1:59:26selfdocumenting because it's so easy to
  2967. 1:59:28read. So that that simplicity that ease
  2968. 1:59:32of use lends itself well to being really
  2969. 1:59:35readable. You can usually just take a
  2970. 1:59:37look at the code, easily read it,
  2971. 1:59:39understand what it's doing, which is
  2972. 1:59:41great, like great for you guys learning,
  2973. 1:59:44great for taking a look at the demos and
  2974. 1:59:46examples that we will do. They're very
  2975. 1:59:48readable.
  2976. 1:59:56Okay. So why has Python really dominated
  2977. 2:00:02AI? So this is a valid question like
  2978. 2:00:04even so it's used for many different
  2979. 2:00:06things. It's a programming language. So
  2980. 2:00:07it can build application and I've given
  2981. 2:00:09you the example of Instagram and there's
  2982. 2:00:10many others um that are built off of
  2983. 2:00:14Python code. Why is it so useful for AI
  2984. 2:00:19in particular?
  2985. 2:00:21mainly
  2986. 2:00:23uh some of the reasons we've already
  2987. 2:00:25talked about mainly how easy it is to
  2988. 2:00:28use lends itself well for AI because um
  2989. 2:00:32that has allowed people to kind of
  2990. 2:00:35quickly get up and running and test out
  2991. 2:00:36their algorithms, test out their models
  2992. 2:00:40just really quickly with Python. That's
  2993. 2:00:42great. The other things listed on here
  2994. 2:00:46are certainly big reasons as well. So
  2995. 2:00:50for example, it has so many community
  2996. 2:00:54libraries, those those packages that um
  2997. 2:00:58have AI models and AI tools that we can
  2998. 2:01:03reuse that people have built these up
  2999. 2:01:05over years and years and years. Um so
  3000. 2:01:08it's to our benefit to reuse those and
  3001. 2:01:11not have to reinvent everything and we
  3002. 2:01:13can get quickly up and running with
  3003. 2:01:14those which would be great.
  3004. 2:01:17The other thing is Python, it lends
  3005. 2:01:18itself very very well to working with
  3006. 2:01:20data in general. Very easy to work with
  3007. 2:01:23data, very easy to load it in from
  3008. 2:01:24external sources, query it, work with
  3009. 2:01:27it, visualize it. Python is so adept at
  3010. 2:01:31that. Um, so that's what makes it really
  3011. 2:01:34nice at doing machine learning and AI
  3012. 2:01:35because so much of it is manipulating
  3013. 2:01:37data. So, um, for that reason alone,
  3014. 2:01:41Python is so popular in the AI community
  3015. 2:01:43just because of its ability to work with
  3016. 2:01:45data. It's so easy. This is something
  3017. 2:01:47we're going to really focus in on like
  3018. 2:01:49in our next course when we talk about
  3019. 2:01:51data science.
  3020. 2:01:53But, um,
  3021. 2:01:55just the ability and the power of it to
  3022. 2:01:58work with data makes lends itself well
  3023. 2:02:00to AI uh, capabilities.
  3024. 2:02:03Um, the other thing is I mentioned the
  3025. 2:02:05rapid prototyping. You can quickly build
  3026. 2:02:07a model in Python because the code is so
  3027. 2:02:09easy. So, and there's so many libraries
  3028. 2:02:11already can quickly prototype. Um,
  3029. 2:02:15it has obviously a big community around
  3030. 2:02:18it that's building out these packages,
  3031. 2:02:20writing documentation, maintaining it
  3032. 2:02:22from an open source level. So, that's
  3033. 2:02:25another reason it's very popular. Um,
  3034. 2:02:27Python's also used with other
  3035. 2:02:29technologies. So, it does have
  3036. 2:02:32capability to integrate with other
  3037. 2:02:34languages. So for instance, Python can
  3038. 2:02:36one of the most popular integrations is
  3039. 2:02:38Python can work with C and C++. So
  3040. 2:02:41sometimes that's necessary to integrate
  3041. 2:02:43with those to do certain things. Um so
  3042. 2:02:47Python has been extended to work with
  3043. 2:02:49other languages. So sometimes there's
  3044. 2:02:51other uh necessary support from other
  3045. 2:02:55like things in other languages that are
  3046. 2:02:56necessary to power something in AI. um
  3047. 2:02:59for example working with GPUs
  3048. 2:03:03and doing things in deep learning. Um
  3049. 2:03:07there's been a lot of integration with
  3050. 2:03:09uh working with um C tools. Now will we
  3051. 2:03:12do that? No, it's already been done for
  3052. 2:03:15us and some of these packages. But um
  3053. 2:03:19the pure ability of Python to do that is
  3054. 2:03:22really powerful and it gets taken for
  3055. 2:03:24granted honestly because you don't see
  3056. 2:03:26that it's underneath the hood and it's
  3057. 2:03:28abstracted away from you when you work
  3058. 2:03:29with those Python packages. But there
  3059. 2:03:32was a lot of work that went into it to
  3060. 2:03:33integrate it with other kind of other
  3061. 2:03:35programming languages.
  3062. 2:03:39Okay.
  3063. 2:03:41So as an example like I mentioned the
  3064. 2:03:43Instagram one. So Netflix for instance,
  3065. 2:03:45all of their recommendation is powered
  3066. 2:03:48by Python. So when you open up Netflix
  3067. 2:03:51or really any streaming service for that
  3068. 2:03:54matter, they're going to use Python to
  3069. 2:03:56deliver those recommendations and
  3070. 2:03:58produce those personalized
  3071. 2:03:59recommendations. Um Spotify as well for
  3072. 2:04:02like music. Um nearly all recommendation
  3073. 2:04:05algorithms are written in Python.
  3074. 2:04:08And in this program, we are actually
  3075. 2:04:11going to learn about recommendation
  3076. 2:04:13systems. So that'll be pretty fun way
  3077. 2:04:16down the road when we get into machine
  3078. 2:04:17learning. We'll talk about how do we
  3079. 2:04:19build a recommendation engine,
  3080. 2:04:22but um they're all done through Python
  3081. 2:04:25for for example. So really cool uh use
  3082. 2:04:28cases there.
  3083. 2:04:32So one of the things I wanted to address
  3084. 2:04:34is how AI itself is changing coding. So
  3085. 2:04:39you guys may be aware of this, but
  3086. 2:04:41obviously there's been a huge um kind of
  3087. 2:04:45explosion in generative AI tools that
  3088. 2:04:48can help write documents and write
  3089. 2:04:50emails and write text and all these
  3090. 2:04:52things. One of the things they can do is
  3091. 2:04:54write code. So um one of the big areas
  3092. 2:04:58where AI is changing coding is it's an
  3093. 2:05:00its ability to generate code for us. And
  3094. 2:05:04so um throughout this program like we
  3095. 2:05:07won't shy away from that necessarily
  3096. 2:05:10and I encourage you guys to use AI tools
  3097. 2:05:13as you see fit to help your own
  3098. 2:05:16understanding and help your own
  3099. 2:05:17productivity. Um
  3100. 2:05:20you know we still will go through the
  3101. 2:05:22fundamentals so you can understand it
  3102. 2:05:24but the AI tools can definitely be a
  3103. 2:05:27supplement to help. Um it's just that I
  3104. 2:05:31think you guys will understand it better
  3105. 2:05:32going through the examples that we do we
  3106. 2:05:34do together and so that when AI
  3107. 2:05:37generates code you will be able to
  3108. 2:05:39understand it and also be able to debug
  3109. 2:05:41it right because it's not always going
  3110. 2:05:43to be perfect. So that's always the
  3111. 2:05:45catch with AI is that you know it
  3112. 2:05:48doesn't always produce perfect answers.
  3113. 2:05:50Um but the at the very least we will be
  3114. 2:05:53able to you know debug things and
  3115. 2:05:56understand things better so that uh we
  3116. 2:05:59can catch those errors.
  3117. 2:06:01Um so obviously like AI is also besides
  3118. 2:06:05flat out generating it it's also
  3119. 2:06:07suggesting what should be there. So, uh,
  3120. 2:06:10some of the code editors really do a
  3121. 2:06:12good job at that, suggesting things, um,
  3122. 2:06:15picking up on what you should produce
  3123. 2:06:17next. That's going to be, um, very
  3124. 2:06:20interesting as we get into, uh, some of
  3125. 2:06:23the platforms that you guys will work
  3126. 2:06:25with to write your Python code. They
  3127. 2:06:27will have that ability. Um, so, uh, the
  3128. 2:06:32other thing is like there's some cloud
  3129. 2:06:35tools that, um, don't require writing
  3130. 2:06:39much code at all and they can just do
  3131. 2:06:41things. So, in other words, you can
  3132. 2:06:42power them by prompts. You're not really
  3133. 2:06:44writing code. You're just writing
  3134. 2:06:45natural language and then they do
  3135. 2:06:47something. Um, they generate the code in
  3136. 2:06:50the background and they execute
  3137. 2:06:52something. Um we will learn about those
  3138. 2:06:55things uh later on in the program
  3139. 2:06:57especially because we we will cover
  3140. 2:06:59generative AI in the future
  3141. 2:07:03um towards the end of our program. So if
  3142. 2:07:05you're wondering like are we going to
  3143. 2:07:07cover LLMs? Are we going to cover how
  3144. 2:07:09these things get generated? Yes. It just
  3145. 2:07:12will be um later on in the program.
  3146. 2:07:18Okay. A lot of votes for B.
  3147. 2:07:20Yeah, pretty unanimous on B. I think I
  3148. 2:07:22agree with it. Yeah, B is definitely the
  3149. 2:07:23right answer. So, all of the
  3150. 2:07:25recommendation systems which we will
  3151. 2:07:27learn how to build ourselves later on
  3152. 2:07:30are written in Python and um they uh are
  3153. 2:07:36machine learning models that make the
  3154. 2:07:37recommendations and that machine
  3155. 2:07:39learning is driven by data um and all of
  3156. 2:07:43that data is manipulated in Python
  3157. 2:07:46um and used to train uh models that do
  3158. 2:07:49the recommendations. That's all
  3159. 2:07:51happening in Python.
  3160. 2:07:54So, we're going to talk about getting
  3161. 2:07:56you guys set up on your own machine and
  3162. 2:07:59talking about the different development
  3163. 2:08:02environments we can use to actually work
  3164. 2:08:04with Python code. Um, before we go into
  3165. 2:08:08that, any questions about anything we
  3166. 2:08:09covered so far?
  3167. 2:08:12Everything's good so far. Yep. And you
  3168. 2:08:15know again if you have experience in
  3169. 2:08:17Python I recognize that it is going to
  3170. 2:08:18be a little slow in beginning. Um it's
  3171. 2:08:21mostly to get us really oriented to some
  3172. 2:08:24background around Python and get us set
  3173. 2:08:26up and then we will be doing you know uh
  3174. 2:08:29getting into the syntax and all that uh
  3175. 2:08:33coming up shortly. So we will actually
  3176. 2:08:36be learning Python specifics coming up
  3177. 2:08:38soon. But you know we're going to um get
  3178. 2:08:41everything set up first.
  3179. 2:08:45All right. So, let's continue then.
  3180. 2:08:47Thank you guys for that.
  3181. 2:08:50So, um it turns out that there are many
  3182. 2:08:53tools in the community for developing
  3183. 2:08:56Python code. And so, um you might hear
  3184. 2:08:59this word ID. It is short for integrated
  3185. 2:09:02development environment. This is a piece
  3186. 2:09:04of software that helps you write and
  3187. 2:09:08test Python code. So, and there's many
  3188. 2:09:12out there. There's a bunch on this list.
  3189. 2:09:14We are going to focus on a few options.
  3190. 2:09:18There's even more than what's on this
  3191. 2:09:20list, but we're going to focus on a few
  3192. 2:09:22options. These IDs are designed to
  3193. 2:09:24really help you write Python. They
  3194. 2:09:26provide many tools in the background
  3195. 2:09:29that make your life easier when you're
  3196. 2:09:31working with Python. So, for example,
  3197. 2:09:33they can provide syntax highlighting.
  3198. 2:09:36They can tell you when you have a syntax
  3199. 2:09:38error. Um, almost like a spell check for
  3200. 2:09:41Python.
  3201. 2:09:43Um, they can help you run Python code
  3202. 2:09:45right within the window. Um, they can
  3203. 2:09:48help you organize your projects. Uh,
  3204. 2:09:50they can do a lot of different things.
  3205. 2:09:53Um, and so there's many tools out there
  3206. 2:09:55that can do it, and it's really a
  3207. 2:09:58personal preference which one you use,
  3208. 2:10:00but in this program, we're really going
  3209. 2:10:02to focus on a few of them to to showcase
  3210. 2:10:05those options because they're very
  3211. 2:10:06popular options. Um, and then, uh, allow
  3212. 2:10:11you guys the flexibility to choose which
  3213. 2:10:13option makes the most sense for you. So,
  3214. 2:10:15generally, that's going to be mostly a a
  3215. 2:10:19preference.
  3216. 2:10:20um mostly a preference as to which one
  3217. 2:10:23you're the most comfortable with, but I
  3218. 2:10:25want to give you guys the option to uh
  3219. 2:10:28explore
  3220. 2:10:30the various options that are available.
  3221. 2:10:36Uh Roberto, is there one that stands out
  3222. 2:10:38as an industry standard? Um there's a
  3223. 2:10:41couple that you see like honestly the
  3224. 2:10:45two of them that we will study uh in
  3225. 2:10:48this coming up in the next few slides
  3226. 2:10:50are the industry standard which are
  3227. 2:10:51going to be VS code Microsoft VS code
  3228. 2:10:54and then Jupyter notebooks. So these two
  3229. 2:11:00are going to be uh ones that we will
  3230. 2:11:03study in particular and use throughout.
  3231. 2:11:07Um
  3232. 2:11:10so so yes we will those will be industry
  3233. 2:11:13standards. PyCharm's also very popular.
  3234. 2:11:16Um so I don't want to rule out PyCharm.
  3235. 2:11:18I know a lot of people who use it. So um
  3236. 2:11:21I would encourage you to explore PyCharm
  3237. 2:11:23as well if you want to but we are not
  3238. 2:11:25going to do that uh in in these slides
  3239. 2:11:28but um I would check it out and see if
  3240. 2:11:32you like it. Um it's another very I'm
  3241. 2:11:35putting a an asterisk next to it because
  3242. 2:11:37I think it's one of the more popular
  3243. 2:11:40uh yes uh yeah we're going to do
  3244. 2:11:43descriptions.
  3245. 2:11:45um requirements uh I'll try my best to
  3246. 2:11:48give those but honestly the requirements
  3247. 2:11:49will be given when you install them. Um
  3248. 2:11:54so the other thing I want to say is we
  3249. 2:11:56will have a couple options that don't
  3250. 2:11:58require you to install anything. So I'm
  3251. 2:12:00going to showcase those as well. So
  3252. 2:12:03there's a couple options that are um we
  3253. 2:12:05won't have to install anything because
  3254. 2:12:07they're going to be cloud-based.
  3255. 2:12:09Okay, I'll show you those.
  3256. 2:12:16Okay. So, but yeah, VS Code, I think VS
  3257. 2:12:18Code and Jupyter notebooks are are
  3258. 2:12:21probably the industry standard most
  3259. 2:12:23popular uh idees.
  3260. 2:12:27Okay. So, what we would recommend in
  3261. 2:12:30this program and the ones that we will
  3262. 2:12:32use the most uh throughout are going to
  3263. 2:12:35be these three. Visual Studio Code, also
  3264. 2:12:37known as VS Code, Jupyter Notebooks, and
  3265. 2:12:40Google Coll Collab, which is Google's
  3266. 2:12:43hosted
  3267. 2:12:45um Google's hosted version of notebooks
  3268. 2:12:48essentially. Um so
  3269. 2:12:52I will showcase each one of these and
  3270. 2:12:54give you some examples of how to set it
  3271. 2:12:56up and examples of how to work with it.
  3272. 2:12:59Um, and that's what we'll do over the
  3273. 2:13:01course of the next few slides and the
  3274. 2:13:03next uh bit of time is I'm going to go
  3275. 2:13:06through each one of these and kind of
  3276. 2:13:07show you what you would need to do to
  3277. 2:13:08get it set up. Um, now that being said,
  3278. 2:13:15excuse me, these two are ones that you
  3279. 2:13:18will install.
  3280. 2:13:20These two you would install locally on
  3281. 2:13:23your on your own machine.
  3282. 2:13:26And this one is um uh cloud hosted
  3283. 2:13:32by Google and it's free. Um all of these
  3284. 2:13:36are free but uh the first two VS code
  3285. 2:13:40and Jupyter notebook you would install
  3286. 2:13:42on your own machine. Collab you would
  3287. 2:13:43just access through your web browser. It
  3288. 2:13:45is hosted by Google. So that's an
  3289. 2:13:47advantage. You don't really need to
  3290. 2:13:48install anything. And for that reason um
  3291. 2:13:51sometimes we will favor Collab. Uh and
  3292. 2:13:53for other reasons too. Collab has some
  3293. 2:13:55really nice features if you've never
  3294. 2:13:57used it. Um, but notebooks, um, Jupyter
  3295. 2:14:03Notebook and Collab are very similar.
  3296. 2:14:06They're very similar. Collab just has
  3297. 2:14:08its own spin-off on on the notebook, um,
  3298. 2:14:11type of file that Jupyter Notebooks work
  3299. 2:14:14with. And it's um, like I said, kind of
  3300. 2:14:16cloud hosted. So, I'm going to I'm going
  3301. 2:14:18to walk us through each one of these and
  3302. 2:14:21explain to you what they do, what they
  3303. 2:14:24look like, and then we will um I'll set
  3304. 2:14:26up each one of them uh kind of in a live
  3305. 2:14:28demo so you guys can see. Um but uh we
  3306. 2:14:34throughout the program, it will really
  3307. 2:14:37be up to you which one of these you want
  3308. 2:14:39to use. There's no hard requirement to
  3309. 2:14:41use any one of them. It's really going
  3310. 2:14:43to be your preference which one of these
  3311. 2:14:46tools you want to use to work with
  3312. 2:14:47Python. Whatever one you feel
  3313. 2:14:49comfortable working with, that's the one
  3314. 2:14:51you should use.
  3315. 2:14:53All three of these are very popular in
  3316. 2:14:55the industry. So, you're not missing out
  3317. 2:14:56by using one versus the other. Um,
  3318. 2:14:59they're all very popular. Even Collab, I
  3319. 2:15:02know it wasn't on the screen, but it is
  3320. 2:15:04widely used in in the community and the
  3321. 2:15:06industry.
  3322. 2:15:10Uh, no system recommendations for
  3323. 2:15:11training LLMs. Um, no, because we don't
  3324. 2:15:13we won't really focus on that until the
  3325. 2:15:15end. When we get to when we get into
  3326. 2:15:17generative AI, we'll talk about that.
  3327. 2:15:21When we get into generative AI, we'll
  3328. 2:15:22talk about that.
  3329. 2:15:28So, yeah, we're not we're not focusing
  3330. 2:15:30on LM in the beginning. That's that's an
  3331. 2:15:33advanced topic for us.
  3332. 2:15:36What is my personal preference? Um, I
  3333. 2:15:40like Visual Studio Code. Um, personally
  3334. 2:15:43I that's what I use for my day-to-day
  3335. 2:15:45work is uh Visual Studio Code. I like
  3336. 2:15:48Visual Studio Code and I like Collab a
  3337. 2:15:50lot. Um, so you know, we'll talk about
  3338. 2:15:54this, but one of the advantages to
  3339. 2:15:56Collab is that it has free access to
  3340. 2:15:58GPUs, which is huge for doing things
  3341. 2:16:02like uh neural nets. Um, so we will lean
  3342. 2:16:06on collab quite a bit later on
  3343. 2:16:10uh later on when we um actually get to
  3344. 2:16:15deep learning and neural nets. We'll
  3345. 2:16:17because collab has free access to GPUs.
  3346. 2:16:20I'll show us that. It's it's really
  3347. 2:16:22nice.
  3348. 2:16:24And when you do anything with neural
  3349. 2:16:25nets, it usually benefits you to have a
  3350. 2:16:27GPU access. Um
  3351. 2:16:30so that'll be nice. But I usually do
  3352. 2:16:33most Python coding inside of VS Code. It
  3353. 2:16:36supports Python pretty pretty well.
  3354. 2:16:42What is more commonly used in the
  3355. 2:16:44industry? Um,
  3356. 2:16:46the two most popular are Visual Studio
  3357. 2:16:49Code and and Notebooks. Jupiter
  3358. 2:16:51notebooks.
  3359. 2:16:52They're both like you can't go wrong
  3360. 2:16:54with either one.
  3361. 2:16:58Those
  3362. 2:17:01two are really popular. Jupyter
  3363. 2:17:02notebooks and Visual Studio Code are
  3364. 2:17:04really popular. There's there's both of
  3365. 2:17:07those you would be okay with. Either
  3366. 2:17:10one.
  3367. 2:17:12Let me start with Visual Stu Studio
  3368. 2:17:13Code. So, um now Visual Studio Code
  3369. 2:17:18is a more general code editor. So, it's
  3370. 2:17:22actually you can edit lots of different
  3371. 2:17:25languages inside of VS Code. Um, so you
  3372. 2:17:29could do Java, you could do C, you could
  3373. 2:17:31do Scala, you can do Go, you can do all
  3374. 2:17:35kinds of languages are supported inside
  3375. 2:17:37of Visual Studio Code. So it's a really
  3376. 2:17:39fantastic product for programming in
  3377. 2:17:41general, not just Python. Um, it has
  3378. 2:17:44built-in terminal support. It has
  3379. 2:17:47co-pilot integrated into it, which is
  3380. 2:17:50nice for AI, like generative AI
  3381. 2:17:53assistance working with your code, which
  3382. 2:17:55is nice. Um, of course it supports
  3383. 2:17:58Python, which is what we are interested
  3384. 2:18:00in. Um, it has it has Python tools. I
  3385. 2:18:04will show us which ones we should
  3386. 2:18:06install as part of VS Code so that we
  3387. 2:18:09can work with Python files and
  3388. 2:18:11notebooks. Um, so it's it's a really
  3389. 2:18:15great code editor in general, which is
  3390. 2:18:17why I like using it. Um, but in
  3391. 2:18:20particular, it's pretty good at working
  3392. 2:18:22with Python. it it supports Python
  3393. 2:18:24pretty uh deeply. Um so and for that
  3394. 2:18:28reason VS code is really really popular.
  3395. 2:18:32But just keep in mind you can actually
  3396. 2:18:33use it for many different types of code
  3397. 2:18:35that uh that people write uh JavaScript
  3398. 2:18:39um Java as I said like many languages
  3399. 2:18:42are supported inside of Visual Studio
  3400. 2:18:43Code. So it's a more general code
  3401. 2:18:46editor. It happens to be really great at
  3402. 2:18:48working with Python though.
  3403. 2:18:52All right. So, I'm going to show us a
  3404. 2:18:53demo on setting up VS Code. Now, we are
  3405. 2:18:57going to do this for each one of these.
  3406. 2:19:01For Jupiter and for Collab, I'm going to
  3407. 2:19:03I'm going to do similar demos. So, um
  3408. 2:19:07don't worry, we'll get to those, but I
  3409. 2:19:09want to start with VS Code to show you
  3410. 2:19:11kind of how to get that set up and what
  3411. 2:19:13it looks like. Um, so where you can find
  3412. 2:19:16this demo
  3413. 2:19:19is inside of the demos that I mentioned
  3414. 2:19:23earlier in the reference material. So
  3415. 2:19:24I'm going to I'm going to jump over to
  3416. 2:19:27that. Let me show you guys.
  3417. 2:19:32So I'm back in the LMS. You guys will
  3418. 2:19:35want to download the demos. I think
  3419. 2:19:37somebody linked it earlier in case this
  3420. 2:19:39didn't show up for you, but we're going
  3421. 2:19:40to be inside of the demos and we're
  3422. 2:19:42going to do demo one for lesson one. We
  3423. 2:19:45do lesson one, demo one, which is going
  3424. 2:19:47to be the VS Code demo.
  3425. 2:19:51So, the main steps that we're going to
  3426. 2:19:53do is just going to be to point you to
  3427. 2:19:57where to install Visual Studio Code. So,
  3428. 2:19:59it is a it is an application is a free
  3429. 2:20:02application you can install on your
  3430. 2:20:03machine. Um, so,
  3431. 2:20:07uh, you will want to follow this link
  3432. 2:20:10that is within the demo file, this
  3433. 2:20:12code.vvisualstudio.com/d
  3434. 2:20:14download and download it for your
  3435. 2:20:16particular platform. So, if you're on
  3436. 2:20:17Windows, obviously, choose the Windows.
  3437. 2:20:20If you're on a Mac, um, choose Mac and
  3438. 2:20:24make sure that you choose the right, one
  3439. 2:20:27of the precautions is to choose the
  3440. 2:20:28right Mac platform. So, if you have like
  3441. 2:20:30an M1, 2, M3, M4 Mac, choose the Apple
  3442. 2:20:34Silicon
  3443. 2:20:36um button. If you're on an older Mac, um
  3444. 2:20:39then you'll want to use the Intel chip
  3445. 2:20:41one. Um
  3446. 2:20:45uh if you're on if you happen to be on
  3447. 2:20:47Linux, which I don't probably most of
  3448. 2:20:50you are not, but if you are, um you want
  3449. 2:20:52to download the right uh distribution uh
  3450. 2:20:55version.
  3451. 2:20:57But, uh follow this link first. So
  3452. 2:20:59that's the first step. Very easy step.
  3453. 2:21:01Just go to that site, pick your right
  3454. 2:21:03platform and uh go ahead and download
  3455. 2:21:06the installer. And mostly we will be
  3456. 2:21:10walking through the steps in the
  3457. 2:21:12installer. And then um I will show us
  3458. 2:21:15what it looks like once it's installed
  3459. 2:21:18and then show you a couple additional
  3460. 2:21:20steps that are actually not mentioned in
  3461. 2:21:21this file that I think are worth doing
  3462. 2:21:23to get you set up.
  3463. 2:21:29Uh yes, we will be doing Jupiter next.
  3464. 2:21:31Yes, we we'll we're going to be covering
  3465. 2:21:34VS Code, Jupiter, and Collab. I'm going
  3466. 2:21:36to show us examples of all of those.
  3467. 2:21:46Okay, let me ask you guys. Were you guys
  3468. 2:21:48able to get to the download page and
  3469. 2:21:50start that download and installation of
  3470. 2:21:52VS Code?
  3471. 2:21:56able to do that.
  3472. 2:21:58Any issues with that?
  3473. 2:22:05Okay. Yeah, it's just like any
  3474. 2:22:08yet I love I love the optimism
  3475. 2:22:12yet.
  3476. 2:22:16Uh already having both of them
  3477. 2:22:18installed. Okay. Yeah. No, if you
  3478. 2:22:20already have it installed, I mean,
  3479. 2:22:21great. I'll show so if you if you
  3480. 2:22:23already have VS Code installed, great.
  3481. 2:22:25You can sit tight. I will show you um a
  3482. 2:22:29couple of extensions that you'll want to
  3483. 2:22:31add for Python support
  3484. 2:22:34if you have it installed already. I'll
  3485. 2:22:37show us how you can use it with Python
  3486. 2:22:39in particular.
  3487. 2:22:41Okay.
  3488. 2:22:44If you already have it installed,
  3489. 2:22:45perfect. Looks like you have it
  3490. 2:22:47launched.
  3491. 2:22:51Still working on it. Okay. So, these
  3492. 2:22:54these instructions um uh show an example
  3493. 2:22:57of someone that would be on a Microsoft
  3494. 2:22:59platform um walking through the
  3495. 2:23:03installation.
  3496. 2:23:07Uh if you're on a Windows, you probably
  3497. 2:23:10want to create a desktop icon. You
  3498. 2:23:11definitely want to add it to your path.
  3499. 2:23:21And this just shows what's being
  3500. 2:23:23installed. So this is all the install
  3501. 2:23:24wizard on Windows. Nothing that exciting
  3502. 2:23:27there. So this if you follow all these
  3503. 2:23:30steps, you will have it installed. I
  3504. 2:23:32hope you have enough disc space. Uh I
  3505. 2:23:34don't think it's too big.
  3506. 2:23:37I don't think it's too too massive. I
  3507. 2:23:39forget how much space it takes. I don't
  3508. 2:23:40think it's that much.
  3509. 2:23:45I don't think it's that much. But um
  3510. 2:23:47yeah, hopefully you have enough.
  3511. 2:23:51So if if you don't
  3512. 2:23:54uh if you do not have enough disc space
  3513. 2:23:57um don't worry because we're going to do
  3514. 2:23:59collab which doesn't require you
  3515. 2:24:01installing anything. So you can always
  3516. 2:24:03use that option. All right. So if if for
  3517. 2:24:06some re let me just say that too just
  3518. 2:24:08even if if it's not a dispace issue if
  3519. 2:24:10you have an in any installation issues
  3520. 2:24:13no worries because we will work with
  3521. 2:24:16collab and Google that is going to be
  3522. 2:24:18cloud hosted that you don't need to
  3523. 2:24:20install anything you just need a Google
  3524. 2:24:22account
  3525. 2:24:24okay a free Google account
  3526. 2:24:27um so no worries at all if you cannot
  3527. 2:24:30get any of these things installed the
  3528. 2:24:33which are going to be Jupiter and
  3529. 2:24:36uh Jupiter and VS Code.
  3530. 2:24:41Where do we go? I haven't said yet. It
  3531. 2:24:42just I'm just making sure it's installed
  3532. 2:24:44for folks.
  3533. 2:24:46I'm going to I'm going to go over to it
  3534. 2:24:47in a second, but did we generally get it
  3535. 2:24:50installed and do you have it open? So,
  3536. 2:24:52if you once you get it installed,
  3537. 2:24:55uh once you get it installed, then open
  3538. 2:24:57it.
  3539. 2:25:04Yeah, you need to get it installed. Uh,
  3540. 2:25:07it should be this first. It should be
  3541. 2:25:09this link here.
  3542. 2:25:13Follow this link to get it installed.
  3543. 2:25:20Oops, I pasted the wrong link.
  3544. 2:25:37Let me find I'll copy and paste the
  3545. 2:25:39link. But yeah, take take a moment to
  3546. 2:25:41get it open. Once you have it open, just
  3547. 2:25:44sit tight
  3548. 2:25:47if you want to.
  3549. 2:25:50What does it say?
  3550. 2:25:54Yeah, feel. So, for you guys seeing the
  3551. 2:25:57co-pilot features, um,
  3552. 2:26:00click click use AI features. I think
  3553. 2:26:03that's okay. Yes. Um, you'll you'll
  3554. 2:26:06likely want copilot. Yes.
  3555. 2:26:09Click click okay on that.
  3556. 2:26:17That's the link, by the way, for the
  3557. 2:26:19download
  3558. 2:26:21in case uh we needed to get to it.
  3559. 2:26:30Okay.
  3560. 2:26:32So, I'm going to go over to VS Code
  3561. 2:26:36and show you what it looks like on uh my
  3562. 2:26:38end.
  3563. 2:26:43Okay. So, you should have something that
  3564. 2:26:44looks roughly like this. I don't have
  3565. 2:26:47anything open. I don't have any files
  3566. 2:26:49open. Uh just kind of have a blank
  3567. 2:26:52screen here. Um, but if you I would
  3568. 2:26:55recommend uh using the AI features if
  3569. 2:26:58you can. Um, I think that'll come in
  3570. 2:27:01handy later on.
  3571. 2:27:04Um, are we
  3572. 2:27:07comfortable uh moving forward? I want to
  3573. 2:27:09show us the extensions that support
  3574. 2:27:11Python. So, right now when you first
  3575. 2:27:14when you first install this, it does not
  3576. 2:27:18work with Python out of the box. We have
  3577. 2:27:20to install a couple extensions inside of
  3578. 2:27:23here to get it to work with Python. I'm
  3579. 2:27:25going to show us how to do that.
  3580. 2:27:30Don't worry about tuning any settings.
  3581. 2:27:32No, don't worry about doing any of that
  3582. 2:27:34at this stage. Don't really need to tune
  3583. 2:27:37anything. We just need to get Python
  3584. 2:27:40support.
  3585. 2:27:47Okay.
  3586. 2:27:49So, you guys with me on this main page?
  3587. 2:27:58You can use your corporate. Sure. Sure.
  3588. 2:28:00Yeah, you can you if you have it. If you
  3589. 2:28:02have co-pilot and want to use your
  3590. 2:28:04corporate, you can use that. That's
  3591. 2:28:05fine.
  3592. 2:28:08But you guys are with me on the main
  3593. 2:28:09page because I'm about to show us uh I'm
  3594. 2:28:12about to show us the extensions we need
  3595. 2:28:14to install to work with Python.
  3596. 2:28:17Okay, really important because this
  3597. 2:28:20isn't this is not in the documentation.
  3598. 2:28:34Um, no, no need to reinstall. Um, you
  3599. 2:28:39can I'll show you how to add that
  3600. 2:28:41through the extensions. No need to
  3601. 2:28:43reinstall.
  3602. 2:28:46You can add it as an extension. Yeah.
  3603. 2:28:52Okay.
  3604. 2:28:54So, let me ask you guys on the left,
  3605. 2:28:58do you see
  3606. 2:29:01this
  3607. 2:29:03little box icon that if you hover over
  3608. 2:29:08it says extensions?
  3609. 2:29:12Do you see that? you. There may be other
  3610. 2:29:14things here too, but at least that one
  3611. 2:29:17with the extensions.
  3612. 2:29:26Okay, so we do see that one. Okay,
  3613. 2:29:32so what we want to do,
  3614. 2:29:38no, I wouldn't I wouldn't uninstall.
  3615. 2:29:40That's okay because we're actually gonna
  3616. 2:29:42install Anaconda to get Jupiter. I
  3617. 2:29:45wouldn't un I wouldn't I would cancel
  3618. 2:29:46that if you can because you're going to
  3619. 2:29:48want that for Jupiter as well. I
  3620. 2:29:51wouldn't uninstall Anaconda.
  3621. 2:29:54I wouldn't uninstall. But I mean, if
  3622. 2:29:56it's already going if it's already doing
  3623. 2:29:57it, that's okay. We'll just reinstall it
  3624. 2:29:59later. All right. So, back to the
  3625. 2:30:01extensions. So, let's click on the
  3626. 2:30:03extensions.
  3627. 2:30:09Okay. So, do we see something like this
  3628. 2:30:11that has a search bar for extensions?
  3629. 2:30:16Do we see the search bar for the
  3630. 2:30:18extensions?
  3631. 2:30:21Okay. What do you think? We're going to
  3632. 2:30:23search for
  3633. 2:30:26Python.
  3634. 2:30:28Python.
  3635. 2:30:30We're going to search for Python. Yeah.
  3636. 2:30:33So, you are going to want to install the
  3637. 2:30:36official Python extension from
  3638. 2:30:39Microsoft. It is this one that has the
  3639. 2:30:41blue check mark next to Python.
  3640. 2:30:45Uh, so there now there are other ones
  3641. 2:30:49here,
  3642. 2:30:50but we just want the one that says
  3643. 2:30:53Python
  3644. 2:30:55from Microsoft. Do we see that extension
  3645. 2:30:57when you type in Python? Do we see that
  3646. 2:31:00one?
  3647. 2:31:02So just so it should just say Python. It
  3648. 2:31:05should be Microsoft.
  3649. 2:31:07Uh it's really popular. It has a lot of
  3650. 2:31:10downloads. Over 192 million downloads as
  3651. 2:31:13an extension.
  3652. 2:31:15It's from Microsoft.
  3653. 2:31:18Okay. Click on that.
  3654. 2:31:21Click on that.
  3655. 2:31:24And then you should see an install
  3656. 2:31:26button. It I already have it installed.
  3657. 2:31:28So it says uninstalled. Right here there
  3658. 2:31:29should be an install button. Install the
  3659. 2:31:32Python extension.
  3660. 2:31:35So out of 192 million installs,
  3661. 2:31:39really popular extension.
  3662. 2:31:44Are you guys able to install it?
  3663. 2:31:47You want to install that? It should be
  3664. 2:31:50pretty quick.
  3665. 2:31:55It should be pretty quick. It's not that
  3666. 2:31:57big of an extension.
  3667. 2:32:02So, what this does is
  3668. 2:32:07just the Python Sherry. It's just a
  3669. 2:32:09Python one. If you go into the
  3670. 2:32:11extensions and then search for Python,
  3671. 2:32:14it is just the one. It's just this one
  3672. 2:32:15that says Python and it's from
  3673. 2:32:17Microsoft.
  3674. 2:32:19Python blue check mark Microsoft.
  3675. 2:32:22You want that one.
  3676. 2:32:25And then you want to click on that one
  3677. 2:32:27and then hit the install.
  3678. 2:32:43Um, Roberto, is that for a co-pilot?
  3679. 2:32:51Is that for a co-pilot? I
  3680. 2:32:54maybe try closing it and reopening it.
  3681. 2:32:57Try closing VS Code, reopening and
  3682. 2:32:59retrying the install.
  3683. 2:33:08Um, no, we're not opening any folders
  3684. 2:33:10right now. We're not opening it. We're
  3685. 2:33:12just installing the extension.
  3686. 2:33:14That's all. We're just installing the
  3687. 2:33:15extension.
  3688. 2:33:18We're not opening any project folders.
  3689. 2:33:22just installing the extension.
  3690. 2:33:25Were we were we able to install that?
  3691. 2:33:46I know there's a lot by Microsoft, but
  3692. 2:33:48there should just be one that that says
  3693. 2:33:50Python.
  3694. 2:33:54there. So see how the name like this
  3695. 2:33:57name is this name here is eyesore. This
  3696. 2:33:59name is Python debugger. This one is
  3697. 2:34:01pilance.
  3698. 2:34:04Just the one that says Python.
  3699. 2:34:10Just that one.
  3700. 2:34:12That's the one we want. Only that one
  3701. 2:34:14right now.
  3702. 2:34:19Okay. Perfect. Perfect.
  3703. 2:34:22Okay, great.
  3704. 2:34:29Okay, so one more extension for you
  3705. 2:34:32guys. So once you install that one, I
  3706. 2:34:34have one more for you that you want to
  3707. 2:34:35install.
  3708. 2:34:38Are we ready for that one? One more we
  3709. 2:34:41want to install.
  3710. 2:34:45Okay, we're ready for the next one. So,
  3711. 2:34:48the next one you want to install
  3712. 2:34:51is the Jupiter extension,
  3713. 2:34:58which is the Jupiter.
  3714. 2:35:01It's this one. It's the very first one
  3715. 2:35:03here on my screen. So, it's it says
  3716. 2:35:05Jupiter
  3717. 2:35:06and it's from Microsoft.
  3718. 2:35:11Okay, we want to install that one.
  3719. 2:35:14Jupiter and it's from Microsoft. want to
  3720. 2:35:17install that one.
  3721. 2:35:21So, this one has 98 million uh installs.
  3722. 2:35:26You want to install this one.
  3723. 2:35:29Did you guys find that one? So, you want
  3724. 2:35:31to type in Jupy
  3725. 2:35:34Ter and it should be the Jupiter
  3726. 2:35:39extension here
  3727. 2:35:42that is uh from Microsoft.
  3728. 2:35:46So you want to install that one.
  3729. 2:35:51Great.
  3730. 2:35:53Now what does this one do? This
  3731. 2:35:55extension will allow you to work with
  3732. 2:36:00Jupiter notebooks inside of VS Code if
  3733. 2:36:03you want to.
  3734. 2:36:06So you Jupiter notebook has its own
  3735. 2:36:10standalone program which we will look at
  3736. 2:36:11next.
  3737. 2:36:14But you can open you can have those
  3738. 2:36:17files, those Jupyter notebook files be
  3739. 2:36:19compatible with VS Code and open them
  3740. 2:36:21and edit them and run them inside of VS
  3741. 2:36:23Code if you want to. So this extension
  3742. 2:36:27gives you the flexibility to work with
  3743. 2:36:29notebooks inside of VS Code. So you
  3744. 2:36:31never have to leave VS Code if you want
  3745. 2:36:32to work with notebooks. Um,
  3746. 2:36:35so this is a good extension if you
  3747. 2:36:38really want to work with notebooks and
  3748. 2:36:39stay inside of VS Code.
  3749. 2:36:48Yes. Uh when you Yeah. When you install
  3750. 2:36:51install an extension, it might it might
  3751. 2:36:53install a couple other dependency
  3752. 2:36:55extensions. Yes. But that's okay. Those
  3753. 2:36:57are required. That's okay.
  3754. 2:37:01That's that's that's okay.
  3755. 2:37:07All right. How do we feel? Good. Uh did
  3756. 2:37:09we get those installed?
  3757. 2:37:12Did we do were we able to get those
  3758. 2:37:13installed?
  3759. 2:37:16Okay, here is how we will test that it
  3760. 2:37:20all worked. So, we're going to do
  3761. 2:37:22something really simple.
  3762. 2:37:25Here's how we will test that it worked.
  3763. 2:37:30Let me go out of here and back to our
  3764. 2:37:32files.
  3765. 2:37:35So, out of the extensions, I just went
  3766. 2:37:37to the top button where it's the little
  3767. 2:37:39file um icon and um I am going to
  3768. 2:37:46um
  3769. 2:37:47go up to the very very top where it um
  3770. 2:37:50so you guys see on your VS Code window
  3771. 2:37:53where it says file, edit, selection,
  3772. 2:37:55view. I'm just going to create um
  3773. 2:38:00I'm just going to create uh a new
  3774. 2:38:04new file.
  3775. 2:38:08So, do you guys see that where where you
  3776. 2:38:09say file edit selection view? Click on
  3777. 2:38:12file and then click on new file.
  3778. 2:38:19You should see what I see on this screen
  3779. 2:38:21right here.
  3780. 2:38:25If you see
  3781. 2:38:28if you see Python and Jupyter notebook
  3782. 2:38:31then you know those are installed
  3783. 2:38:33correctly.
  3784. 2:38:34Do you guys see these options text
  3785. 2:38:36Python and Jupyter notebook?
  3786. 2:38:42Great. So what that means is we we can
  3787. 2:38:44now create those kind of files in the
  3788. 2:38:47future. We can create notebooks. you can
  3789. 2:38:50create Python files and VS Code will be
  3790. 2:38:52able to work with those.
  3791. 2:38:57If you don't see Python, that means your
  3792. 2:38:59Python extension didn't install yet or
  3793. 2:39:03you didn't install it. So, you want to
  3794. 2:39:04go back to you want to go back to your
  3795. 2:39:08extensions and make sure you installed
  3796. 2:39:09Python.
  3797. 2:39:11So go go go to this button over here,
  3798. 2:39:13the extensions,
  3799. 2:39:17type in Python,
  3800. 2:39:21and then make sure you install this
  3801. 2:39:23Python extension.
  3802. 2:39:29Okay. So, you're going to install the
  3803. 2:39:32Python extension and you're going to
  3804. 2:39:35install the Jupiter extension,
  3805. 2:39:40which is this, and install both of
  3806. 2:39:42those.
  3807. 2:39:44Make sure those are installed. If
  3808. 2:39:46they're installed and you still didn't
  3809. 2:39:48see that when you went to file um new
  3810. 2:39:51file,
  3811. 2:39:53if you don't see those, then um try
  3812. 2:39:58exiting VS Code and relaunching it.
  3813. 2:40:02Okay? Try exiting VS Code and reopening
  3814. 2:40:04it and seeing if you can make a new
  3815. 2:40:06file.
  3816. 2:40:11Okay? But it should be under uh at the
  3817. 2:40:14top file and then new file
  3818. 2:40:19and then you should see those options
  3819. 2:40:21Python and Jupiter.
  3820. 2:40:25Once you have those extension installed,
  3821. 2:40:26you may need to close out of VS Code and
  3822. 2:40:29reopen it to see that.
  3823. 2:40:40Okay, perfect. after you relaunched.
  3824. 2:40:42Okay, perfect. Yeah, you may need to
  3825. 2:40:44relaunch so that it can show the it can
  3826. 2:40:47show the extensions.
  3827. 2:40:51Yeah,
  3828. 2:40:54perfect.
  3829. 2:40:57Okay, perfect. So, that's set up for you
  3830. 2:40:59guys. So, um Perfect. It's set up for
  3831. 2:41:02you guys. Uh we will work with it in the
  3832. 2:41:05future, but just wanted to make sure it
  3833. 2:41:07was installed and set up. Once we start
  3834. 2:41:09working with Python, um I will show you
  3835. 2:41:12guys how to how to work with it. Um but
  3836. 2:41:15glad that's set up for now.
  3837. 2:41:23Uh what issue are you having uh Romero?
  3838. 2:41:28Is it not showing? It's not showing
  3839. 2:41:29Python or Jupiter for you when you do
  3840. 2:41:31file new file.
  3841. 2:41:34It's not showing those.
  3842. 2:41:37You may need to exit VS Code and reopen
  3843. 2:41:40it.
  3844. 2:41:48You uh Sil, yeah, you can you can make
  3845. 2:41:51one. We're not going to do anything with
  3846. 2:41:52it right now.
  3847. 2:41:56It's not going to you're not going to do
  3848. 2:41:57anything with it right now, but um
  3849. 2:42:06it's make sure you're searching for it
  3850. 2:42:08with a Y. It's J U P Y T E R.
  3851. 2:42:14You have to search. You have to So when
  3852. 2:42:16you go when you click on the extension,
  3853. 2:42:19search for JUP
  3854. 2:42:23Y. It should be the first thing that
  3855. 2:42:25shows up with Jupy Ter.
  3856. 2:42:29It's this Jupiter one from Microsoft.
  3857. 2:42:36I kernel I'll So the let me show us let
  3858. 2:42:39me show us that later. The kernel you
  3859. 2:42:41have to um you have to have a Python
  3860. 2:42:43interpreter.
  3861. 2:42:46So you may need to install a Python
  3862. 2:42:48interpreter to to be able to run the
  3863. 2:42:50kernel. So, I need to show us that. Um,
  3864. 2:42:54but I I don't want to get into that
  3865. 2:42:55right now.
  3866. 2:43:02Save what to
  3867. 2:43:10Oh, wherever you want. Wherever you want
  3868. 2:43:13on your own machine. It's up to you. It
  3869. 2:43:15doesn't really matter. Just wherever you
  3870. 2:43:17want.
  3871. 2:43:26It doesn't matter. It's up to you.
  3872. 2:43:29All right. So, what I want to do is uh I
  3873. 2:43:33want to take a break. Um because now,
  3874. 2:43:36you know, I said after two hours, we'll
  3875. 2:43:39take a longer break. Um so, we will now
  3876. 2:43:43we'll take a 10-minute break. Now, um if
  3877. 2:43:47you're still having any issues, um we
  3878. 2:43:49can try to get you set up at the end of
  3879. 2:43:51class. Um but we are going to set up.
  3880. 2:43:54So, coming up after our break, we're
  3881. 2:43:56going to take a 10-minute break. Coming
  3882. 2:43:57up after that, we'll we'll go and
  3883. 2:44:00install Jupyter Notebook. And then after
  3884. 2:44:03that, we will look at Collab. So, you're
  3885. 2:44:05going to have multiple options to run
  3886. 2:44:07Python. Not So, if this wasn't working
  3887. 2:44:09for you, that's okay. We'll try a
  3888. 2:44:11different route.
  3889. 2:44:13Okay? I will try a different route. Um I
  3890. 2:44:16I know Collab will work for you because
  3891. 2:44:18that is hosted by Google and really easy
  3892. 2:44:21to get working with. So at the worst
  3893. 2:44:23case scenario, Collab will work for you.
  3894. 2:44:25I know it. Um but we'll try to get
  3895. 2:44:28Jupyter Notebooks installed for you as
  3896. 2:44:30well. But if you're having issues with
  3897. 2:44:31VS Code, let me know at the end of
  3898. 2:44:33class. We'll try to get you set up,
  3899. 2:44:35okay?
  3900. 2:44:37You're still having issues with it.
  3901. 2:44:42But um what we're going to do right now
  3902. 2:44:43is take take a 10-minute break.
  3903. 2:44:48So let's try to be back um in about uh
  3904. 2:44:5210 minutes. Let's call it an even um
  3905. 2:44:56let's call it an even
  3906. 2:45:01uh what will we be covering? Um
  3907. 2:45:03installing the other installing the
  3908. 2:45:05other um Python setups. So Jupyter
  3909. 2:45:07notebook and working with collab. And
  3910. 2:45:10then we will get into the basics of
  3911. 2:45:11Python's the syntax. So we're going to
  3912. 2:45:13talk about indentation, identifiers, um
  3913. 2:45:16maybe if we have time, basic variable
  3914. 2:45:18types, data types. Yep. So we'll get
  3915. 2:45:20into Python.
  3916. 2:45:22We will get into Python today.
  3917. 2:45:25All right. So let's jump over to
  3918. 2:45:28uh Jupiter notebooks. So um what's so
  3919. 2:45:32special about Jupiter?
  3920. 2:45:35Well, it turns out that uh Jupiter is a
  3921. 2:45:39platform for running what are called
  3922. 2:45:42notebook files. So obviously we just
  3923. 2:45:44installed the Jupiter extension in VS
  3924. 2:45:46Code which will allow us to run
  3925. 2:45:48notebooks in VS code but Jupiter has its
  3926. 2:45:52own notebook platform and that's what
  3927. 2:45:54you will install in this setup. Um,
  3928. 2:45:58notebooks are special. They are um
  3929. 2:46:01really great um Python code files that
  3930. 2:46:06give us the ability to execute isolated
  3931. 2:46:09what are called cells of code. So we can
  3932. 2:46:13run one cell at a time and test and
  3933. 2:46:16debug the execution of that single cell
  3934. 2:46:19without affecting any of the other
  3935. 2:46:21cells. So, um, notebooks are great for,
  3936. 2:46:26uh, running code live and interactive.
  3937. 2:46:29When we do a lot of our demos in this
  3938. 2:46:31program, they're all going to be in
  3939. 2:46:33notebooks. Um, so that we can kind of
  3940. 2:46:36run things one cell at a time. Um,
  3941. 2:46:41uh,
  3942. 2:46:42no. So without notebooks you either have
  3943. 2:46:45to run you run like a Python script like
  3944. 2:46:49a Python file um which is a py file and
  3945. 2:46:54usually you have to either run that
  3946. 2:46:55through a debugger or run the entire
  3947. 2:46:58script at once. You don't really get
  3948. 2:47:01code isolated into individual cells
  3949. 2:47:03which is really nice with notebooks. The
  3950. 2:47:06other thing is notebooks are easily
  3951. 2:47:08sharable.
  3952. 2:47:10So you can share a notebook with
  3953. 2:47:12somebody else and they can open it and
  3954. 2:47:13see all of your inputs and outputs in
  3955. 2:47:16the notebook which is really nice like
  3956. 2:47:17all of the outputs get saved into the
  3957. 2:47:20notebook. Um which is nice. So and
  3958. 2:47:25notebooks uh especially in the Jupiter
  3959. 2:47:28platform are going to have all the data
  3960. 2:47:30science libraries available to them. So,
  3961. 2:47:32uh, if you're people usually love doing
  3962. 2:47:34notebooks for working with data, um,
  3963. 2:47:38really easy to work with data inside of
  3964. 2:47:40notebooks and and build things like
  3965. 2:47:42plots. You can display your,
  3966. 2:47:45uh, you can display your graphs really
  3967. 2:47:48easily inside of the notebook and then
  3968. 2:47:50share your notebook so other people can
  3969. 2:47:52see your graphs. Um, so notebooks are
  3970. 2:47:55really awesome like interactive
  3971. 2:47:57environments for running code. Um and we
  3972. 2:48:01will favor notebooks uh as our primary
  3973. 2:48:04way of running code throughout the
  3974. 2:48:06program. Now where you open those
  3975. 2:48:08notebooks is up to you. You can open
  3976. 2:48:11them in VS Code. You can open them in
  3977. 2:48:12the Jupyter notebook platform. Uh
  3978. 2:48:16you can open them inside of Collab and
  3979. 2:48:19run notebooks in Collab. Uh notebooks
  3980. 2:48:23are very very popular.
  3981. 2:48:28Why isn't running in notebooks the
  3982. 2:48:29default? It's because uh not all code
  3983. 2:48:32runs inside of cells. Like applications
  3984. 2:48:34are not going to be well suited for
  3985. 2:48:36notebooks. Like Instagram is not running
  3986. 2:48:39in a notebook. Uh it's more structured
  3987. 2:48:41into actual Python files and actual uh
  3988. 2:48:45more structured programs are going to be
  3989. 2:48:47not in a notebook. Notebook is more for
  3990. 2:48:51prototyping and debugging and uh
  3991. 2:48:55executing small chunks of code to test
  3992. 2:48:58it out. It's not for writing larger
  3993. 2:49:00programs like an like a
  3994. 2:49:03an LLM application like a chatbot would
  3995. 2:49:06generally be in not in a notebook. It'd
  3996. 2:49:08be in like a Python file.
  3997. 2:49:15Uh cells versus class objects. So cells
  3998. 2:49:17are just small uh think of them as small
  3999. 2:49:22little environments to execute our code.
  4000. 2:49:24Um class objects are actual chunks of
  4001. 2:49:27code that define an object. They're
  4002. 2:49:29they're different things.
  4003. 2:49:34Yeah, different things. We'll we'll
  4004. 2:49:35learn about objects. Um and we will
  4005. 2:49:38certainly see what cells are as we go
  4006. 2:49:40through. I'm going to show you an
  4007. 2:49:41example of a cell coming up when we
  4008. 2:49:43install Jupiter.
  4009. 2:49:45But uh let's talk about let's uh go
  4010. 2:49:48through the installation of Jupyter
  4011. 2:49:50notebook so you can see what a notebook
  4012. 2:49:52looks like. I think that'll be helpful
  4013. 2:49:54to orient.
  4014. 2:49:58So let's go over to that demo. So this
  4015. 2:50:01is going to be demo two
  4016. 2:50:04uh demo two inside of um uh lesson one.
  4017. 2:50:09So we're going to go over to that.
  4018. 2:50:17Everybody has this one. Okay, perfect.
  4019. 2:50:19Okay, so you're going to follow this
  4020. 2:50:20instruction. Now, what this is going to
  4021. 2:50:22do is first
  4022. 2:50:25um
  4023. 2:50:27No, this has not this is not going to be
  4024. 2:50:29anything to do with VS Code. This is
  4025. 2:50:31going to be a different platform. This
  4026. 2:50:32is going to be Jupiter.
  4027. 2:50:36Where is this? This is the
  4028. 2:50:38This is the demos.
  4029. 2:50:41This is uh demo two inside of that demos
  4030. 2:50:44folder that we said uh
  4031. 2:50:48to to uh grab all the demos
  4032. 2:50:52from your LMS.
  4033. 2:50:56Does anybody have that uh demo 2 PDF
  4034. 2:50:59they can upload? I I think somebody
  4035. 2:51:01uploaded all of them earlier, but if you
  4036. 2:51:02have demo two, want to upload it real
  4037. 2:51:05quick? I don't have the PDFs.
  4038. 2:51:11if somebody wants to share that.
  4039. 2:51:15So there so they're different. Um VS So
  4040. 2:51:18what I was saying is you can open
  4041. 2:51:21notebooks inside of VS Code and the
  4042. 2:51:23thing that allows you to open notebooks
  4043. 2:51:25in VS Code is the extension.
  4044. 2:51:28So yes, if you're going to work with
  4045. 2:51:29notebooks in VS Code, you need the
  4046. 2:51:30extension installed. But you can use the
  4047. 2:51:34standalone Jupiter platform
  4048. 2:51:37to work with notebooks. It's up to you.
  4049. 2:51:40If you like using VS Code,
  4050. 2:51:43um if you like using VS Code, you can do
  4051. 2:51:45it that way. If you like uh the Jupiter
  4052. 2:51:48platform, you can do it that way. It's
  4053. 2:51:51up to you. It's just a preference. I'm
  4054. 2:51:54giving you guys options. That's my goal
  4055. 2:51:57is to give you options and let you guys
  4056. 2:51:59choose what you're most comfortable
  4057. 2:52:00with.
  4058. 2:52:04Okay. And we're and we're taking time to
  4059. 2:52:06do that now in the beginning of the
  4060. 2:52:08program, right? Because we're going to
  4061. 2:52:10be doing a lot of Python examples coming
  4062. 2:52:12up as we start learning Python. So, it's
  4063. 2:52:15it's valuable to spend that time now. I
  4064. 2:52:17know it can seem a little slow, but I
  4065. 2:52:21promise it'll be worth it so that you
  4066. 2:52:22guys have options for running your
  4067. 2:52:24running your code.
  4068. 2:52:29Yes. Thank you guys for uploading those.
  4069. 2:52:31appreciate it. Those are the demos you
  4070. 2:52:34want to uh follow along with.
  4071. 2:52:38Okay, so the first step here is going to
  4072. 2:52:40be to install Anaconda. Now, you may be
  4073. 2:52:43wondering, what is Anaconda? I thought
  4074. 2:52:44we were talking about Jupiter, and
  4075. 2:52:47that's a valid question. Anaconda is a
  4076. 2:52:51what's called a distribution of Python.
  4077. 2:52:55So Anaconda is a program a software a
  4078. 2:52:59collection of software programs that
  4079. 2:53:02give you a version of Python with a
  4080. 2:53:05bunch of packages
  4081. 2:53:08uh with a bunch of packages already
  4082. 2:53:11installed.
  4083. 2:53:13Um and then
  4084. 2:53:17uh one of those is the Jupiter package
  4085. 2:53:21so that you can run Jupyter notebooks.
  4086. 2:53:24And what Jupyter notebooks will be
  4087. 2:53:28is a uh basically a web browser
  4088. 2:53:32application that will open up a notebook
  4089. 2:53:37editor in your web browser. So that's
  4090. 2:53:40ultimately what we're going to do, but
  4091. 2:53:42we are going to install it via the
  4092. 2:53:44Anaconda distribution
  4093. 2:53:47uh via the Anaconda distribution of
  4094. 2:53:50Python.
  4095. 2:53:51So that's where we're going to start is
  4096. 2:53:53with the initial download of Anaconda.
  4097. 2:53:58Oh, it's no no skipped registration.
  4098. 2:54:01Okay, let me let me uh open the link.
  4099. 2:54:04I think there I think there's a way to
  4100. 2:54:06find it without having to do the
  4101. 2:54:07registration.
  4102. 2:54:11There's a way to get to it without
  4103. 2:54:12having to do that. I'm going to find it
  4104. 2:54:14real quick.
  4105. 2:54:21Oh, you can't. Okay. So, if you can't
  4106. 2:54:22install it, that's okay. We will be able
  4107. 2:54:24to work with notebooks in collab and you
  4108. 2:54:27can work with notebooks inside of VS
  4109. 2:54:29Code. That's fine, too.
  4110. 2:54:38Yeah, I'm getting I I'm going through
  4111. 2:54:40the registration process so I can um I
  4112. 2:54:42can show you that install.
  4113. 2:54:46Okay, let me share my screen.
  4114. 2:54:50Did you guys get to once you go through
  4115. 2:54:52the like setting up your account, do you
  4116. 2:54:55get to this page?
  4117. 2:55:00Do you get to this page for those of you
  4118. 2:55:02going through? Yeah, that looks right
  4119. 2:55:04for you, Ashish. That looks right.
  4120. 2:55:11Do you guys get to this page though when
  4121. 2:55:13you get through your like account setup?
  4122. 2:55:16Okay, you got to this page. Okay, so
  4123. 2:55:18then choose your correct Windows or Mac
  4124. 2:55:21down. You want to be over here on the
  4125. 2:55:23left. You want to do Anaconda
  4126. 2:55:25distribution.
  4127. 2:55:27This is what you want to do. So, choose
  4128. 2:55:30the right one. And if you're on an M1,
  4129. 2:55:31M2, M3, you're going to do the silicon.
  4130. 2:55:35If you're on an older Mac, you're going
  4131. 2:55:37to do the 64. And then obviously, if
  4132. 2:55:39you're on a Windows, you should be
  4133. 2:55:40clicking over here to do Windows. But
  4134. 2:55:43you want to do the Anaconda
  4135. 2:55:44distribution, not Minion. Okay. So,
  4136. 2:55:47click on the installer for Anaconda
  4137. 2:55:51distribution.
  4138. 2:55:56Okay? And then let that install. Now
  4139. 2:56:00while that's installing let me explain
  4140. 2:56:02something about the difference between
  4141. 2:56:05uh I think it was asked earlier what's
  4142. 2:56:07the difference between um Anaconda
  4143. 2:56:11uh as the default Python. So Anaconda
  4144. 2:56:16as I was saying earlier is a version of
  4145. 2:56:19Python that has a bunch of data science
  4146. 2:56:22and machine learning packages already
  4147. 2:56:24installed for you. So uh it comes with a
  4148. 2:56:29bunch of packages that are already
  4149. 2:56:31installed. So if you use that Python
  4150. 2:56:35um that Python has a bunch of packages
  4151. 2:56:38built in with it that you don't need to
  4152. 2:56:39go out and install. So, Anaconda is a
  4153. 2:56:42very popular version of Python for
  4154. 2:56:45people to install that are working in
  4155. 2:56:46data science, AI, ML. Very popular
  4156. 2:56:50version because it already comes with a
  4157. 2:56:52bunch of packages that you would use for
  4158. 2:56:54manipulating data for doing machine
  4159. 2:56:57learning or doing anything in AI. So,
  4160. 2:57:00it's it's a very um popular
  4161. 2:57:02distribution. It also comes with
  4162. 2:57:06Jupiter, which is why we wanted to use
  4163. 2:57:09it because it comes with the notebook
  4164. 2:57:11capability out of the box.
  4165. 2:57:16Okay. So, I'm going to launch. So, when
  4166. 2:57:19this is done installing, you want to
  4167. 2:57:22launch the program that gets installed
  4168. 2:57:25called the uh Anaconda Navigator.
  4169. 2:57:29So, it should install a program on your
  4170. 2:57:31machine called the Anaconda Navigator.
  4171. 2:57:33Do you guys have that? Did anybody get
  4172. 2:57:35through and and you have that program?
  4173. 2:57:38The Anaconda Navigator.
  4174. 2:57:41You don't need any advanced ones. You
  4175. 2:57:43don't need any advanced options.
  4176. 2:57:47Just the just the defaults. All the
  4177. 2:57:49defaults
  4178. 2:57:51should be good.
  4179. 2:57:57Still downloading. Okay. I'm going to
  4180. 2:57:58show you
  4181. 2:58:00I'm going to show you what the navigator
  4182. 2:58:02looks like once you once you have it.
  4183. 2:58:06That's okay if it takes a little bit of
  4184. 2:58:07time to download. That's okay. Um,
  4185. 2:58:09basically once you download it, um, you
  4186. 2:58:12just have to click a couple more buttons
  4187. 2:58:13and then you can access Jupiter.
  4188. 2:58:17Okay, let me share my screen and show
  4189. 2:58:20you what you like. Once it installs,
  4190. 2:58:23this is what it should look like. It's
  4191. 2:58:24okay if it's taking a little bit of
  4192. 2:58:25time.
  4193. 2:58:27You should have something that kind of
  4194. 2:58:30looks like this, which is the um
  4195. 2:58:32dashboard that has the different
  4196. 2:58:35programs available to you to you.
  4197. 2:58:39Um
  4198. 2:58:41do you guys see something like this? If
  4199. 2:58:44you have the navigator,
  4200. 2:58:47do you see something like this?
  4201. 2:58:52which is the which is the like when you
  4202. 2:58:54open the navigator program, you should
  4203. 2:58:56see something like this that has a bunch
  4204. 2:58:58of different um
  4205. 2:59:04you do. Okay.
  4206. 2:59:07It's if it's taking a little bit of time
  4207. 2:59:08that's okay.
  4208. 2:59:11Yes. Na Anaconda Navigator is how you
  4209. 2:59:13launch Yes.
  4210. 2:59:17Anaconda Navigator is how you launch it.
  4211. 2:59:19So yeah, you want to open that. Now the
  4212. 2:59:21the whole reason to come here
  4213. 2:59:24is so we can launch Jupiter notebooks.
  4214. 2:59:28So we can launch Jupyter notebooks. Um
  4215. 2:59:33this is the program we ultimately want
  4216. 2:59:34to launch. This is going to
  4217. 2:59:38uh allow us to open notebook files, edit
  4218. 2:59:41them, run code cells. I'll show you what
  4219. 2:59:44a notebook looks like in a second. But
  4220. 2:59:47but once you have Anaconda installed,
  4221. 2:59:51open the navigator and then launch
  4222. 2:59:54Jupiter notebook. It's just one extra
  4223. 2:59:56step. Launch the Jupiter notebook.
  4224. 3:00:00What that should do is launch the the
  4225. 3:00:04notebook.
  4226. 3:00:06Uh it should launch the web browser of
  4227. 3:00:11your like whatever you have as your
  4228. 3:00:12default web browser. It should open the
  4229. 3:00:14notebook program in every in your web
  4230. 3:00:17browser. So if it's Chrome, Firefox,
  4231. 3:00:19Edge, whatever your default web browser,
  4232. 3:00:21it's going to launch the notebook
  4233. 3:00:23program in the browser.
  4234. 3:00:27Okay.
  4235. 3:00:31So, I'm going to launch it and then I'll
  4236. 3:00:34show you what it looks like. Again, if
  4237. 3:00:35it's taking you a little bit of time,
  4238. 3:00:36that's okay. Whenever it's done,
  4239. 3:00:41how did I get to these icons? Just
  4240. 3:00:43launch. Do you have the Anaconda
  4241. 3:00:44Navigator program?
  4242. 3:00:46It should have got It should be
  4243. 3:00:48installed.
  4244. 3:00:49Open the Anaconda Navigator program.
  4245. 3:00:53It should have been installed with the
  4246. 3:00:54Anaconda installation.
  4247. 3:01:05All right. Was anybody able to get to
  4248. 3:01:06this the Jupiter this? So, it should
  4249. 3:01:09launch in your browser. Anybody
  4250. 3:01:13able to get to that?
  4251. 3:01:16Fantastic. Fantastic. I'm glad some of
  4252. 3:01:18you guys are able to get to it. And if
  4253. 3:01:20it's it's not yet, that's okay.
  4254. 3:01:22Remember, when it's done installing,
  4255. 3:01:23you're going to go to Anaconda Navigator
  4256. 3:01:26and then
  4257. 3:01:28uh launch Jupiter Notebook. That's
  4258. 3:01:32That's what you're going to do.
  4259. 3:01:34That's okay, Roberto. It's okay.
  4260. 3:01:37All right. I do want to I want to show
  4261. 3:01:39you guys a notebook. I just want to show
  4262. 3:01:42you what it looks like. What I'm going
  4263. 3:01:44to do is I'm going to
  4264. 3:01:48um open a notebook by going to new and
  4265. 3:01:52then Python 3 notebook. So you can open
  4266. 3:01:55a folder, you can open a terminal, you
  4267. 3:01:57can open a text file. I'm going to open
  4268. 3:01:59a Python 3 which is a notebook. You so
  4269. 3:02:02it's a it's a a notebook powered by
  4270. 3:02:04Python.
  4271. 3:02:06So I'm going to click on that which will
  4272. 3:02:09launch a new notebook in a new tab.
  4273. 3:02:14And here I am in the notebook editor. So
  4274. 3:02:17now I am in a notebook editor screen. So
  4275. 3:02:20if you go when you first launch Jupiter
  4276. 3:02:23you you can navigate to notice that that
  4277. 3:02:26notebook got created here where I
  4278. 3:02:28currently am on my machine. I could
  4279. 3:02:30navigate to I could navigate to
  4280. 3:02:33documents and then I could um you know
  4281. 3:02:37create a new file there or I could make
  4282. 3:02:39a new folder here and and do it that
  4283. 3:02:42way. Um but uh I am uh just editing this
  4284. 3:02:49notebook right here within this um
  4285. 3:02:51current folder that I'm in.
  4286. 3:02:58Okay. So, do you guys remember when I
  4287. 3:03:00said that code gets executed in a in an
  4288. 3:03:02isolated cell?
  4289. 3:03:05Do you remember that?
  4290. 3:03:08Um,
  4291. 3:03:09this is what a cell looks like. And you
  4292. 3:03:13can make new cells by hitting this plus
  4293. 3:03:15button.
  4294. 3:03:17So, you hit this plus button over here,
  4295. 3:03:19you can make new cells. So, if you hit
  4296. 3:03:23plus++,
  4297. 3:03:25I'm making a bunch of cells.
  4298. 3:03:27Now, what's really cool about cells?
  4299. 3:03:30Yeah, Tim just discovered this. What's
  4300. 3:03:32really cool about cells is you can
  4301. 3:03:34change them to be text or code. So, if
  4302. 3:03:38you change it to markdown,
  4303. 3:03:41I can write markdown text in here to say
  4304. 3:03:44this is my notebook. And then if I run
  4305. 3:03:48this, it's going to display as text.
  4306. 3:03:52So if I run that cell which uh when I'm
  4307. 3:03:56editing it I can click run and it will
  4308. 3:04:00render that as text because I changed
  4309. 3:04:03the cell type to markdown. Markdown is
  4310. 3:04:06just a flavor of text style.
  4311. 3:04:11So otherwise we can write some Python
  4312. 3:04:14code. Now, what I want you guys to type,
  4313. 3:04:16I'll type this in the chat to verify
  4314. 3:04:19everything is working is I want you to
  4315. 3:04:21type print
  4316. 3:04:24hello world.
  4317. 3:04:27I want you to type that
  4318. 3:04:30inside of a cell
  4319. 3:04:35and then
  4320. 3:04:37and then hit run.
  4321. 3:04:43And it should run that code.
  4322. 3:04:47And you should see you should be able to
  4323. 3:04:49see uh you should be able to see that
  4324. 3:04:58Shift enter. Yep. You can whenever
  4325. 3:04:59you're on a cell, you can hit shift
  4326. 3:05:01enter. It'll run the cell.
  4327. 3:05:11You can That's okay. You can always You
  4328. 3:05:12can go back and watch the video. So,
  4329. 3:05:15this is being recorded. You can go back
  4330. 3:05:16and watch the video. I I know it's a
  4331. 3:05:19little frustrating. It's still
  4332. 3:05:20installing for you, but go back and
  4333. 3:05:23watch the video. And I definitely
  4334. 3:05:24encourage once it's installed to go back
  4335. 3:05:26and try this, which would be just
  4336. 3:05:30launching your Anaconda Navigator
  4337. 3:05:33and then launching Jupiter.
  4338. 3:05:41If you don't have, by the way, if you
  4339. 3:05:43don't have Python 3, um you may need to
  4340. 3:05:48uh exit your navigator and reopen it.
  4341. 3:05:53Okay, you may need to exit your your
  4342. 3:05:55navigator, reopen it so that you can
  4343. 3:05:56launch Jupiter again.
  4344. 3:06:05Were you guys able to run this in a
  4345. 3:06:07cell? For those of you that have Jupiter
  4346. 3:06:08running, were you able to run this?
  4347. 3:06:14Nice. And it worked for you. Okay,
  4348. 3:06:16perfect. Perfect.
  4349. 3:06:19So this is what I meant by this is an
  4350. 3:06:22isolated cell. So notice that we can run
  4351. 3:06:24this
  4352. 3:06:26and it doesn't affect
  4353. 3:06:30um
  4354. 3:06:32Sure. Sure. I hear you. I I hear you.
  4355. 3:06:34Update the doc. Uh I can How about I
  4356. 3:06:38post it in our um our Slack channel? By
  4357. 3:06:42the way, do you guys have access to the
  4358. 3:06:44to the Slack channel?
  4359. 3:06:51Okay, I can post it there. I can post
  4360. 3:06:53the instructions to get there.
  4361. 3:06:59Okay, I can post it in our our uh
  4362. 3:07:01cohort's uh channel.
  4363. 3:07:05I hear you. You don't want to search for
  4364. 3:07:07our video. I I hear you.
  4365. 3:07:16Uh I don't have the link on hand, but
  4366. 3:07:19you can get to it through the LMS.
  4367. 3:07:22So if you go to the LMS and go to
  4368. 3:07:27uh there should be
  4369. 3:07:31um
  4370. 3:07:33there should be a link to get to it
  4371. 3:07:35within there. It's should be like over
  4372. 3:07:38here on the right.
  4373. 3:07:43I don't have the link I don't have the
  4374. 3:07:45link to it off hand. Yeah,
  4375. 3:07:51but there should be a way to get to it
  4376. 3:07:53from the LMS.
  4377. 3:07:56Yeah, there should be a banner here. I
  4378. 3:07:59don't know why I don't have it, but
  4379. 3:08:01should be there.
  4380. 3:08:05Okay. So, if you're just getting things
  4381. 3:08:08installed, how do you get to here? Um,
  4382. 3:08:11you open the navigator.
  4383. 3:08:15Open the navigator.
  4384. 3:08:18Syntax is hello world.
  4385. 3:08:22It's just inside of it's just that print
  4386. 3:08:26hello world.
  4387. 3:08:30Um, open the navigator.
  4388. 3:08:34Open the navigator which looks like
  4389. 3:08:36this.
  4390. 3:08:38Let me share my screen.
  4391. 3:08:44Okay. Open the Anaconda Navigator that
  4392. 3:08:47got installed.
  4393. 3:08:48Then click launch on the Jupyter
  4394. 3:08:52notebook program. So you should have
  4395. 3:08:54this at least. You may have other ones.
  4396. 3:08:57Click on this launch. Uh click on this
  4397. 3:09:01launch and then you can launch the uh
  4398. 3:09:04Anaconda Navigator.
  4399. 3:09:14Okay.
  4400. 3:09:16If it if it's a little stuck, that's
  4401. 3:09:17okay. We're going to move on. We're
  4402. 3:09:19going to go to Collab, which can run
  4403. 3:09:21notebooks as well. So, if it seems a
  4404. 3:09:23little stuck, that's okay. I will post
  4405. 3:09:26in our Slack instructions on how to run
  4406. 3:09:28this.
  4407. 3:09:31That's okay.
  4408. 3:09:36All right. But what I wanted to do,
  4409. 3:09:39what I wanted to do before we move on to
  4410. 3:09:41collab is I just wanted to show you I
  4411. 3:09:43wanted to call out a couple things about
  4412. 3:09:45notebooks.
  4413. 3:09:47Um
  4414. 3:09:49is that uh a couple things about
  4415. 3:09:52notebooks. One is that notice that these
  4416. 3:09:54cells are very isolated. Whatever I put
  4417. 3:09:56here
  4418. 3:09:57um does not affect what I had before. So
  4419. 3:10:01I can add numbers like that and it can
  4420. 3:10:04um compute that and this does not affect
  4421. 3:10:07this. So so this is why notebooks are so
  4422. 3:10:09great is you can document the notebooks
  4423. 3:10:12with with mixing in text and code like
  4424. 3:10:16we do here. Um you can run code in its
  4425. 3:10:19own cells.
  4426. 3:10:21uh you can run code in its own cells and
  4427. 3:10:24then you can um have that very isolated.
  4428. 3:10:26So I could jump down here and run
  4429. 3:10:28something and that doesn't matter that
  4430. 3:10:30there's nothing here like it doesn't
  4431. 3:10:32need to be in order. I can um you know I
  4432. 3:10:36can uh run stuff out of I can overwrite
  4433. 3:10:38this
  4434. 3:10:42um and run that and it produces the
  4435. 3:10:44output. Uh
  4436. 3:10:48so you know many things we could do uh
  4437. 3:10:53inside of notebooks that are really
  4438. 3:10:54fantastic for just quickly prototyping
  4439. 3:10:57and running Python code inside of cells.
  4440. 3:11:00So it's very nice that way. So notebooks
  4441. 3:11:03are notebooks are really nice. You can
  4442. 3:11:04also like I could share this file. So
  4443. 3:11:07this produces a file on my machine. Uh
  4444. 3:11:10if I go back to the navigator
  4445. 3:11:13um
  4446. 3:11:15Anaconda navigator Jerry Anaconda
  4447. 3:11:18navigator
  4448. 3:11:22um
  4449. 3:11:24can you read value of variables from
  4450. 3:11:26another cell? Uh you you have to store
  4451. 3:11:29them into variables. So I could I could
  4452. 3:11:31call this uh x
  4453. 3:11:36and then I could refer to x later.
  4454. 3:11:39We'll learn about that. We'll learn
  4455. 3:11:40about that with variables.
  4456. 3:11:44But yes, you can kind of do that with
  4457. 3:11:46variables.
  4458. 3:11:53All right.
  4459. 3:11:56Um,
  4460. 3:12:02what do the numbers after?
  4461. 3:12:07Which numbers? these the ones in the
  4462. 3:12:09brackets.
  4463. 3:12:16Oh th so those are which cells we've uh
  4464. 3:12:20executed. So I executed this one first.
  4465. 3:12:23So it it's it's number one. And then I
  4466. 3:12:26executed um
  4467. 3:12:29uh this I think I did this second. So it
  4468. 3:12:32it's text. It doesn't really get one of
  4469. 3:12:34those. And then I did this one third.
  4470. 3:12:37And then I did um I think I did this one
  4471. 3:12:41fourth and then it got overwritten with
  4472. 3:12:43the fifth. So it just tells you like how
  4473. 3:12:45many executions you've done and what is
  4474. 3:12:47which number execution that was. Then I
  4475. 3:12:49did this one sixth.
  4476. 3:12:51It just keeps track of your executions.
  4477. 3:12:55Okay. So somebody asked about a kernel.
  4478. 3:12:58What is a kernel? So uh the kernel is um
  4479. 3:13:06the kernel is basically the interpreter.
  4480. 3:13:09So it's the thing that that the kernel
  4481. 3:13:11is just a a um a copy of the interpreter
  4482. 3:13:16that the notebook is attaching to in
  4483. 3:13:18order to run. So the notebook can't run
  4484. 3:13:21anything because remember a pi Python
  4485. 3:13:24needs
  4486. 3:13:25Python needs an interpreter to run its
  4487. 3:13:29code. So in notebooks we basically
  4488. 3:13:33create like a virtual copy of the
  4489. 3:13:35interpreter called a kernel. Um and you
  4490. 3:13:38can actually have many kernels based on
  4491. 3:13:41your um your base interpreter. So what's
  4492. 3:13:45nice is Anaconda
  4493. 3:13:49um Anaconda comes with
  4494. 3:13:53uh an interpreter for you and then you
  4495. 3:13:57create kernels that are virtual copies
  4496. 3:13:59of that um that are virtual copies of
  4497. 3:14:03that uh interpreter so that you can run
  4498. 3:14:05your Python code against it. Remember
  4499. 3:14:08you need an interpreter but notebooks
  4500. 3:14:11attach to kernels. Kernels are like
  4501. 3:14:14virtual interpreters.
  4502. 3:14:16Um, and you can have many kernels based
  4503. 3:14:18on the original interpreter. So the
  4504. 3:14:21kernel is literally just think of it
  4505. 3:14:23like the computer that's powering the
  4506. 3:14:25notebook. That's all. It's just the
  4507. 3:14:27compute engine that's allowing you to
  4508. 3:14:29execute your Python code. So every
  4509. 3:14:32notebook has an associated kernel.
  4510. 3:14:36And what's interesting is if you restart
  4511. 3:14:38your kernel, you lose all your data. So
  4512. 3:14:41all of these outputs that we have um you
  4513. 3:14:44would lose if you restarted your kernel.
  4514. 3:14:48So if I restart um now I like since I
  4515. 3:14:52restarted this is not going to know what
  4516. 3:14:54x is. So if I try to print x again it's
  4517. 3:14:57going to say I don't know what x is
  4518. 3:14:58because I restarted my kernel. I lost
  4519. 3:15:01all that data.
  4520. 3:15:05But I can redefine it. And then there it
  4521. 3:15:08is. And notice that my iterations
  4522. 3:15:10restart
  4523. 3:15:12um my iterations restart when I uh
  4524. 3:15:15restart my kernel. So if if I go back
  4525. 3:15:17and restart the kernel again
  4526. 3:15:20and now if I run this, this will be
  4527. 3:15:23first. So notice how that restarts to
  4528. 3:15:25first. This will be second. This will be
  4529. 3:15:28third.
  4530. 3:15:31Try restarting. I don't know what that
  4531. 3:15:33is.
  4532. 3:15:35I don't know why that
  4533. 3:15:40Yeah, choose the Anaconda. Either one.
  4534. 3:15:42Choose. You want to use Anaconda as your
  4535. 3:15:45But what that's saying is what do you
  4536. 3:15:46want to use as your interpreter to to
  4537. 3:15:48build your kernels. So, choose one of
  4538. 3:15:51those. That's fine.
  4539. 3:15:54So, yeah, Anaconda requirements, laptop
  4540. 3:15:56requirements. Um,
  4541. 3:15:58you need a little bit of you need a
  4542. 3:16:00little bit of RAM. Uh, you need a little
  4543. 3:16:04bit of RAM to run the notebooks because
  4544. 3:16:07you need some memory. Um, you don't need
  4545. 3:16:10a lot of it though. I'd be surprised if
  4546. 3:16:12you didn't meet the requirements. It's
  4547. 3:16:13not that much, but you do need some.
  4548. 3:16:17I'm not sure the exact. I'd have to look
  4549. 3:16:20that up on the Anaconda website.
  4550. 3:16:24Oh, it must have been full to start with
  4551. 3:16:26or pretty full. I'd be This doesn't take
  4552. 3:16:28up that much space, I don't think.
  4553. 3:16:33Was it pretty full to begin with?
  4554. 3:16:36I would assume. I don't think this takes
  4555. 3:16:38up that much space.
  4556. 3:16:42Uh, what I want to do then, I want to go
  4557. 3:16:44over to collab. Okay. How do we feel
  4558. 3:16:46about the notebooks? I maybe if it's
  4559. 3:16:48still installing for you, give it a
  4560. 3:16:50little time. Open the open the
  4561. 3:16:52navigator.
  4562. 3:16:54Let's try Coll. I guarantee you Collab
  4563. 3:16:56will work for you if you're still having
  4564. 3:16:58issues with if you're having issues with
  4565. 3:16:59Jupiter.
  4566. 3:17:01No worries. Let's just try collab. I
  4567. 3:17:03promise that will be a lot easier, be a
  4568. 3:17:06million times easier, I think, than than
  4569. 3:17:08working with Jupiter. Okay, great. So,
  4570. 3:17:11we will continue then.
  4571. 3:17:16All right, let me jump over to our
  4572. 3:17:20final
  4573. 3:17:21um
  4574. 3:17:23demo with setting up a collab notebook.
  4575. 3:17:26So I'm just going to jump into doing
  4576. 3:17:28that on in the interest of time.
  4577. 3:17:31Uh
  4578. 3:17:34so
  4579. 3:17:38would you be taking up additional
  4580. 3:17:39sessions too other than uh so we're I'm
  4581. 3:17:42going to be the instructor for all of
  4582. 3:17:45the courses in this program. So you're
  4583. 3:17:48you're stuck with me
  4584. 3:17:50for all of those. Does that make sense?
  4585. 3:17:53like all of the all of the uh AI
  4586. 3:17:57engineer program uh courses.
  4587. 3:18:04Yeah. Yeah. Stuck or be excited. It's
  4588. 3:18:08going to be one or the other. Probably
  4589. 3:18:10not an in between feeling.
  4590. 3:18:13Hopefully. Cool. Hopefully. Hopefully
  4591. 3:18:16good. Yeah. Like I said, I've taught
  4592. 3:18:19this many times. Uh, I think it would be
  4593. 3:18:24uh I think it'll be good.
  4594. 3:18:26You are stuck. Okay. Well, we're going
  4595. 3:18:28to get you unstuck with Collab. I would
  4596. 3:18:30not worry about getting Jupiter set up.
  4597. 3:18:33If it's not working for you, we're going
  4598. 3:18:34to ditch it and we're going to use
  4599. 3:18:35something else that works. I promise
  4600. 3:18:38it's not a I promise getting Jupiter set
  4601. 3:18:40up is not that important relative to
  4602. 3:18:42getting at least one of these options
  4603. 3:18:44that works.
  4604. 3:18:47So, if it's not working over on Jupiter,
  4605. 3:18:50I'm not worried in the slightest about
  4606. 3:18:52it because there's going to be plenty of
  4607. 3:18:53options to run run Python code. In fact,
  4608. 3:18:56we're going to do one next which is
  4609. 3:18:58going to be with um with Collab. So,
  4610. 3:19:02we'll do that. Um so, let me jump into
  4611. 3:19:07that. Let me share my screen here.
  4612. 3:19:11Um,
  4613. 3:19:20learning a lot already. Great. That's
  4614. 3:19:21great. Glad to hear that. Thank you.
  4615. 3:19:26Okay.
  4616. 3:19:29Thank you guys. Appreciate it.
  4617. 3:19:31All right. Let me go to the demo. Demo
  4618. 3:19:35three.
  4619. 3:19:37All right. So, what you want to do
  4620. 3:19:39essentially is to go to this website,
  4621. 3:19:43um, which I have here. I'm going to, uh,
  4622. 3:19:46paste it in the chat. Um, so what you
  4623. 3:19:50want to do is go to Google's website for
  4624. 3:19:52their Collab platform, uh, which is, so
  4625. 3:19:57Collab is a, um, notebook platform that
  4626. 3:20:02Google hosts. So, you don't need to
  4627. 3:20:04install anything. You just go there in
  4628. 3:20:05your favorite web browser, log in with
  4629. 3:20:08your Google account. In fact, I don't
  4630. 3:20:10even think you need to be necessarily
  4631. 3:20:12logged in. You can in order to save your
  4632. 3:20:14notebooks to your drive,
  4633. 3:20:16but um you go there and you basically
  4634. 3:20:20open up a notebook and start working
  4635. 3:20:22with it right away. And it's fantastic.
  4636. 3:20:26Their notebook environment already has a
  4637. 3:20:29lot of packages installed into it for AI
  4638. 3:20:32and machine learning. So that's f that's
  4639. 3:20:34really great. Um
  4640. 3:20:38uh once you get to the page um you
  4641. 3:20:41should log in though if you have a
  4642. 3:20:43Google account. Only reason I say that
  4643. 3:20:46is because it will save your notebooks
  4644. 3:20:48to your drive automatically so that you
  4645. 3:20:50it will automatically save your
  4646. 3:20:51notebooks just like as if you're working
  4647. 3:20:53in a Google doc. So that's great. So
  4648. 3:20:56that um it saves your work
  4649. 3:20:58automatically.
  4650. 3:20:59Um, so please, you know, I would
  4651. 3:21:01recommend getting a Google account if
  4652. 3:21:03you don't have one for free. Logging in
  4653. 3:21:05using Collab is completely free.
  4654. 3:21:08Um, so it's a fantastic platform. Um,
  4655. 3:21:12when you go to that site, uh, assuming
  4656. 3:21:14you've logged in, you want to click on
  4657. 3:21:16that lower left blue button where it
  4658. 3:21:18says new notebook. You can see it in
  4659. 3:21:20this screenshot. And I I'll open up one
  4660. 3:21:23in a moment on on my screen. But do you
  4661. 3:21:26guys see this screen right here that's
  4662. 3:21:28in this screenshot that has the new
  4663. 3:21:31notebook on the bottom in the lower
  4664. 3:21:32left?
  4665. 3:21:36No. From that site, what do you see?
  4666. 3:21:45Oh, so you're already in a notebook. It
  4667. 3:21:47you're already in a notebook. Like it
  4668. 3:21:48says, "Welcome to Collab.
  4669. 3:21:52Oh, okay. So, it already opened the
  4670. 3:21:53notebook for you. Okay, that's that's
  4671. 3:21:55fine. That's fine. I'll show you uh I'll
  4672. 3:21:58show you what that looks like. That's no
  4673. 3:22:00problem. That means you're already
  4674. 3:22:02inside of it.
  4675. 3:22:05Okay.
  4676. 3:22:08Okay. So, then we're pretty much in the
  4677. 3:22:10notebook environment and we can start
  4678. 3:22:12running code. Let me let me hop over to
  4679. 3:22:14Collab and show you guys what it looks
  4680. 3:22:16like.
  4681. 3:22:17Let me stop sharing that and jump over
  4682. 3:22:19to
  4683. 3:22:21collab here.
  4684. 3:22:32Okay.
  4685. 3:22:34So, if you're in the welcome to collab,
  4686. 3:22:36um that's fine or you can start a new
  4687. 3:22:40notebook. Let me assume that we've
  4688. 3:22:41opened up welcome to collab. So, you're
  4689. 3:22:43in this screen. What you want to do if
  4690. 3:22:45you're in this screen is just go to go
  4691. 3:22:47up to file and do new notebook.
  4692. 3:22:52Just go to file, new notebook in drive.
  4693. 3:22:54Just do that. File new notebook
  4694. 3:23:00and this will create a new notebook for
  4695. 3:23:02you which will start fresh a blank
  4696. 3:23:04notebook.
  4697. 3:23:06Okay.
  4698. 3:23:09Were you able to do that? If you guys
  4699. 3:23:13were folks able to get here to this uh
  4700. 3:23:16blank notebook
  4701. 3:23:18one way or the other, you clicked the
  4702. 3:23:19blue button to hit a new one or you went
  4703. 3:23:21up to file and did new notebook.
  4704. 3:23:25Yes. Okay.
  4705. 3:23:28So, there we are. Without doing all the
  4706. 3:23:30Jupiter install steps, we're in a
  4707. 3:23:32notebook. Look how easy that was, right?
  4708. 3:23:34So, why didn't we just start with this?
  4709. 3:23:37Um
  4710. 3:23:39so yeah so we're in Google's notebook
  4711. 3:23:42platform uh which is a fantastic
  4712. 3:23:44platform and uh what's great about this
  4713. 3:23:47is you can export these now these are
  4714. 3:23:50pyb which is which is uh if you're
  4715. 3:23:52curious what that extension means it's
  4716. 3:23:54short for interactive python notebook
  4717. 3:23:58okay IPIB
  4718. 3:24:00so these are the files that you can open
  4719. 3:24:02in Jupiter if you have uh if you notice
  4720. 3:24:05when you open up Jupyter notebook book
  4721. 3:24:07earlier it was a IP YMBB
  4722. 3:24:10um inside of VS code when you work with
  4723. 3:24:13notebooks they are IP YMBB so IPMBB is a
  4724. 3:24:17notebook file and it can be opened in
  4725. 3:24:20any one of these three platforms right
  4726. 3:24:22the Jupiter from Anaconda the uh VS code
  4727. 3:24:26can open IPMB and you can also upload
  4728. 3:24:30your own notebooks here if you have them
  4729. 3:24:32on your own machine you just go to file
  4730. 3:24:34upload notebook and And then it will
  4731. 3:24:37open up a box where you can choose which
  4732. 3:24:39file on your machine to upload. So you
  4733. 3:24:41can upload your own notebooks, which
  4734. 3:24:43will be uh great when we get into um
  4735. 3:24:46demos. We have demo notebooks for you
  4736. 3:24:48guys that we'll work through with our
  4737. 3:24:50code. You can upload those into Collab
  4738. 3:24:52and work with them directly inside of
  4739. 3:24:54here.
  4740. 3:24:56So let's try running something. Let's do
  4741. 3:24:59the print
  4742. 3:25:01hello world.
  4743. 3:25:05So, um, you want to type that in and I
  4744. 3:25:08can paste it in the chat for you guys
  4745. 3:25:11and then you want to you want to hit
  4746. 3:25:13either shift enter or this play button
  4747. 3:25:15right next to the cell.
  4748. 3:25:21Okay, so that might take a moment
  4749. 3:25:23because it's booting up your uh your
  4750. 3:25:25kernel.
  4751. 3:25:27Uh, but then it should run and you
  4752. 3:25:29should see the output. Now, this is
  4753. 3:25:30going to look very similar to Jupiter,
  4754. 3:25:32just slightly different.
  4755. 3:25:36We're inside of Collab
  4756. 3:25:38and we started a new notebook.
  4757. 3:25:42We're just inside of a blank notebook
  4758. 3:25:44for now. And we uh are just within this
  4759. 3:25:47first cell and I I'm doing hello world.
  4760. 3:25:50Were you guys able to run that?
  4761. 3:26:00We didn't. But we could we could open a
  4762. 3:26:03notebook in VS Code because we installed
  4763. 3:26:05the extension. We did that. Remember we
  4764. 3:26:08installed the Jupiter extension. So we
  4765. 3:26:10can open notebooks in VS Code and we can
  4766. 3:26:12run them there. I just didn't show that
  4767. 3:26:14to us. Uh we might do that later down
  4768. 3:26:17the road, but you do have that
  4769. 3:26:19flexibility to run things there if you
  4770. 3:26:22want to.
  4771. 3:26:26Okay. One thing I want to show you guys
  4772. 3:26:29that's really cool. So, um, one thing I
  4773. 3:26:33want to show you is if you go up to
  4774. 3:26:34runtime
  4775. 3:26:36and then go down to change runtime type.
  4776. 3:26:40Do you do you guys have that? Change
  4777. 3:26:41runtime type. So, if you click on
  4778. 3:26:44runtime
  4779. 3:26:46and then change runtime type.
  4780. 3:26:49Do we have that? Click on that. Click on
  4781. 3:26:53change runtime type.
  4782. 3:26:56And look at our options. We can choose a
  4783. 3:26:58GPU for free.
  4784. 3:27:03So we can swap over to a GPU kernel
  4785. 3:27:06which is fantastic for training deep
  4786. 3:27:08learning models and we can use that GPU
  4787. 3:27:11for free. This is one of the reasons
  4788. 3:27:13that uh Collab is so amazing is they
  4789. 3:27:17give free access to a GPU. So if you
  4790. 3:27:21don't have one on your own machine um
  4791. 3:27:24you can use the GPUs from Collab for
  4792. 3:27:26free.
  4793. 3:27:28Yeah, go ahead. I mean, there's no
  4794. 3:27:30nothing wrong with it. So, uh, what it's
  4795. 3:27:32going to ask you to do is, uh, terminate
  4796. 3:27:35your current kernel because you're
  4797. 3:27:37connected to a CPU basic kernel. It
  4798. 3:27:40wants you to disable that so you can
  4799. 3:27:41swap over the GPU. Click okay. That's
  4800. 3:27:44okay.
  4801. 3:27:48All right. And then we are now um, we
  4802. 3:27:50click save. And that will swap us over
  4803. 3:27:53and connect us to a GPU. So, if you how
  4804. 3:27:57you know that you're connected to a GPU
  4805. 3:27:58is if you go over to um
  4806. 3:28:03if you go over to this box on the right.
  4807. 3:28:05Do you guys see that one where it says
  4808. 3:28:07RAM and disk? If we click that,
  4809. 3:28:12it will show us our resource resource
  4810. 3:28:14usage. And you should see GPU RAM
  4811. 3:28:16available of 15 gigs.
  4812. 3:28:20So, you have 15 gigabytes of GPU RAM
  4813. 3:28:22available that you can use.
  4814. 3:28:25So remember, you just click this RAM
  4815. 3:28:28um
  4816. 3:28:31you should click this RAM uh
  4817. 3:28:35uh
  4818. 3:28:36sorry this RAM and disk.
  4819. 3:28:44Roberto, did you swap over the runtime
  4820. 3:28:46to
  4821. 3:28:47uh Yeah, it should pop up. Okay, then
  4822. 3:28:51you should be able to click on this
  4823. 3:28:57Yeah, it might take some time to connect
  4824. 3:28:59to one because what Google has to
  4825. 3:29:01allocate one to you um and then it has
  4826. 3:29:04to like connect it over the cloud. It
  4827. 3:29:06can take a minute. Yeah, it can take a
  4828. 3:29:08minute. It has to allocate one to you.
  4829. 3:29:10So the question is which one is better?
  4830. 3:29:12Um,
  4831. 3:29:14so for the vast majority of things,
  4832. 3:29:17the CPU, the standard CPU runtime, which
  4833. 3:29:20is the default, is going to be better
  4834. 3:29:22for the vast majority of things. The
  4835. 3:29:24only time the GPU is really going to be
  4836. 3:29:26beneficial is when we start doing deep
  4837. 3:29:28learning and training neural networks,
  4838. 3:29:31then the GPU will be really beneficial.
  4839. 3:29:34It will speed up the training time by a
  4840. 3:29:37significant amount.
  4841. 3:29:40I can tell you like I trained a uh
  4842. 3:29:43neural network for images for image
  4843. 3:29:46recognition. Uh it took me it took two
  4844. 3:29:50hours on the CPU and then when I swapped
  4845. 3:29:52it over to GPU it took less than a
  4846. 3:29:54minute
  4847. 3:29:57took less than a minute and it was
  4848. 3:29:58taking two hours on the CPU.
  4849. 3:30:03So yeah, training neural nets on a GPU
  4850. 3:30:07when we get to that is going to be
  4851. 3:30:09beneficial. So if you're not using
  4852. 3:30:12Collab right now, that's okay, but in
  4853. 3:30:14the future when we get into deep
  4854. 3:30:16learning, you're likely going to want to
  4855. 3:30:17use Collab to swap over to the GPU for
  4856. 3:30:20free.
  4857. 3:30:22Now, they do rate limit you,
  4858. 3:30:26so it's free, but you can max it out in
  4859. 3:30:28a day and then they cool you off for 24
  4860. 3:30:31hours, which I have I have done, uh,
  4861. 3:30:34unfortunately. So, like, if you max out
  4862. 3:30:37that RAM and you use it too much in a
  4863. 3:30:4024-hour period, they will, uh, not allow
  4864. 3:30:44you to connect to a free GPU for another
  4865. 3:30:4624 hours.
  4866. 3:30:48So, I doubt you'll run into that
  4867. 3:30:51situation, but I have before
  4868. 3:30:53if you're just if you're just using it
  4869. 3:30:56so much.
  4870. 3:31:02No. So, unfortunately,
  4871. 3:31:04uh, no.
  4872. 3:31:09So, unfortunately, no. You cannot, um,
  4873. 3:31:12Collab doesn't connect to your local
  4874. 3:31:14resources. It only it does the cloud
  4875. 3:31:16Google's cloud resources. So no, you
  4876. 3:31:18can't use your own through collab. But
  4877. 3:31:20yes, you could use your own GPU through
  4878. 3:31:23VS Code. I will show us how to do that
  4879. 3:31:25later when we get into deep learning.
  4880. 3:31:28I will show you that later.
  4881. 3:31:32We we won't need to worry about that
  4882. 3:31:34now, but later on, yes, that'll be
  4883. 3:31:37important.
  4884. 3:31:42Uh it doesn't show GPU. Make sure you
  4885. 3:31:44swap over the runtime to go to change
  4886. 3:31:47runtime type and make sure you pick GPU.
  4887. 3:31:51Make sure you go away from
  4888. 3:31:56No, you should use Collab. I wouldn't
  4889. 3:31:58You don't need to buy anything. You just
  4890. 3:32:01use Collab. Just use Collab for sure.
  4891. 3:32:05Collab's free. It does everything you're
  4892. 3:32:07going to need for the class.
  4893. 3:32:12Yeah, I I highly advocate for Collab. It
  4894. 3:32:15So, by the way, if you're curious, like
  4895. 3:32:17Collab came about because Google wanted
  4896. 3:32:20the the machine learning research
  4897. 3:32:22community to have access to GPUs for
  4898. 3:32:25free to um develop like machine learning
  4899. 3:32:28and uh deep learning models. So, uh it's
  4900. 3:32:33been around for a while. I remember
  4901. 3:32:35using Collab um probably almost 10 years
  4902. 3:32:38ago and it used to be it used to be very
  4903. 3:32:43lucky if you got a GPU. You used to like
  4904. 3:32:46you used to have to click and then hope
  4905. 3:32:48that you would get allocated one and
  4906. 3:32:50sometimes you wouldn't and I would sit
  4907. 3:32:52there and have to refresh and try to
  4908. 3:32:54hope that I would get a GPU but now it's
  4909. 3:32:57it's like readily available which is
  4910. 3:32:59fantastic.
  4911. 3:33:03No, you're But you're joining it at a
  4912. 3:33:04good time because I'm telling you, the
  4913. 3:33:06GPU was very difficult to get. I would
  4914. 3:33:10always try to switch over to that and I
  4915. 3:33:12would rarely be able to.
  4916. 3:33:18So,
  4917. 3:33:22pretty good. But like I said, like if
  4918. 3:33:25the CPU is perfectly fine for everything
  4919. 3:33:28we're going to do, except when we get
  4920. 3:33:30into deep learning, you're going to want
  4921. 3:33:31to swap that over to GPU. But that's
  4922. 3:33:34going to be for we have a while till we
  4923. 3:33:36get to deep learning.
  4924. 3:33:38We have a lot to learn between now and
  4925. 3:33:40then.
  4926. 3:33:45Okay.
  4927. 3:33:47What do we think? Do we like collab?
  4928. 3:33:50We're comfortable with it. Feel free to
  4929. 3:33:51use it. Feel free to use VS Code. Feel
  4930. 3:33:53free to use Jupiter. Whatever you want
  4931. 3:33:56to use, okay? There's options, right? I
  4932. 3:34:00hopefully you have options that work for
  4933. 3:34:02you. Um they are all used in the
  4934. 3:34:05industry. So you're not missing out on
  4935. 3:34:08if whatever you use, people use it of
  4936. 3:34:11these three people use all of them.
  4937. 3:34:15So feel free to use whatever is easiest.
  4938. 3:34:22Yeah, I can show that real quick. Yeah,
  4939. 3:34:28let me go back over to it.
  4940. 3:34:31I'm going to be real quick on it though
  4941. 3:34:32because I want to make sure we get over
  4942. 3:34:33to our other material.
  4943. 3:34:44Okay, let me show you how you can run a
  4944. 3:34:46notebook. Let me show you how to run a
  4945. 3:34:48notebook. So I'm going to go to file,
  4946. 3:34:50new file, and open a Jupyter notebook.
  4947. 3:34:56Okay. So it'll open a new. Now notice
  4948. 3:34:59notice the extension of it
  4949. 3:35:02is
  4950. 3:35:05MB. That should be no surprise. That is
  4951. 3:35:07the universal kind of interactive Python
  4952. 3:35:10notebook file.
  4953. 3:35:12Okay. So the biggest thing you have to
  4954. 3:35:15do when you open a notebook in VS Code
  4955. 3:35:18is you have to
  4956. 3:35:21tell it what kernel to connect to. So
  4957. 3:35:25have to go to select kernel
  4958. 3:35:27and then what you have to select is the
  4959. 3:35:31Python environment. And luckily if you
  4960. 3:35:34installed Anaconda
  4961. 3:35:36you have a built-in Python environment
  4962. 3:35:39which is going to be your uh which is
  4963. 3:35:42going to be um the
  4964. 3:35:47you know which is going to be the uh
  4965. 3:35:49Anaconda that you installed.
  4966. 3:35:51So I have Anaconda here. Now I have a
  4967. 3:35:55lot of other ones but the
  4968. 3:35:58uh Anaconda is here. Say it's this one.
  4969. 3:36:08Does that so when you
  4970. 3:36:11when you uh
  4971. 3:36:16Yeah, you have to install you have to
  4972. 3:36:18install a Python environment. Yes. So
  4973. 3:36:21you want to install Anaconda first and
  4974. 3:36:23then you can run your then you can run
  4975. 3:36:25your notebooks.
  4976. 3:36:33And then you just uh run your code as
  4977. 3:36:36usual
  4978. 3:36:40and then we can run that.
  4979. 3:36:44Yeah, it but like it's working as if you
  4980. 3:36:47know the same kind of notebook that we
  4981. 3:36:49have inside of Collab, the same kind of
  4982. 3:36:52notebook we have in Jupiter.
  4983. 3:36:57You you have to have a Python installed
  4984. 3:37:00for this to work. So you go to you go to
  4985. 3:37:02Python environments
  4986. 3:37:06and then choose a Python environment.
  4987. 3:37:10You could try to create one. I'm not
  4988. 3:37:12sure if that'll work for you. Create
  4989. 3:37:14Python environment. You could try that,
  4990. 3:37:16too.
  4991. 3:37:21Yeah, that's fine. Any anyone will work.
  4992. 3:37:26Any Python will work. You just need to
  4993. 3:37:27pick a Python. Anyone will work.
  4994. 3:38:02Okay. Yeah, if it defaulted to something
  4995. 3:38:04that's fine, too.
  4996. 3:38:07And then we can generate more cells
  4997. 3:38:14and run cells.
  4998. 3:38:19But yeah, that's the thing is you're
  4999. 3:38:20going to want to install um Anaconda
  5000. 3:38:22most likely because you need a Python
  5001. 3:38:25version installed on your machine in
  5002. 3:38:28order to run this.
  5003. 3:38:37Perfect.
  5004. 3:38:42Okay.
  5005. 3:38:44All right. So, what I want to do is jump
  5006. 3:38:45back over to our notes so we can
  5007. 3:38:47continue along. Um again feel free to
  5008. 3:38:50use whatever
  5009. 3:38:52platform works for you. Collab,
  5010. 3:38:55doing notebooks in VS Code, doing
  5011. 3:38:57Jupiter notebooks, whatever works for
  5012. 3:39:00you, please feel free to use that. There
  5013. 3:39:03is no wrong way of using it. Whatever is
  5014. 3:39:06best suited to you and you're most
  5015. 3:39:07comfortable with, please use that
  5016. 3:39:10option.
  5017. 3:39:14Uh, it's lowercase P. That's why
  5018. 3:39:19lowercase P. Capital P is not a function
  5019. 3:39:22in Python.
  5020. 3:39:24Lowerase.
  5021. 3:39:28Yeah, go with Collab. Yeah, if you're if
  5022. 3:39:30you're on a machine, you can't install
  5023. 3:39:32anything, go with Collab. That's totally
  5024. 3:39:34fine. That's why it's there is for the,
  5025. 3:39:38you know, convenient kind of cloud
  5026. 3:39:40aspect to it.
  5027. 3:39:42Um,
  5028. 3:39:43feel free to do collab for everything.
  5029. 3:39:45That's totally fine.
  5030. 3:39:48I will use collab from time to time as
  5031. 3:39:50well.
  5032. 3:39:55All right, let me uh go back to our
  5033. 3:39:59notes then
  5034. 3:40:05and pick up from uh syntax. So, I'm
  5035. 3:40:09going to go back to Let me share my
  5036. 3:40:10screen. Go back to
  5037. 3:40:16Can you use Collab on your phone? I've
  5038. 3:40:18never tried it. I would be surprised.
  5039. 3:40:21Maybe an iPad.
  5040. 3:40:23Maybe like a tablet. It could work
  5041. 3:40:25pretty well.
  5042. 3:40:27Phone. I'm not so sure.
  5043. 3:40:37Yeah. Go ahead. Try it and let me know
  5044. 3:40:39how it works.
  5045. 3:40:45Try it and let me know. I really don't
  5046. 3:40:47know. I'm curious now to try that.
  5047. 3:40:53All right. So, I'm going back over the
  5048. 3:40:54notes. We're going to finish up today uh
  5049. 3:40:56what the time we have left to go through
  5050. 3:40:59some syntax. So really uh
  5051. 3:41:03really getting into um into Python like
  5052. 3:41:07the actual code of it so we can get
  5053. 3:41:10started on that and start working our
  5054. 3:41:12way through it.
  5055. 3:41:14Uh the difference so the the difference
  5056. 3:41:17is um you will be executing py files
  5057. 3:41:22with the within the terminal. So you'll
  5058. 3:41:25be running Python files instead of cells
  5059. 3:41:28in a notebook. you're just you're
  5060. 3:41:30running a you're running a Python script
  5061. 3:41:34rather than individual cells.
  5062. 3:41:39Okay, so there's a difference there.
  5063. 3:41:44And the reason the reason we choose
  5064. 3:41:45notebooks is to run individual cells.
  5065. 3:41:48It's just easier.
  5066. 3:41:50Same syntax,
  5067. 3:41:52same syntax. It's just the code is not
  5068. 3:41:54isolated into cells.
  5069. 3:42:00still Python.
  5070. 3:42:05All right.
  5071. 3:42:08Um, let's go forward into the syntax,
  5072. 3:42:13start learning about it.
  5073. 3:42:15All right. So, something we need to
  5074. 3:42:18learn about is how do we properly write
  5075. 3:42:21Python code? What is the syntax to it?
  5076. 3:42:24So, some things we're going to need to
  5077. 3:42:26learn about are how to write proper
  5078. 3:42:29identifiers, which are names.
  5079. 3:42:32Identifiers are just names for
  5080. 3:42:33variables. So, we need to know what's
  5081. 3:42:36allowed, what's not allowed. We need to
  5082. 3:42:38talk about what the indentation means
  5083. 3:42:40and why do we need it in Python. I want
  5084. 3:42:43to show you guys how to write comments
  5085. 3:42:45because that's really important to
  5086. 3:42:46leaving notes to yourself or others
  5087. 3:42:49about the code and then talk about um
  5088. 3:42:52generally how we produce output and how
  5089. 3:42:54we can accept input um from a user or
  5090. 3:42:58someone interacting with our code. I
  5091. 3:43:01want to talk about all these things.
  5092. 3:43:02We'll see how how much of what we get
  5093. 3:43:04to.
  5094. 3:43:07But let's start with the identifier. So
  5095. 3:43:10what is an identifier in programming?
  5096. 3:43:12This is really for any programming
  5097. 3:43:14language. An identifier is just a name
  5098. 3:43:18we give to something inside of our code.
  5099. 3:43:21So it's a name we give to a variable, a
  5100. 3:43:23name we give to a function, a name we
  5101. 3:43:25give to an object. Um so any name we
  5102. 3:43:30give to something in our code, like when
  5103. 3:43:31we set something equal to x, like x
  5104. 3:43:35equals 3 + 3. um that thing the x is the
  5105. 3:43:41name we are giving or assigning to a
  5106. 3:43:44result or a variable or an object. So
  5107. 3:43:48anytime we write down a name in our code
  5108. 3:43:52of something there are certain rules
  5109. 3:43:54that those names have to follow and
  5110. 3:43:57these are something we will um pick up
  5111. 3:43:59as we go along but I wanted to call them
  5112. 3:44:02out here. So um when we name anything in
  5113. 3:44:06Python
  5114. 3:44:08generally they have to follow these set
  5115. 3:44:10of rules meaning they have to be a combo
  5116. 3:44:13of lowercase or uppercase letters either
  5117. 3:44:16one's allowed it can be even be a
  5118. 3:44:19mixture of lowercase and uppercase
  5119. 3:44:21numerical digits are allowed in the name
  5120. 3:44:24that's okay any digit 0 to nine it's
  5121. 3:44:27okay and underscores are okay
  5122. 3:44:32underscores are Okay. And there's no uh
  5123. 3:44:36minimum or maximum length. So names can
  5124. 3:44:40be really long, they can be really
  5125. 3:44:41short. Um of course they should be
  5126. 3:44:45meaningful. So when we name something,
  5127. 3:44:49it should not remember we want to kind
  5128. 3:44:51of get away from naming everything X or
  5129. 3:44:53Y or A or B um because those names may
  5130. 3:44:58not mean much when we look back at the
  5131. 3:45:00code. So even though those are valid
  5132. 3:45:02names from an identifier perspective, we
  5133. 3:45:05want to be really meaningful when we
  5134. 3:45:07name something. We name a variable, name
  5135. 3:45:09a function, name an object. Um,
  5136. 3:45:13here's one catch.
  5137. 3:45:16The name cannot start with a number. So
  5138. 3:45:20I can't name something uh just the the
  5139. 3:45:23number zero or the number one um because
  5140. 3:45:27I can't start with that. Now, it can
  5141. 3:45:29include that.
  5142. 3:45:31So, if I need to include a number in the
  5143. 3:45:34name, as long as it doesn't start with
  5144. 3:45:37it, that's okay. But names cannot start
  5145. 3:45:40with a digit. That's just one rule of
  5146. 3:45:42Python. Anything that you're assigning a
  5147. 3:45:45name to,
  5148. 3:45:47like a variable, function, whatever,
  5149. 3:45:50cannot start with a number or else it'll
  5150. 3:45:52be invalid.
  5151. 3:45:54Okay?
  5152. 3:45:56So I'll show us examples of that later.
  5153. 3:45:59Yes, they can start with underscores.
  5154. 3:46:00Yes, that's okay. It can start with
  5155. 3:46:03underscores. Of course, it can be lower,
  5156. 3:46:05uppercase. It can start with It cannot
  5157. 3:46:07start with a digit. It can have digits.
  5158. 3:46:10They just can't be the first character
  5159. 3:46:12of the name.
  5160. 3:46:15Yes.
  5161. 3:46:17Um, now special symbols cannot be used
  5162. 3:46:20in the name. So you cannot have a
  5163. 3:46:22percentage, dollar sign, exclamation
  5164. 3:46:24point, hyphen,
  5165. 3:46:27pound symbol, at amperand symbol, at
  5166. 3:46:30symbol. None of those can be used in the
  5167. 3:46:33name. So those symbols are not
  5168. 3:46:35recognized.
  5169. 3:46:36So if you try to include those in the
  5170. 3:46:38name of something,
  5171. 3:46:40that will produce an error. So we don't
  5172. 3:46:43want to do that. The other thing we want
  5173. 3:46:45to avoid is naming something in the same
  5174. 3:46:49name as something that already exists
  5175. 3:46:52internal to Python. So those things are
  5176. 3:46:54called keywords. So there are certain
  5177. 3:46:57keywords that have a meaning in Python.
  5178. 3:47:00They are built into the language. We
  5179. 3:47:03cannot reuse those. They're basically
  5180. 3:47:05reserved. Um so something like class is
  5181. 3:47:09reserved because it means something. It
  5182. 3:47:11means you're declaring a class. We'll
  5183. 3:47:13see. We'll talk about what that means
  5184. 3:47:14later. Or something like global cannot
  5185. 3:47:17be used because it declares something as
  5186. 3:47:19global. Um,
  5187. 3:47:22so
  5188. 3:47:24you know, we'll learn what some of those
  5189. 3:47:26keywords are. There's a list of them
  5190. 3:47:28that are in the Python documentation,
  5191. 3:47:31but we want to avoid naming things after
  5192. 3:47:34builtin
  5193. 3:47:36uh uh functions or builtin keywords. Um
  5194. 3:47:41so so in fact we've already used one
  5195. 3:47:45which is the print function. You know
  5196. 3:47:47when we printed out hello world we would
  5197. 3:47:50want to avoid naming something print
  5198. 3:47:52because print means something. It exists
  5199. 3:47:55as a function. We don't want to name
  5200. 3:47:58something print
  5201. 3:48:00right that would that would produce it
  5202. 3:48:02because it would produce confusion. The
  5203. 3:48:03interpreter would see that and say oh do
  5204. 3:48:05you mean the function print or you
  5205. 3:48:07trying to name something print? it
  5206. 3:48:09wouldn't know. So, we want to avoid
  5207. 3:48:12naming something that already exists
  5208. 3:48:14inside of Python like print or like
  5209. 3:48:17class global. Um, there's many others.
  5210. 3:48:26Okay. Lastly, and this one always throws
  5211. 3:48:29people off, is that when we name
  5212. 3:48:31something uh that is case sensitive. So,
  5213. 3:48:36if you name something lowercase A, that
  5214. 3:48:39is a completely different variable or
  5215. 3:48:41completely different function than if we
  5216. 3:48:42were to name something capital A. These
  5217. 3:48:45are different. They're treated
  5218. 3:48:47differently. So, Python will think that
  5219. 3:48:50those are two different uh objects or
  5220. 3:48:54variables or whatever the case is. So be
  5221. 3:48:56really careful with case sensitivity.
  5222. 3:48:59Python is case sensitive.
  5223. 3:49:03Lowercase A will not be treated the same
  5224. 3:49:05as capital A. And whenever you're naming
  5225. 3:49:07something, so if I have a variable and I
  5226. 3:49:10I I set lowercase A equal to five and
  5227. 3:49:14then I set um uh capital A equals to 10.
  5228. 3:49:19Then if I um print A, that would produce
  5229. 3:49:23five.
  5230. 3:49:25But if I print capital A, that will
  5231. 3:49:27produce 10. It's not the same. So it is
  5232. 3:49:31case sensitive.
  5233. 3:49:33These would be two different names.
  5234. 3:49:35Lowerase A and capital A.
  5235. 3:49:40Okay.
  5236. 3:49:42So, we're going to do examples with
  5237. 3:49:44these, but these are just some rules we
  5238. 3:49:47have to abide by in the syntax if we're
  5239. 3:49:50naming anything like naming our
  5240. 3:49:53variables, naming our functions, naming
  5241. 3:49:54our objects as we go along in in the
  5242. 3:49:57course, right? We just cannot The main
  5243. 3:50:00one that trips people up is we can't
  5244. 3:50:02start with a digit and we can't use
  5245. 3:50:05words that already exist like print.
  5246. 3:50:11Oh, is my video stuck for people?
  5247. 3:50:17Was it stuck?
  5248. 3:50:19Oh, okay. Always let me know because it
  5249. 3:50:22might it might be.
  5250. 3:50:24Always let me know because it definitely
  5251. 3:50:26could be. So, it's better to know than
  5252. 3:50:29not to know.
  5253. 3:50:36Okay.
  5254. 3:50:37Oh, no worries. Like I said, always
  5255. 3:50:39always feel free to to let us know cuz
  5256. 3:50:43um it would if it is, then it's good to
  5257. 3:50:46call it out. So, no worries about that.
  5258. 3:50:49Any questions about these names? Do does
  5259. 3:50:52it make sense about like how we name
  5260. 3:50:55things matters and there just are
  5261. 3:50:57certain rules that we have to follow. Um
  5262. 3:51:00we want to avoid these wacky symbols.
  5263. 3:51:03Um, you know, we want to avoid naming
  5264. 3:51:06things that already exist. We want to
  5265. 3:51:08avoid starting with a digit. Otherwise,
  5266. 3:51:11it's going to be a pretty standard like
  5267. 3:51:13lowercase, uppercase mixture,
  5268. 3:51:16maybe occasionally with an underscore
  5269. 3:51:18mixed in there. Um, or or digits even.
  5270. 3:51:22As long as we don't start with one,
  5271. 3:51:23that's okay.
  5272. 3:51:30Yeah, Tim, that's a good reference. the
  5273. 3:51:32PEP. So PEP
  5274. 3:51:35are the set of guidelines that um the
  5275. 3:51:38Python Foundation has kind of agreed
  5276. 3:51:41upon as um here's what you should use as
  5277. 3:51:45your style guide. Here's what here's
  5278. 3:51:48what the community believes is the best
  5279. 3:51:49style for Python. Those are good to
  5280. 3:51:52read.
  5281. 3:51:56Yeah, those are those are uh good
  5282. 3:51:57references for really like formatting
  5283. 3:52:00and styling your your Python code uh to
  5284. 3:52:03be in line with kind of what the
  5285. 3:52:05community expects.
  5286. 3:52:10Okay.
  5287. 3:52:16Okay. So, let me give you some examples.
  5288. 3:52:19Um so, the ones on the left are going to
  5289. 3:52:21be valid. So, we can name something my
  5290. 3:52:23class. We can name something var_1
  5291. 3:52:26that's okay.
  5292. 3:52:29Count
  5293. 3:52:30um that's okay. Uh
  5294. 3:52:34but if we have
  5295. 3:52:37um like on the right if we have
  5296. 3:52:39something that starts with a digit that
  5297. 3:52:41would be bad. So so this one is no good
  5298. 3:52:44because it starts with this number.
  5299. 3:52:47That's not good. Um this name has this
  5300. 3:52:50wacky at symbol in it. that's no good.
  5301. 3:52:53So, this would be a bad name for
  5302. 3:52:55something that would produce an error.
  5303. 3:52:57Remember, the interpreter is going to
  5304. 3:53:00see that and reject it essentially and
  5305. 3:53:02say, "You can't name something this.
  5306. 3:53:05It's not valid." Um, same thing with
  5307. 3:53:08trying to name something global. This is
  5308. 3:53:09a keyword that already exists in the
  5309. 3:53:11language. The interpreter is going to
  5310. 3:53:13see that and get confused. It's not
  5311. 3:53:15going to know if you're talking about
  5312. 3:53:16the keyword that's built in or you're
  5313. 3:53:19trying to come up with your own name.
  5314. 3:53:21It's not going to know. So, it's just
  5315. 3:53:22going to throw an error.
  5316. 3:53:24Um, so again, we want to we want to keep
  5317. 3:53:27things simple. We want to use
  5318. 3:53:30underscores where it makes sense. We we
  5319. 3:53:32don't want to start with numbers. Um, we
  5320. 3:53:35can use a mixture of lowercase and
  5321. 3:53:36uppercase. That's fine.
  5322. 3:53:39Um, this is a good variable name rather
  5323. 3:53:43than if I just called something X.
  5324. 3:53:46That again, we're trying to avoid that's
  5325. 3:53:49something I always see in the beginning.
  5326. 3:53:51I think is okay in the beginning, but
  5327. 3:53:52it's something we really want to be
  5328. 3:53:54conscious of is naming our variables
  5329. 3:53:57very meaningfully.
  5330. 3:53:59Like count is going to be more
  5331. 3:54:01meaningful if we're keeping track of a
  5332. 3:54:03count of something. We would rather call
  5333. 3:54:06that count than if than if I just called
  5334. 3:54:09it X. Because if you read the code,
  5335. 3:54:11which do you guys believe me? Like when
  5336. 3:54:14you see it, you kind of know exactly
  5337. 3:54:16what it means. X or count? What do we
  5338. 3:54:19think?
  5339. 3:54:22Which one like has more meaning to it
  5340. 3:54:24when you see it? You know exactly what
  5341. 3:54:26it's keeping track of. X or count?
  5342. 3:54:30Yeah, count.
  5343. 3:54:32I would agree with that. Count. Yep.
  5344. 3:54:36So, it's just an like that's just a
  5345. 3:54:37single example of trying to keep track
  5346. 3:54:41of things in a meaningful way. That's a
  5347. 3:54:44good name to give to a variable. That
  5348. 3:54:46would be uh you know keeping track of
  5349. 3:54:49something the count of something
  5350. 3:54:54rather than if we just called it x or y
  5351. 3:54:56or a or b.
  5352. 3:55:02All right, I want to talk to you guys
  5353. 3:55:05about indentation next. So now we know
  5354. 3:55:09we have to name things appropriately and
  5355. 3:55:11the interpreter will give us an error if
  5356. 3:55:12we don't name things appropriately.
  5357. 3:55:15What about indentation?
  5358. 3:55:20So indentation
  5359. 3:55:22is a way for Python to understand what
  5360. 3:55:26code gets executed together.
  5361. 3:55:30Okay.
  5362. 3:55:31So
  5363. 3:55:33and it it also indicates that I am
  5364. 3:55:38breaking the flow of the code from one
  5365. 3:55:41section to the next. So the indentation
  5366. 3:55:43is really important to signify to the
  5367. 3:55:46interpreter there is a new section of
  5368. 3:55:49code that has to be considered
  5369. 3:55:52um before I move on. So um you should
  5370. 3:55:57always use indentation
  5371. 3:56:00whenever you have a colon like we see a
  5372. 3:56:05colon here with if else statements. Now
  5373. 3:56:08we haven't learned about if else
  5374. 3:56:09statements but we will. But notice how
  5375. 3:56:12we have the colon there and the
  5376. 3:56:15interpreter would be okay with this.
  5377. 3:56:18This would work.
  5378. 3:56:20Okay,
  5379. 3:56:22which is a simple statement of saying is
  5380. 3:56:24five greater than two? Yes. So in the
  5381. 3:56:27case that it is, let's run this code.
  5382. 3:56:30But we're only able to run it because
  5383. 3:56:32the interpreter recognizes it's
  5384. 3:56:34indented.
  5385. 3:56:36So the indentation is really really
  5386. 3:56:38critical as it makes the interpreter
  5387. 3:56:41understand what should be next. The
  5388. 3:56:44interpreter understands what should be
  5389. 3:56:46next after this statement. Um like an if
  5390. 3:56:51statement or a loop statement. Um we
  5391. 3:56:56will always have indentation. So this
  5392. 3:56:59would actually uh throw an error because
  5393. 3:57:01there's this is not indented. This is at
  5394. 3:57:04the same level and if you have collab
  5395. 3:57:08open you could try this for yourself.
  5396. 3:57:11Um if you had collab open it you could
  5397. 3:57:13try it for yourself is like this would
  5398. 3:57:16this would throw an error where it says
  5399. 3:57:19I am expecting indentation but you did
  5400. 3:57:21not have have any.
  5401. 3:57:27Does it matter how many spaces?
  5402. 3:57:30Uh you yes you want to use four spaces.
  5403. 3:57:34This this should be four spaces here.
  5404. 3:57:39Spaces or tabs?
  5405. 3:57:41Uh I'm only laughing because it's a
  5406. 3:57:45it's a kind of a controversial question
  5407. 3:57:50in the community. Some people get really
  5408. 3:57:53upset over
  5409. 3:57:55one or the other. I'm not one of those
  5410. 3:57:57people. I don't really care. they so
  5411. 3:58:00most code editors
  5412. 3:58:02uh set the tab automatically as four
  5413. 3:58:08spaces. So a tab will do the same thing
  5414. 3:58:12as if you manually just did four spaces.
  5415. 3:58:14It doesn't really matter in that case.
  5416. 3:58:17So either way,
  5417. 3:58:23yeah, you're so the ide will do that for
  5418. 3:58:26you. The IDs will generally do that for
  5419. 3:58:29you because they know it should go on
  5420. 3:58:30the same line.
  5421. 3:58:35Would it work as on the same line? Yes,
  5422. 3:58:39in some cases it will, but not all. It
  5423. 3:58:42depends on how complex it is. But what
  5424. 3:58:46do you think is easier to read
  5425. 3:58:49from a readability perspective? Which is
  5426. 3:58:51easier
  5427. 3:58:55if it's all in one line or is it more
  5428. 3:58:57readable and easier to follow if it's
  5429. 3:59:00indented?
  5430. 3:59:08Yeah, the that's the purpose. So yes, I
  5431. 3:59:12I agree. indented makes it easier to
  5432. 3:59:14read. So that's another reason Python
  5433. 3:59:18really enforces indentation is to make
  5434. 3:59:20it easier to read. There's a reason they
  5435. 3:59:22do that. It's to make it easier to read.
  5436. 3:59:27Okay?
  5437. 3:59:28And that's one of the best selling
  5438. 3:59:30points of Python is how easy it is to
  5439. 3:59:32read and work with. The indentation
  5440. 3:59:34really helps. So to summarize this, we
  5441. 3:59:38are going to have to use indentation.
  5442. 3:59:40Anytime we have
  5443. 3:59:43uh a statement with a colon. Anytime we
  5444. 3:59:47have a statement with a colon, we're
  5445. 3:59:48going to have to have an indentation
  5446. 3:59:50immediately follow it. And there are
  5447. 3:59:52certain statements that have a colon
  5448. 3:59:53like if, else, else if, and any loop,
  5449. 3:59:59any looping statement. Now, all of these
  5450. 4:00:01things we're going to learn about, we'll
  5451. 4:00:02learn about it in our next lesson.
  5452. 4:00:05But anytime we have a colon, this is
  5453. 4:00:08signaling the interpreter, okay, there
  5454. 4:00:10needs to be a block of code following
  5455. 4:00:13that, which is, yeah, as you say,
  5456. 4:00:15Romero, it's like a a hierarchy. Yes,
  5457. 4:00:18that's a great way of thinking about it.
  5458. 4:00:20It's saying, okay, I should check this,
  5459. 4:00:23then do this if that's true.
  5460. 4:00:26that tells the interpreter this is only
  5461. 4:00:28going to be executed in the event that
  5462. 4:00:30this is actually true. Otherwise, I'm
  5463. 4:00:32going to keep going.
  5464. 4:00:42Okay.
  5465. 4:00:44All right. Any questions about the
  5466. 4:00:45indentation? This is something we're
  5467. 4:00:46going to learn about more as we go
  5468. 4:00:48along. When do we use indentation and
  5469. 4:00:50when do we not? We're going to learn
  5470. 4:00:51about it when we get into the if else
  5471. 4:00:53and the loops which we will study.
  5472. 4:00:57But do we do we let me ask you guys
  5473. 4:00:59this. Do we understand the idea or the
  5474. 4:01:03intent behind indentation?
  5475. 4:01:07Do we roughly get that idea? We don't we
  5476. 4:01:10don't know yet when to use it. I get
  5477. 4:01:12that. But more the intent or the purpose
  5478. 4:01:15of using it is to really like section
  5479. 4:01:18things off.
  5480. 4:01:21Yeah.
  5481. 4:01:23Okay. Good. Good. Good. Good. Glad to
  5482. 4:01:26hear that. Okay.
  5483. 4:01:36Okay.
  5484. 4:01:38Let's wrap up today by talking about
  5485. 4:01:40comments. So, uh what are comments?
  5486. 4:01:43These are like annotations or notes to
  5487. 4:01:47yourself that are completely ignored by
  5488. 4:01:52the interpreter.
  5489. 4:01:53So when your code gets executed, the
  5490. 4:01:56comment does not play any role in what
  5491. 4:01:59gets executed. The interpreter will
  5492. 4:02:00actually just completely ignore it. The
  5493. 4:02:02moment it sees the comment, it will just
  5494. 4:02:04ignore it and go to the next line.
  5495. 4:02:07So the purpose of it is for humans to
  5496. 4:02:10leave a note to other humans reading the
  5497. 4:02:12code and that is very powerful is to be
  5498. 4:02:16able to read those to to leave those
  5499. 4:02:18notes and not have it affect the actual
  5500. 4:02:22code that's uh that's actually executed.
  5501. 4:02:25So there's multiple ways to make
  5502. 4:02:28comments inside of Python. The most
  5503. 4:02:30basic is to use the pound symbol. So the
  5504. 4:02:34remember we cannot use pound symbols to
  5505. 4:02:36name anything.
  5506. 4:02:38This is why because the pound symbol is
  5507. 4:02:41res reserved for making comments. So you
  5508. 4:02:44you when you have a pound symbol like
  5509. 4:02:47this uh that immediately signals to the
  5510. 4:02:50interpreter everything else that follows
  5511. 4:02:53that on this line is a comment. Any
  5512. 4:02:56other text that follows that on that
  5513. 4:02:58line is a comment. And usually your
  5514. 4:03:00editor like VS Code, Jupiter, Collab
  5515. 4:03:05will color that differently, maybe like
  5516. 4:03:08a a like you can see in here, this is a
  5517. 4:03:10Jupiter example. You can see it's kind
  5518. 4:03:12of a light gray,
  5519. 4:03:14greenish gray
  5520. 4:03:16um to signal that this is a comment. Um
  5521. 4:03:21now you may be asking why would we have
  5522. 4:03:23comments? Again, you are going to look
  5523. 4:03:25back at code weeks later,
  5524. 4:03:29especially in this program. You're going
  5525. 4:03:30to look at code in review and be like,
  5526. 4:03:32what what was I doing there? If you
  5527. 4:03:36leave a comment, you'll remember what
  5528. 4:03:38you were doing there, why you had that
  5529. 4:03:40line. Um, and not only that, like when
  5530. 4:03:43you share your code with others, which
  5531. 4:03:45in the real world you would be doing,
  5532. 4:03:47collaborating with others, right? Adding
  5533. 4:03:50in those comments can be really
  5534. 4:03:52beneficial to do. So
  5535. 4:03:55you will see me
  5536. 4:03:58throughout the program. I'm going to
  5537. 4:04:00leave a lot of comments on our demos and
  5538. 4:04:02our notebooks that we work on together
  5539. 4:04:04in the live sessions. I will leave
  5540. 4:04:06comments mainly to call out certain
  5541. 4:04:09things like I will say this is a really
  5542. 4:04:11important step or we are doing this
  5543. 4:04:14because I will leave a lot of comments
  5544. 4:04:17and I encourage you guys to do the same
  5545. 4:04:18in your own code. Um, remember they're
  5546. 4:04:22free. They're they get ignored by the
  5547. 4:04:24interpreter. They don't affect anything.
  5548. 4:04:26They are notes to yourself. So, use them
  5549. 4:04:29accordingly. Um, and you know, there's
  5550. 4:04:32actually multiple ways to leave
  5551. 4:04:34comments, but this is I'll show us those
  5552. 4:04:36as we go along. But this is the uh most
  5553. 4:04:39basic is you just you you type in a
  5554. 4:04:42pound symbol and then everything else
  5555. 4:04:44that follows that is uh is your comment.
  5556. 4:04:50Okay.
  5557. 4:04:52All right, guys. That's it for today.
  5558. 4:04:55Um, what a great first session. Thank
  5559. 4:04:57you guys. Thank you guys for being
  5560. 4:04:58patient. um going through the setup of
  5561. 4:05:01some of those tools. I hope you landed
  5562. 4:05:02on one that worked for you. Um you know,
  5563. 4:05:06use that one going forward, please. If
  5564. 4:05:08it's collab, use that. Jupiter, use
  5565. 4:05:10that. VS Code, use that. Whatever you
  5566. 4:05:13you uh feel most comfortable with,
  5567. 4:05:15please use that. Um we have a lot to
  5568. 4:05:17cover still. You know, we're going to um
  5569. 4:05:20continue on Wednesday. Uh we were we
  5570. 4:05:24will uh continue talking about Python.
  5571. 4:05:26We're just getting our feet wet a little
  5572. 4:05:28bit on on Python. A lot more to cover.
  5573. 4:05:31We're going to get into the the more
  5574. 4:05:33nitty-gritty of the code. So, it'll be
  5575. 4:05:35really fun. We'll cover if else loops,
  5576. 4:05:39um how to control the flow of our
  5577. 4:05:40program. Um we will do that. This is
  5578. 4:05:44where we left off was writing comments
  5579. 4:05:47uh in Python code, which I I did want to
  5580. 4:05:49remind us of how to do that. it is going
  5581. 4:05:52to be using the uh pound symbol to
  5582. 4:05:56initiate a comment. And basically the
  5583. 4:05:58Python interpreter will ignore
  5584. 4:05:59everything else on that line. Uh it it
  5585. 4:06:03treats all of that text as a comment.
  5586. 4:06:05And again like comments are free. You
  5587. 4:06:08might as well use them to your advantage
  5588. 4:06:10to kind of uh leave a note to yourself
  5589. 4:06:12of hey this is what this code is doing.
  5590. 4:06:15Um so that when you come back and read
  5591. 4:06:17it uh you can understand it better. So I
  5592. 4:06:20encourage you guys like when we do demos
  5593. 4:06:24uh and we will do a lot of demos um
  5594. 4:06:27especially today leave comments you know
  5595. 4:06:30put comments in there so so you can make
  5596. 4:06:34a note to yourself what this code is
  5597. 4:06:36doing um so you will get I think it'll
  5598. 4:06:40be good to get in the habit of leaving
  5599. 4:06:41comments uh to kind of mark up the code
  5600. 4:06:44to to kind of remind yourself oh this is
  5601. 4:06:47what it was doing when you look back at
  5602. 4:06:49uh in the future.
  5603. 4:06:52Okay. So we had ended with that.
  5604. 4:06:56What I wanted to do was move on into the
  5605. 4:06:59next slide. So talk about um basically
  5606. 4:07:02how we display output to the screen
  5607. 4:07:05which we've already seen an example of
  5608. 4:07:06when we did the hello world which is the
  5609. 4:07:08print function on the right. So this is
  5610. 4:07:11a by the way this is a Python function
  5611. 4:07:14and you know it's a function because
  5612. 4:07:18of these parentheses. So these
  5613. 4:07:21parenthesis signal that this is a
  5614. 4:07:23function because it expects some sort of
  5615. 4:07:25input to go inside of those parentheses.
  5616. 4:07:28And and the input that would go inside
  5617. 4:07:30of there is going to be text like some
  5618. 4:07:34sort of uh some sort of text that
  5619. 4:07:36belongs inside of quotes. and whatever
  5620. 4:07:39we put there um will display on the
  5621. 4:07:43screen. So that's useful for us to like
  5622. 4:07:45display information
  5623. 4:07:47um print we would say we are printing
  5624. 4:07:49out information to the screen. Um so if
  5625. 4:07:53we want to know the value of a variable
  5626. 4:07:55or the value of something that we are
  5627. 4:07:57doing a calculation with or uh you know
  5628. 4:08:00sanity check something in our code we we
  5629. 4:08:02can print it out which would be using
  5630. 4:08:05the print function and it will display
  5631. 4:08:07that value onto the screen. So we will
  5632. 4:08:10use the print function quite a bit. Um
  5633. 4:08:14you know we haven't learned what
  5634. 4:08:15functions are but uh functions in Python
  5635. 4:08:18are you know um designed to be uh chunks
  5636. 4:08:22of code that execute and do something
  5637. 4:08:25and they take arguments and you know it
  5638. 4:08:28takes an argument because of the
  5639. 4:08:29parenthesis that is um signaling that
  5640. 4:08:31there should be some something inside of
  5641. 4:08:33this parenthesis here which is going to
  5642. 4:08:36be uh text. So whatever text you want to
  5643. 4:08:38display or maybe some variable you want
  5644. 4:08:40to display um that would go inside of
  5645. 4:08:43there. So we'll get the hang of using
  5646. 4:08:45the print function as we go along but
  5647. 4:08:47just wanted to call out that's the
  5648. 4:08:49primary methodology of kind of um
  5649. 4:08:52displaying something on the screen if we
  5650. 4:08:54want to print function. Um now the
  5651. 4:08:58reverse of that is uh asking a user to
  5652. 4:09:03uh input some data. So that would be
  5653. 4:09:05this input function and um this is
  5654. 4:09:09something that uh as you can see an
  5655. 4:09:12example below is we can put some text
  5656. 4:09:15inside of this parenthesis. So again a
  5657. 4:09:17function it has those parenthesis that
  5658. 4:09:19signals it's a it's a function. Um,
  5659. 4:09:23and we can put some text in there which
  5660. 4:09:25would be kind of what displays in to the
  5661. 4:09:29user as kind of a prompt like here. Uh,
  5662. 4:09:33enter your name and that would display
  5663. 4:09:35on the screen and then there would be a
  5664. 4:09:37box next to it. I'm going to show us
  5665. 4:09:39this. I'm going to actually run this
  5666. 4:09:41inside of a notebook in a minute. But
  5667. 4:09:44then there would be a box that displays
  5668. 4:09:46that that would say um hey you know
  5669. 4:09:50enter your name and then you can type
  5670. 4:09:52input in uh and and then when you hit
  5671. 4:09:55enter it will save that input into into
  5672. 4:09:58this variable called name. So um and
  5673. 4:10:01remember name this is a valid identifier
  5674. 4:10:04because it starts with a lowercase n um
  5675. 4:10:08which is fine and it it has all valid
  5676. 4:10:11characters. It doesn't have any wacky,
  5677. 4:10:13you know, uh, pound symbol or at or
  5678. 4:10:17anything crazy. So, it's it's a decent
  5679. 4:10:19identifier. Um,
  5680. 4:10:22so name name would be okay. And so input
  5681. 4:10:25is whenever you want to get whenever you
  5682. 4:10:28want to allow the user to input
  5683. 4:10:30something like it'll bring up a text box
  5684. 4:10:32and they can enter some data. Um, and
  5685. 4:10:35that will be saved in this whatever
  5686. 4:10:38variable you name this you set equal to
  5687. 4:10:40input. Um, and then you can see like as
  5688. 4:10:44soon as we put that in, we immediately
  5689. 4:10:45dis we can display it. So we we print
  5690. 4:10:48hello and then comma name which
  5691. 4:10:50references whatever we stored whatever
  5692. 4:10:53the user input there. So I'm I'm going
  5693. 4:10:54to show us an example of that. Um, but
  5694. 4:10:57input is what get is our primary way of
  5695. 4:11:01getting input from the from the the user
  5696. 4:11:04in a text box so we can use that data in
  5697. 4:11:07our program.
  5698. 4:11:08Print is our primary way of displaying
  5699. 4:11:11data that we already have in our code.
  5700. 4:11:13We can print it which will display it.
  5701. 4:11:16Um
  5702. 4:11:17so we're going to see many examples of
  5703. 4:11:19these as along but just wanted to call
  5704. 4:11:20out those two. These
  5705. 4:11:24functions by the way are built into
  5706. 4:11:26Python. So we don't need to create them
  5707. 4:11:29ourselves. They already exist. They're
  5708. 4:11:31already built into Python. Um nothing
  5709. 4:11:34special we need to do to use them. We
  5710. 4:11:36can just use them right out of the box.
  5711. 4:11:38So again, we'll see this in our in our
  5712. 4:11:39code examples that we're going to do in
  5713. 4:11:41a minute.
  5714. 4:11:45Um, where exactly would an end user be?
  5715. 4:11:47So maybe we ask them for some input. Um,
  5716. 4:11:51and then we do like so we ask them for
  5717. 4:11:53some like their name, their email, their
  5718. 4:11:57uh date of birth, those kind of things.
  5719. 4:12:00We can ask in the input and then we
  5720. 4:12:01maybe we store them in a database or we
  5721. 4:12:03do something with it in the Python code.
  5722. 4:12:05So um whenever you want to accept input
  5723. 4:12:09from a end user that's when you would
  5724. 4:12:11use this input.
  5725. 4:12:14It just depends on the application
  5726. 4:12:17right on the application like what kind
  5727. 4:12:18of input data you you want to accept
  5728. 4:12:20from the from the user.
  5729. 4:12:27Okay. Okay. So, I'm going to show us
  5730. 4:12:28this.
  5731. 4:12:33Um, before we do that demo,
  5732. 4:12:36um, let me ask you guys, which of the
  5733. 4:12:38following do do we remember from Monday?
  5734. 4:12:40Which of the following identifier names
  5735. 4:12:43follows Python's rules
  5736. 4:12:46and best practices for readability?
  5737. 4:12:50So not only so you should be looking for
  5738. 4:12:52the answer choice here that follows the
  5739. 4:12:54rules but also is a meaningful name.
  5740. 4:13:00A lot of different choices. Okay.
  5741. 4:13:04By the way,
  5742. 4:13:06which let me ask let me I'll come back
  5743. 4:13:08to the answer to the original question,
  5744. 4:13:10but let me ask this alternative
  5745. 4:13:11question. Which one of these is not
  5746. 4:13:14valid? Meaning it would Python would
  5747. 4:13:17throw an error how you use it.
  5748. 4:13:20Which one of these is not valid in
  5749. 4:13:22general?
  5750. 4:13:25Cool. Great. It is C. You guys were
  5751. 4:13:27right on top of that. Very good. So C is
  5752. 4:13:28not valid. It names of things cannot
  5753. 4:13:31start with a number. So So that's um C
  5754. 4:13:35is completely invalid in general and
  5755. 4:13:37that would produce uh an error.
  5756. 4:13:41Right. Starts with a number. Exactly.
  5757. 4:13:43Which we cannot do. Now starting with an
  5758. 4:13:46underscore is okay. that's allowed. So
  5759. 4:13:49that's not an issue. And having a number
  5760. 4:13:51be second after the underscore is okay
  5761. 4:13:55as well. So technically A and B would
  5762. 4:13:59follow the rules. So those definitely
  5763. 4:14:01follow the rules. Um now are they
  5764. 4:14:06readable and meaningful is the question.
  5765. 4:14:10I would argue that possibly not. Var 123
  5766. 4:14:15is pretty generic. I would argue that
  5767. 4:14:18even though it's valid, like Python
  5768. 4:14:20would not have any errors with that uh
  5769. 4:14:22variable name, it's not very meaningful.
  5770. 4:14:24It's too generic. It's almost as if we
  5771. 4:14:26just called something X. We just called
  5772. 4:14:28something var 23, that's probably not
  5773. 4:14:32going to be meaningful to us and we're
  5774. 4:14:34not going to understand what that really
  5775. 4:14:36represents. If somebody were to come
  5776. 4:14:37along and read it and see VAR 123,
  5777. 4:14:41that's probably not that great of a
  5778. 4:14:43name. It's not telling us exactly what
  5779. 4:14:45that represents. So, I would say A is
  5780. 4:14:49likely not um
  5781. 4:14:52A is likely not uh a good choice and C
  5782. 4:14:56we know is invalid. So, really I think
  5783. 4:14:57the only two options you could argue are
  5784. 4:15:00B and D. I think D is a really good
  5785. 4:15:02answer. It um it it follows the rules.
  5786. 4:15:07Uh underscores are fine. Everything is
  5787. 4:15:09lowercase. That's fine. Um so it's valid
  5788. 4:15:13but it also is meaningful as a name
  5789. 4:15:15right so final result value um we we
  5790. 4:15:19should probably like in our code we
  5791. 4:15:21would have context we would know what
  5792. 4:15:22that means okay this is our final result
  5793. 4:15:25um so that's a that's a pretty good name
  5794. 4:15:27for something um you know this one is
  5795. 4:15:32okay it's just not that readable 321
  5796. 4:15:35customer details DB table it's okay I
  5797. 4:15:40don't think It's um it's not the worst.
  5798. 4:15:42It it definitely would work, but it's um
  5799. 4:15:46kind of a clunky name. I'm sure we could
  5800. 4:15:48come up with something better, but it
  5801. 4:15:50would work technically. There'd be no
  5802. 4:15:52issues with it.
  5803. 4:15:55Okay. So, I think D is probably the best
  5804. 4:15:58choice, but B is valid, too. I think B
  5805. 4:16:00could work for this. D and B, I think,
  5806. 4:16:01are okay.
  5807. 4:16:05Okay. Good. Good. You guys are right on
  5808. 4:16:07top of that. you have a good I think you
  5809. 4:16:08have a good feel for what the allowed
  5810. 4:16:11names for things are which is good.
  5811. 4:16:15Okay.
  5812. 4:16:17Um
  5813. 4:16:19I'm going to then swap over to this
  5814. 4:16:22demo. So you guys should have the uh
  5815. 4:16:26demos um and we kind of went through
  5816. 4:16:29some of those first few last time to get
  5817. 4:16:31you set up on Cola and Jupiter and VS
  5818. 4:16:34Code. Um, I am going to be using Collab
  5819. 4:16:39for most of these, but feel free to use
  5820. 4:16:42whatever you want to use. If you want to
  5821. 4:16:43use Jupiter, if you want to use VS Code
  5822. 4:16:45and run your notebooks on your own
  5823. 4:16:47machine, feel free to use whatever you
  5824. 4:16:49want to use. I'm going to be using
  5825. 4:16:50Collab just for the simplicity of it.
  5826. 4:16:54Um, so, so this demo will walk us
  5827. 4:16:57through, um, opening up a new Collab
  5828. 4:17:00notebook and then running those input
  5829. 4:17:02and print. So, some examples with input
  5830. 4:17:04and print. Um, so we'll do that
  5831. 4:17:06together. Let me go over to that demo.
  5832. 4:17:13So, if you're following along, we are
  5833. 4:17:15going to be doing um demo 4. So, it
  5834. 4:17:20should be lesson one, demo 4.
  5835. 4:17:27Um, do you guys have this? Give you a
  5836. 4:17:30moment to to pull that up. Lesson one,
  5837. 4:17:32demo 4.
  5838. 4:17:35You guys have access to this one. So, we
  5839. 4:17:37did we did one, two, and three on
  5840. 4:17:39Monday, which were just getting those
  5841. 4:17:41environments set up. So, this is demo 4.
  5842. 4:17:45Um, which again, I know step one says
  5843. 4:17:48open collab. Feel free to open your own
  5844. 4:17:50notebook in VS Code or open your own
  5845. 4:17:52notebook in Jupiter as well, whatever
  5846. 4:17:53you're most comfortable with. Um,
  5847. 4:17:57I'm going to be using the collab to to
  5848. 4:17:58do this, but feel free to use whatever
  5849. 4:18:01works. You're we really just need a
  5850. 4:18:03notebook to be able to run this code.
  5851. 4:18:05So, however you're running notebooks,
  5852. 4:18:07whether that's in VS Code or Jupiter or
  5853. 4:18:09Collab, I any of those, either one is uh
  5854. 4:18:13perfectly fine.
  5855. 4:18:15So, so step one is to open up a
  5856. 4:18:18notebook. I'm going to do it in Collab,
  5857. 4:18:19which is what this says. Um, and then
  5858. 4:18:22you can make a new notebook and then
  5859. 4:18:24rename it to my first program. I'm going
  5860. 4:18:26to do that in a second. And then, um, so
  5861. 4:18:30I'm going to walk through this live with
  5862. 4:18:32you, but just showing you some of the
  5863. 4:18:33steps we're going to do. Um, the first
  5864. 4:18:36thing we're going to do is is just
  5865. 4:18:38practice doing the print hello world
  5866. 4:18:40again so that we can um, execute a print
  5867. 4:18:44statement. So, we'll practice that.
  5868. 4:18:47We're going to make a second cell. Um,
  5869. 4:18:51which we can do in Collab or VS Code or
  5870. 4:18:55Jupiter by hitting the plus button.
  5871. 4:18:56There's usually a plus. Uh, you can see
  5872. 4:18:59it here. Uh, in multiple places in
  5873. 4:19:02Collab, you can do it right below an
  5874. 4:19:03existing cell or there's always a plus
  5875. 4:19:06code here, which is kind of what you
  5876. 4:19:08have in Jupiter. Usually in Jupiter, you
  5877. 4:19:09have a plus button. So, you can just hit
  5878. 4:19:12hit that plus button, it'll make a new
  5879. 4:19:13cell. Um, and so we'll make a new cell
  5880. 4:19:17so we can write some more code.
  5881. 4:19:21Um,
  5882. 4:19:22and then in this new one, we are going
  5883. 4:19:24to practice doing some comments.
  5884. 4:19:27We're going to practice doing some
  5885. 4:19:28comments and then um see how we can do
  5886. 4:19:32uh some more print statements. Okay, so
  5887. 4:19:36let's do that. Let me jump over to
  5888. 4:19:37Collab. Let's walk through these first
  5889. 4:19:39few steps together. Um, and then uh
  5890. 4:19:43we'll come back to this and finish out
  5891. 4:19:45the rest of the steps because we're also
  5892. 4:19:46going to do input. So I'm going to show
  5893. 4:19:48you how to do these uh input which will
  5894. 4:19:51you can see here like the input is going
  5895. 4:19:53to create a text box where you can put
  5896. 4:19:56input and it will you hit enter it will
  5897. 4:19:58save it for you. So input allows you to
  5898. 4:20:01get input from the keyboard
  5899. 4:20:04and save that into a variable to use for
  5900. 4:20:08later.
  5901. 4:20:10Okay. So, let's jump over to
  5902. 4:20:15um let's jump over to
  5903. 4:20:18I'll show us I'll show us in a second.
  5904. 4:20:20How do you rename it?
  5905. 4:20:22I'll show you. Let me jump over to
  5906. 4:20:24collab.
  5907. 4:20:28Um
  5908. 4:20:36okay.
  5909. 4:20:37So, I am over lesson one, demo 4. Yep,
  5910. 4:20:42that's the one we're doing.
  5911. 4:20:45Okay. So, I am in Collab. I'm going to
  5912. 4:20:46start a new notebook.
  5913. 4:20:51Start a new notebook in Collab. Uh, so
  5914. 4:20:53now I'm here. I'm just on a fresh
  5915. 4:20:54notebook. Um, nothing that interesting
  5916. 4:20:57going on. Here's how you rename it. just
  5917. 4:20:59go up to this box on the left
  5918. 4:21:03and almost like a Google doc just just
  5919. 4:21:06uh click into that name and then start
  5920. 4:21:10typing to erase it. So see how I'm like
  5921. 4:21:12hovering over that name and then I'm
  5922. 4:21:14clicking on it and then I can start
  5923. 4:21:16typing to erase it. So I can we can name
  5924. 4:21:18this my
  5925. 4:21:20first program
  5926. 4:21:23and then hit enter and it will save
  5927. 4:21:25that.
  5928. 4:21:33Oh yeah. So if you're in VS Code um do
  5929. 4:21:37to to rename it do file and then save as
  5930. 4:21:42and then you can give a new name to it.
  5931. 4:21:45file, save as.
  5932. 4:21:49Okay, that's how you can rename it in VS
  5933. 4:21:51Code.
  5934. 4:21:58All right, let's do let's do the first
  5935. 4:22:01step. Um, let's do print. So, we're
  5936. 4:22:04going to do print. So, type in print and
  5937. 4:22:07we can uh we can do parenthesis.
  5938. 4:22:11Um, remember this is a function. So, we
  5939. 4:22:14need the we need the parenthesis to
  5940. 4:22:16signal that we want to put some text
  5941. 4:22:18inside of this print function. And then
  5942. 4:22:20you want to do uh you want to do quotes.
  5943. 4:22:24You want to do quotes, the double quotes
  5944. 4:22:26there, in order to allow us to put in
  5945. 4:22:29some text. So, Python will interpret
  5946. 4:22:32what's inside of the quotes as text and
  5947. 4:22:34it will display that text. So we can do
  5948. 4:22:37hello world my first
  5949. 4:22:43Python program.
  5950. 4:22:46Okay. And then we can run it. So feel
  5951. 4:22:49free to put whatever text in here. It
  5952. 4:22:51doesn't really matter exactly what it
  5953. 4:22:52is, but you put some text in there
  5954. 4:22:54between the parenthesis and then hit
  5955. 4:22:56run.
  5956. 4:23:01And the notebook will take a second to
  5957. 4:23:05connect. And then there it is. Right? So
  5958. 4:23:07then you see the the text displayed on
  5959. 4:23:09the screen.
  5960. 4:23:12Try that out. Are you guys able to run
  5961. 4:23:14the print
  5962. 4:23:16in your Jupyter what whether it's
  5963. 4:23:17Collab, whether it's uh Jupyter
  5964. 4:23:19notebook, whether it's VS Code. Can you
  5965. 4:23:21run the print
  5966. 4:23:32install?
  5967. 4:23:35Yeah, I installed that um in VS Code.
  5968. 4:23:38Yep. Try installing that.
  5969. 4:23:48Okay, great. You guys were able to run
  5970. 4:23:49that. Very good. Very good. Okay.
  5971. 4:23:57No, you don't want to save it as a JSON
  5972. 4:23:59file. You want to save it as a pyb just
  5973. 4:24:02like this. See how this one is uh IP
  5974. 4:24:05YMBB?
  5975. 4:24:07That's the format you want. Remember
  5976. 4:24:09that is interactive Python notebook.
  5977. 4:24:13You want that file IPY MB.
  5978. 4:24:25Uh perfect. Yeah, you get you got it to
  5979. 4:24:27run.
  5980. 4:24:32You don't have any extension? No. If
  5981. 4:24:34you're in VS Code, remember from Monday,
  5982. 4:24:36you need to install the the Jupiter
  5983. 4:24:39extension.
  5984. 4:24:45If you're in VS Code, you got to install
  5985. 4:24:47the Jupiter extension.
  5986. 4:24:57You have to manually so manually save
  5987. 4:24:59it.
  5988. 4:25:03You can save it as a py. I would do ipy
  5989. 4:25:06so you can open it in collab.
  5990. 4:25:09Type it yourself.
  5991. 4:25:11Type overwrite what's there and type it.
  5992. 4:25:14Type in um my notebook whatever the name
  5993. 4:25:18is.
  5994. 4:25:20Type it out yourself if you can. like
  5995. 4:25:23save as and then
  5996. 4:25:26type out the full file name yourself.
  5997. 4:25:32Now let's practice a comment. Let's
  5998. 4:25:35practice a comment. So let's build let's
  5999. 4:25:36do a new code cell. So we made a you can
  6000. 4:25:40either do it here. If you hover over
  6001. 4:25:42your cell, you can hit plus to build a
  6002. 4:25:44new code cell or you can hit plus here
  6003. 4:25:46to make a new code cell. So let's do
  6004. 4:25:48that.
  6005. 4:25:56You should be saving.
  6006. 4:25:59Don't worry about the type. Just type in
  6007. 4:26:01the namey imm. I don't think you need
  6008. 4:26:04to.
  6009. 4:26:08Or you could just hit what you could do
  6010. 4:26:09is you could just hit save and then in
  6011. 4:26:12your file explorer you could just rename
  6012. 4:26:15it.
  6013. 4:26:17So if you just save it will save it to
  6014. 4:26:18the default location and then just and
  6015. 4:26:20then just rename it.
  6016. 4:26:23So maybe try that route. Just just do
  6017. 4:26:24save. Just save it. And then it should
  6018. 4:26:30it should try to save it as IPymbb.
  6019. 4:26:38Okay. Okay. Let's practice. Um, so the
  6020. 4:26:42next step in the demo, if you're
  6021. 4:26:43following along in the demo document, it
  6022. 4:26:46wants us to do, so we did the print. We
  6023. 4:26:49want to do um a practice some comments.
  6024. 4:26:57>> Okay, perfect. Uh, let's practice some
  6025. 4:27:00comments. So, um, remember I told you
  6026. 4:27:02that we can do, uh, comments with the
  6027. 4:27:04pound sum. So, this is a
  6028. 4:27:09comment. So practice writing a comment.
  6029. 4:27:11Remember you start a comment with a
  6030. 4:27:13pound symbol.
  6031. 4:27:16Um it will get ignored
  6032. 4:27:19by the interpreter
  6033. 4:27:23interpreter. So feel free to type in
  6034. 4:27:25whatever text you want. I'm just
  6035. 4:27:27reminding us that whatever the comment
  6036. 4:27:29is is going to be ignored and we can
  6037. 4:27:32have whatever code below that that we
  6038. 4:27:34want to have and that comment will get
  6039. 4:27:36completely ignored. So let's do another
  6040. 4:27:38print. So write a comment,
  6041. 4:27:41hit enter. Immediately below that in a
  6042. 4:27:44new line, let's do another print.
  6043. 4:27:48This code
  6044. 4:27:50gets executed.
  6045. 4:27:54So we know this print statement is going
  6046. 4:27:56to get executed, but this comment is
  6047. 4:27:59going to be ignored by the interpreter.
  6048. 4:28:02So let's run that.
  6049. 4:28:04So this code gets executed. This comment
  6050. 4:28:07gets completely ignored,
  6051. 4:28:10right? That comment gets completely
  6052. 4:28:11ignored, which is great. Try writing a
  6053. 4:28:14comment. Are you guys able to write
  6054. 4:28:15comments?
  6055. 4:28:17So write a comment and then try writing
  6056. 4:28:19a print statement right after it.
  6057. 4:28:22And and feel free to put whatever text
  6058. 4:28:24you want inside the comment. And feel
  6059. 4:28:25free to
  6060. 4:28:30uh for the comment, is the space after
  6061. 4:28:32the pound symbol required? No, it's not.
  6062. 4:28:34So, we could test it out. So, I removed
  6063. 4:28:36the space. Doesn't matter. It It's just
  6064. 4:28:39for readability. I usually like doing
  6065. 4:28:42that so that I have some space after it.
  6066. 4:28:44And this is a little It's just a little
  6067. 4:28:46bit more readable, right? It's not like
  6068. 4:28:47mixed together.
  6069. 4:28:52It's just for readability.
  6070. 4:28:55Great. You guys wrote a comment. Okay.
  6071. 4:28:57Perfect. Perfect. We're able to write
  6072. 4:28:59comments. Really great. Okay.
  6073. 4:29:02Okay.
  6074. 4:29:06If you put multiple code lines, do we
  6075. 4:29:09need any separator? Like, no, they just
  6076. 4:29:12go on new lines. So, do you mean like a
  6077. 4:29:14second print statement? Let's We could
  6078. 4:29:16try that. Let's do a secondary print
  6079. 4:29:18statement. So, we can do print.
  6080. 4:29:21Um, this one is on the next line. No
  6081. 4:29:28separator
  6082. 4:29:30needed.
  6083. 4:29:34Do you see that? See how it's on its
  6084. 4:29:35own? I did a print right below this
  6085. 4:29:38other print. And as long as they're on
  6086. 4:29:39their own line, that's okay. They just
  6087. 4:29:42need to be on their own lines. They
  6088. 4:29:43don't need any separator.
  6089. 4:29:46If we run this, then this one gets exe.
  6090. 4:29:50Then see how this is now printed out
  6091. 4:29:51below it. Right there.
  6092. 4:29:58Is there any character limit on the
  6093. 4:29:59comments? Uh, no. There's no character
  6094. 4:30:02limit. Um,
  6095. 4:30:04but
  6096. 4:30:06there's no character limit, but a good
  6097. 4:30:08practice is to not like you don't want
  6098. 4:30:10this to be super long and to to take up
  6099. 4:30:13the whole screen, right? Because then
  6100. 4:30:15it's not really readable.
  6101. 4:30:17So, there's no limit, but you don't want
  6102. 4:30:19to have overly
  6103. 4:30:22long comments. You want to keep them
  6104. 4:30:24kind of concise and short.
  6105. 4:30:26So just so you can read them and they're
  6106. 4:30:28they don't take up a lot of space.
  6107. 4:30:34Not able to add print statement below.
  6108. 4:30:36Why? Why is that?
  6109. 4:30:40You should be able to should be able to
  6110. 4:30:42have a print right below this print.
  6111. 4:30:44Shouldn't be anything that make sure you
  6112. 4:30:46close this parenthesis. Make sure every
  6113. 4:30:48print needs to close the parenthesis
  6114. 4:30:53and they all you also need to close the
  6115. 4:30:55quotes. So close this quote, close this
  6116. 4:30:58quote within within the print
  6117. 4:31:02that needs to be done. So you should be
  6118. 4:31:06able to run I'll paste this for you guys
  6119. 4:31:08in the chat. Should be able to run this
  6120. 4:31:19All right. One thing I want to show you
  6121. 4:31:20guys is just like the demo says in the
  6122. 4:31:22word document, um you can do multi-line
  6123. 4:31:25comments. So if you need to do a lot of
  6124. 4:31:28comments, all you need to do is triple
  6125. 4:31:31quotes. So triple quote,
  6126. 4:31:34then um triple quote, and then
  6127. 4:31:38everything in between.
  6128. 4:31:44That's interesting that it did that.
  6129. 4:31:47Yeah. So, we can do a pound symbol,
  6130. 4:31:50pound symbol, pound symbol,
  6131. 4:31:55pound symbol, and that that should all
  6132. 4:31:57work. So, we can do that.
  6133. 4:32:01Yeah, I think it's a collab thing, but
  6134. 4:32:03normally in in like Jupiter or in
  6135. 4:32:05Python, it it will work just fine. But
  6136. 4:32:08like in collab, I think they don't like
  6137. 4:32:10the triple quotes.
  6138. 4:32:12But yeah, do you guys see do you guys
  6139. 4:32:14see how I just did it like this with the
  6140. 4:32:15pound symbols? That's okay, too.
  6141. 4:32:19Everything between these
  6142. 4:32:23uh pound symbols is a comment and is
  6143. 4:32:26ignored. So now we should be able to run
  6144. 4:32:28that. So there we go. Everything gets
  6145. 4:32:30ignored there. Does that make sense to
  6146. 4:32:32us? The pound symbol comments
  6147. 4:32:41does the does the using the pound
  6148. 4:32:43symbols. So notice how we use that to do
  6149. 4:32:45multiple lines of comments. So we did
  6150. 4:32:48one here, we did one here. We can have
  6151. 4:32:50as many we can have
  6152. 4:32:53um as many
  6153. 4:32:56uh comment lines as we want and they
  6154. 4:32:59will all get ignored.
  6155. 4:33:10What is those? It's supposed to be
  6156. 4:33:11multi-line comments, but for some reason
  6157. 4:33:13it's not working. Um, it so the the
  6158. 4:33:17triple quote is supposed to be like
  6159. 4:33:19representing that you can have a whole
  6160. 4:33:21block of comments.
  6161. 4:33:25I don't know why it's not working in
  6162. 4:33:26collab for me.
  6163. 4:33:31It's working for you. Okay. Okay. I
  6164. 4:33:32don't know why it's not working.
  6165. 4:33:42Single quote.
  6166. 4:33:55It's still It still displays here, which
  6167. 4:33:57I don't get why that's happening.
  6168. 4:34:05It's kind of weird to me.
  6169. 4:34:11Yeah, I don't get why inconsistency. It
  6170. 4:34:13usually It usually works for me. I don't
  6171. 4:34:15get that at all.
  6172. 4:34:41Still still doesn't work for me. I don't
  6173. 4:34:42know why that doesn't
  6174. 4:34:56Yeah.
  6175. 4:34:57I don't get why that's not really liking
  6176. 4:34:59those triple quotes. Oh well. I mean,
  6177. 4:35:02not a big deal. We can just do
  6178. 4:35:05Okay, we can do we can try single.
  6179. 4:35:12Still doesn't work.
  6180. 4:35:25Yeah. Oh well, we can do a pound symbol.
  6181. 4:35:29That will always work. Pound symbol is
  6182. 4:35:31honestly more popular anyway. Most code
  6183. 4:35:34that you see in the wild will have pound
  6184. 4:35:36symbols wherever they're doing um
  6185. 4:35:39wherever they're doing uh comments. So
  6186. 4:35:42that it's fine. Just use a paddle for
  6187. 4:35:44now.
  6188. 4:35:57Uh yeah, that's correct. I don't know
  6189. 4:35:59why that's that's correct. Um I don't
  6190. 4:36:02know why collab doesn't seem to like
  6191. 4:36:03that. It should be ignored
  6192. 4:36:06um generally with the triple quotes, but
  6193. 4:36:09uh that's okay. I'm not too concerned
  6194. 4:36:11about it for now. I guess what you and
  6195. 4:36:14when I do comments, you're usually going
  6196. 4:36:15to see me using the pound symbol
  6197. 4:36:17anyways. It' be very rare that I would
  6198. 4:36:18need to do uh quotes.
  6199. 4:36:27Yeah, it's weird that collab doesn't
  6200. 4:36:29work very consistently. That's okay.
  6201. 4:36:32All right. What I want to show us is I
  6202. 4:36:36want to move on to the input. So, I want
  6203. 4:36:38to I want you guys to see
  6204. 4:36:40I want you guys to type in this code
  6205. 4:36:42here that will take input from a text
  6206. 4:36:45box and save it into a variable called
  6207. 4:36:47name. So, the code we're going to do is
  6208. 4:36:50going to be like this. It's going to be
  6209. 4:36:52name equals input
  6210. 4:36:56and then we'll put um please
  6211. 4:37:00enter your name.
  6212. 4:37:04Okay.
  6213. 4:37:06So this, by the way, I'm going to
  6214. 4:37:09comment this code here. Um, this code
  6215. 4:37:13should
  6216. 4:37:15create
  6217. 4:37:17a text box for us to put in our name.
  6218. 4:37:24Okay, so that's what should happen. So
  6219. 4:37:26when we run this, um, it should pop open
  6220. 4:37:30a text box right below this. And we can
  6221. 4:37:33type in our name and hit enter. And when
  6222. 4:37:35we do that, it will store that result in
  6223. 4:37:38this variable called name, which we can
  6224. 4:37:39use uh wherever we want to in the code.
  6225. 4:37:42So if I hit run, there's that text box.
  6226. 4:37:45Do you guys see that? There's the text
  6227. 4:37:48box. And see how it says, please enter
  6228. 4:37:51your name. And so we can type in our
  6229. 4:37:53name.
  6230. 4:37:56And we hit enter. And there it's stored
  6231. 4:37:59in the name. We can even um display name
  6232. 4:38:03by doing print and then the name which
  6233. 4:38:07will display uh the name that we stored
  6234. 4:38:10when we did the input.
  6235. 4:38:12So try this one out. Try this code out
  6236. 4:38:16for yourself. Try typing input
  6237. 4:38:18parenthesis
  6238. 4:38:20and then you want to have some text
  6239. 4:38:21there. It doesn't matter exactly what it
  6240. 4:38:22is, but something like please enter your
  6241. 4:38:24name or enter your name.
  6242. 4:38:27Try that out. And then it should store
  6243. 4:38:30uh you should be able to type in the box
  6244. 4:38:32that shows up. Hit enter on your
  6245. 4:38:34keyboard. It should save that. And then
  6246. 4:38:36you can um print it out. You can print
  6247. 4:38:39out that name which will um
  6248. 4:38:42display that whatever we typed in
  6249. 4:38:44before.
  6250. 4:38:52What does it look like, Roberto? What
  6251. 4:38:54does it look like? Were
  6252. 4:39:02you Were other people able to run this?
  6253. 4:39:06Oh, yeah. Thank you, Melanie. Yeah, I
  6254. 4:39:08see that. Perfect.
  6255. 4:39:11Perfect. That looks good to me.
  6256. 4:39:20Uh, you don't need a space. um it just
  6257. 4:39:24looks nice, right? It's so that's a good
  6258. 4:39:27practice to have the space so that uh
  6259. 4:39:29this code is um evenly spaced out and it
  6260. 4:39:33looks nicer on the on the screen.
  6261. 4:39:41Um name equals input print hello there
  6262. 4:39:45uh name
  6263. 4:39:52You need Yeah. So, uh, Roberto, you need
  6264. 4:39:55a you need a comma after after the
  6265. 4:39:58quotes.
  6266. 4:40:00After the quotes, you need a comma after
  6267. 4:40:02the quotes to signal to Python that
  6268. 4:40:05you're putting in you have you have some
  6269. 4:40:07text and then an additional input.
  6270. 4:40:12So, it need it needs to be more like it
  6271. 4:40:14needs to be like this. print. Um, hello
  6272. 4:40:17there.
  6273. 4:40:19And then you need an extra comma.
  6274. 4:40:24See how I have an extra comma after the
  6275. 4:40:25quote. You need you need that.
  6276. 4:40:32Sorry. Now, Roberto's uh Kiati.
  6277. 4:40:36Hope I'm pronouncing that right.
  6278. 4:40:40Okay. So, do we feel good about input
  6279. 4:40:42and what it does?
  6280. 4:40:45Perfect. Do we feel good about input and
  6281. 4:40:47what it does? It It brings up a text
  6282. 4:40:49box.
  6283. 4:40:53It Did you hit enter, Roberto? To like
  6284. 4:40:55Were you able to type something in and
  6285. 4:40:57hit It's going to run until you hit
  6286. 4:40:59enter.
  6287. 4:41:00You have to type in the text and then
  6288. 4:41:02hit enter into the box.
  6289. 4:41:06So, let me rerun this. So, it See how
  6290. 4:41:10it's still running? See how this like
  6291. 4:41:12it's going to keep running forever until
  6292. 4:41:15I type something in
  6293. 4:41:21and then when I hit enter it will stop.
  6294. 4:41:28What does your code look like?
  6295. 4:41:52Okay, that looks right.
  6296. 4:41:56Try try stopping it and rerunning it.
  6297. 4:42:03Try try hitting the stop button and then
  6298. 4:42:06rerun it.
  6299. 4:42:18um you should so yeah you should put
  6300. 4:42:20that in a different cell. So if you if
  6301. 4:42:24you separate your code into individual
  6302. 4:42:25cells so you could do like um you could
  6303. 4:42:30do name. So we could we could separate
  6304. 4:42:32this. So this code is the only code
  6305. 4:42:36that's running in this cell.
  6306. 4:42:49That doesn't make sense. Something else
  6307. 4:42:52is
  6308. 4:42:54that doesn't make sense cuz like this
  6309. 4:42:56collab tab is only taking up 235
  6310. 4:43:00megabytes.
  6311. 4:43:02So something is
  6312. 4:43:05chewing up your memory that's not really
  6313. 4:43:07I I can't imagine. Are you using collab?
  6314. 4:43:10You can see like it's not using that
  6315. 4:43:13much. Only 240
  6316. 4:43:15230ish.
  6317. 4:43:17Yeah, I don't think I don't think Collab
  6318. 4:43:19is the culprit unless you loaded in some
  6319. 4:43:21really massive data or something.
  6320. 4:43:25I can't imagine that's the issue.
  6321. 4:43:29You did. You loaded in data. That's
  6322. 4:43:32I mean Yeah. Then it's going to it's
  6323. 4:43:34going to take in memory. Oh, okay. Okay.
  6324. 4:43:37Okay.
  6325. 4:43:43Uh MJ, what are you on? Are you on
  6326. 4:43:49Yeah. Could you screenshot it?
  6327. 4:43:52If it's not working for you, could you
  6328. 4:43:53try collab? If Could you try collab just
  6329. 4:43:56for the sake of like getting it running?
  6330. 4:43:59Things should work in Collab pretty
  6331. 4:44:01easily.
  6332. 4:44:04You're using Collab and nothing's
  6333. 4:44:05working. Uh, are you making sure it's a
  6334. 4:44:07code cell and not a text cell?
  6335. 4:44:12It's not a text cell like this,
  6336. 4:44:15which would be like,
  6337. 4:44:18this is where it will be blue.
  6338. 4:44:29Did you have that? You need to make sure
  6339. 4:44:30it's code. Yeah.
  6340. 4:44:34And when I run that, it's going to be
  6341. 4:44:35it's going to display text. Yeah.
  6342. 4:44:41Okay. Great.
  6343. 4:44:45Glad that it's working. Great.
  6344. 4:44:49Okay.
  6345. 4:44:51Um All right. One more. Uh one more
  6346. 4:44:54example what I want to show you guys is
  6347. 4:44:56how to do how to include the name in a
  6348. 4:44:59print statement. So if we do something
  6349. 4:45:01like print. So um we can include the
  6350. 4:45:06name in a print statement. So if we do
  6351. 4:45:10something like print and then we have um
  6352. 4:45:13hello there and then we have um this and
  6353. 4:45:18then we have welcome to Python.
  6354. 4:45:23um this will
  6355. 4:45:26uh this will display all of that
  6356. 4:45:29together. So notice that we can have as
  6357. 4:45:32many um pieces of information that we
  6358. 4:45:34want to display kind of one after the
  6359. 4:45:36other as long as they're separated by
  6360. 4:45:38these commas.
  6361. 4:45:40So we have this uh text,
  6362. 4:45:44this text because text is stored in that
  6363. 4:45:47variable. Um, then this text and then we
  6364. 4:45:54print that all out and we can have this
  6365. 4:45:56whole collection of text displayed to
  6366. 4:45:58the screen. Try that one out.
  6367. 4:46:11Oh, they do the same thing. They do the
  6368. 4:46:14same thing. So, the comma and the Sorry.
  6369. 4:46:15Yeah, I just noticed the demo does a
  6370. 4:46:17plus. They do the same thing in Python.
  6371. 4:46:20So, we can swap that over to a plus.
  6372. 4:46:21Both of them work.
  6373. 4:46:24They have the same I shouldn't say they
  6374. 4:46:25do the same thing, but they have the
  6375. 4:46:27same effect.
  6376. 4:46:29They have the same effect.
  6377. 4:46:31Actually, there's no You need a little
  6378. 4:46:33bit more spacing here. So, the comma
  6379. 4:46:35gives you a little bit better uh
  6380. 4:46:37spacing.
  6381. 4:46:41So, what the Let me break this down.
  6382. 4:46:43what the so plus
  6383. 4:46:46plus um adds together
  6384. 4:46:51uh text and so what we're doing here
  6385. 4:46:54technically is adding all our text
  6386. 4:46:56together and then displaying it. Um, so
  6387. 4:46:59plus as together text and then the comma
  6388. 4:47:03um,
  6389. 4:47:05uh, prints out multiple pieces of text.
  6390. 4:47:11So they they have the same effect, but
  6391. 4:47:13yeah, you can use you can use either
  6392. 4:47:15one.
  6393. 4:47:25Okay.
  6394. 4:47:27Um, one thing I wanted to show you guys
  6395. 4:47:28too, by the way, in Collab, if you're
  6396. 4:47:31working inside of Collab, I want you to
  6397. 4:47:34hover over your name variable.
  6398. 4:47:37So, if you just take your mouse and
  6399. 4:47:39hover over that,
  6400. 4:47:42do you guys see what it says here?
  6401. 4:47:46Do you see how it says string name and
  6402. 4:47:49then it has the value of that, which is
  6403. 4:47:52which is my name. So, that's something
  6404. 4:47:54cool about Collab is if you hover over
  6405. 4:47:57variables, it will tell you what their
  6406. 4:47:59type is. Now, we haven't learned about
  6407. 4:48:01types, but any text inside of quotes is
  6408. 4:48:05a string. It's it's what we would call a
  6409. 4:48:07string. We're going to learn about that.
  6410. 4:48:09And
  6411. 4:48:11um
  6412. 4:48:13we it also displays what data we
  6413. 4:48:15currently have stored in that variable.
  6414. 4:48:17So all you have to do is hover over a
  6415. 4:48:19variable um to to see what the value is.
  6416. 4:48:24Yeah, that's yeah, that's kind of a
  6417. 4:48:26limitation of VS Code. That's true.
  6418. 4:48:30It doesn't show you immediately on
  6419. 4:48:32hovering.
  6420. 4:48:47Don't see the value on hovering. So, um,
  6421. 4:48:50click into the cell. You have to click
  6422. 4:48:52into the cell and then hover over it.
  6423. 4:48:55Click into the cell and then hover over
  6424. 4:48:57it. It should it should work. Yeah, you
  6425. 4:49:00have to click on the cell or whatever
  6426. 4:49:02cell you're on and then uh hover over
  6427. 4:49:06that and it should work.
  6428. 4:49:10Uh, Mariel asks, "How do we integrate
  6429. 4:49:12that Python code to a client application
  6430. 4:49:14for a user to enter a value?
  6431. 4:49:17um we would likely have a different set
  6432. 4:49:20of code to do that. Um there is Python
  6433. 4:49:23code that can get a UI and uh we we will
  6434. 4:49:26see that um later on in the in the like
  6435. 4:49:30way later on towards the end of the
  6436. 4:49:32program. We'll see that um we can we can
  6437. 4:49:34write Python code to do a UI essentially
  6438. 4:49:38to to make like a almost like a web page
  6439. 4:49:41for someone to enter some input. We'll
  6440. 4:49:44see that uh much later on. So, we're not
  6441. 4:49:47going to get to that right now. It's
  6442. 4:49:48really complex.
  6443. 4:49:55Um what is the purpose of having
  6444. 4:49:58multiple cells? It's so that we can run
  6445. 4:50:00individual pieces of code within those
  6446. 4:50:02cells. It allows us to isolate, right?
  6447. 4:50:05Because I can run I can run code inside
  6448. 4:50:07of these cells and they don't affect any
  6449. 4:50:09other cell. So, it's it's just for like
  6450. 4:50:12debugging and isolation, which is nice,
  6451. 4:50:16right? I don't need to worry about
  6452. 4:50:17running all of it at once. I can run one
  6453. 4:50:19cell at a time.
  6454. 4:50:34Okay. Any other questions?
  6455. 4:50:44Um, can we execute multiple lines
  6456. 4:50:46together? Yes, we did that. Here I had
  6457. 4:50:49multiple. So, I'll I'll show you again.
  6458. 4:50:51I can do um print
  6459. 4:50:54um this is one statement
  6460. 4:50:58and then I can come down and do uh print
  6461. 4:51:02um this is another and then maybe I can
  6462. 4:51:06do some math.
  6463. 4:51:13So you can have as many lines as you
  6464. 4:51:15want
  6465. 4:51:17within a cell.
  6466. 4:51:22Within a cell, you can have as many
  6467. 4:51:24lines of code as you want.
  6468. 4:51:34Is there a way to tell it the order the
  6469. 4:51:36cells execute? Um, you no, if you if you
  6470. 4:51:40go up to um if you go up to run all,
  6471. 4:51:44it's going to run them all in order from
  6472. 4:51:45top to bottom. Uh, in order to tell
  6473. 4:51:49which cells to execute, you can
  6474. 4:51:51rearrange them. You can always like So,
  6475. 4:51:54I could rearrange these cells, by the
  6476. 4:51:56way, by I think there's a way to move it
  6477. 4:51:58down.
  6478. 4:52:00So, I can move it down. So now I'm
  6479. 4:52:03rearranging. So you can move cells. I
  6480. 4:52:05think you can even drag and drop them.
  6481. 4:52:07So notice how I took the one that's at
  6482. 4:52:09the very top and I'm moving it down.
  6483. 4:52:13Otherwise, you have to click, right? You
  6484. 4:52:15just have to like I can run them in any
  6485. 4:52:17order. If I just click like if I click
  6486. 4:52:20here, it will run that one first. If I
  6487. 4:52:21go back up here, it will run that one
  6488. 4:52:24next. So you just click around which
  6489. 4:52:27ones you want to run. Does that make
  6490. 4:52:29sense?
  6491. 4:52:30I can run them in any order as long as I
  6492. 4:52:32click on whatever order I want to do it
  6493. 4:52:34in.
  6494. 4:52:41Okay.
  6495. 4:52:43All right.
  6496. 4:52:45Perfect. So, that that wraps up that
  6497. 4:52:47demo. I hope it was informative. I hope
  6498. 4:52:49you saw the the print statement. Um
  6499. 4:52:51we're going to see that many times. The
  6500. 4:52:53input statement. Um that's pretty cool.
  6501. 4:52:56Um and you got to run you got to run
  6502. 4:52:58some Python. So, if it's your first time
  6503. 4:53:00ever doing programming, congratulations.
  6504. 4:53:02You ran some Python code. That is really
  6505. 4:53:04exciting. Um, so glad we got to do that.
  6506. 4:53:07Um, let's go back to our notes
  6507. 4:53:15and then we'll um
  6508. 4:53:18let let me share the screen.
  6509. 4:53:26Okay.
  6510. 4:53:28So the next thing on our agenda is to
  6511. 4:53:32cover variables and data types. So I
  6512. 4:53:34just said like text is the string data
  6513. 4:53:37type but let's learn about all the
  6514. 4:53:38different data types that are going to
  6515. 4:53:39be available to us inside of Python and
  6516. 4:53:42let's talk about variables. Um it's
  6517. 4:53:44going to be a good discussion. So um I
  6518. 4:53:47think what we'll do is we'll take a
  6519. 4:53:48fivem minute break now and we come back
  6520. 4:53:51and we can start this uh discussion
  6521. 4:53:53about variables and data types. Um, so
  6522. 4:53:56let's take uh a fivem minute break.
  6523. 4:54:00And so let's try to be back um around
  6524. 4:54:06uh 8:30.
  6525. 4:54:13Okay.
  6526. 4:54:21Try to be back around 8 8:30.
  6527. 4:54:34Okay. So, what are variables? These are
  6528. 4:54:38um
  6529. 4:54:40basically our way of storing data to
  6530. 4:54:43make it easier to reference them and
  6531. 4:54:45manipulate uh throughout our program.
  6532. 4:54:48So, we've actually already used a
  6533. 4:54:49variable. We we used one in our uh demo
  6534. 4:54:52we just did where we called the input
  6535. 4:54:54the name. Uh we stored that input into a
  6536. 4:54:58variable called name. And so um
  6537. 4:55:01variables just really are a reference to
  6538. 4:55:05some data. That's all they are. They
  6539. 4:55:07allow us to reference that data
  6540. 4:55:09throughout the program. We can store
  6541. 4:55:11information into a variable and then
  6542. 4:55:13access it throughout our code. Um so on
  6543. 4:55:16this screen are some examples of
  6544. 4:55:17variables. Now, variables have names,
  6545. 4:55:20which is why I said usually we want
  6546. 4:55:23those to be meaningful. Like X is a
  6547. 4:55:25valid name, but it's not that
  6548. 4:55:28interesting of a name. It doesn't give
  6549. 4:55:30us that information much information
  6550. 4:55:32about what it's what it really means.
  6551. 4:55:33So, probably not the best name. Um, but
  6552. 4:55:37we have things like uh we can we can
  6553. 4:55:40store some text inside of this variable
  6554. 4:55:43called name. We can store a number
  6555. 4:55:44inside of this um variable called price.
  6556. 4:55:47we can store uh a true or a false value
  6557. 4:55:50inside of this variable called
  6558. 4:55:52is_active.
  6559. 4:55:54Um and so variables will show up all
  6560. 4:55:58over our code and uh they are basically
  6561. 4:56:01our way to reference some values. Now
  6562. 4:56:06these things over here are basically
  6563. 4:56:09different types of data that we need to
  6564. 4:56:11learn about, right? So we need to learn
  6565. 4:56:13about what is a 10 versus what is in
  6566. 4:56:15something inside of quotes. is it a
  6567. 4:56:17string versus something that has
  6568. 4:56:19decimals which is a floatingoint number
  6569. 4:56:22versus something that is true or false
  6570. 4:56:23which is a boolean value. We need to
  6571. 4:56:25learn about those data types. But notice
  6572. 4:56:28how all of these are being referenced by
  6573. 4:56:31a um by a variable that has some name to
  6574. 4:56:35it. Okay. So the variable is this guy.
  6575. 4:56:38It is our reference to that data. Um and
  6576. 4:56:42we will use variables throughout um so
  6577. 4:56:45that we can have you know references to
  6578. 4:56:47information in our code.
  6579. 4:56:50So variables are fundamental um to to
  6580. 4:56:54working with Python.
  6581. 4:56:56Um now variables can store different
  6582. 4:56:58kinds of data. So I just alluded to
  6583. 4:57:00that. And so the different types of data
  6584. 4:57:02available to us in Python kind of fall
  6585. 4:57:04in these two different categories. one
  6586. 4:57:06being single values or what are known as
  6587. 4:57:09scalar values. So these are things like
  6588. 4:57:12integers. So the number 10, the number
  6589. 4:57:151, the number 2,323,
  6590. 4:57:19those are all whole number integers. Um
  6591. 4:57:22floats, which are anything with a
  6592. 4:57:23decimal.
  6593. 4:57:25So 32.3,
  6594. 4:57:273.14,
  6595. 4:57:29um 1.2, those are all floating point
  6596. 4:57:32numbers. Um, booleans only have two
  6597. 4:57:36options. They only have true or false.
  6598. 4:57:38So, they represent kind of a binary uh
  6599. 4:57:41value um which we say is true or false.
  6600. 4:57:46And um then we also have um complex
  6601. 4:57:50numbers which are which have imaginary
  6602. 4:57:53uh parts to them. We won't really be
  6603. 4:57:55dealing with complex numbers too much so
  6604. 4:57:57I wouldn't worry about them. But in
  6605. 4:57:59reality, Python supports working with
  6606. 4:58:01the uh complex numbers and doing complex
  6607. 4:58:03math. But uh so so complex numbers just
  6608. 4:58:06have kind of a real part and an
  6609. 4:58:08imaginary part to them. Um wouldn't
  6610. 4:58:11worry too much about that. Again, we're
  6611. 4:58:12not really going to work with those ever
  6612. 4:58:14throughout throughout the program, but
  6613. 4:58:15it does exist. Python supports it. So
  6614. 4:58:18scalar data, single values, think
  6615. 4:58:22numbers, think single numbers like
  6616. 4:58:24floats, think integers, um single uh
  6617. 4:58:27true or false values. So these kinds of
  6618. 4:58:30data can be stored into variables.
  6619. 4:58:33On the opposite end of the spectrum are
  6620. 4:58:36aggregated types that we are storing
  6621. 4:58:39multiple things.
  6622. 4:58:41So we're going to learn about all of
  6623. 4:58:43those, but um probably the most common
  6624. 4:58:45and one that we've already dealt with is
  6625. 4:58:47going to be a string. So a string is
  6626. 4:58:49technically an aggregated type because
  6627. 4:58:51it has multiple characters that form,
  6628. 4:58:54you know, an overall uh string, which is
  6629. 4:58:57a string is usually you you know it's a
  6630. 4:59:00string because it's inside of quotes,
  6631. 4:59:02right? It's inside of these double
  6632. 4:59:03quotes or single quotes. Um, Python
  6633. 4:59:07actually doesn't care about quotes
  6634. 4:59:09really in terms of if it's single or
  6635. 4:59:10double as long as you're consistent with
  6636. 4:59:12it. Like if you if you start with double
  6637. 4:59:15quotes, you should end with double
  6638. 4:59:16quotes. If you start with single, you
  6639. 4:59:18should end with single. Python doesn't
  6640. 4:59:20really care either way. Um, so strings
  6641. 4:59:25are going to represent um collections of
  6642. 4:59:27characters. Um we are going to talk
  6643. 4:59:30about sets which are basically like u an
  6644. 4:59:33array of unique values. Um so we'll talk
  6645. 4:59:37about sets we'll talk about lists which
  6646. 4:59:39are a really important structure. It's
  6647. 4:59:41basically an array that can hold many
  6648. 4:59:44different types of data. Um so we'll
  6649. 4:59:46talk about list. We'll talk about
  6650. 4:59:48tupils. So you if you see that word
  6651. 4:59:50tuple e that is um people some people
  6652. 4:59:54pronounce it tuple. I I like to call it
  6653. 4:59:56tupole, but um that is going to be very
  6654. 4:59:59similar to an array. It's just going to
  6655. 5:00:01have slight differences and uh if you
  6656. 5:00:04can change it or not. Tupils you
  6657. 5:00:06actually cannot change once you create
  6658. 5:00:07it. Um versus list you can modify list.
  6659. 5:00:11You can add things to it. You can remove
  6660. 5:00:12things from it. Tupils you cannot. So
  6661. 5:00:15we're going to learn about those
  6662. 5:00:16differences as we go along and start
  6663. 5:00:18working with these different types of
  6664. 5:00:19data.
  6665. 5:00:21Um but they are designed to hold
  6666. 5:00:23multiple values, right? So you can see
  6667. 5:00:25in that example that list has integers,
  6668. 5:00:27it has strings, it can it can hold
  6669. 5:00:29multiple types which is if you're coming
  6670. 5:00:32from other languages is generally not
  6671. 5:00:34the case. Um like arrays in Java, arrays
  6672. 5:00:37in C, they can only hold one type of
  6673. 5:00:40data in the array. They can't hold
  6674. 5:00:42multiple.
  6675. 5:00:43Um
  6676. 5:00:45so uh then finally a dictionary. A
  6677. 5:00:48dictionary is if you're coming from
  6678. 5:00:49other languages, it's like a map, a
  6679. 5:00:51hashmap. Basically it allows you to have
  6680. 5:00:54uh keys mapped to values. So it's a
  6681. 5:00:57really dictionaries are highly useful
  6682. 5:00:59for storing information where we want to
  6683. 5:01:02reference like this value maps to this
  6684. 5:01:06value. So for instance in this
  6685. 5:01:07dictionary the string a maps to one and
  6686. 5:01:11then the string b maps to uh you know
  6687. 5:01:14two and or whatever it maps to. And this
  6688. 5:01:19will allow us to look up values in the
  6689. 5:01:22dictionary. So we could look up, hey,
  6690. 5:01:23what is the value stored at key A or
  6691. 5:01:25what is the value stored at key B? Those
  6692. 5:01:28kind of things. Dictionaries will be
  6693. 5:01:30incredibly useful. We're going to
  6694. 5:01:32explore all of those more as we go along
  6695. 5:01:34in the lesson, but um for right now, it
  6696. 5:01:38should be making sense that there are
  6697. 5:01:40some data types that store multiple
  6698. 5:01:42values like array or sorry, lists, um
  6699. 5:01:45dictionary, strings, and then there are
  6700. 5:01:48some data types that only have a single
  6701. 5:01:49value like a single number like a float,
  6702. 5:01:52integer, um boolean.
  6703. 5:01:55Okay, so more to come on aggregated
  6704. 5:01:57data. We're going to work with those,
  6705. 5:01:58learn about the differences, learn about
  6706. 5:02:00what it looks like in the code to work
  6707. 5:02:02with the set, a dictionary, tupil, list,
  6708. 5:02:05but those generally hold multiple values
  6709. 5:02:08or can hold multiple values whereas um
  6710. 5:02:12scalar data is only going to hold one.
  6711. 5:02:18Okay.
  6712. 5:02:22All right. So uh so as we said earlier
  6713. 5:02:27um you know integers, floats, booleans,
  6714. 5:02:30they only hold a single value. By the
  6715. 5:02:33way, inside of Python, if you ever want
  6716. 5:02:35to see what the type of a variable is.
  6717. 5:02:38So let's say we know we have a variable
  6718. 5:02:39called name. We can always check the
  6719. 5:02:42what data type it is by by using the
  6720. 5:02:45built-in type function. So we can use
  6721. 5:02:48type and then pass in that variable
  6722. 5:02:51and this will display what data type it
  6723. 5:02:54is. So um if we stored the value 42 in
  6724. 5:02:59some variable called int, if we um
  6725. 5:03:02displayed if we did type um if we did
  6726. 5:03:06type of this it would uh produce int
  6727. 5:03:09which would say okay this value is an
  6728. 5:03:12integer versus 3.14 that's going to be a
  6729. 5:03:15float versus capital t true that's going
  6730. 5:03:19to be uh the boolean type bool.
  6731. 5:03:27Okay. So, uh we have integers, we have
  6732. 5:03:30floats, we have booleans, all of which
  6733. 5:03:33we will use throughout and we'll see
  6734. 5:03:34where we will use those one versus the
  6735. 5:03:37other. We'll learn about that.
  6736. 5:03:41Um as I said, complex. So, uh just
  6737. 5:03:44showing you here that those exist
  6738. 5:03:46obviously. Um like I said complex has uh
  6739. 5:03:49a real part and an imaginary part which
  6740. 5:03:52you can access separately. So if you
  6741. 5:03:54store a value as a complex you can uh
  6742. 5:03:57access its real and imaginary parts
  6743. 5:03:59separately which you may need to do for
  6744. 5:04:01some type of uh calculations.
  6745. 5:04:04Um again we won't really work with
  6746. 5:04:07complex numbers in in this program. So
  6747. 5:04:09not a big deal for us but it is
  6748. 5:04:11supported
  6749. 5:04:12and you know a lot of um mathematical
  6750. 5:04:15packages in Python will use complex
  6751. 5:04:18numbers uh if they need to but we won't
  6752. 5:04:21really do it in this program. There's
  6753. 5:04:24not really a need to for us.
  6754. 5:04:28All right. So aggregated data um we have
  6755. 5:04:32those strings which we've already seen.
  6756. 5:04:34Those are the things inside of quotes.
  6757. 5:04:35We have sets which are going to be
  6758. 5:04:37collections of data um that are unique
  6759. 5:04:42basically only allowing one uh copy of
  6760. 5:04:45those elements inside the set. We're
  6761. 5:04:46going to learn about that. Um list which
  6762. 5:04:50is going to be a collection of items
  6763. 5:04:53which we can change, we can add things
  6764. 5:04:54to it, we can remove. Um lists are
  6765. 5:04:56really awesome uh structure in Python.
  6766. 5:05:00Um, what I want you to see right now
  6767. 5:05:02though is you can start to see the
  6768. 5:05:05syntax differences, right? So, like a
  6769. 5:05:07set, um, a a set is where we have, uh,
  6770. 5:05:13this brace. Notice that a set is created
  6771. 5:05:16with a curly brace versus a list which
  6772. 5:05:18is created with a bracket. So, right
  6773. 5:05:20away, like when you see a brace, you
  6774. 5:05:23should be thinking either a set or
  6775. 5:05:25dictionary. Those are the two things
  6776. 5:05:27that are created with a curly brace. Um,
  6777. 5:05:30and you know it's a dictionary because a
  6778. 5:05:31dictionary will have the colon which
  6779. 5:05:33will map I'll show you that on the next
  6780. 5:05:35screen. But that will map things from
  6781. 5:05:37key to value. Um, depending on if you
  6782. 5:05:40know left and right of the colon. Um,
  6783. 5:05:43but do you guys see that like the syntax
  6784. 5:05:45difference of a list? A list has a
  6785. 5:05:48bracket set has a curly brace. Um,
  6786. 5:05:52that's just one small difference. you
  6787. 5:05:53know, we're going to learn like what is
  6788. 5:05:54the actual difference between a set and
  6789. 5:05:56a list, but that's just one I'm pointing
  6790. 5:05:58out right now.
  6791. 5:06:02Um, what does mutable mean? So, mutable
  6792. 5:06:05uh means that we can change it. It's
  6793. 5:06:08it's able to be changed. So, immutable
  6794. 5:06:12would be we cannot change it.
  6795. 5:06:17Yeah. And and one thing about a list
  6796. 5:06:19that's really nice is every list has a
  6797. 5:06:22natural ordering to it which is actually
  6798. 5:06:24really beneficial. So a list has a
  6799. 5:06:27notion of the first item, the second
  6800. 5:06:30item, the third item, the fourth. That's
  6801. 5:06:33really important for accessing data
  6802. 5:06:35within the list. Okay. So lists are
  6803. 5:06:39really powerful. Um
  6804. 5:06:43yeah. So a a set the reason it shows
  6805. 5:06:46it's in a different order is because a
  6806. 5:06:48set does not maintain order. A set never
  6807. 5:06:51maintains order because um it a set is
  6808. 5:06:55not you do not access items by by order.
  6809. 5:07:00So that's just something unique to a set
  6810. 5:07:02is that it doesn't have a natural order.
  6811. 5:07:03So every time you print it out, it will
  6812. 5:07:05display in a different order.
  6813. 5:07:06Potentially it's random. It's random
  6814. 5:07:08order when you when you display it. A
  6815. 5:07:11set is just meant to be a general
  6816. 5:07:13collection. Think of it like a bucket.
  6817. 5:07:15Like here's this bucket of items that I
  6818. 5:07:17have.
  6819. 5:07:18It's just a collection of items. A list
  6820. 5:07:21actually maintains an order, a
  6821. 5:07:23consistent order of items.
  6822. 5:07:27This is different.
  6823. 5:07:30So we'll talk more about that when we
  6824. 5:07:31get into those.
  6825. 5:07:45Okay.
  6826. 5:07:48All right. So I wanted to show you also
  6827. 5:07:49the tupole in the dictionary. So a
  6828. 5:07:51tupole
  6829. 5:07:53is also a collection of items. Now the
  6830. 5:07:55tupil is ordered. So it's like a list.
  6831. 5:07:58It's ordered but it is immutable.
  6832. 5:08:02Meaning you cannot change a tupole. So
  6833. 5:08:04once you create a tupil you cannot
  6834. 5:08:07change it or else you'll get an error.
  6835. 5:08:09Python will tell you hey this is
  6836. 5:08:10immutable I can't change this. So if you
  6837. 5:08:13try changing being if you try to add
  6838. 5:08:15something to the tupil if you try to
  6839. 5:08:17modify one of the entries in the tupil
  6840. 5:08:19like if I try to if I go in and try to
  6841. 5:08:21change this a um to a d
  6842. 5:08:26um this would not be allowed. This would
  6843. 5:08:28this would throw an error. The
  6844. 5:08:30interpreter would say hey you're trying
  6845. 5:08:31to change something that cannot be
  6846. 5:08:32changed. So tupils are immutable but
  6847. 5:08:36they have a benefit beyond a set of
  6848. 5:08:38actually being ordered. So there there's
  6849. 5:08:41a natural ordering to a tupil where this
  6850. 5:08:43is the first item, this is the second
  6851. 5:08:45item, this is the third and every time
  6852. 5:08:46you display a tupil will be in a
  6853. 5:08:48consistent order. But tupils are not
  6854. 5:08:52like a list. You can't change it. So
  6855. 5:08:55tupils are useful for situations where
  6856. 5:08:58you want ordering, but you don't want
  6857. 5:09:00anybody to change any of that data
  6858. 5:09:02that's in the tupil. It's it's not
  6859. 5:09:04changeable.
  6860. 5:09:06Mutable meaning it just means changeable
  6861. 5:09:08like you can modify it. If if something
  6862. 5:09:11is mutable, you can modify it.
  6863. 5:09:15Immutable like a tupole is immutable. We
  6864. 5:09:17cannot modify it once we create it.
  6865. 5:09:19That's what it is.
  6866. 5:09:23Okay.
  6867. 5:09:26Yeah. Okay. So then finally a
  6868. 5:09:29dictionary. Now you by the way um look
  6869. 5:09:32at the tupole. See how it's created with
  6870. 5:09:34a parenthesis.
  6871. 5:09:37So that's different than the curly
  6872. 5:09:38brace. That's different than the
  6873. 5:09:39bracket. Right? So a tupole you know
  6874. 5:09:41it's a tupil because of the parenthesis
  6875. 5:09:44and the items are separated by a comma
  6876. 5:09:46just how just how they are in a set and
  6877. 5:09:47just how they are in a list. Um
  6878. 5:09:51so so the the parenthesis gives it away
  6879. 5:09:53that it's a tupole. Um now look at the
  6880. 5:09:56dictionary and the dictionary is um a
  6881. 5:10:00collection of key value pairs. So this
  6882. 5:10:03is a key value pair. This is a key value
  6883. 5:10:05pair. Um this is a key value pair and on
  6884. 5:10:08and on. We can have as many as we want.
  6885. 5:10:10And one thing I want you to notice about
  6886. 5:10:12this is there is no restriction on the
  6887. 5:10:16data types of the keys and the values.
  6888. 5:10:18So keys can be integers, keys can be
  6889. 5:10:22strings, values can be integers, values
  6890. 5:10:25can be strings, values could be floats,
  6891. 5:10:28values could even be other dictionaries
  6892. 5:10:31or lists. So value like we could have
  6893. 5:10:35what's called a nested dictionary where
  6894. 5:10:37we actually have something mapping over
  6895. 5:10:39to another dictionary.
  6896. 5:10:42That's totally possible in Python. So we
  6897. 5:10:45can have dictionaries that part of the
  6898. 5:10:47values inside of the dictionary actually
  6899. 5:10:49have our dictionaries themselves and
  6900. 5:10:52that would represent kind of a nested
  6901. 5:10:54structure there. So for instance this
  6902. 5:10:57name could map to a dictionary with
  6903. 5:10:59everybody's name in it. Um or it could
  6904. 5:11:03map to a list um you know
  6905. 5:11:07could map to a list it can map to
  6906. 5:11:09whatever it could map to a tupole. Uh so
  6907. 5:11:11you there's really no restriction in
  6908. 5:11:12what the keys and values uh are going to
  6909. 5:11:15be.
  6910. 5:11:18Uh Brent is it more efficient than the
  6911. 5:11:20other uh is what more efficient than the
  6912. 5:11:23other methods? Just want to clarify your
  6913. 5:11:25question so I so I answer it properly.
  6914. 5:11:35The tupole versus using an array.
  6915. 5:11:38Yeah. Yeah. Yeah. So uh these are all
  6916. 5:11:41good questions. So um the tupole
  6917. 5:11:46is guaranteed not to be changed. So it
  6918. 5:11:49is a little faster when we are looking
  6919. 5:11:52up items like when we are referencing
  6920. 5:11:53items. It's a little bit faster because
  6921. 5:11:56uh we know that it's not going to be
  6922. 5:11:58modified ever. So everything is going to
  6923. 5:12:00be consistently in the same spot. So
  6924. 5:12:03like whatever's first is going to stay
  6925. 5:12:05first, whatever's second is going to
  6926. 5:12:06stay second and on and on. So tupil is
  6927. 5:12:09is nice in that sense. A list can be
  6928. 5:12:12changed. So whatever is first may not
  6929. 5:12:15guarantee to be first in the future. We
  6930. 5:12:17can modify it. We can remove things. We
  6931. 5:12:19can add things to the list. So we can
  6932. 5:12:22expand. The list is very like dynamic.
  6933. 5:12:24The list. So the list is less efficient
  6934. 5:12:28because it's way more dynamic. Does that
  6935. 5:12:30make sense? Like it can change. You can
  6936. 5:12:31add you can keep expanding the list by
  6937. 5:12:34adding things to it. You can shrink the
  6938. 5:12:35list by removing things from it.
  6939. 5:12:38So list is way more dynamic which for a
  6940. 5:12:41lot of scenarios is useful,
  6941. 5:12:44right? We want to be able to add and
  6942. 5:12:45remove and modify things.
  6943. 5:12:48Um but a tupole is more rigid in the
  6944. 5:12:51sense that once you create it, you
  6945. 5:12:52cannot change anything about it.
  6946. 5:13:00Yeah. Yeah. So a dictionary is good for
  6947. 5:13:03Yeah. Like a phone book would be a good
  6948. 5:13:05example of a dictionary because you with
  6949. 5:13:08a dictionary you're usually looking up
  6950. 5:13:10things. So you so like a diction in a
  6951. 5:13:12phone book you have a name that maps to
  6952. 5:13:15a phone number.
  6953. 5:13:17Um so yes you have a you have that a
  6954. 5:13:21dictionary will map a key to a value
  6955. 5:13:23just like a name would be mapped to a
  6956. 5:13:25phone number. So yeah a phone book makes
  6957. 5:13:28a lot of sense.
  6958. 5:13:31Um,
  6959. 5:13:34a list, a list is like any is like a
  6960. 5:13:36normal like like your grocery list. Like
  6961. 5:13:38you may add things to it, you may remove
  6962. 5:13:39things from it, you may change things on
  6963. 5:13:41it. It's very dynamic. Um, a tupole is
  6964. 5:13:45kind of like a fixed um set of data
  6965. 5:13:49that's ordered in some way. So maybe
  6966. 5:13:52like um what you would see on on a on a
  6967. 5:13:55letter like you have your name, you have
  6968. 5:13:57your address, you have um your zip code,
  6969. 5:14:00like you kind of have those and it it
  6970. 5:14:02should stay that way in order to mail
  6971. 5:14:04the letter kind of thing.
  6972. 5:14:07Uh can you convert a tupil to a list?
  6973. 5:14:09Yes, you can do vice versa. You can
  6974. 5:14:11convert a tupil to a list and you can
  6975. 5:14:13convert um you can convert a list to a
  6976. 5:14:15tupole. Yes, you can convert between
  6977. 5:14:18them.
  6978. 5:14:20I'll show us examples of that later.
  6979. 5:14:29Okay. So, just to recap there,
  6980. 5:14:33tupil is not changeable, but it has an
  6981. 5:14:38order. So, it has a natural ordering to
  6982. 5:14:40it. Whatever is first is first. Whatever
  6983. 5:14:43second is second, third and third. So,
  6984. 5:14:44you can access things based on their
  6985. 5:14:46position within the tupole. That's
  6986. 5:14:48really nice. But you cannot modify
  6987. 5:14:50anything about a tupil once you create
  6988. 5:14:51it.
  6989. 5:14:52Okay. A list has an ordering to it. You
  6990. 5:14:57can access things based on their
  6991. 5:14:58position. But a list is dynamic. It is
  6992. 5:15:01mutable. Meaning you can change it. You
  6993. 5:15:03can change values. You can add things to
  6994. 5:15:06it. You can remove things from it. Okay?
  6995. 5:15:08So very dynamic. That's what a list is.
  6996. 5:15:10Um dictionary. It maps keys to values.
  6997. 5:15:14No restrictions on what those keys and
  6998. 5:15:16values can be.
  6999. 5:15:18All right. And then a set. A set is
  7000. 5:15:20think of it like a bucket. It just has
  7001. 5:15:22things in it. A set has no order to it.
  7002. 5:15:25So you cannot access things based on
  7003. 5:15:26their order. And every time you uh
  7004. 5:15:29display the set, you can get a different
  7005. 5:15:31ordering. Um
  7006. 5:15:34but a a set only is special in that it
  7007. 5:15:38only allows unique items. So if you try
  7008. 5:15:40to put multiple copies of a piece of
  7009. 5:15:42data, it's only going to keep one of
  7010. 5:15:43them. So a a set is like a bucket with
  7011. 5:15:47only unique things in it.
  7012. 5:15:50Okay. So and sometimes that's really
  7013. 5:15:52useful is to know like what are the
  7014. 5:15:54unique values? Uh a set would help us
  7015. 5:15:57maintain that.
  7016. 5:16:02Any questions about those? You know, we
  7017. 5:16:04have to we have to work with this and
  7018. 5:16:05see this in the code and we will. But
  7019. 5:16:07just any questions right now about these
  7020. 5:16:09different types of data that we're
  7021. 5:16:10talking about.
  7022. 5:16:21Okay.
  7023. 5:16:23Very good.
  7024. 5:16:26All right. Let's talk about assignment.
  7025. 5:16:28So, what that means is um Oops. Let's
  7026. 5:16:31talk about assignment which means that
  7027. 5:16:34we will be um taking a variable name and
  7028. 5:16:38assigning data to it. Now we've already
  7029. 5:16:40seen this. We already saw it in our demo
  7030. 5:16:42where we did input. We did name equals
  7031. 5:16:45input.
  7032. 5:16:46So the equals symbol is how we assign
  7033. 5:16:51values to a variable.
  7034. 5:16:54That makes sense, right? It's very like
  7035. 5:16:56self-explanatory.
  7036. 5:16:58But um what we should think about with a
  7037. 5:17:02variable is really the fact that a
  7038. 5:17:04variable is is a reference to that data.
  7039. 5:17:09Okay. So when we say x= 34, we are
  7040. 5:17:13assigning 34 to the name x. So x becomes
  7041. 5:17:18a variable which is referencing the data
  7042. 5:17:21which is an integer 34. Right?
  7043. 5:17:25What's really interesting about that and
  7044. 5:17:26this is how you can kind of test your
  7045. 5:17:28intuition of the fact that this is a
  7046. 5:17:31reference is if we come along and have
  7047. 5:17:33another name Y and we set that equal to
  7048. 5:17:36X.
  7049. 5:17:38This is just saying that we are creating
  7050. 5:17:41another reference that is equal to the
  7051. 5:17:43reference we already have. Now, why
  7052. 5:17:46would we ever do that? Probably we
  7053. 5:17:48wouldn't. That's kind of redundant. But
  7054. 5:17:50this just proves that they're ultimately
  7055. 5:17:52references because when we display x, we
  7056. 5:17:56get 34. Of course, that's what we stored
  7057. 5:17:59the value 34
  7058. 5:18:01uh referenced by x.
  7059. 5:18:04And then when we print y, we get the
  7060. 5:18:06same number, right? We get 34. And why
  7061. 5:18:09does that happen? Because we we
  7062. 5:18:11literally declared y equal to x. Meaning
  7063. 5:18:15y should reference the same data that x
  7064. 5:18:17does. Okay, so as variables they are
  7065. 5:18:21equal meaning that um X is being
  7066. 5:18:25assigned to Y meaning Y should reference
  7067. 5:18:28the same data that X does. So they they
  7068. 5:18:31uh contain the same data. Now what's
  7069. 5:18:34interesting is if you print out the ID.
  7070. 5:18:37So the ID is the internal
  7071. 5:18:40um the internal memory address
  7072. 5:18:46of of the reference.
  7073. 5:18:49Um now it usually we don't care about
  7074. 5:18:51that but this is just to prove the point
  7075. 5:18:53is that you can see these are the same
  7076. 5:18:55address. These are the same. That's by
  7077. 5:18:57design because we're saying okay I have
  7078. 5:19:00this reference X which is referencing
  7079. 5:19:01this data 34. it's stored at this
  7080. 5:19:03address. Um, and then when I come along
  7081. 5:19:06and say, okay, y equals x, that's just
  7082. 5:19:10the same reference. You see how it's the
  7083. 5:19:12same exact address,
  7084. 5:19:14same reference.
  7085. 5:19:16So, just proving that variables are
  7086. 5:19:19literally just references to data. They
  7087. 5:19:21allow us to reference that data, which
  7088. 5:19:23is really, really, you know, nice. So,
  7089. 5:19:25we can reuse x throughout the code. Um,
  7090. 5:19:28we can reuse name. we can you know
  7091. 5:19:30whatever we create we can reuse.
  7092. 5:19:35Um if you look over to the right we have
  7093. 5:19:38an alternative example which um now
  7094. 5:19:42resets y to store a new value. So
  7095. 5:19:45instead of saying y equals to x we
  7096. 5:19:47actually overwrite y and reassign it to
  7097. 5:19:49the integer 78. That's a new piece of
  7098. 5:19:52data right 78. So now if you look at
  7099. 5:19:56their their uh references, they're
  7100. 5:19:59different. These are different. And that
  7101. 5:20:02makes sense because now they're pointing
  7102. 5:20:03to two different uh pieces of data,
  7103. 5:20:06right? X is pointing to 34. Y is
  7104. 5:20:09referencing to 78. So of course they're
  7105. 5:20:12going to be different uh different
  7106. 5:20:13addresses. And this is a bit of a typo.
  7107. 5:20:16This should say ID of Y
  7108. 5:20:19because we're ref we're talking about Y.
  7109. 5:20:21It's a bit of a typo there.
  7110. 5:20:24Okay, so hopefully this now this example
  7111. 5:20:27is just to reinforce the fact that when
  7112. 5:20:29we use the equal sign, we're setting
  7113. 5:20:31equal we're setting a variable name
  7114. 5:20:33equal to a piece of data, right? And
  7115. 5:20:36that is creating a reference to that
  7116. 5:20:38piece of data.
  7117. 5:20:40That's all we're that's all we're saying
  7118. 5:20:41with this. So we are assigning a piece
  7119. 5:20:45of data to that reference X or Y or
  7120. 5:20:47whatever it is.
  7121. 5:20:53Okay.
  7122. 5:20:56All right. Let me ask you guys. Um, what
  7123. 5:20:59is the default data type of a variable
  7124. 5:21:02assigned using the input function? This
  7125. 5:21:04is an interesting question. We didn't
  7126. 5:21:05actually cover this, so I'm really
  7127. 5:21:06curious to see what you guys think about
  7128. 5:21:08this.
  7129. 5:21:24A lot of votes for for string.
  7130. 5:21:28Let's get a few more.
  7131. 5:21:38Perfect. Yeah. So, water votes receipt
  7132. 5:21:40it is a string. So, that that begs the
  7133. 5:21:44question like what happens if we input a
  7134. 5:21:47number? Like what happens if we put in a
  7135. 5:21:49two? What happens to that? You know that
  7136. 5:21:53two will actually be read in as the
  7137. 5:21:56string two. So it would be So if we use
  7138. 5:21:59the input and we it pulls up that text
  7139. 5:22:01box and we put in a number like two
  7140. 5:22:06um and we set that equal to the variable
  7141. 5:22:09x, whatever we name that name x,
  7142. 5:22:11whatever. What that really means is x is
  7143. 5:22:14going to be um equal to the the um x is
  7144. 5:22:19going to be equal to the
  7145. 5:22:22uh string 2. So that's something to be
  7146. 5:22:25cautious about with the input is it
  7147. 5:22:28always assumes the input data is going
  7148. 5:22:30to be a string. So luckily there's a way
  7149. 5:22:33to convert between strings and numbers.
  7150. 5:22:37So if we wanted to turn this into the
  7151. 5:22:39actual number, what we would do is use
  7152. 5:22:41the the data type function int, which
  7153. 5:22:44would convert uh this would convert it
  7154. 5:22:48over to the numerical two. Would
  7155. 5:22:50actually convert it from a string to an
  7156. 5:22:52integer. We just use int. Or we could
  7157. 5:22:54use like if we if somebody put in a
  7158. 5:22:56decimal like 2.5
  7159. 5:22:59then um we could do a float
  7160. 5:23:04of 2.5
  7161. 5:23:06and that would convert that over to uh
  7162. 5:23:09the the number.
  7163. 5:23:12Okay, let me actually show you guys
  7164. 5:23:15this. Let me go over to Collab real
  7165. 5:23:16quick and show you guys this. I know
  7166. 5:23:18it's not in a demo, but I think it'll be
  7167. 5:23:19better if I just show you what I mean by
  7168. 5:23:21this because this is an important point
  7169. 5:23:23with input.
  7170. 5:23:26So, let me uh stop sharing there. Let me
  7171. 5:23:30go over to Collab for a second so I can
  7172. 5:23:32show you literally what this means.
  7173. 5:23:38So, go back into the notebook here. So
  7174. 5:23:41what I want to show you is that um when
  7175. 5:23:45we do input
  7176. 5:23:47the default type
  7177. 5:23:51is string.
  7178. 5:23:53So for instance when I do um
  7179. 5:23:59when I do uh uh value equals to input
  7180. 5:24:05and let's try um enter your age.
  7181. 5:24:12Oops. Enter your age.
  7182. 5:24:15And then we uh run this.
  7183. 5:24:21So we enter the age. Now this is going
  7184. 5:24:24to be read in as a string. So even
  7185. 5:24:28though I'm putting a number there, it's
  7186. 5:24:30actually going to be read in as a
  7187. 5:24:31string. So now
  7188. 5:24:34look at what the type of value is.
  7189. 5:24:39It's a string. Do we see that? So
  7190. 5:24:42string. So this number even though we
  7191. 5:24:45put in a number it gets it the the input
  7192. 5:24:49function always converts it to a string
  7193. 5:24:52no matter what we put there. If we put a
  7194. 5:24:53decimal if we put a a large number it's
  7195. 5:24:56always going to assume it's a it's a
  7196. 5:24:58string. So luckily
  7197. 5:25:01um we can convert to an integer
  7198. 5:25:05by using int the int function.
  7199. 5:25:09So um we can print sorry we can say
  7200. 5:25:12value
  7201. 5:25:14uh or we can do int value which which
  7202. 5:25:18will convert that 32 string because
  7203. 5:25:23right now if I were to um just display
  7204. 5:25:27value it's a string 32. You can see it
  7205. 5:25:30inside of the quotes. But now when I do
  7206. 5:25:33this uh and I can run that now it's an
  7207. 5:25:37integer. Do we see that now it's
  7208. 5:25:40actually a number
  7209. 5:25:42which is great. It no longer has those
  7210. 5:25:43quotes. It's actually going to be
  7211. 5:25:45treated as an actual integer which which
  7212. 5:25:47may be useful for calculations or
  7213. 5:25:49storing it or whatever whatever we need
  7214. 5:25:51to do with it. So that's just one piece
  7215. 5:25:53of caution with the input is if you're
  7216. 5:25:56working with numerical data it's going
  7217. 5:25:58to treat it as a string. We have to
  7218. 5:26:00convert it.
  7219. 5:26:02Okay
  7220. 5:26:05questions on that. Does that make sense
  7221. 5:26:07to us? like the input's always going to
  7222. 5:26:09accept the input as a string. So if we
  7223. 5:26:12want to work with it alternatively
  7224. 5:26:15um we should convert it.
  7225. 5:26:20Uh you can yeah so like you could
  7226. 5:26:23convert um if I did this if I wrapped
  7227. 5:26:27this around in the int function that
  7228. 5:26:30would automatically
  7229. 5:26:32take whatever we put whatever this
  7230. 5:26:34returns would automatically be um cast
  7231. 5:26:37over to an int. So we could do that. So
  7232. 5:26:41let me show you that. So when I run
  7233. 5:26:43this, I can put in 32
  7234. 5:26:48and it it's like automatically going to
  7235. 5:26:51be casted to an integer. So there now
  7236. 5:26:53it's an integer. Does that make sense?
  7237. 5:26:55Like when I wrap this int around the
  7238. 5:26:57input, it's going to automatically
  7239. 5:26:59convert
  7240. 5:27:17Uh, what did you put in the input box?
  7241. 5:27:19So, yes, you'll get an error if you
  7242. 5:27:22don't put in a valid integer.
  7243. 5:27:24So, let's put in like if I put in my
  7244. 5:27:28name,
  7245. 5:27:30this is going to be this should be an
  7246. 5:27:31error because I don't know how to
  7247. 5:27:34convert this string over to a number. It
  7248. 5:27:36doesn't make sense to do that, right?
  7249. 5:27:38So, this should be an error.
  7250. 5:27:42Right? That will be an error because
  7251. 5:27:44it's a string.
  7252. 5:27:53So why did you get an error? Uh input
  7253. 5:27:55enter your age value int value.
  7254. 5:28:04Uh did the did the text box show up?
  7255. 5:28:06Maybe try separating it into a different
  7256. 5:28:08cell.
  7257. 5:28:10Try try putting the other two lines in a
  7258. 5:28:12different cell. Um, you need the text
  7259. 5:28:14box to show up and then you need to
  7260. 5:28:16enter something.
  7261. 5:28:18Yeah.
  7262. 5:28:22Okay.
  7263. 5:28:24All right. Does this all make sense? Any
  7264. 5:28:26questions about this? About the input
  7265. 5:28:28function.
  7266. 5:28:33Okay.
  7267. 5:28:36Good. Okay, let me go over back to the
  7268. 5:28:38notes then.
  7269. 5:28:46Okay.
  7270. 5:28:53All right. So, we have another demo.
  7271. 5:28:55We'll do that now. Uh I was just kind of
  7272. 5:28:57doing one, but let's go back over to
  7273. 5:28:59this will be demo five. Let's do that.
  7274. 5:29:01So, we're going to practice assigning
  7275. 5:29:03different values um to variables and
  7276. 5:29:06displaying them just so you get in the
  7277. 5:29:08habit of being able to create your own
  7278. 5:29:10variables and just go through that kind
  7279. 5:29:12of one more time. We'll do this one
  7280. 5:29:14relatively quickly um and then uh move
  7281. 5:29:17on.
  7282. 5:29:20So, this will be uh demo five.
  7283. 5:29:29So, let me pull that one up for you
  7284. 5:29:31guys.
  7285. 5:29:51Okay, let me share my screen.
  7286. 5:29:57All right. So, this is going to be demo
  7287. 5:29:59five. Um,
  7288. 5:30:03now again, like feel free to use
  7289. 5:30:05whatever platform you've been using.
  7290. 5:30:07Collab, Jupyter Notebook. I know this
  7291. 5:30:09instruction says set up a Jupyter
  7292. 5:30:11notebook. Feel free to use whatever you
  7293. 5:30:12want. You can use Collab. Um, whatever's
  7294. 5:30:15been working for you to build your build
  7295. 5:30:17your notebooks. So obviously this this
  7296. 5:30:20looks a little different than collab but
  7297. 5:30:21it's because it's the Jupiter. Um so we
  7298. 5:30:25create a notebook.
  7299. 5:30:27Now what I want you to see
  7300. 5:30:31is this takes the approach of everything
  7301. 5:30:33we just did. Let me zoom in on this. Uh,
  7302. 5:30:37I know that's a little small,
  7303. 5:30:41but this is doing everything we just
  7304. 5:30:43said we could do where we
  7305. 5:30:46um essentially take
  7306. 5:30:49So, I just want to zoom in on this. Um,
  7307. 5:30:52notice that we
  7308. 5:30:55uh take the um input and this will be
  7309. 5:31:00saved as a string.
  7310. 5:31:01Um,
  7311. 5:31:03so this will be saved as a string and
  7312. 5:31:06this will be saved into this name. And
  7313. 5:31:08for instance, this will be saved as a
  7314. 5:31:11string, but we convert it over to an
  7315. 5:31:12int, which is exactly the kind of
  7316. 5:31:14example I just did, right? Where we take
  7317. 5:31:17take an input, we convert it over to to
  7318. 5:31:21uh int.
  7319. 5:31:24Does somebody have Yeah. Does somebody
  7320. 5:31:26have the demos available? like if if
  7321. 5:31:29somebody doesn't mind sharing those in
  7322. 5:31:30the chat. I again I don't have the PDFs.
  7323. 5:31:33They should be from your LMS. They
  7324. 5:31:35should be in the reference material.
  7325. 5:31:37There should be a demos folder that you
  7326. 5:31:39can download. If somebody has those and
  7327. 5:31:41doesn't mind sharing them.
  7328. 5:31:45They have that folder of them, like a
  7329. 5:31:47zip folder of them, that'd be fantastic.
  7330. 5:31:54Yeah. Thanks. Thanks. This is This is
  7331. 5:31:58the demo we're going through currently.
  7332. 5:32:01Perfect. So, for you guys having trouble
  7333. 5:32:04navigating the demos, please download
  7334. 5:32:07this zip folder.
  7335. 5:32:10Download the zip folder that that these
  7336. 5:32:13guys are uploading. Thank you so much.
  7337. 5:32:14Download the zip folder so you have all
  7338. 5:32:16of them.
  7339. 5:32:18Please take a moment to do that.
  7340. 5:32:25Okay.
  7341. 5:32:26Uh,
  7342. 5:32:36copy the code and got an error at height
  7343. 5:32:38value. Use foot, not meter. I mean, it
  7344. 5:32:40shouldn't matter. It, you know, you
  7345. 5:32:42should just be the point of that one is
  7346. 5:32:44to put in a decimal.
  7347. 5:32:50How to create a new file. Um, what
  7348. 5:32:52platform are you on? Collab.
  7349. 5:32:56I don't know what platform you're on.
  7350. 5:32:58Collab. Uh, just go to file, new
  7351. 5:33:01notebook.
  7352. 5:33:03New notebook in drive, I think is what
  7353. 5:33:05it's called.
  7354. 5:33:13Do you see that? It should be like it
  7355. 5:33:15should be at the top. There should be a
  7356. 5:33:16file and then new notebook.
  7357. 5:33:22Let me go over to it.
  7358. 5:33:26Uh,
  7359. 5:33:30this one. You don't see this
  7360. 5:33:34file. It's at the top. The top of the
  7361. 5:33:36notebook. Do file and then new notebook.
  7362. 5:33:41You don't see new notebook.
  7363. 5:33:56Uh if you if you don't see that, just go
  7364. 5:33:58to a new tab. Just go to a new tab and
  7365. 5:34:01go to um Google Collab.
  7366. 5:34:05You can always do that. Just go to just
  7367. 5:34:07start a new um just go to Google Collab
  7368. 5:34:10and then it will let you like launch a
  7369. 5:34:12new notebook. So just just do that. Just
  7370. 5:34:13do a new tab if it doesn't work.
  7371. 5:34:22Okay.
  7372. 5:34:24So, by the way, one of those examples
  7373. 5:34:26was entering a float. So, it looked kind
  7374. 5:34:28of like this. So we had um our our
  7375. 5:34:31height is equal to float and then we had
  7376. 5:34:36uh input and then we had um enter your
  7377. 5:34:41height and then this was um uh some sort
  7378. 5:34:47of uh this should be some sort of
  7379. 5:34:49decimal value. So let's say it is um I
  7380. 5:34:54don't know uh 5.7
  7381. 5:34:57whatever that is uh feet it doesn't it's
  7382. 5:35:00just some decimal um and then we hit uh
  7383. 5:35:04we hit enter that will store the height
  7384. 5:35:07as a float so that when we um display
  7385. 5:35:10the height uh it will be rendered as a
  7386. 5:35:13float appropriately right that's what
  7387. 5:35:15that that's what should happen
  7388. 5:35:20that's the point of that It just needs
  7389. 5:35:21to be some decimal. It should work.
  7390. 5:35:31All right, let me go back to the demo
  7391. 5:35:34document.
  7392. 5:35:38All right, were you guys able to run
  7393. 5:35:40some of these? Like, were you able to
  7394. 5:35:42run some of the inputs and change them?
  7395. 5:35:44So, try these out on your own real
  7396. 5:35:46quick. like try doing int and then input
  7397. 5:35:48for enter your age. It should convert
  7398. 5:35:52that. You should be putting in a number
  7399. 5:35:54or else you'll get an error and it
  7400. 5:35:56should convert that over.
  7401. 5:36:04I by the way I wouldn't worry about this
  7402. 5:36:06last one uh because we haven't learned
  7403. 5:36:08about the comparison operator yet which
  7404. 5:36:11is this equals equals. So we'll learn
  7405. 5:36:13about that in a in a little bit in a few
  7406. 5:36:15minutes. So, don't worry about that one
  7407. 5:36:17too much right now. But at least these
  7408. 5:36:18first few should make some sense and we
  7409. 5:36:21should be able to do.
  7410. 5:36:25Were you guys able to run one of those
  7411. 5:36:27and convert over the the float or int
  7412. 5:36:32and do the input and convert it?
  7413. 5:36:36Did that work for you?
  7414. 5:36:41Give it a try.
  7415. 5:36:54Let me clear that. Any questions about
  7416. 5:36:57that?
  7417. 5:37:05Should look something like this.
  7418. 5:37:10Good. We're good on that on converting
  7419. 5:37:11over the input. Okay, perfect. Sounds
  7420. 5:37:14like Sounds like we're able to run that
  7421. 5:37:15and uh it was okay.
  7422. 5:37:23What are you entering for the feet?
  7423. 5:37:28Like, are you literally entering like
  7424. 5:37:30quotes?
  7425. 5:37:32Yeah, that's not going to work when you
  7426. 5:37:34do that because it's going to um there's
  7427. 5:37:37a string f, there's a character there.
  7428. 5:37:39it's not going to be able to convert
  7429. 5:37:40over to.
  7430. 5:37:42So if you did if you did 6.4 that would
  7431. 5:37:46work.
  7432. 5:37:48Any decimal should work. But like the f
  7433. 5:37:50is a character. So the the float doesn't
  7434. 5:37:54know how to convert over a character,
  7435. 5:37:57right? Yeah. So so that's not going to
  7436. 5:38:00work. You need to put in a decimal to to
  7437. 5:38:03be able to convert over to the number.
  7438. 5:38:07Okay.
  7439. 5:38:09Very good. Very good. Let's go back over
  7440. 5:38:12to our notes so we can continue along.
  7441. 5:38:291.7. Yeah. If you have any if you have
  7442. 5:38:33any character, it's not going to work.
  7443. 5:38:36It's not going to work. You need to put
  7444. 5:38:37in you need to put in a decimal.
  7445. 5:38:44All right, let's talk about operators.
  7446. 5:38:47So, these are going to be really
  7447. 5:38:48important. Um,
  7448. 5:38:54let's talk about operators so that we
  7449. 5:38:57can uh
  7450. 5:39:00uh be able to compare things and work
  7451. 5:39:03with things. Um so let's let's talk
  7452. 5:39:06about Python operators.
  7453. 5:39:08So what are operators? What do we mean
  7454. 5:39:10by that? In Python, operators are
  7455. 5:39:13special symbols or keywords that perform
  7456. 5:39:16operations. So as the name suggests,
  7457. 5:39:18it's performing some level of operation.
  7458. 5:39:21Um which means that the interpreter
  7459. 5:39:24should do some sort of logical
  7460. 5:39:25operation, mathematical operation,
  7461. 5:39:27relational operation to produce a
  7462. 5:39:30result. Um, so usually that means
  7463. 5:39:33there's going to be multiple variables
  7464. 5:39:35that are going to be used to do some
  7465. 5:39:36operation between. So an example of an
  7466. 5:39:39operation would be like adding,
  7467. 5:39:41subtracting, multiplying. That's an
  7468. 5:39:42operation. But we can have logical
  7469. 5:39:45operations like taking the um logical
  7470. 5:39:49and or logical or of things. We'll see
  7471. 5:39:51what that means. But um in Python,
  7472. 5:39:54there's many situations where we want to
  7473. 5:39:56we want to be able to do operations
  7474. 5:39:59between variables. Whether that's simple
  7475. 5:40:01mathematical or maybe some type of
  7476. 5:40:03relational like testing if a value is in
  7477. 5:40:06a list. That's an important operation.
  7478. 5:40:09Is 10 in my list? Is five in my list? Um
  7479. 5:40:13those are important operations. So we
  7480. 5:40:15want to learn about these operators and
  7481. 5:40:17they're going to be really important for
  7482. 5:40:18us going forward is because these will
  7483. 5:40:21be very standard. um things we will use
  7484. 5:40:24as we uh go along. So we're going to
  7485. 5:40:27spend some time talking about operators.
  7486. 5:40:30Um so it turns out in Python um you can
  7487. 5:40:33kind of group operators into many
  7488. 5:40:35different categories. Um there's going
  7489. 5:40:38to be standard arithmetic operators.
  7490. 5:40:40Those are your everyday things like
  7491. 5:40:41plus, minus, um division,
  7492. 5:40:44multiplication. Um, assignment
  7493. 5:40:46operators, which we've already seen, is
  7494. 5:40:48things like equals, where we're setting
  7495. 5:40:50a reference equal to something. We've
  7496. 5:40:52already seen that. That's an assignment.
  7497. 5:40:54Comparison, which is things like greater
  7498. 5:40:56than or less than. Those are important
  7499. 5:40:58for comparing values, comparing
  7500. 5:41:00variables. Um, logical operators are
  7501. 5:41:03going to be something like and and or,
  7502. 5:41:06which will um do a logical operation
  7503. 5:41:08between two two boolean values. That'll
  7504. 5:41:11be important. And then we have a
  7505. 5:41:13collection of miscellaneous operators.
  7506. 5:41:15Um those will be things like is
  7507. 5:41:17something in a collection like is five
  7508. 5:41:20in a list? That's an operator. So we'll
  7509. 5:41:23talk we're going to talk about all of
  7510. 5:41:24these but just pointing out that there's
  7511. 5:41:26many different categories of operators
  7512. 5:41:28in Python.
  7513. 5:41:35Okay, let's first talk about the
  7514. 5:41:36arithmetic operators. So these are going
  7515. 5:41:39to be your standard everyday um uh
  7516. 5:41:42operations between numbers. So if we
  7517. 5:41:45have numerical values like integers or
  7518. 5:41:47floats, we can do math between them.
  7519. 5:41:49That makes sense. Like that should be a
  7520. 5:41:50capability of Python and it certainly
  7521. 5:41:52is. We can add things, we can subtract
  7522. 5:41:55things, we can multiply things, we can
  7523. 5:41:57divide things. So um here are all those
  7524. 5:42:01operators. We have plus minus the
  7525. 5:42:03asterisk is a multiplication. So x
  7526. 5:42:06asterisk y will multiply those together.
  7527. 5:42:09So if we have two variables, one of them
  7528. 5:42:11is 50, one of them is four, we do x
  7529. 5:42:14asterisk y, that's going to multiply
  7530. 5:42:16them together to get 200. Pretty pretty
  7531. 5:42:18straightforward. Um
  7532. 5:42:21division is one that we should be
  7533. 5:42:22careful of. Of course, like we don't
  7534. 5:42:24want to divide by zero. So if you I if
  7535. 5:42:28the uh this secondary value that we end
  7536. 5:42:32up dividing by is zero, that'll give us
  7537. 5:42:34an error. Um the interpreter will say,
  7538. 5:42:36"Hey, you're trying to divide by zero."
  7539. 5:42:38We can't do that. It'll it'll produce an
  7540. 5:42:40error. So that's the only thing we have
  7541. 5:42:42to be on the lookout for with division.
  7542. 5:42:43Just don't want to divide by zero.
  7543. 5:42:46Um
  7544. 5:42:48so all these are pretty standard. I
  7545. 5:42:50think they all make sense.
  7546. 5:42:52Hopefully they do to you. I think
  7547. 5:42:54they're all pretty standard. you know,
  7548. 5:42:55the kinds of things you'd see on a on a
  7549. 5:42:57basic calculator. They all make sense.
  7550. 5:42:59They should exist. Now, here's some more
  7551. 5:43:02exotic ones. Um, I don't know if you
  7552. 5:43:05guys have ever seen the the modulus
  7553. 5:43:07operator, also known as modulo. This is
  7554. 5:43:10one that returns the remainder of a
  7555. 5:43:13division. Okay? So the the percentage
  7556. 5:43:15sign is a mathematical operation between
  7557. 5:43:18two numbers that returns not the
  7558. 5:43:21quotient like not the actual division
  7559. 5:43:23result but the remainder. So 50 divided
  7560. 5:43:27by four
  7561. 5:43:29um you know four goes into 50 um it goes
  7562. 5:43:34in there uh uh 12 times evenly but it
  7563. 5:43:38has two left over right. So there the
  7564. 5:43:41remainder there is two. So the result of
  7565. 5:43:44x mod we would read this as x mod y or
  7566. 5:43:48modulo y um returns two. So if you're if
  7567. 5:43:53you're unfamiliar with the modulo
  7568. 5:43:54operation that seems a little bizarre
  7569. 5:43:55that you take these two numbers
  7570. 5:43:58um oops it seems a little bizarre that
  7571. 5:44:01you take these two numbers and you like
  7572. 5:44:02do this operation and you get a
  7573. 5:44:06remainder result but it's actually a
  7574. 5:44:07very powerful operation. Um the reason
  7575. 5:44:11being is that sometimes we want to know
  7576. 5:44:13what the remainder is more than we want
  7577. 5:44:14to know what the quotient is. For
  7578. 5:44:16instance, things that are very like
  7579. 5:44:18cyclic in nature. Um so maybe we cycle
  7580. 5:44:21through a collection and we want to know
  7581. 5:44:24like how many times do we cycle through
  7582. 5:44:26and then we have something left over
  7583. 5:44:28which is the remainder. Um so the modulo
  7584. 5:44:31operation is pretty useful. You could
  7585. 5:44:33also check like if a number is even or
  7586. 5:44:35odd using this. Like so if you modulo by
  7587. 5:44:38two and it returns zero, that means it's
  7588. 5:44:40even, right? Because that means there's
  7589. 5:44:42there's nothing left over when I divide
  7590. 5:44:44by two. So modulo is kind of a nice way
  7591. 5:44:46to check if a number is even or odd. Um
  7592. 5:44:50so modulo is a pretty nice uh operation.
  7593. 5:44:53We'll use it from time to time. Uh but
  7594. 5:44:55that is the percent operator. So x
  7595. 5:44:59percent y will look for that remainder
  7596. 5:45:01of the division. Um now there is also a
  7597. 5:45:05double slash operator which is the
  7598. 5:45:08integer division operator. This is kind
  7599. 5:45:11of the reverse of modulo. It takes the
  7600. 5:45:14largest integer quotient that that uh we
  7601. 5:45:17can do from a division perspective. So
  7602. 5:45:20remember I said 50 / 4. We can divide 4
  7603. 5:45:23into 50 12 times evenly and we have two
  7604. 5:45:27left over. So the integer division will
  7605. 5:45:29just return to us an integer always
  7606. 5:45:32which will be that quotient.
  7607. 5:45:35So this is the quotient
  7608. 5:45:38um and this is the uh remainder of 50 /
  7609. 5:45:414. So the integer division returns to
  7610. 5:45:45you that whole number like the largest
  7611. 5:45:47number of times that that number goes
  7612. 5:45:50into the other. So 12 times evenly
  7613. 5:45:53obviously there's a remainder there but
  7614. 5:45:55um but but yeah so integer division that
  7615. 5:45:59one's useful if we want to know like how
  7616. 5:46:01many times can I fit a value into
  7617. 5:46:03another value a whole number of times
  7618. 5:46:06and that happens from from time to time
  7619. 5:46:07we may need to know that.
  7620. 5:46:11Okay last operation here is exponent. So
  7621. 5:46:15the exponent is the asterisk asterisk
  7622. 5:46:18operator. Um so that raises a number to
  7623. 5:46:23a power. Um so for instance like x star
  7624. 5:46:27y or asteris y would mean that we are
  7625. 5:46:30doing an operation like 5 to the 4th
  7626. 5:46:33power um which is 625.
  7627. 5:46:37Okay. So asterisk pretty useful. Like
  7628. 5:46:39probably the most common asterisk would
  7629. 5:46:41be squaring something which would be x
  7630. 5:46:43um star star 2 which would would would
  7631. 5:46:47be um x squared. So I mean that's a
  7632. 5:46:50pretty common operation there is to
  7633. 5:46:52raise something to the second power
  7634. 5:46:54maybe the third raising something to the
  7635. 5:46:56fourth probably less common but um the
  7636. 5:47:00the asterisk asterisk operator is is how
  7637. 5:47:03we do exponents in Python.
  7638. 5:47:07Okay,
  7639. 5:47:08so these are all basic arithmetic
  7640. 5:47:11operations we can do between variables
  7641. 5:47:13in Python. All right, any questions on
  7642. 5:47:17those? Do those kind of make sense to us
  7643. 5:47:19from a syntax perspective?
  7644. 5:47:23Pretty straightforward, I think.
  7645. 5:47:25Hopefully nothing too surprising there.
  7646. 5:47:27Um,
  7647. 5:47:29do you used to use module all the time
  7648. 5:47:31for date date calculation? Yeah. Yeah.
  7649. 5:47:34like when you uh find out how many like
  7650. 5:47:36days how many weeks or where you are in
  7651. 5:47:38the week, you cycle through like uh
  7652. 5:47:41modulo 7 or something.
  7653. 5:47:44That make sense?
  7654. 5:47:52Okay,
  7655. 5:47:54very good.
  7656. 5:47:56Okay, I want to talk about assignment
  7657. 5:47:58operators now. Now we've already seen
  7658. 5:48:00this which is the basic equal sign that
  7659. 5:48:04is a data assignment operator right so
  7660. 5:48:07that means that we are setting a value
  7661. 5:48:10equal to a reference so we are storing
  7662. 5:48:12data inside of this reference variable a
  7663. 5:48:15we use the basic equal sign as our
  7664. 5:48:18assignment operator so that equal sign
  7665. 5:48:20is called the assignment operator now
  7666. 5:48:23what's really awesome is we can combine
  7667. 5:48:26this basic assignment operator with our
  7668. 5:48:28arithmetic ones to update values
  7669. 5:48:33um and modify them uh as kind of a
  7670. 5:48:37shortcut to say uh so so for example
  7671. 5:48:40like a plus= 5 really represents the
  7672. 5:48:43fact that I want to reassign a to the
  7673. 5:48:47result of a + 5. So this means take
  7674. 5:48:52whatever it is add five to it and
  7675. 5:48:54reassign it to the value of a. So this
  7676. 5:48:57is the same thing as if we just shortcut
  7677. 5:49:00it in and Python will recognize if we do
  7678. 5:49:02plus equals 5 it's the same thing. So
  7679. 5:49:07and actually we can do that with any of
  7680. 5:49:09these arithmetic operators. So if we
  7681. 5:49:11want to take a variable multiply it by
  7682. 5:49:14two and reassign it to that variable we
  7683. 5:49:16can use star equals. So like a asterisk
  7684. 5:49:21equals 2 is the same thing as if we were
  7685. 5:49:25to reassign a to the value of a * 2.
  7686. 5:49:30Does that make sense on the
  7687. 5:49:33reassignment portion of that? So plus
  7688. 5:49:35equals divide equal modulo equals star
  7689. 5:49:39star equals would exponent something and
  7690. 5:49:41reset it back to the variable.
  7691. 5:49:45um minus equals we'll subtract and
  7692. 5:49:48reassign that back to the variable. So
  7693. 5:49:51you know x minus equ= 3 we'll subtract
  7694. 5:49:54three from x and re and basically update
  7695. 5:49:56it right reassign it back to x.
  7696. 5:50:01So so that's pretty useful like whenever
  7697. 5:50:03we need to do an operation and add like
  7698. 5:50:07um you know a a very typical
  7699. 5:50:11reassignment is to do like a plus equals
  7700. 5:50:131
  7701. 5:50:16That's a very typical reassignment
  7702. 5:50:18because what this is the same as is a
  7703. 5:50:21equals a + one. So that's like a single
  7704. 5:50:25increment of a. We're just updating it
  7705. 5:50:26by one.
  7706. 5:50:29So plus equals 1. We may see that from
  7707. 5:50:31time to time.
  7708. 5:50:33A loop coming on. Yeah. Yeah. These are
  7709. 5:50:35used in like while loops. Yeah. Like you
  7710. 5:50:38do plus equals and you increment it
  7711. 5:50:41until you reach a certain condition.
  7712. 5:50:43Yeah.
  7713. 5:50:44Now, if you're coming from other
  7714. 5:50:45languages, if you have programming
  7715. 5:50:47programming experience, you're coming
  7716. 5:50:48from other languages, Python does not
  7717. 5:50:50have an increment operator like plus+. I
  7718. 5:50:54wish it did, but it doesn't. So, like I
  7719. 5:50:56know in in Java and I think C they have
  7720. 5:50:59um you can do like uh a plus+ or
  7721. 5:51:03actually reverse you can do plus a but
  7722. 5:51:06um that does not exist in Python
  7723. 5:51:09unfortunately. You have to do the plus
  7724. 5:51:11equals reassignment. So they don't have
  7725. 5:51:13an increment operator. You'd have to
  7726. 5:51:15you'd have to do just plus equals one to
  7727. 5:51:18do the same effect as plus+.
  7728. 5:51:21So I I know some people ask about that,
  7729. 5:51:23but yeah, doesn't exist unfortunately.
  7730. 5:51:29All right. Any questions about
  7731. 5:51:30assignment? It's just really the equal
  7732. 5:51:32sign and we can tack on the arithmetic
  7733. 5:51:34to do some type of basic math and
  7734. 5:51:37reassign to the variable.
  7735. 5:51:39Hopefully the fact that we're using a
  7736. 5:51:41single equals makes sense. Where people
  7737. 5:51:43get confused all the time is the
  7738. 5:51:46difference between a single equal sign
  7739. 5:51:47and multi and two equal signs which
  7740. 5:51:50we're going to see. Two equal signs
  7741. 5:51:52means something completely different
  7742. 5:51:54than a single equal sign. Single equal
  7743. 5:51:57sign is an assignment. We are taking
  7744. 5:52:00data and storing it in a reference
  7745. 5:52:02variable,
  7746. 5:52:04right?
  7747. 5:52:06But multiple equal signs, we're going to
  7748. 5:52:07learn about what that means. That's
  7749. 5:52:08actually a comparison.
  7750. 5:52:10It's something different.
  7751. 5:52:13All right, we'll continue. Thank you
  7752. 5:52:14guys. All right, so we're talking about
  7753. 5:52:17uh comparison. So, uh we're going to
  7754. 5:52:20talk about a few operators that allow us
  7755. 5:52:22to compare two values. Now, this is
  7756. 5:52:24going to be useful as we go forward
  7757. 5:52:26because sometimes we want to know when
  7758. 5:52:27is a value bigger than something or less
  7759. 5:52:29than something or equal to something,
  7760. 5:52:31not equal to something. Those
  7761. 5:52:33comparisons are going to be useful.
  7762. 5:52:35um and we have a collection of operators
  7763. 5:52:38to do that for us. So again, one that I
  7764. 5:52:41think a lot of people get confused on is
  7765. 5:52:43the um equals comparison operator which
  7766. 5:52:47is uh the double equals symbol. So a lot
  7767. 5:52:52of people get confused on that. What is
  7768. 5:52:54the difference between a single equal
  7769. 5:52:55sign and a double? This double equal
  7770. 5:52:58sign is checking if two values are
  7771. 5:53:02equal.
  7772. 5:53:03Um
  7773. 5:53:05so for instance we have uh these two
  7774. 5:53:08numbers x and y they're both integers
  7775. 5:53:10that are 20. We check if x equals equals
  7776. 5:53:13to y and that returns true because
  7777. 5:53:18uh these two values are the same. They
  7778. 5:53:20both equal 20. So when x equals equals y
  7779. 5:53:23that is a true statement. So these
  7780. 5:53:26that's something to realize is that
  7781. 5:53:27these comparisons are things that return
  7782. 5:53:29booleans true or false because a number
  7783. 5:53:32is going to be bigger than another yes
  7784. 5:53:34you know true or false they they are a a
  7785. 5:53:38uh comparison that gives us a kind of a
  7786. 5:53:40yes or no answer. Um
  7787. 5:53:44so the equals equals checks if two
  7788. 5:53:46values are the same and then the uh not
  7789. 5:53:50equals operator which is uh an
  7790. 5:53:52exclamation point with an equals um
  7791. 5:53:55checks to see if two values are
  7792. 5:53:57different. So they are not equal. So for
  7793. 5:54:00instance if we had um uh 45 and 24 we we
  7794. 5:54:04uh do x not equals y that would return
  7795. 5:54:07true.
  7796. 5:54:08Um now if we had these two values as
  7797. 5:54:11before and we checked here x not equals
  7798. 5:54:14to y um this would be false because they
  7799. 5:54:18are equal right so um
  7800. 5:54:21not equals to checks if values are
  7801. 5:54:24different so that's a simple comparison
  7802. 5:54:26are they not equal um so so in this case
  7803. 5:54:29that would return true
  7804. 5:54:32so these are pretty useful if we want to
  7805. 5:54:34compare directly is a value equal to
  7806. 5:54:36another we use the equals equals If
  7807. 5:54:38they're different, we use the not
  7808. 5:54:40equals. And we're going to have
  7809. 5:54:42different scenarios where we will use
  7810. 5:54:43those.
  7811. 5:54:45I also want to call out the basic, you
  7812. 5:54:48know, greater than and less than. So the
  7813. 5:54:50this first one is the less than
  7814. 5:54:52operator. It is uh going to be obviously
  7815. 5:54:54returning true when a number is less
  7816. 5:54:56than another number. So when we have
  7817. 5:54:59things like 20 and uh 30, this x less
  7818. 5:55:04than y would return true because 20 is
  7819. 5:55:06definitely smaller than 30. So this
  7820. 5:55:08returns true. Um
  7821. 5:55:11and then greater than checks if a number
  7822. 5:55:13is bigger than another. So that
  7823. 5:55:15comparison uh x bigger than y in this
  7824. 5:55:18case would uh return true as a
  7825. 5:55:21comparison. So again these are all
  7826. 5:55:23operators that check uh comparison
  7827. 5:55:26between two numbers that will be uh
  7828. 5:55:30really useful as we go forward and start
  7829. 5:55:32to work with data and numbers and we do
  7830. 5:55:35comparisons.
  7831. 5:55:36uh we will do those all the time later
  7832. 5:55:38on.
  7833. 5:55:42Now there's also scenarios when when we
  7834. 5:55:44want to know is it less than or equal
  7835. 5:55:46to. So that operator just tacks on an
  7836. 5:55:49equal sign. So less than equals
  7837. 5:55:52is the less than or equal to operator.
  7838. 5:55:55So for instance 10 less than or equal to
  7839. 5:55:5730 that is true um because 10 is
  7840. 5:56:00certainly smaller than 30. But um it
  7841. 5:56:03would have been true even if x was 30.
  7842. 5:56:06That would also be true because 30 uh 30
  7843. 5:56:10equals to um 30 would equal to 30. That
  7844. 5:56:14would be a true statement.
  7845. 5:56:16Um greater than or equal to same same
  7846. 5:56:19scenario. We have a greater than and
  7847. 5:56:21then we have an equal sign right after
  7848. 5:56:22it. This returns true if something is
  7849. 5:56:24bigger than or equal to another number.
  7850. 5:56:27So here's an interesting one. We do 30
  7851. 5:56:29bigger than or equal to 30. That returns
  7852. 5:56:31true because 30 equals to 30. That that
  7853. 5:56:34makes sense. So less than or equal to
  7854. 5:56:36bigger than or equal to we can we can do
  7855. 5:56:38with these simple operators.
  7856. 5:56:42Uh is greater than greater than similar
  7857. 5:56:44to the usage of brackets? No.
  7858. 5:56:47Uh so greater than or greater than is
  7859. 5:56:49what's called a uh a bit shift operator.
  7860. 5:56:54It's a little bit different. Um I I
  7861. 5:56:57would I'm going to save any explanation
  7862. 5:57:00that just just look that one up is what
  7863. 5:57:02I'll say. It it does like a bit um a bit
  7864. 5:57:05manipulation which is um a bit of a bit
  7865. 5:57:10of a hassle to deal with but we we won't
  7866. 5:57:12ever use greater than or we won't ever
  7867. 5:57:14use greater than greater than. It's it
  7868. 5:57:16does some sort of a shifting operation
  7869. 5:57:18like a bit mathematics which we we don't
  7870. 5:57:21need to do.
  7871. 5:57:29Okay. So those are comparisons. Um let's
  7872. 5:57:32look at our logical operators. So now
  7873. 5:57:35these ones are going to be really really
  7874. 5:57:36interesting and useful when we get into
  7875. 5:57:38controlling the flow of our program. Um
  7876. 5:57:42so logical operators are used for
  7877. 5:57:44combining conditional statements. So
  7878. 5:57:47conditional statements are things that
  7879. 5:57:49return these are statements that return
  7880. 5:57:52um true or false. So they return a
  7881. 5:57:54boolean and we can it's it's like we are
  7882. 5:57:57combining them together in certain ways.
  7883. 5:58:00Okay,
  7884. 5:58:02so the and operator, let's look at that
  7885. 5:58:06one first, which in Python is the
  7886. 5:58:08literal word and. So that's very nice.
  7887. 5:58:11It's it's literally the the keyword and.
  7888. 5:58:14Um, and what this does is it takes the
  7889. 5:58:18result of some boolean comparison and
  7890. 5:58:20some other boolean comparison and
  7891. 5:58:22returns true if both of them are true.
  7892. 5:58:26So and will only return true as a
  7893. 5:58:30combination if both individual
  7894. 5:58:32statements are true. They both have to
  7895. 5:58:34be true. The moment one of them is false
  7896. 5:58:36and will return false.
  7897. 5:58:38So this is useful for doing a
  7898. 5:58:41combination of things where we want
  7899. 5:58:43every individual thing to be true. So a=
  7900. 5:58:471. This is a true statement because a
  7901. 5:58:50equals 1 and then b= 2 is a true
  7902. 5:58:52statement. So both of these would be
  7903. 5:58:54true. So therefore when we combine them
  7904. 5:58:57with the and this overall combination is
  7905. 5:59:00true.
  7906. 5:59:02So keep that in mind. These operators
  7907. 5:59:05are ones that combine individual logical
  7908. 5:59:09statements or conditional statements.
  7909. 5:59:11Right?
  7910. 5:59:13Okay. Now the one that is less
  7911. 5:59:16restrictive than and is the or statement
  7912. 5:59:19which um is used when you only want at
  7913. 5:59:24least one of the statements to be true.
  7914. 5:59:27So if we want to combine these things
  7915. 5:59:28and only require at least a minimum of
  7916. 5:59:31one to be true, we use the or statement.
  7917. 5:59:34So for instance, a= 1 is true because a=
  7918. 5:59:391. So that's true. and then B equals
  7919. 5:59:42equals to 2 is false. So this one is
  7920. 5:59:45false. But that doesn't matter from the
  7921. 5:59:47perspective of or because we have a
  7922. 5:59:49minimum of one of these statements being
  7923. 5:59:51true. So or when we use the logical
  7924. 5:59:53combination of or um we just need either
  7925. 5:59:57or to be true. So a= 1 is true. So this
  7926. 6:00:01overall returns true.
  7927. 6:00:04So or is something that will combine
  7928. 6:00:06conditional statements and return uh
  7929. 6:00:08return true if at least one of them is
  7930. 6:00:11true. If all of them are false or it
  7931. 6:00:13would return false because none of them
  7932. 6:00:15are none of them would be true.
  7933. 6:00:20Okay. So we have and we have or and then
  7934. 6:00:23we have not. So not is an interesting
  7935. 6:00:26one. Not essentially reverses a boolean.
  7936. 6:00:31So if we have a statement that is
  7937. 6:00:33inherently true and we put a not in
  7938. 6:00:35front of it, it will invert that to be
  7939. 6:00:38false. If we if we have something that
  7940. 6:00:40is false and we put a not in front of
  7941. 6:00:42it, it will return uh true.
  7942. 6:00:45One of the interesting examples in
  7943. 6:00:47Python and this trips up people all the
  7944. 6:00:49time is the fact that Python treats zero
  7945. 6:00:54very specially. So the integer zero
  7946. 6:00:58is
  7947. 6:00:59oops the integer zero is inherently
  7948. 6:01:02treated by Python as false.
  7949. 6:01:06So Python treats zero as false and then
  7950. 6:01:10every other integer as true. Basically
  7951. 6:01:12being Python is indicating that it is
  7952. 6:01:15something that is not zero. uh anything.
  7953. 6:01:19So like um B equals to one would be
  7954. 6:01:22treated as true because it's as long as
  7955. 6:01:25it's something that's not zero
  7956. 6:01:28then Python treats that integer as as a
  7957. 6:01:32true boolean essentially. Um so why
  7958. 6:01:36that's interesting is if you put a not
  7959. 6:01:38in front of this this would actually
  7960. 6:01:40return true because not false what is
  7961. 6:01:44the opposite of false? It is true right?
  7962. 6:01:46So not false it would return true. So we
  7963. 6:01:51we will see not from time to time. Uh
  7964. 6:01:54not shows up when we want to negate
  7965. 6:01:56something. So when we um you know you
  7966. 6:01:59know maybe we have a an iteration an
  7967. 6:02:02iterative loop and we say while not
  7968. 6:02:05finished and we you know then we will
  7969. 6:02:08execute a bunch of statements while we
  7970. 6:02:10continue to not be finished and then the
  7971. 6:02:12moment that that it finishes then it
  7972. 6:02:14then the loop would be over. So not is
  7973. 6:02:17powerful to kind of invert uh trus to
  7974. 6:02:20falses and falses to true.
  7975. 6:02:23Um so maybe we want to check if
  7976. 6:02:25something is not empty. Meaning that um
  7977. 6:02:28if it's empty
  7978. 6:02:30uh if it's not empty that would be
  7979. 6:02:32false. Not empty um you know maybe it
  7980. 6:02:36would return true. So not is something
  7981. 6:02:39we will uh see from time to time as a
  7982. 6:02:42negation operator logical negation.
  7983. 6:02:48Okay.
  7984. 6:02:50Any questions on
  7985. 6:02:52uh any questions on these operators?
  7986. 6:02:56These now these we're going to use these
  7987. 6:02:58in the control of the flow of our
  7988. 6:03:00program.
  7989. 6:03:05One more example for not. Yeah. So a
  7990. 6:03:07pretty typical case for not would be
  7991. 6:03:09something like this where um
  7992. 6:03:12uh maybe we have some code oops I always
  7993. 6:03:16forget to swap over to this maybe we
  7994. 6:03:18have some code that checks so if we have
  7995. 6:03:20a list
  7996. 6:03:23if we have a list and it's uh empty
  7997. 6:03:27okay and it has nothing in it let's say
  7998. 6:03:29it has nothing in it we could we could
  7999. 6:03:30have some code that says like if um if
  8000. 6:03:35not
  8001. 6:03:38uh list
  8002. 6:03:41um then we then do something. So then if
  8003. 6:03:47not list uh meaning that it's not empty
  8004. 6:03:50then check then grab the first value.
  8005. 6:03:52Let's say that grab the first value. So
  8006. 6:03:54we'd have some code like this. So uh
  8007. 6:03:58this this not is used to like we can
  8008. 6:04:01negate the fact that this is going to be
  8009. 6:04:03empty and then this would be true and
  8010. 6:04:06then we can continue to access something
  8011. 6:04:08because that would mean it's not empty.
  8012. 6:04:11So not empty is a pretty standard use
  8013. 6:04:15case for not like to check that
  8014. 6:04:16something is not empty.
  8015. 6:04:22Uh can we also use not for checking
  8016. 6:04:24value in the list? Yeah, that's that's
  8017. 6:04:26what we're doing here to say like is it
  8018. 6:04:28not empty?
  8019. 6:04:42Okay.
  8020. 6:04:47All right.
  8021. 6:04:49Let's go to some miscellaneous
  8022. 6:04:51operators. So we have now some of these
  8023. 6:04:54are going to be incredibly useful. The
  8024. 6:04:56one on this page not that useful. The is
  8025. 6:05:00mainly because it's very rare that we
  8026. 6:05:03would uh that we would check these. So
  8027. 6:05:07is is what we call the identity operator
  8028. 6:05:10and this is something that um checks to
  8029. 6:05:12see if two references are the same.
  8030. 6:05:16Okay, two references are the same. um
  8031. 6:05:20meaning that they're referencing the
  8032. 6:05:21same piece of data. Um so now this is a
  8033. 6:05:26very interesting case where we have a
  8034. 6:05:28equals to a list b equals to the list.
  8035. 6:05:32However, when we ask the question a is b
  8036. 6:05:35this would actually return false. Now
  8037. 6:05:37that seems very counterintuitive but the
  8038. 6:05:39reason that's the case is because we are
  8039. 6:05:41creating two different references. We're
  8040. 6:05:44saying A equals to this list, B equals
  8041. 6:05:47to this list, which is a whole new piece
  8042. 6:05:49of data.
  8043. 6:05:51It's a whole new piece of data. So
  8044. 6:05:53therefore, we can't claim that they're
  8045. 6:05:54the same reference even though they're
  8046. 6:05:56the Now what would be true is A equals
  8047. 6:06:00equals B because their data is the same.
  8048. 6:06:04That would be true, but their their
  8049. 6:06:06references are different because they're
  8050. 6:06:09different variables, right? A and B are
  8051. 6:06:10different variables.
  8052. 6:06:12Different memory location. Exactly.
  8053. 6:06:14Different references. So A and A is B is
  8054. 6:06:18the same as checking um if ID A equals
  8055. 6:06:23equals IDB. Does that make sense? That's
  8056. 6:06:26basically checking that that logical uh
  8057. 6:06:29comparison if their addresses are the
  8058. 6:06:32same. It's the same check. So is is
  8059. 6:06:35basically a shorthand for doing this.
  8060. 6:06:39And so th those would be false because
  8061. 6:06:41they're going to be two different
  8062. 6:06:42references, A and B.
  8063. 6:06:45Um, however, we can use our not. So A is
  8064. 6:06:49not B. That's actually true because it's
  8065. 6:06:51the inverse of is, right? So that that
  8066. 6:06:54actually would invert the false and this
  8067. 6:06:56would be true. A is not B. That is true.
  8068. 6:07:02Behind the scenes, yeah, like the
  8069. 6:07:04memory, yeah, the location in the
  8070. 6:07:06computer memory is different. Yes,
  8071. 6:07:08because they are different variables,
  8072. 6:07:09different references.
  8073. 6:07:15Yeah, the data is equal. The data is
  8074. 6:07:17equal, but the references are different,
  8075. 6:07:20which is what this checks. You know, we
  8076. 6:07:23have two different names, A and B. Those
  8077. 6:07:25are different.
  8078. 6:07:27Now, take a look at this last example.
  8079. 6:07:29This is saying A is a list. B equals to
  8080. 6:07:32A. Now, remember what that does? That is
  8081. 6:07:35the assignment of we're saying B is the
  8082. 6:07:38same reference as A. That's what this
  8083. 6:07:40does here. The same reference. So does
  8084. 6:07:44it make sense to us that when we ask now
  8085. 6:07:47A is B. This should be true. And it is
  8086. 6:07:50like this is true because um
  8087. 6:07:56this is true because they are literally
  8088. 6:07:57the same reference. We're setting B
  8089. 6:07:59equals to A. So they are referring to
  8090. 6:08:03the same data. Now
  8091. 6:08:05um in terms of a variable reference they
  8092. 6:08:08so now their ids are the same
  8093. 6:08:09essentially their memory locations are
  8094. 6:08:12the same.
  8095. 6:08:15If you want to compare only data then uh
  8096. 6:08:19what operators have we looked at that
  8097. 6:08:20are comparison
  8098. 6:08:22we should think about that I mean we
  8099. 6:08:24just saw if we go back a couple slides
  8100. 6:08:26we have a bunch of operators to compare
  8101. 6:08:28data that's these guys right like equals
  8102. 6:08:31equals greater than less than so what we
  8103. 6:08:35could do is say does the data equal the
  8104. 6:08:38other data which would be something like
  8105. 6:08:39this equals equals operator
  8106. 6:08:42when two values are equal not the
  8107. 6:08:45reference ES. Does that make sense? This
  8108. 6:08:47equals equals is checking if two values
  8109. 6:08:49are the same, which is the the data, not
  8110. 6:08:52the not the reference.
  8111. 6:09:02Okay.
  8112. 6:09:05Now I want to show you a really powerful
  8113. 6:09:08um I want to show you a really powerful
  8114. 6:09:11miscellaneous operator which is going to
  8115. 6:09:14be uh which is going to be the in or
  8116. 6:09:18what's called the membership operator.
  8117. 6:09:21So this is an operator that checks if a
  8118. 6:09:25value is a member of a collection.
  8119. 6:09:29So this could be like uh this could be
  8120. 6:09:33like um you know where we have a list, a
  8121. 6:09:38tupil, a dictionary, just a collection
  8122. 6:09:40of data and we want to know is a value a
  8123. 6:09:44member of that collection which is
  8124. 6:09:46really useful for testing you know do we
  8125. 6:09:49have membership of something inside of
  8126. 6:09:51something else. So for instance, let's
  8127. 6:09:54say we have a list and we have a list A
  8128. 6:09:57equals and then we have 20 45 and 10
  8129. 6:09:59inside that list. So if we ask the
  8130. 6:10:01question
  8131. 6:10:0310 in A, this actually would return true
  8132. 6:10:06because 10 is a member of A. 10 is a
  8133. 6:10:09member of that list. So that is true.
  8134. 6:10:12Now that's useful to know. So the N
  8135. 6:10:15operator is a really powerful uh really
  8136. 6:10:18powerful operator.
  8137. 6:10:20Um, same same thing with not. So, we can
  8138. 6:10:23use not in. So, 10 not NA would be false
  8139. 6:10:27because there it is. We know it's a
  8140. 6:10:28member of A. So, that would be false. Of
  8141. 6:10:31course, it's NA. We can see it right
  8142. 6:10:33there. It's a member of that list.
  8143. 6:10:36Um, but if we check the different value
  8144. 6:10:38that's completely not inside of a, 30,
  8145. 6:10:40not NA, that would return true because
  8146. 6:10:4230 is not a member of that list. And by
  8147. 6:10:46the way, this in operator works for all
  8148. 6:10:49kinds of collections. So it would work
  8149. 6:10:50for a tupole, it would work for a set,
  8150. 6:10:53it would work for a dictionary.
  8151. 6:10:56Um, it would work it would work for all
  8152. 6:10:59kinds of collections.
  8153. 6:11:03Yeah, Roberto. Um, it's the fact that
  8154. 6:11:06there are different variables. So we may
  8155. 6:11:09sometimes it makes sense to have
  8156. 6:11:11different variables that are that maybe
  8157. 6:11:14they have the same value but they're
  8158. 6:11:16different variables altogether different
  8159. 6:11:18references
  8160. 6:11:20and the reason that is is because maybe
  8161. 6:11:21we have a copy of that data and then
  8162. 6:11:25maybe we manipulate the this one. Maybe
  8163. 6:11:28this is a copy of it and we manipulate
  8164. 6:11:30this guy and we leave this guy the same
  8165. 6:11:33to check the differences later.
  8166. 6:11:38Yes, references are tied to the variable
  8167. 6:11:41like A is a reference, B is a reference
  8168. 6:11:45even though their data that they're
  8169. 6:11:46pointing to is the same value.
  8170. 6:11:50Maybe we just have a copy of it that and
  8171. 6:11:53we manipulate one of those copies.
  8172. 6:11:59Yeah, that's why
  8173. 6:12:12What is the reason to check for what? If
  8174. 6:12:14they're equal, like as references, A is
  8175. 6:12:17B.
  8176. 6:12:21Honestly, there's not many good reasons.
  8177. 6:12:24Um, maybe if you want to know if
  8178. 6:12:26something is a copy of something else,
  8179. 6:12:27like you want to know that, like let's
  8180. 6:12:29say you're checking later down the code
  8181. 6:12:31and you have an A and a B and you want
  8182. 6:12:32to know if one is a copy of the other,
  8183. 6:12:34you can check and see if they're the
  8184. 6:12:35same reference.
  8185. 6:12:37That's the only reason I could think of
  8186. 6:12:39why you would do that. It's rarely used.
  8187. 6:12:42Rarely used, but it is an operator that
  8188. 6:12:44I wanted to show you in case you do
  8189. 6:12:46stumble across it um somewhere and
  8190. 6:12:49you're reading about Python or something
  8191. 6:12:50and you see the is
  8192. 6:12:54Like that's the only reason I can really
  8193. 6:12:55think of is to check if something is a
  8194. 6:12:57copy of another meaning it's the same
  8195. 6:12:59like maybe it's it's uh the same
  8196. 6:13:02reference
  8197. 6:13:04B equals to A then we could check A is B
  8198. 6:13:06and we know that then they're the same
  8199. 6:13:08reference.
  8200. 6:13:11Yes, the is operator is comparing
  8201. 6:13:12references not exact not the values.
  8202. 6:13:14Yes, that's true.
  8203. 6:13:23All right. How do we feel about this in
  8204. 6:13:25operator? Like the the membership
  8205. 6:13:28operator. Does that make sense? If
  8206. 6:13:30you're checking if an if a value is a
  8207. 6:13:33member of a collection,
  8208. 6:13:36that's going to be highly useful later
  8209. 6:13:38on. Highly useful. This one we will use
  8210. 6:13:41quite a bit. The is we will probably
  8211. 6:13:43rarely ever use, but but this one we
  8212. 6:13:46will definitely use.
  8213. 6:13:57Okay. So, I wanted to quiz you guys. Um,
  8214. 6:14:01what is the main difference between the
  8215. 6:14:02equals equals and the is operator? What
  8216. 6:14:06is the main difference?
  8217. 6:14:08We were we've just been discussing this,
  8218. 6:14:10so hopefully this this one is easy. Been
  8219. 6:14:14discussing it quite a bit.
  8220. 6:14:33Yeah, Roberto, now now you know the
  8221. 6:14:35answer. Perfect. Yeah, it is B. All you
  8222. 6:14:37guys answering B. Perfect. It is B. Good
  8223. 6:14:39job. You guys are right on top of that.
  8224. 6:14:41Good job. So, just wanted to point that
  8225. 6:14:43out. Like equals equals compares the
  8226. 6:14:44values. We're doing a comparison. Um is
  8227. 6:14:48checks if the references are the same,
  8228. 6:14:51which is the variable reference like the
  8229. 6:14:53memory location.
  8230. 6:14:58Yes, it is. Identity is the same as Yep,
  8231. 6:15:00that's what we mean. The references are
  8232. 6:15:02the same.
  8233. 6:15:06All right. So, wanted to do a short demo
  8234. 6:15:08on the operators on comparison, etc.
  8235. 6:15:11Wanted to just show off that demo um so
  8236. 6:15:14you can see and practice with it a
  8237. 6:15:16little bit. Uh so, let's hop over to
  8238. 6:15:18demo six inside of the lesson one. I'm
  8239. 6:15:22going to hop over to that.
  8240. 6:15:30which will be our last thing we will do
  8241. 6:15:32in lesson one and we'll move on to
  8242. 6:15:33lesson two.
  8243. 6:15:36Um, let me pull up demo six here.
  8244. 6:15:44Give me a moment.
  8245. 6:15:49All
  8246. 6:16:01right,
  8247. 6:16:05let me share my screen.
  8248. 6:16:13All right, so demo six. Uh, hopefully
  8249. 6:16:16you guys have access to this. This is
  8250. 6:16:17the last one inside of lesson one. Um,
  8251. 6:16:20now again, this one says try to use VS
  8252. 6:16:23Code. You can if you want. Again, Collab
  8253. 6:16:25works fine. You can use whatever you've
  8254. 6:16:28been using, Jupiter, Collab, VS Code,
  8255. 6:16:30whatever works for you. No big deal on
  8256. 6:16:32which one you use.
  8257. 6:16:37So, no worries on any of this.
  8258. 6:16:48Oh, yeah. So, um, F is So, yeah, this
  8259. 6:16:53that's a good question. What is F? So, F
  8260. 6:16:55tells Python to format. F is short for
  8261. 6:16:59format. It's basically format the the
  8262. 6:17:02string which we're going to print by
  8263. 6:17:04having some placeholders.
  8264. 6:17:06And um
  8265. 6:17:09we have variables called A and B. And
  8266. 6:17:12this fills in the blank of these
  8267. 6:17:14placeholders with whatever the values of
  8268. 6:17:16A and B are. So F just allows us to
  8269. 6:17:19format and fill in the blanks. Does that
  8270. 6:17:22make sense? Like wherever these um
  8271. 6:17:24braces are, we have a variable name
  8272. 6:17:26inside of it A and B. And we are um
  8273. 6:17:29going to fill in the blanks of those A
  8274. 6:17:32and B um by uh you know by just
  8275. 6:17:37replacing them whenever we do the print
  8276. 6:17:39function.
  8277. 6:17:42So there think of it as F is short for
  8278. 6:17:45format.
  8279. 6:17:47and we have a couple placeholders and
  8280. 6:17:49those will be filled in by our variables
  8281. 6:17:51A and B.
  8282. 6:17:55Okay, so this demo uh does a bunch of
  8283. 6:17:59operators. So it's going to do a bunch
  8284. 6:18:01of comparisons where we input one number
  8285. 6:18:04and we turn that into an integer. Input
  8286. 6:18:06another number, turn that into an
  8287. 6:18:08integer, and then do a bunch of
  8288. 6:18:09comparisons. So I'm going to jump over
  8289. 6:18:10to the notebook. I'll do that for us to
  8290. 6:18:12to show that off. But that's all we're
  8291. 6:18:15doing in in really in the beginning of
  8292. 6:18:17this uh demo. Um so let me jump over to
  8293. 6:18:23the notebook and show off that so we can
  8294. 6:18:25see it
  8295. 6:18:30but uh should be straightforward to
  8296. 6:18:31follow because we've done a lot of that
  8297. 6:18:32already.
  8298. 6:18:38Okay. So hopping over to notebook. Again
  8299. 6:18:40feel free to use whatever you want to
  8300. 6:18:41use. You can use collab, you can use um
  8301. 6:18:45uh Jupiter, you can use VS Code,
  8302. 6:18:47whatever you use. Um let's store a
  8303. 6:18:50variable as a and let's make it an
  8304. 6:18:52integer
  8305. 6:18:54and let's do enter your first number.
  8306. 6:18:59So we will do that.
  8307. 6:19:03Let's run that. So let's enter our first
  8308. 6:19:05number. Let's put in 10 or whatever you
  8309. 6:19:07want really, but I'm going to put in 10.
  8310. 6:19:10So that gets stored as a. Now let's do a
  8311. 6:19:13second number and let's do int
  8312. 6:19:17input um enter your second number
  8313. 6:19:24and let's input that.
  8314. 6:19:28So now I'm going to put in a second
  8315. 6:19:30number. Let's do 20.
  8316. 6:19:35So now we have a and b. Now let's do
  8317. 6:19:38some comparisons. So let's do print.
  8318. 6:19:42Um then we can do uh let's do the f
  8319. 6:19:46formatting like they had in there. Now
  8320. 6:19:48what this is going to be is we are going
  8321. 6:19:51to check a
  8322. 6:19:54um greater than or let's do yeah let's
  8323. 6:19:56do greater than b
  8324. 6:20:00um is
  8325. 6:20:02and then let's do uh comma a greater
  8326. 6:20:06than b.
  8327. 6:20:08Let's compare those two numbers. So now
  8328. 6:20:10this is doing the comparison. A greater
  8329. 6:20:13than b is going to compare those two
  8330. 6:20:15integers. What should this return? What
  8331. 6:20:18should a greater than b return? Should
  8332. 6:20:21it be true or should it be false?
  8333. 6:20:27Should be false. Right? So we should we
  8334. 6:20:30should uh display false here.
  8335. 6:20:36And that's what it is. 10 greater than
  8336. 6:20:3820 is false.
  8337. 6:20:41Okay, so that is false.
  8338. 6:20:44Let's do another one.
  8339. 6:20:47Let's do um let's try equals equals. So
  8340. 6:20:50let's say um a
  8341. 6:20:54equals equals to b is now what do we
  8342. 6:20:58think this one's going to be?
  8343. 6:21:03a equals equals to b.
  8344. 6:21:06What should this one be?
  8345. 6:21:10Yes, very good. This one should also be
  8346. 6:21:14false.
  8347. 6:21:18Let's run that.
  8348. 6:21:20And that one will be false. Very good.
  8349. 6:21:26Okay.
  8350. 6:21:28So, I think we get that. Let me give you
  8351. 6:21:29guys a Let me uh show you something
  8352. 6:21:31interesting. Let me uh go a little off
  8353. 6:21:34script from that demo document and let's
  8354. 6:21:36introduce a third number called C. Let's
  8355. 6:21:40do a third integer.
  8356. 6:21:43So enter your third number.
  8357. 6:21:47Let's do a third number.
  8358. 6:21:52Let's enter um another value of 10.
  8359. 6:22:03Okay. So, we have another number of 10.
  8360. 6:22:07Now, what I want to do is let's just do
  8361. 6:22:12uh multiple comparisons and do a logical
  8362. 6:22:16operator between them. So let's do um a
  8363. 6:22:21not equal to b
  8364. 6:22:24and
  8365. 6:22:27a less than c
  8366. 6:22:31or sorry b
  8367. 6:22:33less than c.
  8368. 6:22:37What do we think this is going to
  8369. 6:22:38return? This might be a little
  8370. 6:22:39challenging. What do you think this is
  8371. 6:22:41going to return?
  8372. 6:23:01We should use F. We can. I'm just not
  8373. 6:23:03printing. I'm just going to I'm just
  8374. 6:23:05going to run the cell. I'm I'm kind of
  8375. 6:23:07doing a shortcut and just print. I'm not
  8376. 6:23:08going to print. I'm just going to run
  8377. 6:23:10the cell and it should display what it
  8378. 6:23:11is.
  8379. 6:23:12Do we think it's going to be false?
  8380. 6:23:14Yeah, you guys are right on top of it.
  8381. 6:23:17Should be false. Now, let's break that
  8382. 6:23:20down. Why is that false? So, A not equal
  8383. 6:23:22to B is checking if A is not equal to B,
  8384. 6:23:25which is true.
  8385. 6:23:27A has a value of 10.
  8386. 6:23:32So, A has a value of 10. So, 10 is not
  8387. 6:23:35equal to 20. That makes sense. But 20 is
  8388. 6:23:39not less than 10. So this part is false.
  8389. 6:23:43So let's make a comment.
  8390. 6:23:46So the reason reason this is false
  8391. 6:23:50is because B is not less than C. So and
  8392. 6:23:58returns false.
  8393. 6:24:00Yes. Very good.
  8394. 6:24:04Now,
  8395. 6:24:06what if I take this code
  8396. 6:24:10and do this?
  8397. 6:24:16What's this?
  8398. 6:24:26Perfect. Yeah, you guys are right on top
  8399. 6:24:28of it. This should be true. And it is.
  8400. 6:24:31Now that's because the moment we have at
  8401. 6:24:34least one true which is going to be this
  8402. 6:24:38that makes sense right that is going to
  8403. 6:24:41be true.
  8404. 6:24:54Okay, one more and then we can wrap up
  8405. 6:24:57this demo. So I want to create a list.
  8406. 6:25:00I'm going to call it X and I'm going to
  8407. 6:25:02create a list of three numbers 10 20 30.
  8408. 6:25:08Okay. Now, what do you think uh this
  8409. 6:25:12result is?
  8410. 6:25:14What what should this be?
  8411. 6:25:23It should be true.
  8412. 6:25:25Very good. Should be true.
  8413. 6:25:32Now what should what should this be?
  8414. 6:25:41This is also true. Very good. This
  8415. 6:25:43should this should be true because this
  8416. 6:25:45is not going to be inside of the list.
  8417. 6:25:48So that makes sense. That is not true.
  8418. 6:25:52Now what I want you to notice is nothing
  8419. 6:25:54will change if I change this to a
  8420. 6:25:55tupole.
  8421. 6:25:57Nothing will change. We can still check.
  8422. 6:26:00So we can still check if uh 10 is a
  8423. 6:26:03member of this tupole and we can still
  8424. 6:26:04check if 40 is not a member of this
  8425. 6:26:06tupole. So nothing really changes,
  8426. 6:26:08right? It's still in membership operator
  8427. 6:26:11in checks is it a member of any
  8428. 6:26:14collection?
  8429. 6:26:19Okay,
  8430. 6:26:21very good. Any questions about the these
  8431. 6:26:23examples?
  8432. 6:26:25Any questions? Do we feel comfortable
  8433. 6:26:28with some of these operators? You guys
  8434. 6:26:29were right on top of it. It was very
  8435. 6:26:30impressive. You guys got those right
  8436. 6:26:32away.
  8437. 6:26:35Uh, you change print
  8438. 6:26:39a= b and a is b. And I entered four both
  8439. 6:26:44times
  8440. 6:26:46and I got the same result. True.
  8441. 6:26:53Uh so you had your A and your B. So you
  8442. 6:26:56you had um
  8443. 6:26:59you had A equals to 4
  8444. 6:27:02and B equals to 4 and you checked
  8445. 6:27:06uh you checked this.
  8446. 6:27:16Yes.
  8447. 6:27:20Yeah. So that is a little confusing and
  8448. 6:27:23I can understand why. So the reason this
  8449. 6:27:25ends up being true is because this is a
  8450. 6:27:28scalar data. So for scalar data it's
  8451. 6:27:32going to optimize in the memory to point
  8452. 6:27:35because four the integer four occupies
  8453. 6:27:39the same memory address always. But when
  8454. 6:27:43we create an array, when we create a
  8455. 6:27:45list, that is um a new object.
  8456. 6:27:49So yeah, that one's a little confusing,
  8457. 6:27:51but it's it's only because this is
  8458. 6:27:54scalar data that um Python kind of
  8459. 6:27:57knows, okay, four is the same like
  8460. 6:28:01integer in the me in memory always
  8461. 6:28:05regardless of if we're referencing it
  8462. 6:28:07from this from two different variables.
  8463. 6:28:09That's that's the reason I um yeah I
  8464. 6:28:12forgot to mention that example but it's
  8465. 6:28:15purely so the the reason this is true is
  8466. 6:28:18because
  8467. 6:28:20this is true because of scalar
  8468. 6:28:25data optimization. Essentially it's it's
  8469. 6:28:29not going to waste creating a new object
  8470. 6:28:31when it's just a single integer four. it
  8471. 6:28:33basically occupies the same memory
  8472. 6:28:35address
  8473. 6:28:37uh as as a as a reference.
  8474. 6:28:42But when we build a list like that is a
  8475. 6:28:44different object.
  8476. 6:28:52Okay.
  8477. 6:28:54Very good.
  8478. 6:28:58All right. So,
  8479. 6:29:00um, that will wrap up lesson one. What
  8480. 6:29:05I'm going to do is go into lesson two.
  8481. 6:29:07I'm going to pull up the lesson two
  8482. 6:29:09notes, the slides for lesson two. So, if
  8483. 6:29:11you have those, let's pull those up. Um,
  8484. 6:29:14I do encourage you now, there is a
  8485. 6:29:16guided practice at the end of lesson
  8486. 6:29:17one, and that is for your yourself. I
  8487. 6:29:21would encourage you as as kind of
  8488. 6:29:22homework between now and and the next
  8489. 6:29:24time we meet um to do the guided
  8490. 6:29:27practice for lesson one. Try that out on
  8491. 6:29:30your own. It is um you guys should have
  8492. 6:29:32access to it from your LMS and the
  8493. 6:29:34reference materials. There's a guided
  8494. 6:29:36practice. Try that. Try those out. Okay?
  8495. 6:29:40Try out the guided practice for lesson
  8496. 6:29:41one. It's just it's just some additional
  8497. 6:29:44practice of the things we just covered.
  8498. 6:29:49Okay?
  8499. 6:29:50Let me open up
  8500. 6:29:53um
  8501. 6:29:55lesson two.
  8502. 6:30:04Give me a moment here.
  8503. 6:30:10Okay.
  8504. 6:30:12So, let me share my screen.
  8505. 6:30:24Okay, so now we're going to move on to
  8506. 6:30:27looking more closely at those things
  8507. 6:30:29like lists, tupils, dictionaries, and
  8508. 6:30:31then looking at control flow with
  8509. 6:30:33conditional statements and loops. So
  8510. 6:30:35we're going to get into the fun stuff, I
  8511. 6:30:37would think, um that you guys may may
  8512. 6:30:39have been waiting for.
  8513. 6:30:43Okay, so we've talked about uh we just
  8514. 6:30:45finished talking about lesson one where
  8515. 6:30:46we have you know Python as a really
  8516. 6:30:48important thing to learn and study
  8517. 6:30:50because it's used all over the place
  8518. 6:30:52with data science and a IML. So one of
  8519. 6:30:54the things is we need to continue
  8520. 6:30:56learning about it with things like loops
  8521. 6:30:59things like if else statements to
  8522. 6:31:01control the flow of our programs and
  8523. 6:31:04these basic data structures.
  8524. 6:31:08Let's continue forward. Um so this
  8525. 6:31:11lesson we're going to talk about list
  8526. 6:31:13tupils dictionary sets uh we're going to
  8527. 6:31:15you know talk about the differences how
  8528. 6:31:18we can access data within things like
  8529. 6:31:20list how we can modify them um how we
  8530. 6:31:22can access things from tupils what are
  8531. 6:31:24the differences we'll review all that
  8532. 6:31:26one of the big topics is going to be to
  8533. 6:31:29um look at how we can control the flow
  8534. 6:31:32meaning control the flow is using like
  8535. 6:31:34decision logic like if this is true then
  8536. 6:31:37do this else do this we'll talk about
  8537. 6:31:39those kind of statements. We'll talk
  8538. 6:31:41about iteration. So how we can do loops
  8539. 6:31:44to repeat um statements of code that we
  8540. 6:31:47want to do. Um we'll also talk about
  8541. 6:31:49organizing our code a bit into
  8542. 6:31:51functions. Um which is going to be
  8543. 6:31:53really helpful to for our own
  8544. 6:31:55organization and reuse and
  8545. 6:31:57maintainability.
  8546. 6:32:02Okay.
  8547. 6:32:05So let's jump into it. Let's talk about
  8548. 6:32:08some of those data structures.
  8549. 6:32:11So, we've already talked about that
  8550. 6:32:13these aggregate data structures exist um
  8551. 6:32:17and they allow us to manipulate data
  8552. 6:32:19inside of Python. Lists, tupil, sets,
  8553. 6:32:22dictionaries are the main ones we're
  8554. 6:32:23going to focus on.
  8555. 6:32:26Let's start with lists. And I think
  8556. 6:32:28lists are going to be something we're
  8557. 6:32:30going to use quite a bit of throughout.
  8558. 6:32:33So, they're going to be a really good
  8559. 6:32:34place to start with. They are super
  8560. 6:32:36popular in Python. A lot of people um
  8561. 6:32:38use them to do to work with data. They
  8562. 6:32:41are kind of the most basic um data most
  8563. 6:32:45basic and useful data structure that
  8564. 6:32:46there there is.
  8565. 6:32:50So what is a list? A list is a an
  8566. 6:32:53ordered mutable meaning it can be
  8567. 6:32:56modified data structure that can hold
  8568. 6:32:59elements of different data types. So
  8569. 6:33:00it's a collection of data and there's no
  8570. 6:33:04requirement that all the members of the
  8571. 6:33:06list be the same type. In fact, you can
  8572. 6:33:07have different types. You can have
  8573. 6:33:09integers, you can have floats, you can
  8574. 6:33:10have strings,
  8575. 6:33:12um you can have even more abstract
  8576. 6:33:14objects be members of a list. Um so any
  8577. 6:33:18kind of data can live within a list. But
  8578. 6:33:20the big thing is that it is modifiable.
  8579. 6:33:23It's dynamic. You can change you can add
  8580. 6:33:26things to it. You can remove you can
  8581. 6:33:27change things. Um and it has an inherent
  8582. 6:33:31ordering which is nice. So you can you
  8583. 6:33:33can be reassured that there is some
  8584. 6:33:36inherent position of items and it will
  8585. 6:33:39maintain that order. So we can access
  8586. 6:33:41things based on the order like we can
  8587. 6:33:43access the first can access the last we
  8588. 6:33:45can access anything in between.
  8589. 6:33:48So that's nice. Um so what are some key
  8590. 6:33:52characteristics? So uh lists support
  8591. 6:33:55multiple data types. We talked about
  8592. 6:33:57that. There's no requirement that
  8593. 6:33:58they're all the same. They can have
  8594. 6:33:59multiple. They allow for indexing, which
  8595. 6:34:03we're going to talk about. This means
  8596. 6:34:05that we can access things based on their
  8597. 6:34:08index, which is their position within
  8598. 6:34:10the list. So, we can always access the
  8599. 6:34:12first thing, the last thing, anything in
  8600. 6:34:14between, based on its position. That's
  8601. 6:34:17another word for uh index because lists
  8602. 6:34:21have natural ordering to them which is
  8603. 6:34:24um really really powerful to ensure that
  8604. 6:34:27there's um one you know there's a first
  8605. 6:34:30position a second position third etc. So
  8606. 6:34:34lists are really nice for that.
  8607. 6:34:36Um
  8608. 6:34:38uh they are modifiable which is really
  8609. 6:34:41nice. So we can add things into the
  8610. 6:34:43list. It's very dynamic. So once we
  8611. 6:34:44create a list, we can throughout our
  8612. 6:34:46program, we can add data to it, we can
  8613. 6:34:48remove it, we can change. Um, lists also
  8614. 6:34:51allow duplicates, which may be
  8615. 6:34:53desirable. Like maybe we add something
  8616. 6:34:55into our list that already exists.
  8617. 6:34:57That's okay. List allow duplicates. This
  8618. 6:35:00is going to be different than sets. Sets
  8619. 6:35:03do not allow duplicates. Sets are just a
  8620. 6:35:05bucket of unique things. So if we added
  8621. 6:35:08a duplicate into a set, it would reject
  8622. 6:35:10it. we wouldn't have any errors, but it
  8623. 6:35:13just wouldn't um show up as a copy. We
  8624. 6:35:16would just have a set is only going to
  8625. 6:35:18maintain one copy of an item. It it only
  8626. 6:35:21allows unique items. Lists allow you can
  8627. 6:35:23have as many copies of data as you want
  8628. 6:35:25inside of a list. Um so it does allow
  8629. 6:35:28duplicates.
  8630. 6:35:30Now, we already saw in terms of syntax,
  8631. 6:35:33lists are um defined by brackets. So
  8632. 6:35:37when you see those brackets um it
  8633. 6:35:39defines a list and its items are
  8634. 6:35:41separated by uh commas.
  8635. 6:35:45So
  8636. 6:35:47we can have a list that looks like this.
  8637. 6:35:53Yeah, I was going to explain slice. We
  8638. 6:35:54have a a couple slides about slicing
  8639. 6:35:56coming up, but slicing just means that
  8640. 6:35:59we can slicing means that we can grab a
  8641. 6:36:02section of elements at a time from a
  8642. 6:36:04list. So for instance we can uh actually
  8643. 6:36:08let me use this example down here. A
  8644. 6:36:10slice would mean we can grab like these
  8645. 6:36:12first three slice of the of the list or
  8646. 6:36:16we can grab the last five elements or
  8647. 6:36:19whatever like this is a slice. It's just
  8648. 6:36:22a a subset of the list that we can grab
  8649. 6:36:25we can access.
  8650. 6:36:27So slicing just means taking a subset,
  8651. 6:36:31taking a smaller section of the list and
  8652. 6:36:34we can grab all those elements at a
  8653. 6:36:36time. And what that's called is within a
  8654. 6:36:38slice.
  8655. 6:36:42Yeah, like a slice of pizza. We're
  8656. 6:36:44taking the whole thing and we're taking
  8657. 6:36:45a small section of it.
  8658. 6:36:50So that's and actually that's going to
  8659. 6:36:53be possible because the list has a
  8660. 6:36:56natural order to it.
  8661. 6:37:01Does the data have to be sequential? No,
  8662. 6:37:03it doesn't have to be. In fact, you it
  8663. 6:37:05can be completely different types.
  8664. 6:37:08Does that make sense? Like so look at
  8665. 6:37:10this example down here. Like we have 10
  8666. 6:37:122 5 hello. That's a valid list. You can
  8667. 6:37:16have different types of data in there
  8668. 6:37:17which isn't sequential at all
  8669. 6:37:29in a slice. No, it doesn't have to be.
  8670. 6:37:32So you can have like you can have a
  8671. 6:37:34slice that picks every third element
  8672. 6:37:38uh every other element. Um yeah, it
  8673. 6:37:42doesn't have to be sequential. No, it
  8674. 6:37:45can be customizable.
  8675. 6:37:49What's also nice is you can slice from
  8676. 6:37:51the beginning or you can slice from the
  8677. 6:37:52end as well. So you can you can go from
  8678. 6:37:54the end and slice backwards. Um or you
  8679. 6:37:58can go from the beginning and slice
  8680. 6:37:59forwards. So you can grab like every
  8681. 6:38:01other element from the beginning. You
  8682. 6:38:03can grab every other element from the
  8683. 6:38:04back and work your way forward and stop
  8684. 6:38:06at a certain point.
  8685. 6:38:10Slicing is very nice. Yeah. So, I'm
  8686. 6:38:12going to show us how to do that.
  8687. 6:38:16Can you slice in the middle? Yeah, you
  8688. 6:38:17can slice anywhere you want.
  8689. 6:38:21Can you slice a pizza in the middle?
  8690. 6:38:23Sure. Would you do that? Maybe not. But
  8691. 6:38:26yeah, you can slice anywhere in the
  8692. 6:38:28list. You can slice.
  8693. 6:38:37Does slicing change the original list?
  8694. 6:38:39No, it's just selection of a subset.
  8695. 6:38:42No, it just it just extracts elements.
  8696. 6:38:45It doesn't like permanently change it in
  8697. 6:38:47any way. It just gives you a view. Think
  8698. 6:38:51of it as like giving you a view of that
  8699. 6:38:53subset.
  8700. 6:39:00Okay, so you may be wondering when would
  8701. 6:39:03I ever use list? So normally you use
  8702. 6:39:06lists whenever you want an ordered
  8703. 6:39:09collection that is dynamic meaning it
  8704. 6:39:12may you may want to add things from it.
  8705. 6:39:15Um we may want to modify things from it
  8706. 6:39:19frequently. Um but we want something
  8707. 6:39:22that is dynamic and has an ordering to
  8708. 6:39:24it. Lists are perfect for that reason.
  8709. 6:39:26So they can contain data that we can add
  8710. 6:39:29to remove from change. So lists are very
  8711. 6:39:32versatile. I think most use cases with
  8712. 6:39:34manipulating
  8713. 6:39:36um a collection of items would fall
  8714. 6:39:39under a list. A list would be a very
  8715. 6:39:41good choice. Um
  8716. 6:39:46we can slice elements and use them.
  8717. 6:39:48Yeah. Yeah. You can slice you can store
  8718. 6:39:51a slice inside of a variable and use use
  8719. 6:39:54the use that resulting variable. Yeah,
  8720. 6:39:57for sure.
  8721. 6:40:11order information like
  8722. 6:40:16Yeah. Yeah, that's a good example. Yeah,
  8723. 6:40:18order like the collection of orders
  8724. 6:40:22would probably be in a list because it
  8725. 6:40:23can change. Um, but the prices would be
  8726. 6:40:26probably static. So, that would be more
  8727. 6:40:29uh suitable for a tupole. Yeah, a tupole
  8728. 6:40:32or maybe a dictionary. A dictionary is
  8729. 6:40:34probably better because you can have
  8730. 6:40:35like a product ID or a name that maps to
  8731. 6:40:38a price. Probably a dictionary would be
  8732. 6:40:40more appropriate for a price list. But
  8733. 6:40:43but yeah, tupole maybe makes sense too.
  8734. 6:40:51By the way, look at this list. Do you
  8735. 6:40:52guys see how it has different types of
  8736. 6:40:55data in it, right? It has so like the
  8737. 6:40:57first element is an integer, the next is
  8738. 6:40:59a string, the next is a float, the last
  8739. 6:41:02is a boolean. That's totally valid in
  8740. 6:41:04Python, which is kind of unique to
  8741. 6:41:07Python. Like a lot of languages don't
  8742. 6:41:09support that. A mixtyped array.
  8743. 6:41:13Don't really they don't really have that
  8744. 6:41:15notion of that.
  8745. 6:41:26All right. So, what I wanted to get to
  8746. 6:41:28was the positions. So, this is going to
  8747. 6:41:31be this is really important to pay
  8748. 6:41:33attention to because this is going to be
  8749. 6:41:35something that we will be using
  8750. 6:41:38throughout the program is how to access
  8751. 6:41:41data by its position.
  8752. 6:41:44Okay, which so there's another word for
  8753. 6:41:46that. The position um sometimes you will
  8754. 6:41:48hear called the index. So the index in
  8755. 6:41:51the list of where these where data
  8756. 6:41:54members live is their order like their
  8757. 6:41:56position amongst amongst the list.
  8758. 6:42:00So what's special about Python is that
  8759. 6:42:04it has the first position
  8760. 6:42:08is index zero which trips people up all
  8761. 6:42:13the time. The very classic trip up of
  8762. 6:42:17Python is that the first element of a
  8763. 6:42:20list is at position zero. The next
  8764. 6:42:23element is at position one. The next
  8765. 6:42:26element is at position two. On and on
  8766. 6:42:28and on. The last element is at position
  8767. 6:42:32n minus one where n is the size of the
  8768. 6:42:38um size of list. So, however many
  8769. 6:42:41elements we have in our list.
  8770. 6:42:43Um,
  8771. 6:42:45so
  8772. 6:42:47if you want to access the first element,
  8773. 6:42:49you would be looking for the element
  8774. 6:42:51that's at position zero. If you want to
  8775. 6:42:54access the second element, that's this
  8776. 6:42:56this name Bob, that is at position
  8777. 6:42:59number one or index one. So that's
  8778. 6:43:02something that trips up people is that
  8779. 6:43:04it actually starts
  8780. 6:43:08starts at zero
  8781. 6:43:11which is really important to understand
  8782. 6:43:12is that the positions start at zero.
  8783. 6:43:18Why is that? Um that's a good question.
  8784. 6:43:21I mean
  8785. 6:43:23different so so different languages
  8786. 6:43:26treat that differently. Um,
  8787. 6:43:29it's more historical reasons that it was
  8788. 6:43:32created that way. Uh, but I think it's
  8789. 6:43:36based on your like
  8790. 6:43:40think of it as like how far into the
  8791. 6:43:42list you are. So, if you're in the
  8792. 6:43:44beginning, that means you're basically
  8793. 6:43:46at at um position zero because you
  8794. 6:43:49haven't made any progress like
  8795. 6:43:50traversing the list. I think that was
  8796. 6:43:53the intuition.
  8797. 6:43:57Oh, yeah. which is kind of what Tim says
  8798. 6:43:58there is like yeah if you're at the very
  8799. 6:44:01beginning of the list you're at position
  8800. 6:44:02zero because you haven't made any you
  8801. 6:44:04haven't made any forward progress in
  8802. 6:44:06traversing so it's like you're at step
  8803. 6:44:08zero you're at the beginning
  8804. 6:44:17okay but okay so aside from why it was
  8805. 6:44:20that way does it make sense that that
  8806. 6:44:22the first like what I'm saying is the
  8807. 6:44:26position zero is the first element,
  8808. 6:44:27position one is the second element,
  8809. 6:44:29position two is the third element, and
  8810. 6:44:31on and on and on.
  8811. 6:44:33That's that's how it works in Python.
  8812. 6:44:37Okay.
  8813. 6:44:39Now,
  8814. 6:44:40what I'm telling you is the positions if
  8815. 6:44:43you were to view it going left to right.
  8816. 6:44:45So, in other words, in the forward
  8817. 6:44:47direction, we start from zero and go up
  8818. 6:44:49to the to the n minus one in terms of
  8819. 6:44:52the position. Now, what's really
  8820. 6:44:55convenient is the positions can also be
  8821. 6:44:58indexed from back to front, meaning they
  8822. 6:45:02can also be indexed going this way.
  8823. 6:45:05And what's really nice is the very last
  8824. 6:45:08element starts at position minus one.
  8825. 6:45:14The very last element on the right
  8826. 6:45:16starts at minus1 and then goes all the
  8827. 6:45:19way up to minus n.
  8828. 6:45:22Okay, now that's really convenient
  8829. 6:45:24because if we want to access the very
  8830. 6:45:26last element, I don't need to know how
  8831. 6:45:28big the list is. I just need to access
  8832. 6:45:30the position minus one. That guarantees
  8833. 6:45:34to access the very last element. The
  8834. 6:45:38minus one position is the very last
  8835. 6:45:40element.
  8836. 6:45:42And then it it so this is the very last
  8837. 6:45:45element. This is the second to last,
  8838. 6:45:48third to last, fourth to last, on and on
  8839. 6:45:51and on. And then by the time you get to
  8840. 6:45:53the front, it is minus n, which is like
  8841. 6:45:57the
  8842. 6:45:59uh number of elements in the list minus.
  8843. 6:46:03So in this case, minus 6 is the very
  8844. 6:46:05beginning.
  8845. 6:46:06Now, why why in the world do we care
  8846. 6:46:08about negative position? It's for that
  8847. 6:46:10exact reason.
  8848. 6:46:11we can um we can think about the
  8849. 6:46:14positions from the end of the list going
  8850. 6:46:17backward which is very powerful like if
  8851. 6:46:19I want to access things from the very
  8852. 6:46:21end I don't need to know exactly how
  8853. 6:46:23many there are I just need to know minus
  8854. 6:46:25one is the end minus two is the second
  8855. 6:46:27to last minus3 is the third from last um
  8856. 6:46:31which is pretty convenient
  8857. 6:46:36okay
  8858. 6:46:38so does that make sense on the negative
  8859. 6:46:40index index. It's the negative index is
  8860. 6:46:43going from right to left. It's from the
  8861. 6:46:46back to the front. Minus one is the last
  8862. 6:46:49element.
  8863. 6:46:51And then you go second to last, third to
  8864. 6:46:53last, right? And that is minus 2, -3,
  8865. 6:46:57-4.
  8866. 6:47:06So from right to left. Yeah. Right.
  8867. 6:47:08Right to left. minus 6 is actually the
  8868. 6:47:10first element
  8869. 6:47:12because it's six back from the end which
  8870. 6:47:14is the f which is the front. There's
  8871. 6:47:16only six elements.
  8872. 6:47:22Yeah.
  8873. 6:47:28So are there any question about this is
  8874. 6:47:30really important to understand really
  8875. 6:47:33important to understand because this is
  8876. 6:47:35how we are going to slice and access
  8877. 6:47:37data is based on these positions.
  8878. 6:47:45Uh is there syntax to get the number of
  8879. 6:47:47elements in a list? Yes, it's the length
  8880. 6:47:49function which is len. So length of list
  8881. 6:47:55would give you uh the number of elements
  8882. 6:47:58which in this case is six. So ln the
  8883. 6:48:01length function gives you how many
  8884. 6:48:03elements are in the list.
  8885. 6:48:06len or length. So this guy this
  8886. 6:48:11function.
  8887. 6:48:17Uh when are you counting backwards?
  8888. 6:48:19You're usually counting backwards when
  8889. 6:48:22you want to know what's at the end of
  8890. 6:48:24the list. So when you only care what has
  8891. 6:48:26been added at the end and you want to
  8892. 6:48:29maybe you want to get the last five
  8893. 6:48:31elements.
  8894. 6:48:32So you slice backwards from the from the
  8895. 6:48:35end.
  8896. 6:48:37That's that may be useful like maybe you
  8897. 6:48:39want to know like it imagine a list is
  8898. 6:48:42holding in your orders
  8899. 6:48:44and so you want the last five orders. So
  8900. 6:48:47you can just go from the back and go
  8901. 6:48:49towards the front. Min -1, -2, -3, -4,
  8902. 6:48:52-5
  8903. 6:48:55would be the last five orders. There's
  8904. 6:48:58going to be many scenarios where we want
  8905. 6:48:59to count from the end.
  8906. 6:49:02uh when we're manipulating data
  8907. 6:49:05later on when we're when we're working
  8908. 6:49:07with um bigger sets of data um in
  8909. 6:49:10something like an like a matrix, it's
  8910. 6:49:13going to be useful to grab like the last
  8911. 6:49:15five rows, last 10 rows.
  8912. 6:49:19So going from the end makes more sense.
  8913. 6:49:21So you imagine if we have like a larger
  8914. 6:49:23matrix of data um maybe we want to slice
  8915. 6:49:27out these last 10 rows in which case we
  8916. 6:49:30want to count from the minus like the
  8917. 6:49:32last row is minus one and we want to go
  8918. 6:49:35back towards the front.
  8919. 6:49:43Can you sort? Yes, you can sort. Uh I
  8920. 6:49:46I'll have an example of that in a
  8921. 6:49:47second. Yeah, you can sort.
  8922. 6:49:50There's a built-in sorting function in
  8923. 6:49:53Python that that will allow you to sort
  8924. 6:49:54data in a list. Yes,
  8925. 6:49:57there's a built-in function for that.
  8926. 6:50:07Okay. What I wanted to do is show you
  8927. 6:50:11guys an example of accessing elements
  8928. 6:50:13from the list. So assuming we know the
  8929. 6:50:18position which is that index all we have
  8930. 6:50:20to do is use brackets to access items of
  8931. 6:50:24a list. So imagine we have this list
  8932. 6:50:26called fruits which has some strings in
  8933. 6:50:29it apple banana cherry mango and we want
  8934. 6:50:33to access the first element. So that is
  8935. 6:50:36just this code here fruits bracket zero.
  8936. 6:50:40So the bracket tells the interpreter,
  8937. 6:50:43hey, I want to access something within
  8938. 6:50:45this list. And then all we have to do is
  8939. 6:50:47give it the position that we want to
  8940. 6:50:49access. So this is position zero, which
  8941. 6:50:53is going to be the first element. This
  8942. 6:50:54would retrieve the first element. Now,
  8943. 6:50:58it's not removing it. It's just
  8944. 6:51:02accessing it so we can view what that
  8945. 6:51:05is. So, it's actually going to give us a
  8946. 6:51:07copy of what that value is under the
  8947. 6:51:10hood. Basically, a copy of it. It's not
  8948. 6:51:12going to permanently. It's not going to
  8949. 6:51:13delete it. It's not going to remove it.
  8950. 6:51:16It's going to give us a
  8951. 6:51:19copy of what is at position zero. In
  8952. 6:51:22this case, apple.
  8953. 6:51:24Now if we put in a position two in that
  8954. 6:51:27bracket that should be remember the
  8955. 6:51:29indexing is 0 1 2 3
  8956. 6:51:34because there's four items. So the last
  8957. 6:51:37element is n minus one which is three.
  8958. 6:51:40So the item that's at position two is
  8959. 6:51:43going to be cherry. So this should
  8960. 6:51:45return the string cherry.
  8961. 6:51:48Okay. So we use we always use this kind
  8962. 6:51:52of syntax
  8963. 6:51:55position
  8964. 6:51:57to access the element that is at that
  8965. 6:52:00position.
  8966. 6:52:03Okay, pretty simple. And what's really
  8967. 6:52:06nice by the way about this is this the
  8968. 6:52:09same exact thing works for tupils
  8969. 6:52:12because tupils also are ordered. So
  8970. 6:52:15we're going to see that when we get into
  8971. 6:52:16tupils. But the same exact things works
  8972. 6:52:19where a tupole has a first item, a
  8973. 6:52:21second item, a third which are index
  8974. 6:52:22zero, 1, two, three. And we can access
  8975. 6:52:25things in the exact same way. So we can
  8976. 6:52:27access a tupole by its position as well.
  8977. 6:52:31Same exact thing will happen position.
  8978. 6:52:36So we can get the first item of a tupole
  8979. 6:52:38with with uh by doing um zero and we can
  8980. 6:52:42get the the last by doing minus one.
  8981. 6:52:45and on and on.
  8982. 6:52:55Any questions about the accessing?
  8983. 6:53:01Can we get position based on value? Yes.
  8984. 6:53:04So that is a special function called
  8985. 6:53:06index.
  8986. 6:53:08So if you do like list dot dot index
  8987. 6:53:16and then you pass in a value like apple,
  8988. 6:53:21this would return to you the index of
  8989. 6:53:24the first occurrence. Not every
  8990. 6:53:26occurrence, but just the first. So like
  8991. 6:53:28because we could have multiple copies of
  8992. 6:53:30Apple in our list, but the first time we
  8993. 6:53:33stumble upon Apple, this would return
  8994. 6:53:35this would return uh zero because apple
  8995. 6:53:39occurs at index zero.
  8996. 6:53:43So index function
  8997. 6:53:46uh would return um the index
  8998. 6:53:50which is which is the same word as
  8999. 6:53:52position.
  9000. 6:53:58Okay.
  9001. 6:54:01All right. Any other questions? We can
  9002. 6:54:03uh I think what we'll do is we can take
  9003. 6:54:05if unless there's any other questions we
  9004. 6:54:07can take a short break and then come
  9005. 6:54:08back and continue.
  9006. 6:54:11Uh, I want to talk about slicing.
  9007. 6:54:22No, it would return none. It would
  9008. 6:54:25return null. Basically none
  9009. 6:54:36as opening closing braces are square on
  9010. 6:54:40it if it's a list. No, it's actually the
  9011. 6:54:42same as a tupole. Uh in terms of
  9012. 6:54:44accessing it's the same
  9013. 6:54:47uh
  9014. 6:54:51it's it's in terms of accessing it's the
  9015. 6:54:52same. Um, for creating a list, yes, it
  9016. 6:54:56is square brackets. This creates a list.
  9017. 6:54:59The brackets always create a list. For
  9018. 6:55:02accessing elements, this is the same as
  9019. 6:55:04a tupole. The the square brackets for
  9020. 6:55:07accessing.
  9021. 6:55:09Actually, in most things in Python, it's
  9022. 6:55:12the same where we uh access data using
  9023. 6:55:15the square brackets. A dictionary, same
  9024. 6:55:18thing. We access keys by using the
  9025. 6:55:21square brackets.
  9026. 6:55:23So square brackets in Python is is
  9027. 6:55:25basically like an access operator.
  9028. 6:55:31Uh no, so list.index doesn't only work
  9029. 6:55:34for the first item. What I'm saying is
  9030. 6:55:36you can have lists that have multiple
  9031. 6:55:38I'm saying it returns to you the first
  9032. 6:55:41occurrence,
  9033. 6:55:42the index of the first occurrence
  9034. 6:55:44because we could have a copy of Apple
  9035. 6:55:46later on in the list, right? So there's
  9036. 6:55:49nothing that stops us from having a
  9037. 6:55:51duplicate. So, we could have another
  9038. 6:55:53apple down here. Like, let's say we had
  9039. 6:55:55another apple at the end of the list. If
  9040. 6:55:57I did list.index apple, it's only going
  9041. 6:56:00to return to me this first one, even
  9042. 6:56:03though there's another copy of it in the
  9043. 6:56:04list.
  9044. 6:56:07So, it only returns to you the first
  9045. 6:56:09occurrence.
  9046. 6:56:12But, you know, there's nothing stopping
  9047. 6:56:13me from doing list.index of banana or
  9048. 6:56:15cherry or whatever.
  9049. 6:56:18Are there data types by this for all
  9050. 6:56:20Unicode characters? Um, I think there's
  9051. 6:56:23Yeah, I think there's special strings
  9052. 6:56:25you can make that do Unicode,
  9053. 6:56:28but I'm honestly not 100% sure.
  9054. 6:56:32I would research that. I don't really
  9055. 6:56:35know. I think you can do that with
  9056. 6:56:37strings,
  9057. 6:56:39special strings.
  9058. 6:56:42Yeah, I don't think there's anything
  9059. 6:56:43like I don't think there's anything
  9060. 6:56:44inherently special about it that Python
  9061. 6:56:48can't handle. It's just you sometimes
  9062. 6:56:51have to like escape characters
  9063. 6:56:54uh
  9064. 6:56:56to distinguish them. But
  9065. 6:57:00yeah.
  9066. 6:57:02All right. Going back to our example,
  9067. 6:57:07um we had the uh going back to that
  9068. 6:57:12fruits list.
  9069. 6:57:14We can access things from the end. So
  9070. 6:57:16here's an example of us using the
  9071. 6:57:18negative index, right? So minus one
  9072. 6:57:21grabs the element that's at the last uh
  9073. 6:57:24that's the last member of the list. So
  9074. 6:57:26mango minus3
  9075. 6:57:29um would be the uh third from last,
  9076. 6:57:32which is going to be banana.
  9077. 6:57:35So negative index really handy to access
  9078. 6:57:39from the end of the list going backward.
  9079. 6:57:42Um pretty useful there.
  9080. 6:57:45There's an example of it.
  9081. 6:57:51All right.
  9082. 6:57:56All right. Let's talk about slicing.
  9083. 6:57:58Okay. Let's talk about slicing. So,
  9084. 6:58:00slicing allows us to extract a subset of
  9085. 6:58:04items from a list using a specific range
  9086. 6:58:08of indices. And so the syntax to do
  9087. 6:58:11slicing
  9088. 6:58:13is going to be using our brackets again,
  9089. 6:58:16but um we use a colon to specify where
  9090. 6:58:22we are starting and stopping our
  9091. 6:58:24positions. And not only that, but we
  9092. 6:58:27also use uh a colon to signal how many
  9093. 6:58:30we want to step by. So do we want to do
  9094. 6:58:33every other in which case we would step
  9095. 6:58:35by two. Do you want to do every third
  9096. 6:58:37element which would step by three? So
  9097. 6:58:41most slicing is going to follow um this
  9098. 6:58:44sort of syntax where we do a list and
  9099. 6:58:48then we do like a start index and then
  9100. 6:58:52we do colon
  9101. 6:58:54and then we do stop index
  9102. 6:58:57um and then we do colon uh step. Now
  9103. 6:59:02what you will see is that the the step
  9104. 6:59:06defaults to a step size of one meaning
  9105. 6:59:09we grab every element in between
  9106. 6:59:11starting and stop. Um so step size of
  9107. 6:59:15one is the default. So we actually
  9108. 6:59:18typically will not include the step size
  9109. 6:59:20unless we specifically want to get every
  9110. 6:59:23other which would be a step size of two
  9111. 6:59:25or every third or every fourth or every
  9112. 6:59:27fifth. Um so usually we leave off this
  9113. 6:59:31step size and we just we just have a
  9114. 6:59:33starting and a stop as part of our
  9115. 6:59:35slice. Um and what what that does is it
  9116. 6:59:39tells the interpreter to access
  9117. 6:59:43everything between this start and stop.
  9118. 6:59:47So um with one catch which is a very
  9119. 6:59:51important catch that trips everyone up
  9120. 6:59:54which is that Python um is very annoying
  9121. 6:59:57and that it uh when you do slicing it
  9122. 7:00:01allows you to include the starting
  9123. 7:00:04index. So uh if we start somewhere we
  9124. 7:00:09will guarantee that the slice will
  9125. 7:00:10include that but it will not include the
  9126. 7:00:14stopping index. it will do everything
  9127. 7:00:17between there up to the stop index but
  9128. 7:00:20not actually including what's what's at
  9129. 7:00:23the stop position.
  9130. 7:00:26So for example,
  9131. 7:00:28this slice that you see um on the screen
  9132. 7:00:33is a slice that would be starting at
  9133. 7:00:36position two
  9134. 7:00:39because remember um let me draw this
  9135. 7:00:42out. This is position zero. This is
  9136. 7:00:44position one, position two, position
  9137. 7:00:47three, position four, five, and six. And
  9138. 7:00:51so this is a slice that would start at
  9139. 7:00:54position two
  9140. 7:00:56is our start.
  9141. 7:00:59And meaning we're guaranteed to get 34
  9142. 7:01:02because we're starting there. But the
  9143. 7:01:04stop for this slice would have to be
  9144. 7:01:08here. This is our stop because we are
  9145. 7:01:12going to include everything in between
  9146. 7:01:16there. We're going to include all of
  9147. 7:01:19this this uh slice as part of our uh
  9148. 7:01:23what we can access. Um so that would be
  9149. 7:01:27positions 2, three, and four. So this
  9150. 7:01:29slice would be um basically like this
  9151. 7:01:33list. Um it would start at two and go to
  9152. 7:01:37five.
  9153. 7:01:39And then it technically would be a step
  9154. 7:01:41size of one, but remember we don't
  9155. 7:01:43really need to include that. So this
  9156. 7:01:45would really be list um two to five.
  9157. 7:01:53That does feel annoying. Yes, it's it's
  9158. 7:01:55because they you have to know where to
  9159. 7:01:58start and stop. And so they cho Python
  9160. 7:02:01chooses to to be um not inclusive of the
  9161. 7:02:06stopping index, but it it includes the
  9162. 7:02:09starting index. Um and and that's just a
  9163. 7:02:14choice. That's a design choice of
  9164. 7:02:15Python.
  9165. 7:02:19Okay. So the colon gives us the slice.
  9166. 7:02:22Um, so if you see if you see uh uh if
  9167. 7:02:29you see
  9168. 7:02:31a colon inside of a brackets, that
  9169. 7:02:35signals you're grabbing a collection of
  9170. 7:02:37items. So we So this this actually
  9171. 7:02:40returns a smaller list. This returns a
  9172. 7:02:43list of 34, 20, and 80. So it's a it's a
  9173. 7:02:49slice meaning we get multiple items
  9174. 7:02:51rather than just a single item from from
  9175. 7:02:53the selection.
  9176. 7:02:57No, 54 would not make the cut. 67
  9177. 7:02:59doesn't make the cut either because
  9178. 7:03:01remember we don't include the stopping
  9179. 7:03:03index.
  9180. 7:03:08Can you do two to four plus one? Yes,
  9181. 7:03:11you can do arithmetic in there and
  9182. 7:03:12Python will evaluate the arithmetic
  9183. 7:03:14first. So it will do 4 + 1 first and
  9184. 7:03:17determine that is five.
  9185. 7:03:20Yes, you can do that.
  9186. 7:03:31Okay. So before
  9187. 7:03:34before I move on, uh does the slicing
  9188. 7:03:39idea make sense? It is grabbing a
  9189. 7:03:42collection of items from the larger list
  9190. 7:03:45and we set it up with this syntax.
  9191. 7:03:52Uh what do you think? What do you think
  9192. 7:03:530 to six would return? What would be
  9193. 7:03:55your guess?
  9194. 7:04:09What's included in the slice though?
  9195. 7:04:11It's not just 67. If you go 0 to six,
  9196. 7:04:16how many? Like, you should be getting
  9197. 7:04:17more than that. Yeah, you should get
  9198. 7:04:20everything, right? Exactly. You should
  9199. 7:04:22get all of those numbers up to
  9200. 7:04:26uh 54. You would not include 54. So you
  9201. 7:04:29would get 76, you get 12, you get 34,
  9202. 7:04:3220, 80, and 67.
  9203. 7:04:41Yes, that's true.
  9204. 7:04:49Yeah. Yeah. So the the step relevance is
  9205. 7:04:52that we can so from our slice we can
  9206. 7:04:57choose um our step size of how many
  9207. 7:05:00element like how much we want to skip
  9208. 7:05:04positions within that slice. So step of
  9209. 7:05:06one which is a default means we get
  9210. 7:05:09every position we go we increment by one
  9211. 7:05:13one position to the next to the next to
  9212. 7:05:15the next. A step of two
  9213. 7:05:19uh
  9214. 7:05:21a step of two would be that uh I'm going
  9215. 7:05:25to grab every other element from the
  9216. 7:05:27start. So a step of two like if I sliced
  9217. 7:05:31this
  9218. 7:05:33and changed this to a step size of two,
  9219. 7:05:36then that would only grab this and this
  9220. 7:05:39because that would step over. It would
  9221. 7:05:41take two steps to get to the next
  9222. 7:05:43element of the slice.
  9223. 7:05:47Whereas a step size of one is going to
  9224. 7:05:49grab everything
  9225. 7:05:53because it's going to go one index to
  9226. 7:05:54the next. So think of the step size as
  9227. 7:05:57how many positions are we incrementing?
  9228. 7:06:05Yes. 2 to 7 would include 54. Yes,
  9229. 7:06:08that's right.
  9230. 7:06:102 to 7 would include 54.
  9231. 7:06:20Okay.
  9232. 7:06:23All right. Let's see some more examples.
  9233. 7:06:27So if we have a list like this and we
  9234. 7:06:30slice it from 1 to 4, the output is
  9235. 7:06:34going to start at index one
  9236. 7:06:38and go all the way up to index 4 but not
  9237. 7:06:42include four. So index 4 is this guy and
  9238. 7:06:46it's not going to include that. So it's
  9239. 7:06:48going to be these three elements here
  9240. 7:06:52would be our slice. Yep. 20 30 40. Very
  9241. 7:06:55good. That's what it would be.
  9242. 7:06:58So that's pretty useful
  9243. 7:07:01to be able to slice. Let me show you
  9244. 7:07:03another example.
  9245. 7:07:08Oops, don't have another. Let me go
  9246. 7:07:10back.
  9247. 7:07:13Let me show you another example with
  9248. 7:07:15this. Um, so what we can do is we can
  9249. 7:07:20actually slice backwards as well. So,
  9250. 7:07:25what do you Let me ask you guys this.
  9251. 7:07:26What do you think this slice would be?
  9252. 7:07:38Actually, let me erase this. Let me do
  9253. 7:07:40minus
  9254. 7:07:42or
  9255. 7:07:47what do you think this would return?
  9256. 7:07:51We can use negative index index uh
  9257. 7:07:53indices in our slices.
  9258. 7:07:56What do you think that would be?
  9259. 7:08:04Very good. 30 40 50. So it's going to
  9260. 7:08:07go. So remember -4 is the fourth from
  9261. 7:08:11the last. So it's going to be this is
  9262. 7:08:13minus4
  9263. 7:08:15and then this is minus one. So we're not
  9264. 7:08:19going to include minus one. So we should
  9265. 7:08:21be doing this slice here.
  9266. 7:08:24That should be the slice. 30 40 50 would
  9267. 7:08:26be that.
  9268. 7:08:29Okay. One other a couple other examples
  9269. 7:08:32I wanted to give you is that you can act
  9270. 7:08:36in certain special cases you can
  9271. 7:08:38actually leave off the starting and
  9272. 7:08:42stop. And what that would signal the
  9273. 7:08:43interpreter is that you want to go all
  9274. 7:08:45the way to the end or start all the way
  9275. 7:08:47from the beginning. So if you do
  9276. 7:08:50something like this,
  9277. 7:08:52let me show you an example. If you do
  9278. 7:08:54something like this and you do not
  9279. 7:08:57include, you leave the the start blank.
  9280. 7:09:01You leave the start blank and you go all
  9281. 7:09:03the way up to minus one. What that would
  9282. 7:09:05signal to the interpreter is by default
  9283. 7:09:09um start at the beginning. So start at
  9284. 7:09:11zero. Essentially start at zero. If you
  9285. 7:09:14leave off a slice
  9286. 7:09:17uh as your start that the interpreter
  9287. 7:09:19assumes you want to start at the
  9288. 7:09:21beginning.
  9289. 7:09:23So what do you think this slice would be
  9290. 7:09:25knowing that?
  9291. 7:09:32Yes, exactly. You guys you guys are
  9292. 7:09:33right on top of it. 10 to 50. Perfect.
  9293. 7:09:35So it's going to be everything but the
  9294. 7:09:37last
  9295. 7:09:39everything but the last would be
  9296. 7:09:40included in that slice.
  9297. 7:09:43Perfect. And so the other thing is we
  9298. 7:09:47can leave off the end which would signal
  9299. 7:09:49that we want to go all the way to the
  9300. 7:09:50end. So what do you guys think this is
  9301. 7:09:53going to be?
  9302. 7:10:00What would that be?
  9303. 7:10:10Yes. So this is So this is actually
  9304. 7:10:12going to include the end. So I know
  9305. 7:10:15that's a little counterintuitive, but
  9306. 7:10:16it's actually this this guarantees we
  9307. 7:10:18include the end. So if it's blank, it's
  9308. 7:10:22going to go all the way to the end,
  9309. 7:10:23including the end. I know that's that's
  9310. 7:10:26annoying. I don't blame you for thinking
  9311. 7:10:28it should be 50,
  9312. 7:10:31but it basically goes to n. Basically
  9313. 7:10:34goes to n, which would mean that we
  9314. 7:10:36remember the index is n minus one is the
  9315. 7:10:38last index.
  9316. 7:10:40So, so if it's blank, this means we
  9317. 7:10:43should go we should end at n
  9318. 7:10:47which is the length of the list meaning
  9319. 7:10:49that um the last element is at n minus
  9320. 7:10:52one. So we should include the n minus
  9321. 7:10:54one. So yeah that would be 20 to 60
  9322. 7:11:00actually sorry 30 to 60 because we uh
  9323. 7:11:03index two is 30. So that would be So
  9324. 7:11:06that slice would be this one all the way
  9325. 7:11:09to the end.
  9326. 7:11:13Good. I have one more example for you
  9327. 7:11:16that's really going to throw you for a
  9328. 7:11:17loop is this one. So what happens if we
  9329. 7:11:20have
  9330. 7:11:23this case?
  9331. 7:11:31This probably won't for loop, but I'll
  9332. 7:11:32have one more following that.
  9333. 7:11:37All yeah, 10 to 60 everything. Perfect.
  9334. 7:11:41You guys are right on top of that
  9335. 7:11:42because we're leaving the starting blank
  9336. 7:11:44meaning that we should start at the
  9337. 7:11:46front. We're leaving the end blank
  9338. 7:11:48meaning we should go all the way to the
  9339. 7:11:49end. So that would be everything
  9340. 7:11:50everything in between. So at that point
  9341. 7:11:53we're not really slicing anything,
  9342. 7:11:55right? We're not really slicing much.
  9343. 7:11:56We're just taking the whole list.
  9344. 7:11:59Now, one example I want to give you
  9345. 7:12:03is
  9346. 7:12:06what if we sliced
  9347. 7:12:13and we had a step size
  9348. 7:12:17of minus one.
  9349. 7:12:22Any ideas what that would do?
  9350. 7:12:27What does this do?
  9351. 7:12:32What is a step size of minus one?
  9352. 7:12:36So minus negative index goes from the
  9353. 7:12:39back, right?
  9354. 7:12:41So if we're step sizing minus one, what
  9355. 7:12:44should we be doing effectively?
  9356. 7:12:53So, so this is signaling we basically
  9357. 7:12:54want to have the whole list but step by
  9358. 7:12:57minus one.
  9359. 7:13:04Yeah. So this this would be the reverse.
  9360. 7:13:13So this would be the reverse list. So, I
  9361. 7:13:17know that seems wacky, but that actually
  9362. 7:13:18is a way to validly reverse a list in
  9363. 7:13:21Python is to index to step size by minus
  9364. 7:13:24one. Because what that means is you're
  9365. 7:13:26slicing the entire array, but you're
  9366. 7:13:29stepping by minus one. We know minus one
  9367. 7:13:34um we know minus one goes
  9368. 7:13:38backwards, right? Effectively, because
  9369. 7:13:40it starts from the end. So, step size by
  9370. 7:13:42minus one would be go back this way.
  9371. 7:13:45each element going back this way. So
  9372. 7:13:48that would that would effectively
  9373. 7:13:49reverse the list.
  9374. 7:13:54Yeah, there now there is a reverse
  9375. 7:13:57function. A list has a reverse function.
  9376. 7:14:00So that is in English. But this is like
  9377. 7:14:04this is an alternative to to reversing
  9378. 7:14:07minus one
  9379. 7:14:09step size of minus one.
  9380. 7:14:17Uh, no, Roberto, they're not quite the
  9381. 7:14:19same because remember when you do two
  9382. 7:14:22colon and then you leave out the blank,
  9383. 7:14:25that means you're you're going all the
  9384. 7:14:26way to the end, including the blank,
  9385. 7:14:28including the end.
  9386. 7:14:30So, -4 to minus one would be the same as
  9387. 7:14:332 to
  9388. 7:14:35uh 2 to six or 2 to 7. Two to six.
  9389. 7:14:40Sorry. Two to six.
  9390. 7:14:4560 to 10. Yep. It would be it. So this
  9391. 7:14:47reverses it. Meaning this would be this
  9392. 7:14:50would return to 60 then 50 then 40. It's
  9393. 7:14:54the reverse of the list when you step
  9394. 7:14:57size by minus one.
  9395. 7:15:18I am sure.
  9396. 7:15:21How about slice from position one?
  9397. 7:15:26Slice from position one to second to
  9398. 7:15:27last and reverse it.
  9399. 7:15:31Uh what do you think that would be?
  9400. 7:15:34You're starting at one going to second
  9401. 7:15:36to last.
  9402. 7:15:39And then rever like we already know
  9403. 7:15:41what's reversing is step size minus one
  9404. 7:15:43that will always reverse.
  9405. 7:15:52What is colon colon?
  9406. 7:15:57Colon is is the fact that we're leaving
  9407. 7:16:01colon colon is not anything special.
  9408. 7:16:03It's the fact that we're leaving the
  9409. 7:16:04starting. It's just the syntax of
  9410. 7:16:06slicing, right? colon colon is because
  9411. 7:16:09we are we're leaving the starting and
  9412. 7:16:11stopping blank
  9413. 7:16:13which we can do. We're allowed to do in
  9414. 7:16:15slicing. So that would signal that we're
  9415. 7:16:16doing everything but we're going step
  9416. 7:16:18size minus one.
  9417. 7:16:25Yeah.
  9418. 7:16:30Okay.
  9419. 7:16:33Any
  9420. 7:16:38other any other questions?
  9421. 7:16:43How do we feel about slicing? Do you
  9422. 7:16:45feel okay with it? Are we going to we're
  9423. 7:16:48going to practice it more as we go
  9424. 7:16:50along? Uh because we're going to use
  9425. 7:16:52slicing quite a bit when we work with
  9426. 7:16:54data, but does the concept of slicing
  9427. 7:16:57make sense? Yeah, you need you need more
  9428. 7:16:59p We'll do more of it. We're going to do
  9429. 7:17:01slicing throughout the program.
  9430. 7:17:07Yes, we'll do multi-dimensional uh in
  9431. 7:17:10our next course.
  9432. 7:17:12Not right now, but in our in our data
  9433. 7:17:14science course, we'll do
  9434. 7:17:15multi-dimensional.
  9435. 7:17:32that help?
  9436. 7:17:37Yeah. Step can be a very Yeah, sure.
  9437. 7:17:40Sure. So, there's there's nothing that's
  9438. 7:17:42stopping you from, you know, there's
  9439. 7:17:44nothing that's stopping you from doing
  9440. 7:17:45like let's say x is two and then we do
  9441. 7:17:49um numbers
  9442. 7:17:52numbers and then we have uh two to to
  9443. 7:17:58six and then x
  9444. 7:18:01Yeah, that's fine. There's nothing that
  9445. 7:18:03would stop us from doing that.
  9446. 7:18:07I do you mean that I think that's I
  9447. 7:18:10think that's fine. There there would be
  9448. 7:18:12nothing wrong with that.
  9449. 7:18:16But you're right that it should be an
  9450. 7:18:18integer. If it's if it's like a float,
  9451. 7:18:19Python will complain. It needs to be an
  9452. 7:18:22integer step size and it needs to be it
  9453. 7:18:24needs to be uh
  9454. 7:18:26um in order to get any meaningful data.
  9455. 7:18:28We wouldn't want that step size to be
  9456. 7:18:30too big or like it, you know, if we pick
  9457. 7:18:32it to be like 20 and there's only five
  9458. 7:18:34elements, that's not going to make any
  9459. 7:18:36sense. The step size needs to be
  9460. 7:18:38reasonable.
  9461. 7:18:44All right. So, what I want to do is show
  9462. 7:18:48you some functions that lists have. So
  9463. 7:18:52probably one of the most useful
  9464. 7:18:54functions a list has is the ability to
  9465. 7:18:56append items to the list. Now this will
  9466. 7:19:00add items to the list and particularly
  9467. 7:19:03it will add it at the end. So this is
  9468. 7:19:06this is something we will use quite a
  9469. 7:19:08bit is the list.append
  9470. 7:19:11function.
  9471. 7:19:12So this will uh this will add this will
  9472. 7:19:17modify the list and add a new element at
  9473. 7:19:20the end. So append always appends to the
  9474. 7:19:23end. Um
  9475. 7:19:26and so this will uh allow us to um take
  9476. 7:19:31this this string cherry and now the when
  9477. 7:19:34we append it this list is now
  9478. 7:19:36permanently been changed to have cherry
  9479. 7:19:39at the very end. So there it is. It is
  9480. 7:19:42now at the back of the list and it is it
  9481. 7:19:45is now at the kind of end position when
  9482. 7:19:49we do append. So here you see the list
  9483. 7:19:52being really dynamic allowing us to add
  9484. 7:19:55elements to it through this append
  9485. 7:19:57function.
  9486. 7:20:02So append really really useful allows us
  9487. 7:20:05to to add we just pass in we pass in an
  9488. 7:20:08element inside of the append uh function
  9489. 7:20:11here that we want to add to the list and
  9490. 7:20:14it will it will go to the back of the
  9491. 7:20:16list.
  9492. 7:20:23Um, lists also have a pop function
  9493. 7:20:28um, which you pass in a position and it
  9494. 7:20:33will remove that item that's at that
  9495. 7:20:35position. And not only will it
  9496. 7:20:38permanently remove the item that's at
  9497. 7:20:40that position, but it will return it
  9498. 7:20:42back to you. So, pop is really useful if
  9499. 7:20:46you want to remove things um from the
  9500. 7:20:50from the list. um if you don't provide a
  9501. 7:20:54position. So if you don't provide any
  9502. 7:20:56index that you want to remove from and
  9503. 7:20:58you just do if you if you just do um pop
  9504. 7:21:02without any uh thing in there that will
  9505. 7:21:05always remove the last element by
  9506. 7:21:08default. So always just so if we just
  9507. 7:21:11did this it would remove the 40.
  9508. 7:21:18Can we append in a specific position?
  9509. 7:21:21Um, yes. You would use the insert
  9510. 7:21:24function and then give it the index you
  9511. 7:21:26want to insert into. So,
  9512. 7:21:29would would uh be every list has ainsert
  9513. 7:21:34function to to and then you put in a
  9514. 7:21:36position you want to add it to.
  9515. 7:21:40Is it common use for append and remove
  9516. 7:21:42during? Yes. So append is really common
  9517. 7:21:45to add new things to the list which
  9518. 7:21:47maybe we're doing like data aggregation.
  9519. 7:21:49We want to add things to a list and then
  9520. 7:21:51take the average of the list. That's
  9521. 7:21:53very common. Um remove. Yes. Maybe we're
  9522. 7:21:57working our way through a collection of
  9523. 7:21:59things and when we process it we want to
  9524. 7:22:00remove it. So we can do pop to remove it
  9525. 7:22:03from the list.
  9526. 7:22:05Yes.
  9527. 7:22:12Uh yes. So when so when we pop it
  9528. 7:22:16permanently affects the list. So um
  9529. 7:22:20everything gets shifted. Yes. All their
  9530. 7:22:22positions get shifted according to what
  9531. 7:22:24we removed. So like in this example um
  9532. 7:22:2830 is uh 30 is index um two but when we
  9533. 7:22:34pop it now becomes uh the last element.
  9534. 7:22:38So it would now be eligible to be index
  9535. 7:22:40minus one, right? Cuz when we remove 40,
  9536. 7:22:4430 is now the end of the list. Um,
  9537. 7:22:48for example,
  9538. 7:22:54so yeah, everything shifts
  9539. 7:22:57and you know this this example here um
  9540. 7:23:01pops from index two. So we would go to
  9541. 7:23:04index two, which is 30, and remove that.
  9542. 7:23:07And so 40 now shifts up to be at index 2
  9543. 7:23:11whereas previously it was at index 3.
  9544. 7:23:19Insert as well. Yep. When you insert
  9545. 7:23:21everything shifts. Yep.
  9546. 7:23:29Okay. Here's a here's a really useful
  9547. 7:23:31function as well. So we have the extend
  9548. 7:23:34function
  9549. 7:23:35um which would allow us to add in
  9550. 7:23:38multiple elements. So this is the same
  9551. 7:23:40as if we appended every individual item
  9552. 7:23:43in this collection to the list. So
  9553. 7:23:47extend takes a list and adds its
  9554. 7:23:51elements to the other list. So notice
  9555. 7:23:54that we have um
  9556. 7:23:56we have a list of colors here, red and
  9557. 7:23:59blue, and we're extending it with a list
  9558. 7:24:02of green and yellow, which will result
  9559. 7:24:05in the colors list now having all of
  9560. 7:24:08those elements. So extend is really
  9561. 7:24:10helpful if we want to add in multiple
  9562. 7:24:12pieces of data to an existing list.
  9563. 7:24:21a lot of questions. Um, can we get the
  9564. 7:24:23index and values with a print command?
  9565. 7:24:25Uh, yeah.
  9566. 7:24:28Yeah. I mean, you can use a I'm not sure
  9567. 7:24:30what example you have in mind, but yes.
  9568. 7:24:37Yes. Append is one item. Extend is is
  9569. 7:24:39taking an entire list and and adding all
  9570. 7:24:42of those elements to the existing list.
  9571. 7:24:45Yes. Append is for only one value at a
  9572. 7:24:48time. Yes. Extend is when you're adding
  9573. 7:24:51multiple values.
  9574. 7:24:53Append is one value at a time. Yes.
  9575. 7:25:10Uh okay. So let me ask you guys what do
  9576. 7:25:12you think of this? Um,
  9577. 7:25:15which of the following method adds a
  9578. 7:25:17single element at the end of the list?
  9579. 7:25:23So, adding a single element at the end
  9580. 7:25:25of the list.
  9581. 7:25:42Very good. It should be a Yeah, we
  9582. 7:25:44append. Append adds and append always
  9583. 7:25:47adds to the end.
  9584. 7:25:53Very good.
  9585. 7:25:59Okay. So, now we're going to have a
  9586. 7:26:02demo. Um, and by the way, this is um
  9587. 7:26:06this is a demo that uh is an existing
  9588. 7:26:10notebook. So, um, what you would want to
  9589. 7:26:14do is, especially if you're working in
  9590. 7:26:15collab, is take the notebook. Now, this
  9591. 7:26:18is within lesson two. So, we're going to
  9592. 7:26:20do demo one and lesson two. You would
  9593. 7:26:22want to take that notebook if you're
  9594. 7:26:23working in Collab and upload it. I'll
  9595. 7:26:26show you how to do that, but we're going
  9596. 7:26:28to do we're going to do the demo that's
  9597. 7:26:30inside of uh the first demo inside of
  9598. 7:26:33lesson two.
  9599. 7:26:36Um,
  9600. 7:26:44so let me share my screen.
  9601. 7:26:51Okay. So, if you're inside of Collab,
  9602. 7:26:53what you're going to want to do is go to
  9603. 7:26:55file and then upload notebook. So,
  9604. 7:26:58you're going to want to go to upload
  9605. 7:26:59notebook and then um pick the hopefully
  9606. 7:27:03you've downloaded the demos in which
  9607. 7:27:05case you have the notebook from from
  9608. 7:27:08lesson two. There's a bunch of notebook
  9609. 7:27:10files, the IP YMBs. You want to upload
  9610. 7:27:14um those demo those demo notebooks.
  9611. 7:27:18Okay. If you're working in collab, if
  9612. 7:27:19you're working in Jupiter, um, or you're
  9613. 7:27:22working in, uh, VS Code, you can just
  9614. 7:27:24open that file, uh, within VS Code or
  9615. 7:27:27Jupiter, um,
  9616. 7:27:30and, uh, you should be good to go from
  9617. 7:27:32there. So, I've I've already uh, done
  9618. 7:27:34that. This is this is demo one inside of
  9619. 7:27:38lesson two. Do you guys have access to
  9620. 7:27:39that notebook?
  9621. 7:27:42Demo one and lesson two.
  9622. 7:27:47There should be lesson two has a bunch
  9623. 7:27:50of notebooks that we're going to work
  9624. 7:27:51through. Um,
  9625. 7:27:57okay.
  9626. 7:27:58Very good.
  9627. 7:28:01So, if we run if we run this piece of
  9628. 7:28:04code um that's in this first set or
  9629. 7:28:08sorry first cell, it's going to um
  9630. 7:28:11create this list which has different mix
  9631. 7:28:13types. So this list has integers, it has
  9632. 7:28:16strings, it has floats, but we can
  9633. 7:28:20create this list. If we just hit run,
  9634. 7:28:23um, we now have a list. And what I want
  9635. 7:28:26to show you is if we were to check the
  9636. 7:28:29type of this my list, um, of course,
  9637. 7:28:32this should be a list, which it is.
  9638. 7:28:38Okay. Uh, thank you for uploading that.
  9639. 7:28:41Perfect.
  9640. 7:28:44Okay. So let's go ahead and access a few
  9641. 7:28:48elements. So we can access the first
  9642. 7:28:50element here. We can access the element
  9643. 7:28:53the fourth which would be at position
  9644. 7:28:55three and then the seventh which would
  9645. 7:28:57be at position six. We can access all of
  9646. 7:29:00those
  9647. 7:29:01and we put those into a new list here by
  9648. 7:29:05putting them inside of the brackets. So
  9649. 7:29:06that that means that we're accessing
  9650. 7:29:08this first one. That's the first element
  9651. 7:29:11of this list that we're creating. It's
  9652. 7:29:1325, which is here.
  9653. 7:29:17By the way, what do you guys think
  9654. 7:29:18happens if we try to if we try to use an
  9655. 7:29:20index that is too big for this list?
  9656. 7:29:25What do you think would happen? Like if
  9657. 7:29:27we if we tried to do if we tried to use
  9658. 7:29:30code that would be like um my list and
  9659. 7:29:34then we put in the index like 20. What
  9660. 7:29:37do you think would happen?
  9661. 7:29:40because there's definitely not 20 items
  9662. 7:29:41in this list.
  9663. 7:29:45Yeah, it'll be an error. So, let's try
  9664. 7:29:46running that. This will give me an error
  9665. 7:29:49that says it's out of range. Yeah, an
  9666. 7:29:52exception, right? It would be an
  9667. 7:29:54exception, which would say, uh, we have
  9668. 7:29:56an index error. Um, we're trying to use
  9669. 7:29:58an index that's too big for our list
  9670. 7:30:01essentially.
  9671. 7:30:05So, just pointing that out. Um, let me
  9672. 7:30:07make a comment there.
  9673. 7:30:10Um, this
  9674. 7:30:13uh index is out of range for our list.
  9675. 7:30:21So, we should get an error.
  9676. 7:30:25See how I'm making a comment? Making a
  9677. 7:30:28comment there to remind myself of why I
  9678. 7:30:30got this error. So, remember, comments
  9679. 7:30:33are useful.
  9680. 7:30:38All right. So, we access things and we
  9681. 7:30:42can uh put those inside of a list. Now,
  9682. 7:30:44let's do negative index. So, we know
  9683. 7:30:46negative -1 should give us the item
  9684. 7:30:48that's at the very end of the list,
  9685. 7:30:50which would be this 2.718.
  9686. 7:30:54Um, so that should be there. And then
  9687. 7:30:56minus 4 would be fourth from the back.
  9688. 7:30:58Minus 7 would be seventh from the back.
  9689. 7:31:02So, we can uh get those values. Not too
  9690. 7:31:04bad.
  9691. 7:31:06Then we have a slicing example.
  9692. 7:31:10So 2 to 7 we know as a slice. This
  9693. 7:31:13should um this is a slice that uh slice
  9694. 7:31:18that starts at index 2 and goes to index
  9695. 7:31:237 but doesn't include index 7.
  9696. 7:31:31Right? So that should be the slice. Um,
  9697. 7:31:34so if we run this code, um, this would
  9698. 7:31:38extract everything starting at position
  9699. 7:31:40two, which should be the third item of
  9700. 7:31:42the list, all the way up to, uh,
  9701. 7:31:45position 7.
  9702. 7:31:48And one other thing I wanted to show you
  9703. 7:31:50is we can extract how many elements are
  9704. 7:31:53in the list.
  9705. 7:31:55So I wanted to show you guys that this
  9706. 7:31:58code tells us the length of the list
  9707. 7:32:03which would be if we did length of my
  9708. 7:32:07list.
  9709. 7:32:08Yeah. Len. So we pass that in the the
  9710. 7:32:11length function. We pass in my list um
  9711. 7:32:14which should give us 10. So there's 10
  9712. 7:32:16items in this list.
  9713. 7:32:25So just wanted to call out that there is
  9714. 7:32:28this length function that we can do with
  9715. 7:32:30the list.
  9716. 7:32:35Can you show printing index numbers for
  9717. 7:32:38the list?
  9718. 7:32:43Like do you mean an uh every number in
  9719. 7:32:45its index?
  9720. 7:32:50Do you mean that every number and its
  9721. 7:32:52index? Yeah. So the code that does that
  9722. 7:32:54is the enumerate function and we we
  9723. 7:32:57would use a loop. Um so it would be
  9724. 7:33:00something like for index
  9725. 7:33:04um value in enumerate
  9726. 7:33:08uh my list and then we could do um print
  9727. 7:33:14uh index
  9728. 7:33:16and value
  9729. 7:33:24like that
  9730. 7:33:27you know we haven't learned this We
  9731. 7:33:29haven't learned any of this yet, but
  9732. 7:33:30that's that's what it's doing.
  9733. 7:33:33Yeah.
  9734. 7:33:50Okay,
  9735. 7:33:52cool. Uh let's see. So um finally what I
  9736. 7:33:57wanted to show is that we can append. So
  9737. 7:34:00if we uh take our list this is what it
  9738. 7:34:03currently is
  9739. 7:34:05and then we append a new element we can
  9740. 7:34:08print out the list and you can see how
  9741. 7:34:09it ends up at the end. So we take that
  9742. 7:34:12original list and we just add a new
  9743. 7:34:14element at the end. Um and then this by
  9744. 7:34:18the way we didn't we didn't uh explain
  9745. 7:34:20this but remove will find that value
  9746. 7:34:24find the value 100 and remove it from
  9747. 7:34:28the list. It will find the first
  9748. 7:34:31occurrence.
  9749. 7:34:33Yeah. So remove will find the first
  9750. 7:34:36occurrence of this value. Pop is index
  9751. 7:34:39based. Remove is value based. So remove
  9752. 7:34:44will look for the 100 and and take that
  9753. 7:34:47out of the list. But pop will um be
  9754. 7:34:52index based. So if we do um
  9755. 7:34:56if we do so pop is index based. So we
  9756. 7:35:02could uh do my list.pop
  9757. 7:35:05pop and we could pass in a zero which
  9758. 7:35:09should remove um this will remove the
  9759. 7:35:13first element
  9760. 7:35:15and then we can uh print my list.
  9761. 7:35:21So that removed the 25.
  9762. 7:35:30Does remove all? No, I think it's just
  9763. 7:35:33the first occurrence.
  9764. 7:35:35This is the first occurrence. You'd have
  9765. 7:35:36to do it multiple times if you have
  9766. 7:35:38duplicates.
  9767. 7:35:42I think I have to double check that, but
  9768. 7:35:44I think it's just the first occurrence.
  9769. 7:35:52Okay. All right. So, I know we're a
  9770. 7:35:54minute over.
  9771. 7:35:56Uh, thank you guys so much. What a great
  9772. 7:35:59first couple of sessions. Um, I think
  9773. 7:36:01we're picking up this really well. So
  9774. 7:36:03very good job. A lot of great questions,
  9775. 7:36:05a lot of good um back and forth. So I
  9776. 7:36:08appreciate that. Hope you guys are
  9777. 7:36:09learning and picking up this Python as
  9778. 7:36:12we go along. Um we have a lot more to
  9779. 7:36:14cover. So uh you know next time we meet
  9780. 7:36:18um you know we will uh continue talking
  9781. 7:36:21about the other data structures. So we
  9782. 7:36:22have to talk about sets, tupils,
  9783. 7:36:24dictionaries and then we have to get
  9784. 7:36:26into uh loops and if else statements and
  9785. 7:36:30then we'll eventually work our way to
  9786. 7:36:32functions. So, a lot more to cover, but
  9787. 7:36:34we'll get there. Um, and but hopefully
  9788. 7:36:38you guys are learning a lot. Any
  9789. 7:36:40homework to do? Not formally, but I
  9790. 7:36:42would request that you guys work on the
  9791. 7:36:45guided practice for lesson one. Work on
  9792. 7:36:48the guided practice for lesson one if
  9793. 7:36:51you can.
  9794. 7:36:53Okay? So, go into your reference
  9795. 7:36:55materials, find the guided practices,
  9796. 7:36:58work on the lesson one guided practice
  9797. 7:37:00between now and our next session.
  9798. 7:37:02Thank you guys. Thank you so much. Uh
  9799. 7:37:04thank you for I know these 4hour
  9800. 7:37:06sessions are a lot. Appreciate your
  9801. 7:37:08patience. Thank you so much.
  9802. 7:37:11Have a great rest of your week.
  9803. 7:37:13>> So you might be wondering what's changed
  9804. 7:37:15in machine learning and why is it the
  9805. 7:37:17best time to get into it now. Well,
  9806. 7:37:19let's go back a few years. In the past,
  9807. 7:37:22machine learning was more about building
  9808. 7:37:23models based on historical data. It was
  9809. 7:37:26about training algorithms to predict
  9810. 7:37:28specific outcomes like classifying
  9811. 7:37:30emails, spam or not spam, predicting
  9812. 7:37:32house prices based on past data. But
  9813. 7:37:34fast forward to 2026 and the landscape
  9814. 7:37:37has changed dramatically. Today, machine
  9815. 7:37:39learning isn't just about making
  9816. 7:37:40predictions. It's about building systems
  9817. 7:37:42that can learn, adapt, and improve over
  9818. 7:37:44time. We're no longer creating
  9819. 7:37:46algorithms to just run experiments
  9820. 7:37:48offline. Now, machine learning systems
  9821. 7:37:50are integrated into real world
  9822. 7:37:52operations and are capable of making
  9823. 7:37:54decisions that impact the business
  9824. 7:37:56immediately. Here's an example to make
  9825. 7:37:58it clearer. In the past, an e-commerce
  9826. 7:38:00website might use machine learning to
  9827. 7:38:02predict what products a customer might
  9828. 7:38:04want based on their past purchases. Now,
  9829. 7:38:07the systems can constantly learn from
  9830. 7:38:09new customer data, continuously refining
  9831. 7:38:11those predictions in real time as
  9832. 7:38:13customers preferences are changing. The
  9833. 7:38:15world of machine learning has evolved
  9834. 7:38:17from theory to practice and this has
  9835. 7:38:19created a huge demand for machine
  9836. 7:38:21learning engineers who can build
  9837. 7:38:23scalable systems and make them work in
  9838. 7:38:25real world environments. The impact of
  9839. 7:38:27machine learning is now directly tied to
  9840. 7:38:29business outcomes and machine learning
  9841. 7:38:31engineers are at the center of that
  9842. 7:38:32transformation. You might be thinking
  9843. 7:38:35okay I get it machine learning is
  9844. 7:38:36impactful but what exactly does a
  9845. 7:38:39machine learning engineer do compared to
  9846. 7:38:40other roles in tech? That's a great
  9847. 7:38:42question. In the world of machine
  9848. 7:38:44learning, you'll hear about a few key
  9849. 7:38:46roles such as data scientist, machine
  9850. 7:38:48learning, and AI engineer. Let's break
  9851. 7:38:50them down so you know exactly where you
  9852. 7:38:52fit in. Data scientists are like the
  9853. 7:38:54detectives of data. They spend their
  9854. 7:38:57time analyzing large data sets, finding
  9855. 7:38:59trends, and trying to extract meaningful
  9856. 7:39:01insights. They build models, but their
  9857. 7:39:04main focus is usually on data
  9858. 7:39:05exploration, and experimenting with
  9859. 7:39:07various algorithms. They don't typically
  9860. 7:39:09focus on deploying those models into
  9861. 7:39:12production environments. Machine
  9862. 7:39:13learning engineers on the other hand
  9863. 7:39:15these are architects. They take the
  9864. 7:39:17models built by data scientists and
  9865. 7:39:19build scalable deployable systems. They
  9866. 7:39:22work on creating solutions that will not
  9867. 7:39:24only work in the short term but can also
  9868. 7:39:26scale to handle real world data in
  9869. 7:39:28massive volumes. The machine learning
  9870. 7:39:30engineer is responsible for ensuring
  9871. 7:39:32that machine learning systems are
  9872. 7:39:34integrated into businesses that can work
  9873. 7:39:36seamlessly with existing technologies.
  9874. 7:39:38AI engineers focus more on the
  9875. 7:39:40application side of things. They build
  9876. 7:39:42AI powered products like chatbots, voice
  9877. 7:39:45assistants, and real-time systems. While
  9878. 7:39:47their work often overlaps with ML
  9879. 7:39:49engineers, they are typically more
  9880. 7:39:51focused on the userfacing product and
  9881. 7:39:53how machine learning fits into it. As an
  9882. 7:39:55ML engineer, your primary focus is to
  9883. 7:39:57take models and turn them into
  9884. 7:39:59actionable solutions that are deployed
  9885. 7:40:01in real world systems. We shall now move
  9886. 7:40:03on to why 2026 is the right time to
  9887. 7:40:05enter machine learning. Now that you
  9888. 7:40:07know the role of an ML engineer, let's
  9889. 7:40:09talk about why 2026 is the perfect time
  9890. 7:40:11for you to jump into the field. You've
  9891. 7:40:14probably heard that machine learning is
  9892. 7:40:15a hot topic, but what does that mean for
  9893. 7:40:17you as someone starting out in this
  9894. 7:40:19field? So, I'll help you break that down
  9895. 7:40:21for you. First, the demand for ML
  9896. 7:40:23engineers has skyrocketed. The world is
  9897. 7:40:26full of problems that needs solving, and
  9898. 7:40:28machine learning has proven to be one of
  9899. 7:40:30the most effective tools to solve them.
  9900. 7:40:32From predicting customer preferences to
  9901. 7:40:34automating critical business functions,
  9902. 7:40:36machine learning is changing how
  9903. 7:40:38businesses operate. Secondly, the tools
  9904. 7:40:40used to build machine learning systems
  9905. 7:40:42are more accessible than ever before. In
  9906. 7:40:44the past, machine learning was viewed as
  9907. 7:40:46something experimental, something that
  9908. 7:40:48required a lot of effort just to set up.
  9909. 7:40:50But today machine learning platforms and
  9910. 7:40:52frameworks such as TensorFlow, PyTorch
  9911. 7:40:54and Scikitlearn have matured
  9912. 7:40:56significantly. These tools make it
  9913. 7:40:58easier to build and deploy models that
  9914. 7:41:00can scale to handle real world data. We
  9915. 7:41:03shall now move on to why is this the
  9916. 7:41:04right time for you to get started out as
  9917. 7:41:06a machine learning engineer. So how do
  9918. 7:41:08you get started? The first step is to
  9919. 7:41:10build a strong foundation. You might be
  9920. 7:41:12excited to start building models and
  9921. 7:41:14diving into algorithms. But before that
  9922. 7:41:16you need to understand the core concepts
  9923. 7:41:18that drive all the machine learning
  9924. 7:41:19systems. These include mathematics,
  9925. 7:41:22programming and data handling. You don't
  9926. 7:41:24need to be an expert in all of these
  9927. 7:41:26areas, but you do need to understand the
  9928. 7:41:28basics. Think of these as building
  9929. 7:41:29blocks of everything that you will need
  9930. 7:41:31to learn machine learning. Let's start
  9931. 7:41:32with mathematics. You don't need to be a
  9932. 7:41:34math genius, but you do need to
  9933. 7:41:36understand the basics. There are three
  9934. 7:41:38main areas of math that will help you
  9935. 7:41:40get started out as an ML engineer, and
  9936. 7:41:42those are linear algebra. This is the
  9937. 7:41:45study of vectors and matrices which are
  9938. 7:41:47used to manipulate and process data in
  9939. 7:41:50machine learning models. If you have
  9940. 7:41:52heard terms like feature vectors or
  9941. 7:41:54matrix operations, that's linear algebra
  9942. 7:41:56play. This is the core of how machine
  9943. 7:41:59learning algorithms operate. This is the
  9944. 7:42:01math behind optimizing models. You'll
  9945. 7:42:03use calculus to adjust the parameters of
  9946. 7:42:05machine learning models and minimize the
  9947. 7:42:08error between predicted and the actual
  9948. 7:42:10values. Specifically, derivatives are
  9949. 7:42:12used to find the best way to fit a model
  9950. 7:42:14into the data. Statistics. Machine
  9951. 7:42:16learning is all about working with
  9952. 7:42:18uncertaintity, and statistics will help
  9953. 7:42:20you make sense of it. Whether you're
  9954. 7:42:22dealing with probability distributions
  9955. 7:42:23or hypothesis testing, statistics can
  9956. 7:42:26help you understand patterns in data and
  9957. 7:42:28make decisions with uncertain
  9958. 7:42:30information. These mathematical concepts
  9959. 7:42:32will help you build a more accurate
  9960. 7:42:34model and optimize it effectively. Let's
  9961. 7:42:37move on to programming stack. If you're
  9962. 7:42:39new to programming, don't worry. Python
  9963. 7:42:41is the language that you will want to
  9964. 7:42:43learn. It's simple to get started with
  9965. 7:42:45and has a huge ecosystem of libraries
  9966. 7:42:47especially designed for machine
  9967. 7:42:48learning. The key libraries that you
  9968. 7:42:50will need to master include NumPy. This
  9969. 7:42:53library is used to handle large arrays
  9970. 7:42:55of matrices and data which is the
  9971. 7:42:57backbone of most machine learning
  9972. 7:42:59algorithms. If you plan on working with
  9973. 7:43:01large data sets, you'll be using NumPy a
  9974. 7:43:03lot. Pandas. This library is great for
  9975. 7:43:06data manipulation and analysis. You'll
  9976. 7:43:08use pandas to clean, organize, and
  9977. 7:43:11transform data, making it ready for
  9978. 7:43:12machine learning models. Scikitlearn.
  9979. 7:43:15This library provides simple, easy to
  9980. 7:43:17use tools for building machine learning
  9981. 7:43:19models. It covers everything from data
  9982. 7:43:21prep-processing models like regression
  9983. 7:43:24and classification. Once you're
  9984. 7:43:26comfortable with these tools, you'll be
  9985. 7:43:27able to start building machine learning
  9986. 7:43:29models and working with real world data.
  9987. 7:43:31Next up is SQL, which stands for
  9988. 7:43:33structured query language. As an ML
  9989. 7:43:36engineer, you will be working on lots of
  9990. 7:43:38data, and SQL will help you query
  9991. 7:43:40databases to retrieve information that
  9992. 7:43:42you need. You'll use SQL to extract
  9993. 7:43:44data, filter it, and join tables
  9994. 7:43:46together, making sure that you have the
  9995. 7:43:48right data to your models. Once you have
  9996. 7:43:50your data, the next step is data
  9997. 7:43:52wrangling. Data wrangling is a huge part
  9998. 7:43:55of your job, and if you master it, it
  9999. 7:43:56will save you a lot of time and
  10000. 7:43:58frustration while building models. We
  10001. 7:44:00shall now move on to types of machine
  10002. 7:44:02learning. Let's talk about the different
  10003. 7:44:04types of machine learning. There are
  10004. 7:44:05three main categories which are
  10005. 7:44:07supervised learning, unsupervised
  10006. 7:44:09learning and reinforcement learning.
  10007. 7:44:11Speaking of supervised learning, this is
  10008. 7:44:14when you have label data. You train your
  10009. 7:44:16model on data where the answers are
  10010. 7:44:18already known. The goal is for the model
  10011. 7:44:20to learn the relationship between inputs
  10012. 7:44:22and outputs so it can predict the future
  10013. 7:44:24outcomes. Unsupervised learning. In this
  10014. 7:44:27case, the model works with unlabelled
  10015. 7:44:29data and it tries to find hidden
  10016. 7:44:31patterns and groupings in the data. This
  10017. 7:44:33is useful for tasks like clustering or
  10018. 7:44:35anomaly detection. Reinforcement
  10019. 7:44:37learning. This type of learning involves
  10020. 7:44:39training an agent to make decisions by
  10021. 7:44:42interacting with its environment and
  10022. 7:44:43receiving rewards or penalties based on
  10023. 7:44:46its actions. This is often used in game
  10024. 7:44:48AI or robotics. We shall now move on to
  10025. 7:44:51feature engineering and evaluation.
  10026. 7:44:53After cleaning your data, the next step
  10027. 7:44:55is feature engineering. This is the
  10028. 7:44:57process of transforming raw data into
  10029. 7:44:59meaningful features that help your model
  10030. 7:45:01make better predictions. Once you've
  10031. 7:45:03engineered your features, it's time to
  10032. 7:45:05evaluate the model. You'll use metrics
  10033. 7:45:07like accuracy, precision, and recall to
  10034. 7:45:10see how well your model is performing.
  10035. 7:45:12Proper evaluation ensures that your
  10036. 7:45:14model is ready for real world
  10037. 7:45:15applications and that it can handle data
  10038. 7:45:18in production. We shall now move on to
  10039. 7:45:20the six-month learning plan in order to
  10040. 7:45:21become an ML engineer in 2026. So, how
  10041. 7:45:24do you actually get started? Here's your
  10042. 7:45:26six-month learning plan to guide your
  10043. 7:45:28journey. Firstly, focus on learning
  10044. 7:45:30Python, mathematics, and SQL. Dive into
  10045. 7:45:33machine learning algorithms and hands-on
  10046. 7:45:35projects. Then, learn deep learning and
  10047. 7:45:37work with frameworks like TensorFlow or
  10048. 7:45:39PyTorch. By the sixth month, you can
  10049. 7:45:42work on a capstone project that covers
  10050. 7:45:44the full pipeline from data collection
  10051. 7:45:46to deployment. We shall now speak about
  10052. 7:45:48the ML engineer tool stack. So let's
  10053. 7:45:51talk about the tools that you will need
  10054. 7:45:52as a machine learning engineer. You must
  10055. 7:45:54be wondering what tools should I be
  10056. 7:45:56learning to become a successful ML
  10057. 7:45:58engineer in 2026. So let's simplify
  10058. 7:46:00this. Git and GitHub are the first
  10059. 7:46:03things on your list. At first glance,
  10060. 7:46:05version control might seem like
  10061. 7:46:06something that only coders need to worry
  10062. 7:46:08about. However, it can be really
  10063. 7:46:10crucial. Why? Because version control
  10064. 7:46:13allows you to track every change that
  10065. 7:46:15you make to your code, collaborate with
  10066. 7:46:17teams, and roll back changes if
  10067. 7:46:19something goes wrong. Imagine you're
  10068. 7:46:21working on a huge project and you mess
  10069. 7:46:22something up. Without version control,
  10070. 7:46:25you could lose hours of work. But with
  10071. 7:46:27Git and GitHub, you could go back and
  10072. 7:46:29fix to any point and time. You might be
  10073. 7:46:32thinking, okay, I get that Git is
  10074. 7:46:34useful, but what about the tools that
  10075. 7:46:35can actually help me build and deploy
  10076. 7:46:37machine learning models? Now, that's a
  10077. 7:46:39great question. So let's talk about the
  10078. 7:46:41cloud platforms like AWS, Google Cloud,
  10079. 7:46:44and Azure. In the past, machine learning
  10080. 7:46:46models were often built and tested upon
  10081. 7:46:48local machines, but we quickly realized
  10082. 7:46:51that it wasn't scalable. These cloud
  10083. 7:46:53platforms give you the ability to handle
  10084. 7:46:55large data sets, run models on
  10085. 7:46:57highowered servers, and scale them as
  10086. 7:46:59your projects grow. Instead of relying
  10087. 7:47:02on your personal laptop to do all of the
  10088. 7:47:04heavy lifting, you can leverage the
  10089. 7:47:05cloud to train models much faster and
  10090. 7:47:08handle massive data sets without
  10091. 7:47:09worrying about memory limitations. We
  10092. 7:47:12shall now move on to the next topic
  10093. 7:47:13which is on MLOps life cycle. You have
  10094. 7:47:16now got all of your tools in place. But
  10095. 7:47:18how do you go from all of these tools
  10096. 7:47:20coming together in the MLOps life cycle.
  10097. 7:47:22So you must be wondering what does all
  10098. 7:47:24of the MLOps life cycle tools even look
  10099. 7:47:26like and how do I manage the process
  10100. 7:47:29from starting to finish? Well, here's
  10101. 7:47:31the thing. MLOps is like a welloiled
  10102. 7:47:33machine. It involves a series of stages
  10103. 7:47:36that ensure that your models remain
  10104. 7:47:37reliable, scalable, and adaptable to new
  10105. 7:47:40data over time. Think of it like a car
  10106. 7:47:42assembly line. First, you train the car
  10107. 7:47:45and then you deploy it into the real
  10108. 7:47:47world. After that, you need to monitor
  10109. 7:47:49it and see if it runs smoothly. If it
  10110. 7:47:51breaks down, you will have to retrain
  10111. 7:47:52it. Let's break it down into key steps.
  10112. 7:47:55Training your model. This is where you
  10113. 7:47:57create the model using your training
  10114. 7:47:59data. This is a very experimental phase
  10115. 7:48:01where you try different algorithms,
  10116. 7:48:03tweak parameters, and optimize the model
  10117. 7:48:05for better performance. Deploying it to
  10118. 7:48:08production. Now that your model is
  10119. 7:48:10ready, it's time to put it into action.
  10120. 7:48:12This means making the model available to
  10121. 7:48:15users or clients. Whether it's in a
  10122. 7:48:17mobile app or a web service, deployment
  10123. 7:48:19is where the magic happens. Monitoring
  10124. 7:48:22performance. Once your model is live,
  10125. 7:48:24you can't just forget about it. It's
  10126. 7:48:26like checking the car tire pressure
  10127. 7:48:28after it's been on the road for a while.
  10128. 7:48:30You need to continuously track how well
  10129. 7:48:32your model is doing. If it starts to
  10130. 7:48:34slip or underperform, it's time for
  10131. 7:48:36tweaks and adjustments. Lastly, we have
  10132. 7:48:39retraining. Over time, your model may
  10133. 7:48:41need to be retrained with new data to
  10134. 7:48:43stay accurate. This is especially true
  10135. 7:48:45in industries where the environment is
  10136. 7:48:47constantly changing like e-commerce or
  10137. 7:48:50finance. Restraining ensures that the
  10138. 7:48:52model stays relevant and continues to
  10139. 7:48:54provide value. You may be wondering that
  10140. 7:48:57this sounds like a lot of work and
  10141. 7:48:59that's true. That's where MLOps tools
  10142. 7:49:01like MLS help you streamline this
  10143. 7:49:03process. By automating and managing this
  10144. 7:49:06life cycle, you can spend less time
  10145. 7:49:08dealing with the back end and more time
  10146. 7:49:09focusing on building innovative
  10147. 7:49:11solutions. We shall now move on to
  10148. 7:49:13experiment tracking. As you dive deeper
  10149. 7:49:15into machine learning, you'll quickly
  10150. 7:49:17realize how important it is to track
  10151. 7:49:19your experiments. At first, this might
  10152. 7:49:21seem a little overwhelming. You might be
  10153. 7:49:23thinking, I can't just run a model and
  10154. 7:49:25hope for the best, right? But trust me,
  10155. 7:49:28tracking your experiments is one of the
  10156. 7:49:29best habits that you can develop early
  10157. 7:49:31on. Think of it like logging your
  10158. 7:49:33workout progress. Even if you don't
  10159. 7:49:35track the results, how will you know
  10160. 7:49:37whether you're improving or not? The
  10161. 7:49:39same goes for machine learning. By using
  10162. 7:49:41experiment tracking tools flow and
  10163. 7:49:43weights and biases, you can keep a
  10164. 7:49:45detailed record of every model and every
  10165. 7:49:47hyperparameter and every evaluation
  10166. 7:49:50metric. Now imagine you're building a
  10167. 7:49:52recommendation system for an online
  10168. 7:49:54store. You try different adjustments,
  10169. 7:49:56algorithms, and get different results.
  10170. 7:49:59With experiment tracking, you can easily
  10171. 7:50:01compare which configurations work best.
  10172. 7:50:03You can easily compare which
  10173. 7:50:04configurations work best and learn from
  10174. 7:50:06the past mistakes. You'll always know
  10175. 7:50:09which experiment gives you the best
  10176. 7:50:10results and which needs tweaking.
  10177. 7:50:12Tracking experiments also helps in
  10178. 7:50:14collaboration. If you're working with
  10179. 7:50:16the team, being able to see everyone's
  10180. 7:50:18experiments in one place makes it easier
  10181. 7:50:20to understand their approach and build
  10182. 7:50:22on each other's work. We shall now move
  10183. 7:50:24on to projects and portfolio. Now that
  10184. 7:50:26you've got your tools and workflow in
  10185. 7:50:28place, it's time to focus on building a
  10186. 7:50:30strong portfolio. You might be thinking,
  10187. 7:50:32how do I make my portfolio stand out to
  10188. 7:50:34potential employees? That's a great
  10189. 7:50:37question. The answer is simple. Real
  10190. 7:50:39world projects. Think about it.
  10191. 7:50:42Employers want to see what you can do in
  10192. 7:50:44practice. They don't just want to see
  10193. 7:50:46theoretical knowledge. They want to see
  10194. 7:50:47that you can solve problems and build
  10195. 7:50:49working systems that can make a
  10196. 7:50:51difference. So, for that reason, we have
  10197. 7:50:53mini project examples. At this point,
  10198. 7:50:56you're probably itching to start
  10199. 7:50:57building something yourself. Well,
  10200. 7:51:00hands-on projects are the best way to
  10201. 7:51:02solidify what you've already learned.
  10202. 7:51:04For example, you could create a customer
  10203. 7:51:06churn prediction model or a fraud
  10204. 7:51:09detection system. These are practical
  10205. 7:51:11real world projects that demonstrate
  10206. 7:51:13your ability to build solutions from
  10207. 7:51:15start to finish. Make sure you document
  10208. 7:51:18your projects clearly on GitHub and
  10209. 7:51:20always include a detailed explanation of
  10210. 7:51:23your approach, challenges, and results.
  10211. 7:51:25Remember, the goal is to show that you
  10212. 7:51:27understand the problem and build a
  10213. 7:51:29working solution. We'll now move on to
  10214. 7:51:31Kaggle and open source. As you continue
  10215. 7:51:34to build your portfolio, I highly
  10216. 7:51:36recommend diving into Kaggle
  10217. 7:51:37competitions and contributing to open
  10218. 7:51:39source projects. Kaggle is an amazing
  10219. 7:51:42platform where you can work on real
  10220. 7:51:43world data sets and solve problems that
  10221. 7:51:46companies and research institutions are
  10222. 7:51:48facing. Not only will you improve your
  10223. 7:51:50skills, but you'll also have the
  10224. 7:51:52opportunity to see how top data
  10225. 7:51:53scientists and machine learning
  10226. 7:51:55engineers approach similar problems.
  10227. 7:51:57Contributing to open-source projects is
  10228. 7:51:59another excellent way to showcase your
  10229. 7:52:01skills. It shows that you can work well
  10230. 7:52:03with others understanding existing
  10231. 7:52:05systems and contribute to the community.
  10232. 7:52:08Plus, it's a great way to gain
  10233. 7:52:09visibility and make connections with
  10234. 7:52:11other engineers. We shall now speak
  10235. 7:52:13about the résumés which work for you.
  10236. 7:52:15Speaking about your resume, when it
  10237. 7:52:17comes to landing a job as a machine
  10238. 7:52:19learning engineer, your resume needs to
  10239. 7:52:21be focused on real world projects and
  10240. 7:52:23practical experience. So, be sure to
  10241. 7:52:26highlight the projects that you've
  10242. 7:52:27worked on, the tools you've used, and
  10243. 7:52:29most importantly, the impact that your
  10244. 7:52:31work has had. If you build a
  10245. 7:52:33recommendation engine that boosted
  10246. 7:52:35product sales by 20%, make sure that you
  10247. 7:52:37include that. Employers want to see how
  10248. 7:52:39your work contributes to solving real
  10249. 7:52:41business problems. Don't forget to
  10250. 7:52:43include links to Kaggle profile or any
  10251. 7:52:45other open source contributions that you
  10252. 7:52:47have made. Employers love seeing code
  10253. 7:52:50and showing them your projects is the
  10254. 7:52:51best way to stand out. We shall now
  10255. 7:52:53speak about what interviewers test in
  10256. 7:52:552026. By now you must be wondering what
  10257. 7:52:58do employers actually look for in an ML
  10258. 7:53:00engineer. So in 2026 interviews are not
  10259. 7:53:03just about technical knowledge.
  10260. 7:53:05Employers just want to see how well you
  10261. 7:53:07can communicate your thought process and
  10262. 7:53:09real world problems. You'll likely face
  10263. 7:53:11practical tests that challenge you to
  10264. 7:53:13build a model or analyze data in real
  10265. 7:53:16time. You'll also be tested on how well
  10266. 7:53:18you explain your approach and justify
  10267. 7:53:20the decisions that you made during the
  10268. 7:53:22project. Prepare to discuss things like
  10269. 7:53:25why you choose a specific algorithm for
  10270. 7:53:27a problem or how you handle issues like
  10271. 7:53:30data imbalance or overfitting along with
  10272. 7:53:33what metrics you use to evaluate your
  10273. 7:53:35model's performance. Being able to
  10274. 7:53:38communicate clearly your process and how
  10275. 7:53:40you arrived at your solution is a skill
  10276. 7:53:42that will set you apart from other
  10277. 7:53:43candidates. We shall now speak about the
  10278. 7:53:46ML trends that you will need to follow.
  10279. 7:53:48Machine learning is evolving fast and
  10280. 7:53:50staying up to date is key to remaining
  10281. 7:53:52relevant in this field. So here are a
  10282. 7:53:54few trends to keep an eye on in 2026.
  10283. 7:53:57Generative models. These models can
  10284. 7:54:00generate new data based on patterns they
  10285. 7:54:02learn and they're used for things like
  10286. 7:54:04text generation, image creation, and
  10287. 7:54:06even music composition. AutoML automated
  10288. 7:54:10machine learning tools are making it
  10289. 7:54:11easier for non-experts to build machine
  10290. 7:54:13learning models. As a result, more
  10291. 7:54:16people will be able to contribute to
  10292. 7:54:18this field without needing to become
  10293. 7:54:19experts. Privacy first models. As
  10294. 7:54:22privacy concerns grow, machine learning
  10295. 7:54:24models are being designed to work
  10296. 7:54:26securely and ethically and these don't
  10297. 7:54:28compromise on user privacy. Staying on
  10298. 7:54:31top of these trends will help you remain
  10299. 7:54:33competitive and innovative in the ever
  10300. 7:54:35evolving field of machine learning. So
  10301. 7:54:38in conclusion, I will say consistency is
  10302. 7:54:40key in machine learning. You don't need
  10303. 7:54:42to know everything right away, but you
  10304. 7:54:44must stay committed to learning and
  10305. 7:54:46building. Start small, keep
  10306. 7:54:48experimenting, and keep improving. The
  10307. 7:54:50road to become a machine learning
  10308. 7:54:52engineer may seem long, but with the
  10309. 7:54:53right tools and mindset, you can get
  10310. 7:54:55there. Thank you for watching and I'll
  10311. 7:54:57see you in the next one. Thanks for
  10312. 7:54:59watching the machine learning engineer
  10313. 7:55:00road map for 2026. We hope that this
  10314. 7:55:02video gave you a clear path to becoming
  10315. 7:55:04a successful ML engineer. Ready to take
  10316. 7:55:07the next step? Explore the Simply Lance
  10317. 7:55:09professional certificate in AI and
  10318. 7:55:10machine learning in partnership with
  10319. 7:55:12Purdue University. Gain hands-on
  10320. 7:55:14experience and industry recognized
  10321. 7:55:15credentials to boost your career.
  10322. 7:55:17>> Machine learning is which is a subset of
  10323. 7:55:20artificial intelligence, right? That's
  10324. 7:55:22uh basically um machines learning from
  10325. 7:55:26data
  10326. 7:55:28in order to uh make decisions
  10327. 7:55:30essentially. Um, so this was a big
  10328. 7:55:34departure from the rules-based systems
  10329. 7:55:37at the time, right, that were explicitly
  10330. 7:55:39programmed to make decisions. So just
  10331. 7:55:42think of an example like a really big
  10332. 7:55:45kind of if this, then that, then that,
  10333. 7:55:47then that, and and else if this, this,
  10334. 7:55:50this, right? So bunch of rules that had
  10335. 7:55:52to be pre-programmed in order to um come
  10336. 7:55:56out with some final answer. Uh, with
  10337. 7:55:58machine learning, it's the exact
  10338. 7:56:00opposite of that. we're actually
  10339. 7:56:01training something from examples from
  10340. 7:56:03existing data um in order to predict
  10341. 7:56:06something or um
  10342. 7:56:10make some type of decision. Uh and so
  10343. 7:56:12we're going to learn about the various
  10344. 7:56:14ways we can do machine learning. But if
  10345. 7:56:16you guys remember we
  10346. 7:56:18um talked about some of this like the
  10347. 7:56:20differences and the uh basically
  10348. 7:56:23rules-based approaches to learning from
  10349. 7:56:26data approach. Um and in included in
  10350. 7:56:30that is going to be uh complex
  10351. 7:56:32unstructured data. So things like
  10352. 7:56:34images, text, audio. What handles those
  10353. 7:56:37really well is uh deep learning which we
  10354. 7:56:40will get to in the course after this.
  10355. 7:56:43But uh those are certainly in there as
  10356. 7:56:46learning from data even complex data.
  10357. 7:56:51So we had this picture uh and I think
  10358. 7:56:53this is kind of around where we left off
  10359. 7:56:55last time was uh just distinguishing
  10360. 7:56:59between those three terms. We see
  10361. 7:57:01artificial intelligence, deep learning
  10362. 7:57:03and machine learning kind of used
  10363. 7:57:04interchangeably, but this is really how
  10364. 7:57:05they fit in. Artificial intelligence is
  10365. 7:57:08kind of a broad anything mimicking human
  10366. 7:57:11intelligence. Um which doesn't have to
  10367. 7:57:15be learning from data, but uh machine
  10368. 7:57:17learning is part of that. And then um
  10369. 7:57:20one way to accomplish machine learning
  10370. 7:57:22is to use neural nets which is the focus
  10371. 7:57:24of uh deep learning. Um and so deep
  10372. 7:57:28learning has been has found a lot of
  10373. 7:57:29success especially recently with uh
  10374. 7:57:32those complex data types like images,
  10375. 7:57:35speech, text, right? So deep learning
  10376. 7:57:38used all over the place. Even in um
  10377. 7:57:40modern like generative AI, we see deep
  10378. 7:57:43learning used quite a bit. Um it really
  10379. 7:57:46anything that's using neural nets is uh
  10380. 7:57:49going to be deep learning.
  10381. 7:57:52Um
  10382. 7:57:53and again we'll focus on that later but
  10383. 7:57:56we're going to be mainly focused on
  10384. 7:57:57machine learning for this course
  10385. 7:57:59primarily machine learning that does not
  10386. 7:58:01use neural networks. Okay so just models
  10387. 7:58:04that are not necessarily neural networks
  10388. 7:58:07be our focus.
  10389. 7:58:11So in machine learning we had an example
  10390. 7:58:14of a game uh essentially um learning
  10391. 7:58:19what decisions to make uh based on the
  10392. 7:58:23uh kind of current um state of the
  10393. 7:58:26board. This could be a um you know
  10394. 7:58:29machine learning example that uh learns
  10395. 7:58:32from many previous examples. So a lot of
  10396. 7:58:35data around these games are used to
  10397. 7:58:38train these um kind of robots that can
  10398. 7:58:42play these games and play them at a very
  10399. 7:58:43high level. Um so there's been a lot of
  10400. 7:58:46successes actually in machine learning
  10401. 7:58:48and deep learning um around
  10402. 7:58:51uh playing games like chess or go
  10403. 7:58:56um using machine learning algorithms. So
  10404. 7:58:58pretty cool.
  10405. 7:59:01All right. So I think this is where we
  10406. 7:59:02ended. We last time we said there's a
  10407. 7:59:03bunch of different use cases for machine
  10408. 7:59:05learning. So um recommendation system is
  10409. 7:59:08going to be a big one and we will
  10410. 7:59:10actually study that uh in one of our
  10411. 7:59:12final lessons of this course. Um chat
  10412. 7:59:15bots like generative AI doing sentiment
  10413. 7:59:18analysis chat bots we'll study later but
  10414. 7:59:20those are certainly an application of
  10415. 7:59:22learning from data in order to uh
  10416. 7:59:25generate responses to text prompts
  10417. 7:59:28right. Um spam filtering that's a good
  10418. 7:59:30example like classifying an email as
  10419. 7:59:33spam or not spam. Um that that gets
  10420. 7:59:37trained from examples and uh learning
  10421. 7:59:40from data such as previous emails. Um
  10422. 7:59:44social media posts analysis is another
  10423. 7:59:46kind of text data um use case but you uh
  10424. 7:59:52can do a lot with that text like you can
  10425. 7:59:54predict the sentiment um you can predict
  10426. 7:59:58uh the category of what what the post is
  10427. 8:00:01talking about um those kind of things
  10428. 8:00:04all can be done with machine learning
  10429. 8:00:07>> and many other use cases not on this
  10430. 8:00:09list that we will uh cover
  10431. 8:00:11>> you know as we as as we go further.
  10432. 8:00:22Okay, so this is where we kind of left
  10433. 8:00:24off. Um, so what's doing all the hard
  10434. 8:00:28work here is
  10435. 8:00:31>> uh machine learning algorithms. So these
  10436. 8:00:32are things that will um these are things
  10437. 8:00:36that will learn from the data. So they
  10438. 8:00:39are uh they they are basically um
  10439. 8:00:44algorithms or sets of rules that uh or
  10440. 8:00:48mathematical rules I should say not
  10441. 8:00:49formal rules like in the in the sense of
  10442. 8:00:51a rule system but mathematical um
  10443. 8:00:54formulas and mathematical uh rules
  10444. 8:00:57essentially that help us learn from the
  10445. 8:01:01data. So they correlate the data to some
  10446. 8:01:03type of outcome. So some type of
  10447. 8:01:06prediction uh whether that's going to be
  10448. 8:01:09as we will see whether that could be
  10449. 8:01:10like a number like we're predicting a
  10450. 8:01:13price or demand or sales
  10451. 8:01:16um or it could be a category like is
  10452. 8:01:18this transaction fraud or not fraud or
  10453. 8:01:21what's the probability that this is
  10454. 8:01:23fraud um so we have different kinds of
  10455. 8:01:27predictions we can make with machine
  10456. 8:01:28learning.
  10457. 8:01:30Um
  10458. 8:01:31but uh we will study the kind of the
  10459. 8:01:34differences of those coming up. Um but
  10460. 8:01:37machine learning algorithms are really
  10461. 8:01:39what power they're kind of the models,
  10462. 8:01:41right? They're the models that help
  10463. 8:01:42power uh machine learning to actually
  10464. 8:01:46learn from data.
  10465. 8:01:48So we're going to spend a lot of time in
  10466. 8:01:50this course studying those algorithms
  10467. 8:01:53like the different models that we can
  10468. 8:01:54build and what their differences are,
  10469. 8:01:56what their strengths are, what their
  10470. 8:01:58weaknesses are. we'll we'll learn a lot
  10471. 8:02:00about those.
  10472. 8:02:04Okay. So, I guess you can imagine like
  10473. 8:02:06everything is so data dependent, right?
  10474. 8:02:09Um we're learning from data. So, uh it
  10475. 8:02:12makes sense that the quality of data
  10476. 8:02:14really really matters here in
  10477. 8:02:16determining how strong the model can be.
  10478. 8:02:19Um so you see this graph here charting
  10479. 8:02:22kind of the um high quality data um
  10480. 8:02:27versus just uh any old data but a decent
  10481. 8:02:31enough quantity of it. Um you can see
  10482. 8:02:34that performance and the performance is
  10483. 8:02:36measured by some evaluation metric. Um,
  10484. 8:02:40so think of it as uh something like an
  10485. 8:02:42accuracy. Like if we were predicting
  10486. 8:02:44fraud or not fraud, how accurate can our
  10487. 8:02:47model get at actually detecting fraud
  10488. 8:02:50um it gets better and better and better
  10489. 8:02:53the graph shows that the higher quality
  10490. 8:02:56of data that we have. So there's kind of
  10491. 8:02:59that there's a there's a saying in
  10492. 8:03:00machine learning um called garbage in
  10493. 8:03:03garbage out. What that means is if you
  10494. 8:03:06have poor data, even the best model in
  10495. 8:03:08the world, poor data is not going to
  10496. 8:03:11result in having a good model that can
  10497. 8:03:13be accurate and perform well. Um, so it
  10498. 8:03:16needs to be high quality, meaning um
  10499. 8:03:20there needs to be a decent amount of it
  10500. 8:03:21and it needs to be labeled appropriately
  10501. 8:03:24as we will will talk about
  10502. 8:03:27um and it needs to not have any, you
  10503. 8:03:30know, significant outliers. it needs to
  10504. 8:03:32be clean, not have those missing values,
  10505. 8:03:35all of those things. Um, you can you
  10506. 8:03:38have a good chance at deriving good
  10507. 8:03:40predictions from higher quality data
  10508. 8:03:44as this kind of shows.
  10509. 8:03:49Okay.
  10510. 8:03:52So, one thing we're going to learn um as
  10511. 8:03:54we go along is
  10512. 8:03:56quantity matters as well. So, not only
  10513. 8:03:58quality, but a decent amount of it. And
  10514. 8:04:01um we're going to learn those kind of
  10515. 8:04:02rules of thumb like how much data do I
  10516. 8:04:04need for certain algorithms. Um one
  10517. 8:04:08thing that we will see is that uh the
  10518. 8:04:11the basic machine learning models that
  10519. 8:04:13we'll study don't need as much as a
  10520. 8:04:16neural network would. It you know neural
  10521. 8:04:19networks are going to require a lot more
  10522. 8:04:22um than a basic machine learning model
  10523. 8:04:25learning model. So uh that's something
  10524. 8:04:28we will see as we go along. But uh this
  10525. 8:04:30is something we'll talk about and
  10526. 8:04:32discuss with each model that we study is
  10527. 8:04:34kind of how much data do we actually
  10528. 8:04:36need to produce a high quality model.
  10529. 8:04:44Okay, any questions uh so far?
  10530. 8:04:54Okay, let's talk about the different
  10531. 8:04:57types of machine learning that we're
  10532. 8:04:59going to discuss. Prim, there's going to
  10533. 8:05:01be two primary ones that we will study
  10534. 8:05:03in this course and then a couple others
  10535. 8:05:06that'll be a little bit more advanced
  10536. 8:05:07that we won't get to but worth knowing
  10537. 8:05:10about. Um, so there's going to be four
  10538. 8:05:11total that we'll study or talk about and
  10539. 8:05:14they'll be on this list here, which is
  10540. 8:05:18um supervised learning and unsupervised
  10541. 8:05:20learning. Now, I'd say the majority of
  10542. 8:05:23our focus will probably be on supervised
  10543. 8:05:26learning, and we'll talk about what that
  10544. 8:05:28means, but we'll also cover unsupervised
  10545. 8:05:32learning as well. And so, we'll look at
  10546. 8:05:34the most popular techniques in each of
  10547. 8:05:36these types of machine learning.
  10548. 8:05:39Um,
  10549. 8:05:41and then we'll talk about these two, but
  10550. 8:05:43not really study them because they're
  10551. 8:05:45more advanced topics. Um, that that will
  10552. 8:05:48be beyond the scope of what we'll do.
  10553. 8:05:50But uh these are going to be um
  10554. 8:05:54different styles of machine learning
  10555. 8:05:58that are going to be characterized by um
  10556. 8:06:02what kinds of predictions they make,
  10557. 8:06:03what kind of data they need and require.
  10558. 8:06:06Um and uh what kind of outcomes they're
  10559. 8:06:10actually producing. Um, so let's let's
  10560. 8:06:14get into each of these, but uh the the
  10561. 8:06:17one that we'll probably spend the
  10562. 8:06:18majority of our time on is going to be
  10563. 8:06:20supervised learning, but we will study
  10564. 8:06:23unsupervised learning as well. We'll
  10565. 8:06:25study both and we're going to talk about
  10566. 8:06:26we're going to define both of those um
  10567. 8:06:28coming up. And again, these will be a
  10568. 8:06:31little bit more advanced topics that we
  10569. 8:06:32won't spend too much time on.
  10570. 8:06:35Um, but but we'll discuss their
  10571. 8:06:37relevancy in machine learning. Um, and
  10572. 8:06:41give a good definition to it.
  10573. 8:06:48Okay.
  10574. 8:06:51All right. Let's start with supervised
  10575. 8:06:54learning. Now this is going to be uh a
  10576. 8:06:57term that really refers to
  10577. 8:07:01using examples. So using labeled
  10578. 8:07:06examples. So here we say labeled data to
  10579. 8:07:10help our model train. In other words,
  10580. 8:07:14help our model be able to predict guided
  10581. 8:07:18by specific input output pairs. So
  10582. 8:07:21supervised really refers to the fact
  10583. 8:07:23that we have answers. We have examples.
  10584. 8:07:29We have answers with those and we use
  10585. 8:07:32that collection of data to build our
  10586. 8:07:35model off of so that we can predict
  10587. 8:07:39um those kinds of things like a price,
  10588. 8:07:42like a category, like a spam, not spam.
  10589. 8:07:47in this in this slide like we would be
  10590. 8:07:49predicting if this shape is a square, a
  10591. 8:07:52triangle or a circle.
  10592. 8:07:54Um but but when we build a model for
  10593. 8:07:58that, we have data that has an answer
  10594. 8:08:02attached to it, right? We've talked
  10595. 8:08:04about this before a little bit with
  10596. 8:08:05labels. So there's a guide there that
  10597. 8:08:10can guide us towards building our model.
  10598. 8:08:12there's an actual every every example
  10599. 8:08:15has an answer and that answer is really
  10600. 8:08:18critical to help build our model off of.
  10601. 8:08:21So, um that's it's almost like you have
  10602. 8:08:26um a you have a bunch of exercises
  10603. 8:08:31in let's say like a math textbook. You
  10604. 8:08:33have a bunch of exercises and you have
  10605. 8:08:35the answers and that way you can kind of
  10606. 8:08:38check your work. you think about model
  10607. 8:08:40training um that is the really a lot of
  10608. 8:08:44that process of model training as we are
  10609. 8:08:46going to discover is um basically
  10610. 8:08:49checking our work against these answers
  10611. 8:08:51in our data in our training data.
  10612. 8:08:55Okay. So supervised learning is any type
  10613. 8:08:58of machine learning that involves
  10614. 8:09:01learning from labeled data in order to
  10615. 8:09:03predict outcomes. Okay. Predict outcomes
  10616. 8:09:06like now the the outcomes can be
  10617. 8:09:09numerical. They can be like a price,
  10618. 8:09:12temperature, demand, sales, revenue.
  10619. 8:09:15They can be numerical, but they can also
  10620. 8:09:17be categorical. So they can be like spam
  10621. 8:09:20not spam, fraud, not fraud, cancer, not
  10622. 8:09:22cancer. Um dog, cat, giraffe, those kind
  10623. 8:09:27of categories. Um we could predict
  10624. 8:09:30those. It's some type of outcome. Okay,
  10625. 8:09:32some type of outcome. The key is we're
  10626. 8:09:34using labeled examples to guide our
  10627. 8:09:38model building. That's why it's called
  10628. 8:09:40supervised learning.
  10629. 8:09:42So we know in our data we know what the
  10630. 8:09:46inputs are. Of course, those are going
  10631. 8:09:48to be think of the inputs as like all of
  10632. 8:09:49our columns and then we have a special
  10633. 8:09:53label column that represents the output
  10634. 8:09:55we're trying to predict. So if you think
  10635. 8:09:57about that housing price data, the label
  10636. 8:10:00could be the price. And that's something
  10637. 8:10:02we would build a model to predict, but
  10638. 8:10:04we have answers for all of our examples
  10639. 8:10:07in our rows. We have answers to help
  10640. 8:10:10guide our model building.
  10641. 8:10:12They help tweak our model because we
  10642. 8:10:14know the answer ahead of time. So
  10643. 8:10:17they're they're really good examples to
  10644. 8:10:19build our model off of.
  10645. 8:10:23Okay. So that's that's supervised
  10646. 8:10:25learning.
  10647. 8:10:32Uh in this example, is circle not in the
  10648. 8:10:34prediction because it's not part of the
  10649. 8:10:36test data even though it's in the
  10650. 8:10:38labeled data?
  10651. 8:10:40Um no, it just not necessarily. It just
  10652. 8:10:43means that like we learn against all of
  10653. 8:10:47these examples that have these answers
  10654. 8:10:50and then when we observe new examples um
  10655. 8:10:53we can try to predict what those would
  10656. 8:10:55be based on what we've seen before. So I
  10657. 8:10:58if there was a you know it's just a
  10658. 8:11:00coincidence we only have two two
  10659. 8:11:02examples in our test data like we could
  10660. 8:11:04have a circle here in which case we
  10661. 8:11:06would predict circle
  10662. 8:11:08that's fine or at least we would hope
  10663. 8:11:10our model would predict circle right
  10664. 8:11:12that's what we're hoping may or may not
  10665. 8:11:14get it right
  10666. 8:11:16um but it's it's only not there because
  10667. 8:11:21we only like we're just assuming that we
  10668. 8:11:23only have two examples we're testing
  10669. 8:11:24against but in reality we would probably
  10670. 8:11:26do a lot more than two.
  10671. 8:11:29It's just it's just a coincidence
  10672. 8:11:30really.
  10673. 8:11:37In reality, we would test against a lot
  10674. 8:11:39more data. And we're actually going to
  10675. 8:11:41see why we would do that. Like why would
  10676. 8:11:44we train our model and then kind of use
  10677. 8:11:48additional data um to to evaluate it?
  10678. 8:11:52It's actually really important that we
  10679. 8:11:53do that step to get a sense of how good
  10680. 8:11:56our model is before we take it out in
  10681. 8:11:58the real world. So if we apply our model
  10682. 8:12:01that we build on our label data to
  10683. 8:12:05um this kind of set of test data that we
  10684. 8:12:09haven't been exposed to before. It helps
  10685. 8:12:12give us a sense of how good is our
  10686. 8:12:14model. So it's test is usually used for
  10687. 8:12:17evaluation.
  10688. 8:12:20So that's something that's something
  10689. 8:12:21we'll study.
  10690. 8:12:24How do we train? Uh it depends on the
  10691. 8:12:26model. Um so training will be a sense uh
  10692. 8:12:31will be an algorithm that will um
  10693. 8:12:34basically update the model according to
  10694. 8:12:36the data. These labeled examples. Um
  10695. 8:12:39every model is going to be different in
  10696. 8:12:40exactly how it trains. So we're going to
  10697. 8:12:44we're going to talk about that when we
  10698. 8:12:45get to the individual models that we'll
  10699. 8:12:47study.
  10700. 8:12:49But uh loosely speaking, they're going
  10701. 8:12:51to use the data to adjust itself. Like
  10702. 8:12:54imagine adjust like tuning a bunch of
  10703. 8:12:57knobs. Um, like the best example I can
  10704. 8:13:00give you is we I think I did this one
  10705. 8:13:03last week where you have kind of a
  10706. 8:13:04function
  10707. 8:13:06that predicts the price and let's say it
  10708. 8:13:09has
  10709. 8:13:11um weights like weight one with feature
  10710. 8:13:14one, weight two with feature two,
  10711. 8:13:18weight three with feature three. So
  10712. 8:13:20imagine we had three input features and
  10713. 8:13:22we we built an answer according to that.
  10714. 8:13:25Essentially what we would do to train
  10715. 8:13:27the model is adjust these
  10716. 8:13:32um in order to get this correct based on
  10717. 8:13:35our our labeled examples.
  10718. 8:13:39Okay.
  10719. 8:13:41So that's something we're going to learn
  10720. 8:13:42about coming up shortly when we when we
  10721. 8:13:44actually dive into model. Every model is
  10722. 8:13:45going to be slightly different in how it
  10723. 8:13:47trains, but at a high level it's going
  10724. 8:13:49to use the training data with those
  10725. 8:13:53examples, right? the labeled examples to
  10726. 8:13:55help guide the formula essentially to
  10727. 8:13:58adjust to generate the proper kind of
  10728. 8:14:02model here. The these things are going
  10729. 8:14:05to be adjusted according to the data
  10730. 8:14:08in order to produce the correct output.
  10731. 8:14:12So think about these as knobs that will
  10732. 8:14:14turn.
  10733. 8:14:19Okay.
  10734. 8:14:22Uh which type of machine learning is
  10735. 8:14:23used? Uh probably supervised um which is
  10736. 8:14:26what we're talking about now. So
  10737. 8:14:28probably supervised because most people
  10738. 8:14:29want to
  10739. 8:14:31um build some type of model to predict
  10740. 8:14:34something. Uh so yeah, I'd say I'd say
  10741. 8:14:38supervise.
  10742. 8:14:44Yes, we're are we are definitely going
  10743. 8:14:46to learn how to train. Yeah, we'll see.
  10744. 8:14:48We'll do the code. Um, I'll tell you
  10745. 8:14:51about how it's done. Yeah, we're
  10746. 8:14:53definitely going to learn it. But what I
  10747. 8:14:54was saying is it's kind of on a model by
  10748. 8:14:56model basis.
  10749. 8:14:58So, I want to wait till we get into the
  10750. 8:15:00individual models, then we'll talk about
  10751. 8:15:01how they're trained.
  10752. 8:15:04But yeah, we'll we'll learn how to do
  10753. 8:15:05that.
  10754. 8:15:12But yeah, supervisor is used all over
  10755. 8:15:13the place. Even even for uh generative
  10756. 8:15:17models, they use supervised learning
  10757. 8:15:19because um like an LLM
  10758. 8:15:23is going to use labeled examples in
  10759. 8:15:26order to train, right? In order to train
  10760. 8:15:29how to generate responses according to
  10761. 8:15:32prompts. Um it needs to learn against a
  10762. 8:15:35lot of text examples.
  10763. 8:15:38So that supervised learning is what um
  10764. 8:15:41results in that model,
  10765. 8:15:44right? Learning from those labeled
  10766. 8:15:45examples.
  10767. 8:15:56Okay.
  10768. 8:16:02It is yeah image image uh a lot of um
  10769. 8:16:07yeah a lot of image processing is
  10770. 8:16:08supervised like object detection. So the
  10771. 8:16:11yolo model is an object detection model.
  10772. 8:16:14Yes. Um because it has to be trained
  10773. 8:16:17right it has to be trained on uh it has
  10774. 8:16:20to be trained on images
  10775. 8:16:24with labels such as this is what object
  10776. 8:16:26is in this image. This is the box around
  10777. 8:16:29the object.
  10778. 8:16:31Um yes. So if if it's if it ever uses
  10779. 8:16:35label data to train and build the model,
  10780. 8:16:38it is supervised. So YOLO is definitely
  10781. 8:16:41supervised and we actually we will we
  10782. 8:16:44will cover the YOLO model later on in in
  10783. 8:16:46our deep learning course. We talk about
  10784. 8:16:49object detection.
  10785. 8:16:51So we'll we'll study that.
  10786. 8:16:54But yeah, it's supervised
  10787. 8:17:04Okay. So on the slide we have some
  10788. 8:17:08common supervised learning algorithms
  10789. 8:17:10that are we will study. So all of these
  10790. 8:17:13we will study and understand what they
  10791. 8:17:15do and how they work but just giving you
  10792. 8:17:18some to name them. linear regression is
  10793. 8:17:20kind of the one I just drew out which is
  10794. 8:17:22the um this is the prototypical like
  10795. 8:17:26easiest to understand model that is kind
  10796. 8:17:28of the um exactly like this where we
  10797. 8:17:31have a weight times a feature um a
  10798. 8:17:34weight times a feature and then a weight
  10799. 8:17:37times a feature
  10800. 8:17:39and on and on and on. You can have as
  10801. 8:17:41many as you want.
  10802. 8:17:43um that is a linear regression. And so
  10803. 8:17:46that is um that's a supervised model
  10804. 8:17:49because we need this value here and we
  10805. 8:17:52need all of our inputs in order to um
  10806. 8:17:56actually train this model and generate
  10807. 8:17:58all those weights
  10808. 8:18:00um that that is uh that uses um labeled
  10809. 8:18:05examples to help tune all those knobs.
  10810. 8:18:07Um same with all these other models. So,
  10811. 8:18:09we're going to talk about decision
  10812. 8:18:10trees. We're going to talk about
  10813. 8:18:11logistic regression and and SVMs, which
  10814. 8:18:13are support vector machines. We'll talk
  10815. 8:18:15about all of those, but they're all
  10816. 8:18:17examples of supervised uh supervised
  10817. 8:18:19learning.
  10818. 8:18:21Okay, we'll talk about all of these.
  10819. 8:18:25They're all supervised because they all
  10820. 8:18:28require labeled examples in order to
  10821. 8:18:30train them and and then subsequently use
  10822. 8:18:33them. Okay.
  10823. 8:18:45Okay. So what are some use case
  10824. 8:18:46examples? So for for instance in uh
  10825. 8:18:49supervised learning we may be predicting
  10826. 8:18:50temperature based on yearly temperature
  10827. 8:18:53trends. So we would have that yearly
  10828. 8:18:56data as our um as our labeled examples
  10829. 8:18:59and those would supervise the learning
  10830. 8:19:02of a model that predicts temperature.
  10831. 8:19:04Um, same thing with predicting crop
  10832. 8:19:06yield based on um, seasonal crop quality
  10833. 8:19:10changes. So maybe we have a bunch of
  10834. 8:19:11features relating to crop quality. We
  10835. 8:19:14could predict crop yield. Um, we would
  10836. 8:19:18just need historical examples with those
  10837. 8:19:21labels, right? What the crop yield is
  10838. 8:19:23for each time period. Let's say we would
  10839. 8:19:27just need those uh, supervised examples
  10840. 8:19:29and we could easily build a model off of
  10841. 8:19:32it.
  10842. 8:19:33Um
  10843. 8:19:35uh this this last one sorting waste
  10844. 8:19:37based on known waste items and their
  10845. 8:19:39corresponding waste types. Um that's
  10846. 8:19:42kind of like spam. It's like filtering
  10847. 8:19:44basically like a spam filtering. Um so
  10848. 8:19:47think of it like the the shapes example.
  10849. 8:19:49We sorting things into squares, circles,
  10850. 8:19:52triangles. Um, same kind of idea here
  10851. 8:19:54where we have a bunch of examples on
  10852. 8:19:56what those um what those waste items
  10853. 8:20:00should uh should belong to, like what
  10854. 8:20:03waste bins they would go to, for
  10855. 8:20:04example. Um, and those could be labeled
  10856. 8:20:09and therefore then we could um
  10857. 8:20:13understand what category of waste they
  10858. 8:20:15belong to.
  10859. 8:20:17Um, same thing with spam. something is
  10860. 8:20:19fraud or not fraud, spam or not spam,
  10861. 8:20:22cancer or not cancer. All of those are
  10862. 8:20:23going to be supervised learning examples
  10863. 8:20:25because they're going to require in
  10864. 8:20:27order to train them, they're going to
  10865. 8:20:29require data that has those labels.
  10866. 8:20:32Okay? So, anything that has labels is
  10867. 8:20:35going to be supervised learning.
  10868. 8:20:39So, again, this is where we will spend
  10869. 8:20:42probably the the majority of our time is
  10870. 8:20:46doing supervised learning problems. ones
  10871. 8:20:48that we have labeled data. We're
  10872. 8:20:51building a model and we're going to
  10873. 8:20:52predict those those uh labels
  10874. 8:20:55essentially.
  10875. 8:21:04Okay, before we go to unsupervised, any
  10876. 8:21:07questions about uh supervised
  10877. 8:21:24Okay.
  10878. 8:21:27All right. So, supervised requires
  10879. 8:21:29labels
  10880. 8:21:31in order to have an example to go off of
  10881. 8:21:34to build your model. And that's because
  10882. 8:21:36you're predicting those kind of outcomes
  10883. 8:21:39like spam or not spam, cancer not
  10884. 8:21:41cancer. Now unsupervised learning is
  10885. 8:21:45completely different. It's the opposite.
  10886. 8:21:48So unsupervised learning is where we do
  10887. 8:21:52not use labels whatsoever. So we're not
  10888. 8:21:55using any labels at all. So it's it it
  10889. 8:21:58can be completely unlabeled or even if
  10890. 8:22:00it's labeled, we're not using labels in
  10891. 8:22:02any way. But um we primarily would say
  10892. 8:22:05it's unlabeled data. We have no guidance
  10893. 8:22:08because we're not using the labels in
  10894. 8:22:10any way. we have no guidance to um
  10895. 8:22:14predict anything but that's because
  10896. 8:22:16we're not really predicting anything in
  10897. 8:22:17unsupervised learning. Generally what
  10898. 8:22:19we're doing is looking for some
  10899. 8:22:21structure or pattern.
  10900. 8:22:23Okay, with unsupervised learning we're
  10901. 8:22:25looking for some structure or pattern.
  10902. 8:22:27So um one type of example that's very
  10903. 8:22:32very popular is going to be this second
  10904. 8:22:34one which is um identification
  10905. 8:22:37identification of user groups based on
  10906. 8:22:40similarities or commonalities. Now this
  10907. 8:22:42is going to be a problem basically known
  10908. 8:22:45as clustering
  10909. 8:22:49and it's a problem we will study quite a
  10910. 8:22:51bit. There's going to turn out to be
  10911. 8:22:53lots of different algorithms that can
  10912. 8:22:55accomplish clustering. So what
  10913. 8:22:57clustering attempts to do is basically
  10914. 8:22:59say um we have data that's like this and
  10915. 8:23:02then data over here and then data over
  10916. 8:23:06here. Let's just group these together.
  10917. 8:23:08So like this should be one group. This
  10918. 8:23:10should be one group and this should be
  10919. 8:23:11one group. And we can find those
  10920. 8:23:14structures and say okay this is group
  10921. 8:23:17one this is group two and this is group
  10922. 8:23:19three.
  10923. 8:23:22One 2 3. And we can basically build what
  10924. 8:23:26we would call clusters of data um based
  10925. 8:23:29on how close together the points are
  10926. 8:23:32kind of located in these kind of cluster
  10927. 8:23:34zones like these boxes I've drawn.
  10928. 8:23:38Okay. Now that doesn't require any label
  10929. 8:23:40to do which is really fascinating. So
  10930. 8:23:42unsupervised you don't need any label at
  10931. 8:23:44all to accomplish the algorithm. Um so
  10932. 8:23:47clustering is one good example. Um
  10933. 8:23:50finding outliers or anomalies is
  10934. 8:23:52another. So we don't necessarily have
  10935. 8:23:54any label of what is an outlier or what
  10936. 8:23:57is an anomaly. We are deriving that from
  10937. 8:24:00the features alone. There's no guidance.
  10938. 8:24:02There's no label um to doing like
  10939. 8:24:05outlier detection or anomaly detection.
  10940. 8:24:08Okay. So that's another good example.
  10941. 8:24:10One that's not listed on here um but is
  10942. 8:24:15also really important that we will study
  10943. 8:24:16is something known as dimensionality
  10944. 8:24:19reduction.
  10945. 8:24:21So dim reduction and what that what this
  10946. 8:24:25focuses on is basically compressing the
  10947. 8:24:28data set a bit. So we take our data and
  10948. 8:24:31basically compress it. Um so that but we
  10949. 8:24:34do it in such a way that we retain as
  10950. 8:24:37much information as we can. It's a very
  10951. 8:24:40like smart compression and what it does
  10952. 8:24:43is it lowers the dimension.
  10953. 8:24:46um dimension. Think of the dimension as
  10954. 8:24:48like number of columns.
  10955. 8:24:52Number of columns.
  10956. 8:24:54So imagine we had 100 columns in a data
  10957. 8:24:57frame. What we could do is actually
  10958. 8:24:58reduce that down to 10. So like 10% of
  10959. 8:25:02that. So we reduce it down to 10. And um
  10960. 8:25:06but those 10 are it's not like we
  10961. 8:25:09chopped out um 90 other columns. we um
  10962. 8:25:14smartly kind of compressed all that
  10963. 8:25:15information into these 10 new columns um
  10964. 8:25:19that are compressed versions of the
  10965. 8:25:21hundred that we used to have. Um so
  10966. 8:25:24dimensionality reduction is is another
  10967. 8:25:26unsupervised technique. It requires no
  10968. 8:25:28guidance, no label to do, but is um a
  10969. 8:25:33really useful technique to reduce the
  10970. 8:25:35size of your data if you're doing things
  10971. 8:25:37with it. Um so this is another one that
  10972. 8:25:40we will we'll study how to do it and
  10973. 8:25:43basically more details behind it what
  10974. 8:25:44the algorithms are.
  10975. 8:25:46Um we'll so probably those two in
  10976. 8:25:49unsupervised will spend the most amount
  10977. 8:25:51of time on clustering and dimensionality
  10978. 8:25:53reduction.
  10979. 8:26:02Uh and supervised if some data is
  10980. 8:26:03present but we didn't label it means in
  10981. 8:26:06example we had circle triangle square in
  10982. 8:26:09the training data we add pentagon
  10983. 8:26:18but we didn't label that in that case.
  10984. 8:26:23Uh yeah. So every um in supervised
  10985. 8:26:27learning, every row, think about it as
  10986. 8:26:30like every row in our data frame needs
  10987. 8:26:32to have a label
  10988. 8:26:34uh associated to it. It needs to have a
  10989. 8:26:36a column that represents the label.
  10990. 8:26:41So if we've never seen Pentagon before,
  10991. 8:26:43I can't use that as a label.
  10992. 8:26:48So it has to the pentagon has to exist
  10993. 8:26:52in the data. if I'm going to be able to
  10994. 8:26:54predict it,
  10995. 8:27:01right? So, it can't predict, right? If
  10996. 8:27:03we've never seen it before, we have no
  10997. 8:27:05examples to go off. We have no guidance.
  10998. 8:27:07So, how could we predict that?
  10999. 8:27:09Right? We can't predict it
  11000. 8:27:25if it's if it's in there. So if if we
  11001. 8:27:28have labels of Pentagon, let's say, then
  11002. 8:27:31yeah, we could predict Pentagon.
  11003. 8:27:33We could
  11004. 8:27:49remove. Remove what?
  11005. 8:27:56We wouldn't if it was talking about the
  11006. 8:27:57Pentagon, we wouldn't remove that. No,
  11007. 8:27:59let me go back to that page. We wouldn't
  11008. 8:28:02remove it. Um, it's just if it's not in
  11009. 8:28:06our labels, we're not going to be able
  11010. 8:28:07to predict it. So, Pentagon's a good
  11011. 8:28:10example here. Uh, Pentagon is not one of
  11012. 8:28:15our labels. So, it currently is not in
  11013. 8:28:18our data set as one of the labels. We
  11014. 8:28:20only have data that's either a triangle,
  11015. 8:28:22circle, or
  11016. 8:28:24square. We don't have pentagon. So, I
  11017. 8:28:27would never be able to predict pentagon.
  11018. 8:28:30I'll never be able to do that if I
  11019. 8:28:31haven't seen examples of it before.
  11020. 8:28:35Okay. But let's say we had that in
  11021. 8:28:37there.
  11022. 8:28:39So, we had Pentagon.
  11023. 8:28:44So if we had Pentagon, um we could have
  11024. 8:28:47an example of it in our labels
  11025. 8:28:56and then Yeah, we it could be then we
  11026. 8:28:58could predict it.
  11027. 8:29:06Yeah. Yeah. The the don't get worried
  11028. 8:29:08don't worry about the test data. So the
  11029. 8:29:10test data is just saying here's a new
  11030. 8:29:13here's a shape what is it okay that's a
  11031. 8:29:15square here's a shape what is it okay
  11032. 8:29:17that's a triangle and we could have as
  11033. 8:29:19many of those examples as we want in our
  11034. 8:29:21test data so we could have a circle and
  11035. 8:29:24say okay what's this should be circle
  11036. 8:29:29right the test data can be whatever it
  11037. 8:29:31whatever it wants but yeah if if we've
  11038. 8:29:33never seen pentagon before we're never
  11039. 8:29:34going to be able to predict
  11040. 8:29:45These are the the label data and labels
  11041. 8:29:48are basically the talking about the same
  11042. 8:29:51thing. The labels just mean what are the
  11043. 8:29:54categories
  11044. 8:29:56that are present in our data. So in this
  11045. 8:30:00data we only have three labels that are
  11046. 8:30:01present.
  11047. 8:30:08So the labels is are relative to our
  11048. 8:30:11label data, right? It's saying
  11049. 8:30:14what labels,
  11050. 8:30:18excuse me, what labels uh do we have
  11051. 8:30:23in our data and we only have those three
  11052. 8:30:25circle, triangle, square. So so Pentagon
  11053. 8:30:28would not be part of those labels. We
  11054. 8:30:30couldn't predict it.
  11055. 8:30:44No. So unsupervised is not going to make
  11056. 8:30:47a prediction. That's the big difference
  11057. 8:30:49with unsupervised. They're not going to
  11058. 8:30:51make a prediction like this. Um so
  11059. 8:30:54unsupervised is not going to make a
  11060. 8:30:56prediction. It's going to do something
  11061. 8:30:57different like um basically say like
  11062. 8:31:00these guys are similar, these are
  11063. 8:31:02similar, these are similar, this is a
  11064. 8:31:04cluster, this is a cluster, this is a
  11065. 8:31:06cluster. It's not going to make a
  11066. 8:31:09prediction. That's what supervised
  11067. 8:31:12learning does.
  11068. 8:31:16Clustering, yes, which is unsupervised.
  11069. 8:31:19Yes, clustering does not require any
  11070. 8:31:21labels. Unsupervised just means we don't
  11071. 8:31:23have any labels. We don't require any
  11072. 8:31:25labels.
  11073. 8:31:43So the other thing unsupervised might do
  11074. 8:31:46is it might say
  11075. 8:31:49and again without the labels it might
  11076. 8:31:51say that this is an outlier.
  11077. 8:31:56it might say that this guy is an outlier
  11078. 8:31:58because there's only there's only one of
  11079. 8:32:00those and they're not like the other. So
  11080. 8:32:02that that's something that um that's
  11081. 8:32:05something that uh unsupervised could do.
  11082. 8:32:14Um it it yeah and no. It kind of labels
  11083. 8:32:19a cluster in the sense that um it would
  11084. 8:32:23basically assign a number to it like
  11085. 8:32:25this is cluster one, this is cluster
  11086. 8:32:28two, this is cluster three.
  11087. 8:32:33It'll assign a number to it, but it's
  11088. 8:32:35not a very meaningful it doesn't assign
  11089. 8:32:37like a prediction label in in the
  11090. 8:32:40traditional sense of a label.
  11091. 8:32:42It does provide like a numerical index
  11092. 8:32:44for the cluster to because what we want
  11093. 8:32:46to know is like okay this guy has the
  11094. 8:32:49cluster of one. This guy belongs to
  11095. 8:32:51cluster one. This guy belongs to cluster
  11096. 8:32:53one. This guy belongs to cluster two.
  11097. 8:32:55This guy belongs to cluster two. Does
  11098. 8:32:57that make sense? So there needs to be
  11099. 8:32:58some like index of what cluster you
  11100. 8:33:00belong to.
  11101. 8:33:03So it's kind of like a label but not in
  11102. 8:33:06the traditional like prediction sense.
  11103. 8:33:24Okay.
  11104. 8:33:35Very good. So again, unsupervised, no
  11105. 8:33:38labels. You're doing things like
  11106. 8:33:42identifying clusters,
  11107. 8:33:44um identifying outliers, doing
  11108. 8:33:47dimensionality reduction. These are all
  11109. 8:33:49like structure and pattern oriented
  11110. 8:33:52things. They're not predictions of a
  11111. 8:33:54label. Okay? They're not which is what
  11112. 8:33:57we would see in supervised learning.
  11113. 8:34:06Okay. So an example would be that we
  11114. 8:34:09take we put in the data um we can group
  11115. 8:34:13together uh data such as images into
  11116. 8:34:17categories based on similarities um
  11117. 8:34:20which would be like those clusters. So
  11118. 8:34:21there's no these would be groups that we
  11119. 8:34:24don't have any label on ahead of time
  11120. 8:34:26like we don't have we don't say that
  11121. 8:34:27this image should belong to this this
  11122. 8:34:29image should belong to this we derive
  11123. 8:34:32that from the characteristics of the
  11124. 8:34:34data. Um so think like a good example is
  11125. 8:34:38um customer groups. So we would identify
  11126. 8:34:41customers based on like okay do they
  11127. 8:34:44have similar spending levels? How many
  11128. 8:34:46days do they go shopping in a week? How
  11129. 8:34:49much money do they spend? And we can
  11130. 8:34:51kind of group together customers based
  11131. 8:34:53on similar qualities.
  11132. 8:34:56Clustering will find those groups that
  11133. 8:34:58should exist.
  11134. 8:35:00um it will discover those groups based
  11135. 8:35:03on um the similarities in the data, but
  11136. 8:35:07there's no labels that that say like
  11137. 8:35:09this person should be in this group,
  11138. 8:35:12this person should be in this ahead of
  11139. 8:35:14time. There's no labels of that. It gets
  11140. 8:35:16derived during the algorithm. It's
  11141. 8:35:18unsupervised,
  11142. 8:35:22right? There's no unsupervised really
  11143. 8:35:24literally means no guidance. There's no
  11144. 8:35:27guidance to doing it. We just derive
  11145. 8:35:29that from the structure of the data
  11146. 8:35:31which is the similarities.
  11147. 8:35:47Okay.
  11148. 8:35:53All right. So,
  11149. 8:35:56a couple more for you. So we had um
  11150. 8:35:59supervised which uses the labels. We
  11151. 8:36:04have unsupervised which uses no labels
  11152. 8:36:07looking for structure. And then we have
  11153. 8:36:09something that's kind of in between
  11154. 8:36:12which is um what is known as
  11155. 8:36:16semiupervised learning. And this is
  11156. 8:36:19where you use a combination of a little
  11157. 8:36:23bit of label data, but most of your data
  11158. 8:36:25is actually unlabeled data. Um, and you
  11159. 8:36:29try to get some use out of that label
  11160. 8:36:32data in order to um build a model out of
  11161. 8:36:37it. And so, uh, it uses the, um, it uses
  11162. 8:36:42that label data to, um, generally
  11163. 8:36:46provide some guidance on usually what
  11164. 8:36:49happens with semi-supervised learning is
  11165. 8:36:51you use your label data to kind of
  11166. 8:36:54predict what the label should be for the
  11167. 8:36:57unlabelled data and then you can go from
  11168. 8:36:59there. So you can create artificial
  11169. 8:37:01labels on this unlabeled data and then
  11170. 8:37:05you can use all of it once it's all been
  11171. 8:37:07labeled kind of like a supervised
  11172. 8:37:09learning uh approach. So but but this is
  11173. 8:37:12semi-supervised basically refers to the
  11174. 8:37:14fact that you start out with most of
  11175. 8:37:17your data not being labeled but you do
  11176. 8:37:20have some labeled examples and what you
  11177. 8:37:23can do is basically extrapolate those
  11178. 8:37:25labels into the unlabeled data set and
  11179. 8:37:28then provide some artificial labels and
  11180. 8:37:32then now everything has a label you can
  11181. 8:37:34do supervised learning.
  11182. 8:37:36Okay. So, it falls kind of between um
  11183. 8:37:40supervised and and unsupervised.
  11184. 8:37:43Uh and there So, this is this is kind of
  11185. 8:37:46rare. Most of the time you're not going
  11186. 8:37:48to do that. You're actually just going
  11187. 8:37:50to um prefer to just start with all
  11188. 8:37:53label data. That's usually the preferred
  11189. 8:37:55approach. Most of the time you'll
  11190. 8:37:57actually just be doing supervised
  11191. 8:37:58learning, not really semi-supervised
  11192. 8:38:01learning. So, it's pretty rare, but um
  11193. 8:38:09it it could like if Yeah, it could if
  11194. 8:38:12the if we had a lot of examples of
  11195. 8:38:14Pentagon and we wanted and so they were
  11196. 8:38:16unlabeled and then we tried to guess
  11197. 8:38:18what kind of shape they were um and
  11198. 8:38:21provide an artificial label uh and then
  11199. 8:38:25um then use that whole data set to build
  11200. 8:38:27a model off of then then yeah, it could
  11201. 8:38:29it could fall into this category. Okay.
  11202. 8:38:38They Oh, going back to the question,
  11203. 8:38:40they still use some kind of label data
  11204. 8:38:41like age, gender. They use uh that's
  11205. 8:38:44those aren't those aren't really labels.
  11206. 8:38:46That's the features. So, yeah, they
  11207. 8:38:48still use the core features of the data.
  11208. 8:38:52They just don't have any like labels in
  11209. 8:38:54the traditional sense of a label. Like
  11210. 8:38:56you should think of a label as something
  11211. 8:38:58we are trying to predict.
  11212. 8:39:01So whether that's a price, whether
  11213. 8:39:03that's like a category like spam, not
  11214. 8:39:05spam, cancer, not cancer, it's something
  11215. 8:39:07we'd be interested in kind of
  11216. 8:39:09predicting. And so um in our data, we
  11217. 8:39:12would have an answer for every row. We'd
  11218. 8:39:15have one of our columns would be like
  11219. 8:39:16the the result like the outcome answer
  11220. 8:39:19that we're trying to predict. That's the
  11221. 8:39:21label.
  11222. 8:39:23So in unsupervised, we don't have any of
  11223. 8:39:25the labels.
  11224. 8:39:27We do have just the regular features
  11225. 8:39:29like gender, age, income, square
  11226. 8:39:33footage,
  11227. 8:39:37bedrooms, bathrooms, all those things.
  11228. 8:39:44Okay.
  11229. 8:39:48So, we have semi-supervised that falls
  11230. 8:39:50in between supervised. Now, the reason
  11231. 8:39:52it falls between is be is because
  11232. 8:39:54there's a decent amount of data that's
  11233. 8:39:57unlabeled. In fact, a majority of it
  11234. 8:39:59unlabeled. But what we can do is try to
  11235. 8:40:03label it. We can try to take what we
  11236. 8:40:05know from our existing labels and
  11237. 8:40:08predict an artificial label and then use
  11238. 8:40:11all that data together in kind of a
  11239. 8:40:14supervised fashion for a model down the
  11240. 8:40:16road.
  11241. 8:40:22So that's kind of what this picture uh
  11242. 8:40:24says is we can try to take um you know
  11243. 8:40:28maybe we try to infer some labels based
  11244. 8:40:30on we have some some labelled data here.
  11245. 8:40:33We have most of our data is unlabeled
  11246. 8:40:36and we try to supply some labels to it.
  11247. 8:40:40Um like maybe we have a baby's category
  11248. 8:40:42of teens, a tween, uh you know youth and
  11249. 8:40:47um adults. Um and then we try so we we
  11250. 8:40:51take our our labels and we try to
  11251. 8:40:55extrapolate those into artificial labels
  11252. 8:40:57for this unlabelled data so that we can
  11253. 8:40:59use it now because then everything has a
  11254. 8:41:02label at this point and then we can just
  11255. 8:41:04go ahead and do supervised learning from
  11256. 8:41:06there.
  11257. 8:41:11So we can do supervised from there. What
  11258. 8:41:13we would prefer to do and what we'll do
  11259. 8:41:15in this course
  11260. 8:41:17um is just start with supervised. We'll
  11261. 8:41:20just start with the labels. We won't try
  11262. 8:41:22to derive artificial labels usually.
  11263. 8:41:24We'll just start with labels.
  11264. 8:41:35So one example in the real world is
  11265. 8:41:38something like Google photos which um
  11266. 8:41:42whenever you take a picture it can
  11267. 8:41:44provide uh uh labels based on previous
  11268. 8:41:48uh images in your library. So it can it
  11269. 8:41:52can produce tags or um labels on those.
  11270. 8:41:56Uh generally when you take that picture
  11271. 8:41:58it's kind of unlabeled unless you go in
  11272. 8:42:01and specifically provide some tags and
  11273. 8:42:03some labels. But um if you don't do that
  11274. 8:42:06it can still it can still uh make it can
  11275. 8:42:11artificially create one of those based
  11276. 8:42:12on the other label data that you already
  11277. 8:42:15have.
  11278. 8:42:17So that's um
  11279. 8:42:20that's an example.
  11280. 8:42:28Okay.
  11281. 8:42:32All right. Last one in terms of machine
  11282. 8:42:34learning. So we have supervised, we have
  11283. 8:42:38unsupervised.
  11284. 8:42:39Uh then we had semi-supervised which is
  11285. 8:42:42somewhere in between a mixture of having
  11286. 8:42:43some unlabelled data and label data. Um
  11287. 8:42:46now we're going to talk about
  11288. 8:42:47reinforcement learning which is
  11289. 8:42:49completely different. Um it's it's
  11290. 8:42:52completely different than the other
  11291. 8:42:53three. It's a type of machine learning
  11292. 8:42:55where we uh basically learn from
  11293. 8:42:58interaction with the environment. And
  11294. 8:43:01you might ask what are we learning? We
  11295. 8:43:04are learning what actions to take in the
  11296. 8:43:08environment. Um and the way we do that
  11297. 8:43:11is by reinforcing
  11298. 8:43:13positive actions that lead to a a
  11299. 8:43:16reward. Um, so that's where the word
  11300. 8:43:20reinforcement comes from is we we
  11301. 8:43:22basically uh imagine like a child
  11302. 8:43:25that's, you know, learning from trial
  11303. 8:43:27and error. Like they're trying to crawl,
  11304. 8:43:28they're trying to walk and they keep
  11305. 8:43:30falling down. um eventually they learn
  11306. 8:43:33how to do it through trial and error and
  11307. 8:43:35they might get a reward
  11308. 8:43:38or they might um reinforce some of those
  11309. 8:43:41positive movements that lead them to
  11310. 8:43:43walk or crawl um or they might learn
  11311. 8:43:48from the penalties, right? They might
  11312. 8:43:49learn from uh some type of feedback. So
  11313. 8:43:53they might learn from falling down like,
  11314. 8:43:55"Oh, that hurts. I should uh support
  11315. 8:43:57myself a little bit better, right?" Or
  11316. 8:43:58be a little more coordinated. Um
  11317. 8:44:02and so they they learn from those
  11318. 8:44:04actions and their interaction with the
  11319. 8:44:06environment. Um
  11320. 8:44:09uh so this is a complex um algorithm
  11321. 8:44:15essentially uh it's it deals a lot with
  11322. 8:44:19um again taking actions. Usually when
  11323. 8:44:22you take an action something changes in
  11324. 8:44:24the environment um then you kind of
  11325. 8:44:28observe some type of feedback. So, think
  11326. 8:44:30about like a a board game where you're
  11327. 8:44:33trying to figure out what move you
  11328. 8:44:35should make. Or another good example is
  11329. 8:44:37like with a robot um trying to navigate
  11330. 8:44:40a maze. So, like what route should it
  11331. 8:44:42take? Should it move forward? Should it
  11332. 8:44:44move backward? Should it move left or
  11333. 8:44:45right? Those are different actions it
  11334. 8:44:47can take. Also, like a self-driving car,
  11335. 8:44:50should it should it turn? Should it
  11336. 8:44:52speed up? Should it slow down? Those are
  11337. 8:44:54all good examples of things that have
  11338. 8:44:56been trained from reinforcement
  11339. 8:44:58learning.
  11340. 8:45:05Uh yeah. So real world examples would be
  11341. 8:45:08like in a board game, uh a a reward
  11342. 8:45:10would be like if you win the game. Um or
  11343. 8:45:14if you like capture a piece like in
  11344. 8:45:17checkers or chess, that's a reward. A
  11345. 8:45:20penalty would be like if you lose the
  11346. 8:45:21game or lose one of your pieces, that
  11347. 8:45:23could be a a penalty.
  11348. 8:45:26um in a board game or sorry in like a a
  11349. 8:45:31robot navigation task, it could get
  11350. 8:45:33rewards for um moving in the right
  11351. 8:45:36direction
  11352. 8:45:38um towards the exit or like when it like
  11353. 8:45:41let's say you wanted to train a robot on
  11354. 8:45:42how to open the door and navigate a
  11355. 8:45:45room. Um you would penalize it for
  11356. 8:45:47bumping into the wall.
  11357. 8:45:49Um you would give it a reward for moving
  11358. 8:45:57usually oh like oh the algorithm
  11359. 8:45:59themselves usually it's like a a step
  11360. 8:46:02function um it's usually it's like a
  11361. 8:46:04discrete function that kind of is based
  11362. 8:46:07on the state so the reward it could be
  11363. 8:46:10like um like depending on the let's
  11364. 8:46:13let's go back to the board game example
  11365. 8:46:15like the reward could be like or even
  11366. 8:46:18the maze let's say like a navigating the
  11367. 8:46:20maze like getting to this let's say this
  11368. 8:46:22was the exit
  11369. 8:46:25and this was the entrance.
  11370. 8:46:29Then if they make it to here, they get a
  11371. 8:46:31numerical like if they make it to the
  11372. 8:46:33exit, they get a numerical reward of
  11373. 8:46:34like plus 100, let's say. So it's just a
  11374. 8:46:37number. And then if they uh like if they
  11375. 8:46:41bump if they go into here, like let's
  11376. 8:46:43say this is kind of like a death trap or
  11377. 8:46:45like a pit, this this would be like a
  11378. 8:46:47minus 100. So it could be like discrete
  11379. 8:46:50numerical values could be the reward. If
  11380. 8:46:53they're moving in the right direction
  11381. 8:46:54like let's say we want to encourage
  11382. 8:46:56going this way then we could give
  11383. 8:46:57smaller intermediate rewards like this
  11384. 8:46:59should be a plus like if you move
  11385. 8:47:01forward this is a plus five this is a
  11386. 8:47:04plus 10 this is a plus 15 if you're
  11387. 8:47:07moving in the wrong direction away from
  11388. 8:47:09the exit. Um that would be like a minus5
  11389. 8:47:12or a minus 10. Does that make sense? So
  11390. 8:47:15they're they're numerical in nature and
  11391. 8:47:17what you're trying to do is collect the
  11392. 8:47:19most reward. You're trying to get the
  11393. 8:47:21largest reward you can through trial and
  11394. 8:47:25error. So you you try this out many many
  11395. 8:47:27many times. You basically simulate
  11396. 8:47:30running through this maze many many
  11397. 8:47:32times. And what dictates it what
  11398. 8:47:36dictates like where I should go is based
  11399. 8:47:39on what I've observed in the past. It's
  11400. 8:47:41almost like you're a child remembering
  11401. 8:47:42like, okay, what move should I make from
  11402. 8:47:44this space? Like, if I'm here, if I'm
  11403. 8:47:47here, which way should I go? Should I go
  11404. 8:47:49down? Should I go right? Should I go
  11405. 8:47:51left? You kind of know that from
  11406. 8:47:53experience.
  11407. 8:47:55Does that make sense? Based on the
  11408. 8:47:56reward that I've seen in the past, like
  11409. 8:47:58when I've moved down, I've gotten a
  11410. 8:48:00higher reward than moving left or right.
  11411. 8:48:04Does that make sense? So, yeah, it's
  11412. 8:48:05it's a numerical value
  11413. 8:48:08as a reward.
  11414. 8:48:17Yeah, that's a great question. Um, how
  11415. 8:48:20does it differentiate rewards based on
  11416. 8:48:21gain and loss i.e. chess? So it's it's a
  11417. 8:48:25very comp complicated uh answer but
  11418. 8:48:27essentially every so in the chess board
  11419. 8:48:32you can think of the board as like every
  11420. 8:48:34every um
  11421. 8:48:37space is a state.
  11422. 8:48:41So I could be in this state I could be
  11423. 8:48:43in this state and then it's not not only
  11424. 8:48:46is every every uh space but where all
  11425. 8:48:48the other pieces are. So there's lots of
  11426. 8:48:50states that are possible.
  11427. 8:48:53Um, so
  11428. 8:48:55the way there's a way to quantify
  11429. 8:48:59essentially what's the value of taking a
  11430. 8:49:03certain action like moving my piece
  11431. 8:49:04left, moving it right, moving it up or
  11432. 8:49:07down um given the rest of the state. So
  11433. 8:49:11you're you're right, it may be
  11434. 8:49:12beneficial to sacrifice. Um, but we
  11435. 8:49:16would learn that through experience that
  11436. 8:49:18okay, the best move in this situation is
  11437. 8:49:20to sacrifice.
  11438. 8:49:22We would we would have to learn that
  11439. 8:49:24through trial and error many many many
  11440. 8:49:25times which is to say like okay if I'm
  11441. 8:49:28in this current state of the world right
  11442. 8:49:31all these pieces are distributed in this
  11443. 8:49:33way the best move for me right now in
  11444. 8:49:36the long run to get the most reward in
  11445. 8:49:40the long run is to actually sacrifice my
  11446. 8:49:42piece and move it right move it into
  11447. 8:49:44like a bad position theoretically but we
  11448. 8:49:47know from experience that's actually the
  11449. 8:49:49most long-term reward is from that
  11450. 8:49:51position
  11451. 8:49:52like moving it right may be the best
  11452. 8:49:54action for me. So what you learn is how
  11453. 8:49:58to take actions
  11454. 8:50:00and actions are usually like move right,
  11455. 8:50:03move left, move up, move down. You think
  11456. 8:50:05about like a self-driving car though,
  11457. 8:50:07that's going to be like slow down, speed
  11458. 8:50:09up, turn your wheel 10°,
  11459. 8:50:13um those kind of actions.
  11460. 8:50:19So the the short answer is it's there's
  11461. 8:50:22a calculation there that you learn what
  11462. 8:50:26the long-term value of every state is
  11463. 8:50:30every unique state
  11464. 8:50:32and then you're trying to basically say
  11465. 8:50:35what action should I take from that
  11466. 8:50:37state given that current state of the
  11467. 8:50:40world.
  11468. 8:50:49Okay.
  11469. 8:50:51And I really I really like reinforcement
  11470. 8:50:53learning. It's actually probably my
  11471. 8:50:55favorite field of machine learning.
  11472. 8:50:57Unfortunately, we won't be covering it
  11473. 8:51:00um in our main uh course. We have
  11474. 8:51:03offered uh electives around
  11475. 8:51:05reinforcement learning in the past. So,
  11476. 8:51:07um stay tuned. Maybe when we get to the
  11477. 8:51:09end of this program, uh we'll offer an
  11478. 8:51:11elective on it and if enough people sign
  11479. 8:51:13up for it, we'll we'll run it. But, um
  11480. 8:51:17we it's not part of our we don't really
  11481. 8:51:19cover reinforcement learning as part of
  11482. 8:51:20our main topics. It's it is an advanced
  11483. 8:51:23uh more advanced topic than than what
  11484. 8:51:25we'll cover. But, um I I really enjoy
  11485. 8:51:28it. Find it very fascinating.
  11486. 8:51:37Okay. So, all of this is kind of um
  11487. 8:51:40illustrating what I was saying, which is
  11488. 8:51:42um you think of like uh the thing that's
  11489. 8:51:45interacting in the environment like the
  11490. 8:51:46robot or the car or the human moving a
  11491. 8:51:50chest piece is known as the agent. It's
  11492. 8:51:54interacting with the environment by
  11493. 8:51:55taking actions which updates the state
  11494. 8:51:58um of of the environment. So that's
  11495. 8:52:01that's why you see this word state here.
  11496. 8:52:03This gets updated constantly every time
  11497. 8:52:05you take an action. Um ultimately what
  11498. 8:52:07reinforcement learning is trying to do
  11499. 8:52:09is learn the best action like what would
  11500. 8:52:12be the best action to take. Um
  11501. 8:52:16and the best action is is the one that
  11502. 8:52:18leads to the most long-term reward.
  11503. 8:52:21That's the best action. Um, so you have
  11504. 8:52:24to uh you have to learn what you know
  11505. 8:52:29what leads to a good reward by kind of
  11506. 8:52:31experiencing this over and over and over
  11507. 8:52:33through trial and error. So there's a
  11508. 8:52:36lot of um kind of simulation or letting
  11509. 8:52:38the robot try something a lot um in
  11510. 8:52:42order to kind of learn what's rewarding
  11511. 8:52:44and what's not. Think about it again
  11512. 8:52:46like I think a good example is like with
  11513. 8:52:47children, right? you kind of have to let
  11514. 8:52:49them try things until they learn on
  11515. 8:52:52their own what's what can they do and
  11516. 8:52:54what can they not do
  11517. 8:52:56what's the best actions right
  11518. 8:53:00so reinforcement learning has made its
  11519. 8:53:03way into other places so I I said like a
  11520. 8:53:05good example is self-driving cars or ro
  11521. 8:53:08robotics a lot of reinforce
  11522. 8:53:10reinforcement learning is used there one
  11523. 8:53:11place it's found its way into recently
  11524. 8:53:14is recommendation systems have kind of
  11525. 8:53:17merged with reinforcement learning
  11526. 8:53:18learning. Um, and this is because you
  11527. 8:53:23you can imagine there's kind of a
  11528. 8:53:25built-in reward for you clicking on a
  11529. 8:53:28video and kind of watching it.
  11530. 8:53:31Um, so that kind of reinforces that
  11531. 8:53:33recommendation and then uh that's where
  11532. 8:53:37um you can then kind of recommend a
  11533. 8:53:40similar thing and see if that's
  11534. 8:53:42rewarding and generates a click or
  11535. 8:53:45generates some view time or watch time
  11536. 8:53:47or whatever. Um so reinforcement
  11537. 8:53:50learning has found its way into a lot of
  11538. 8:53:52areas. Um recommendations being one of
  11539. 8:53:55them because it's just natural for the
  11540. 8:53:57idea of like what um should I recommend
  11541. 8:54:00next to generate the most reward. In
  11542. 8:54:03this case the reward is kind of
  11543. 8:54:04correlated to did they click on it or
  11544. 8:54:06not or did they how long did they watch
  11545. 8:54:09for longer it's more rewarding.
  11546. 8:54:12um those kind of things but uh place
  11547. 8:54:15places where reinforcement learning have
  11548. 8:54:17been used I said self-driving cars um
  11549. 8:54:21games so uh one of the most famous
  11550. 8:54:24examples if you want to look it up is
  11551. 8:54:26the um Alph Go this was in 2016 um the
  11552. 8:54:31Alph Go uh algorithm was a reinforcement
  11553. 8:54:34learning bot that beat um some of the
  11554. 8:54:38world's best Go players which go if
  11555. 8:54:41you're not familiar Go is a um board
  11556. 8:54:44game
  11557. 8:54:45that is a little bit more uh complex
  11558. 8:54:48than chess. It has more more uh it's a
  11559. 8:54:52larger board um more pieces to it. Um
  11560. 8:54:57but they there was a reinforcement
  11561. 8:54:58learning powered bot that actually um
  11562. 8:55:01learned how to play the game so
  11563. 8:55:03effective it could beat um world kind of
  11564. 8:55:06masters at the games was pretty amazing.
  11565. 8:55:08Um that's the alpha go and that was by
  11566. 8:55:11deep mind Google and deep mind in 2016.
  11567. 8:55:15That was pretty that was only in 10
  11568. 8:55:17years ago not that long.
  11569. 8:55:19Um so certain uh we said recommendation
  11570. 8:55:24uh even autocorrect um learning to
  11571. 8:55:27predict like what is the best correction
  11572. 8:55:30uh to generate a reward which would be
  11573. 8:55:32like you accept that correction or you
  11574. 8:55:34reject it would be a penalty. Um so
  11575. 8:55:37reinforced learning has been adapted to
  11576. 8:55:40these kind of problems very
  11577. 8:55:41successfully. Let's take a look at the
  11578. 8:55:44packages that we will use throughout. So
  11579. 8:55:47um of course we will rely on these three
  11580. 8:55:50which we've already relied on to do a
  11581. 8:55:53lot of things like numpy to do numerical
  11582. 8:55:56manipulations and calculations.
  11583. 8:55:59Uh mapplot lib to do any plotting and
  11584. 8:56:02not only mapp but maybe seabour as well.
  11585. 8:56:05both of those to do plotting. Um, pandas
  11586. 8:56:08is a big one because
  11587. 8:56:11that's where all of our data is going to
  11588. 8:56:12be manipulated and prepped before it
  11589. 8:56:15goes into modeling.
  11590. 8:56:17So, all of that stuff we learned from
  11591. 8:56:19pandis is definitely going to be applied
  11592. 8:56:21here in this course uh as we actually
  11593. 8:56:24build models. Um, so of course like
  11594. 8:56:27these old ones that we've been working
  11595. 8:56:29with quite a bit um still going to be
  11596. 8:56:31useful here in the modeling stage. Um,
  11597. 8:56:35mainly for different reasons though,
  11598. 8:56:37mostly to get our data prepared to do
  11599. 8:56:40some type of modeling or maybe to
  11600. 8:56:41visualize it before we do modeling to
  11601. 8:56:43get a sense of what it looks like, those
  11602. 8:56:46kind of things.
  11603. 8:56:48Um,
  11604. 8:56:50sci is sometimes useful for certain uh
  11605. 8:56:55um processing like in unsupervised
  11606. 8:56:58learning. We'll actually use scyp a
  11607. 8:56:59little bit to do dimensionality
  11608. 8:57:01reduction or help us do that. Um so
  11609. 8:57:04scypi will be used here and there and
  11610. 8:57:07we've seen it before with hypothesis
  11611. 8:57:08testing we use scypi like the test and z
  11612. 8:57:11test came from there. Um some of the
  11613. 8:57:14unsupervised learning stuff will come
  11614. 8:57:16out of there but the package we will use
  11615. 8:57:19by far the most in this course is going
  11616. 8:57:22to be scikitlearn
  11617. 8:57:25which is here. Um and we've already seen
  11618. 8:57:29a little bit about scikitlearn in terms
  11619. 8:57:31of its pre-processing capability. So we
  11620. 8:57:35use the uh minmax scaler and the
  11621. 8:57:38standard scaler from there from the
  11622. 8:57:40pre-processing module in scikitlearn.
  11623. 8:57:43But it has um many different models
  11624. 8:57:47built into it that we can use to help uh
  11625. 8:57:50do our training and predictions. Um so
  11626. 8:57:54it's a incredibly useful machine
  11627. 8:57:56learning library. It is the industry
  11628. 8:57:58standard machine learning library. Um if
  11629. 8:58:01you're going to do anything in machine
  11630. 8:58:03learning, it would be expected that you
  11631. 8:58:05know how to use scikitlearn.
  11632. 8:58:08Now what's really lucky about that is
  11633. 8:58:10that scikitlearn is a really easy
  11634. 8:58:13package to get used to. Nearly
  11635. 8:58:15everything we do in scikitlearn will
  11636. 8:58:17mostly follow the same pattern and so um
  11637. 8:58:20the code will be extremely simple. They
  11638. 8:58:22did a great job with that package of
  11639. 8:58:24making things really user friendly,
  11640. 8:58:26really simple. Um, it's a really
  11641. 8:58:29fantastic package and we're going to get
  11642. 8:58:30a lot of practice with it uh as we go
  11643. 8:58:33along. Every model we build will
  11644. 8:58:34essentially be from scikitlearn
  11645. 8:58:37and not only like the models but um
  11646. 8:58:40doing the training, doing the
  11647. 8:58:41predictions and then doing the
  11648. 8:58:43evaluation will all come from different
  11649. 8:58:45uh scikitlearn u modules. So that'll be
  11650. 8:58:49really nice and we'll get um good
  11651. 8:58:52exposure to that package throughout the
  11652. 8:58:53course. So if anything will come away
  11653. 8:58:56from this course as um scikitlearn uh uh
  11654. 8:59:01experts that'll be very nice. So this is
  11655. 8:59:04this will be the new one for us learn
  11656. 8:59:06but we'll get a lot of practice with it.
  11657. 8:59:12Okay.
  11658. 8:59:16All right. So just to recap that lesson
  11659. 8:59:18before we move on to lesson three. Um we
  11660. 8:59:21talked about machine learning as
  11661. 8:59:22learning from data um which is included
  11662. 8:59:25underneath the AI umbrella but deep
  11663. 8:59:28learning is also included under machine
  11664. 8:59:30learning because it's still learning
  11665. 8:59:31from data but it's learning using neural
  11666. 8:59:34networks.
  11667. 8:59:35Um we talked about the four different
  11668. 8:59:37types of machine learning. We had
  11669. 8:59:38supervised, unsupervised,
  11670. 8:59:41semi-supervised and reinforcement. So
  11671. 8:59:44those are the the different types of
  11672. 8:59:45machine learning that are out there. Um
  11673. 8:59:48and then we talked about some of the pi
  11674. 8:59:49Python packages uh that we will use. The
  11675. 8:59:53main one being scikitlearn and of course
  11676. 8:59:55we'll use our older like pandas to
  11677. 8:59:57manipulate our data and get uh pass it
  11678. 8:59:59into our model training etc. But
  11679. 9:00:02scikitlearn will be uh our goto for
  11680. 9:00:06anything machine learning.
  11681. 9:00:10All right. So, I have some questions for
  11682. 9:00:11you guys, some checks.
  11683. 9:00:14So, let me know in the chat. What do you
  11684. 9:00:16guys think? Uh, which of the following
  11685. 9:00:19best describes machine learning?
  11686. 9:00:26Which choice do you think makes the best
  11687. 9:00:29is the best for this?
  11688. 9:01:12Very good. Very good. I see I see a lot
  11689. 9:01:14of choices for A and A would be the
  11690. 9:01:16correct choice. So machine learning is
  11691. 9:01:19definitely um a a subset of AI. It's
  11692. 9:01:23underneath that AI umbrella, but of
  11693. 9:01:25course we're learning from experience
  11694. 9:01:27and of course that experience is
  11695. 9:01:28recorded in the data um without being
  11696. 9:01:32explicitly programmed. Uh so it's the
  11697. 9:01:34exact opposite of BNC. We're definitely
  11698. 9:01:36not learning from rules and it's
  11699. 9:01:38definitely not just used for image and
  11700. 9:01:40speech speech recognition. It can be
  11701. 9:01:42used for many other things beyond those.
  11702. 9:01:46So yeah, A is the best choice there.
  11703. 9:01:49What we say here?
  11704. 9:01:53Okay. What do you guys think about this?
  11705. 9:01:55Which example illustrates the use of
  11706. 9:01:57machine learning to enhance customer
  11707. 9:01:58experience in an ecommerce company?
  11708. 9:02:13In other words, what would be some what
  11709. 9:02:14would be some uh typical use cases of
  11710. 9:02:17machine learning?
  11711. 9:02:45Good. So I think uh C is going to be the
  11712. 9:02:48best answer here. Definitely C. So it's
  11713. 9:02:51using machine learning to do uh fraud
  11714. 9:02:54transactions. So so that would be a
  11715. 9:02:56prediction probably a supervised
  11716. 9:02:58learning, right? If if this is fraud or
  11717. 9:03:00not fraud. Um, and then maybe some
  11718. 9:03:03customer behavior uh that might be
  11719. 9:03:06unsupervised. So maybe grouping together
  11720. 9:03:08customers uh clustering them based on
  11721. 9:03:10their data like their shopping behavior
  11722. 9:03:13and characteristics. Um that that might
  11723. 9:03:16be unsupervised but either way it's
  11724. 9:03:18machine learning.
  11725. 9:03:21Okay.
  11726. 9:03:26Okay. Final one. What distinguishes deep
  11727. 9:03:28learning from machine learning in
  11728. 9:03:30artificial intelligence? So what's
  11729. 9:03:32unique about deep learning?
  11730. 9:04:03Oh, very good. Yep. So, deep learning
  11731. 9:04:05uses neural networks as so you guys are
  11732. 9:04:09right on top of that. Neural deep
  11733. 9:04:10learning uses neural nets. That's what
  11734. 9:04:12makes it unique. So, machine learning
  11735. 9:04:14would be part A. Machine learning is
  11736. 9:04:17focused on learning from data.
  11737. 9:04:18Underneath of that is learning from data
  11738. 9:04:20using neural networks which is what uh
  11739. 9:04:23deep learning is.
  11740. 9:04:27Very good.
  11741. 9:04:30All right. Let's go to lesson three.
  11742. 9:04:34And lesson 3 has two notebooks. We're
  11743. 9:04:36going to be starting with 3.1.
  11744. 9:04:40So, you'll want to open up that
  11745. 9:04:41notebook. I'm going to go over to it
  11746. 9:04:43now. Give you a moment to open that up.
  11747. 9:04:52So, we're going to open the 3.1
  11748. 9:04:54notebook. Um, there's two of them. We'll
  11749. 9:04:57see how far if we can get into the
  11750. 9:04:59second one today. probably will.
  11751. 9:05:02Um, but we're going to do the uh we're
  11752. 9:05:04going to start with 3.1 notebook. Do you
  11753. 9:05:06guys have this notebook? Should be in
  11754. 9:05:08your materials for for this course.
  11755. 9:05:13Let me give you a moment to open that
  11756. 9:05:14one.
  11757. 9:05:30Do you guys have it?
  11758. 9:05:44Okay.
  11759. 9:05:46All right. So, we're going to start by
  11760. 9:05:50talking about uh supervised learning.
  11761. 9:05:54um in our machine learning journey. So
  11762. 9:05:56remember we're going to talk about uh
  11763. 9:05:58supervised and unsupervised after we do
  11764. 9:06:00supervised
  11765. 9:06:02um and there's going to be a lot to
  11766. 9:06:03cover with supervised mainly because um
  11767. 9:06:07there are uh two different types of
  11768. 9:06:09problems we can tackle uh which will be
  11769. 9:06:13uh we'll talk about in a moment
  11770. 9:06:14predicting different kinds of values. Um
  11771. 9:06:17but let's talk about the kind of what
  11772. 9:06:18we're hoping to learn here which is um
  11773. 9:06:21talk about the different kinds of
  11774. 9:06:22problems that we'll study which are
  11775. 9:06:24these these categories of supervised
  11776. 9:06:26learning. Um those two categories are
  11777. 9:06:28going to be called classification and
  11778. 9:06:29regression. We'll talk about those and
  11779. 9:06:31their differences and then talk about
  11780. 9:06:33some applications and some uh example
  11781. 9:06:36algorithms
  11782. 9:06:38and that's just within this notebook. Um
  11783. 9:06:403.2 two we'll get into uh regression in
  11784. 9:06:44particular um which will be uh very very
  11785. 9:06:48interesting. Okay. So that'll be our
  11786. 9:06:50first models that we'll build will be
  11787. 9:06:51over there in 3.2.
  11788. 9:06:56Okay. So if you guys remember um
  11789. 9:06:58supervised learning is where we learn
  11790. 9:07:00from labeled data. So we have input and
  11791. 9:07:03outputs in our in our data set. Um and
  11792. 9:07:07so you train a model on this data that
  11793. 9:07:10includes input features and
  11794. 9:07:13corresponding outputs that are that are
  11795. 9:07:16the labels, right? So um the goal is to
  11796. 9:07:21learn a relationship between the input
  11797. 9:07:24and the output. Of course, that's what
  11798. 9:07:25any model is trying to do. Um, and what
  11799. 9:07:28this allows us to do is then take that
  11800. 9:07:32model and use it to make predictions on
  11801. 9:07:34never-beforeseen
  11802. 9:07:36uh data. Right? So then we have a
  11803. 9:07:39predictive model out of that that we can
  11804. 9:07:41use um going forward on new examples. Um
  11805. 9:07:46so
  11806. 9:07:47remember we will have in our data a
  11807. 9:07:50bunch of features which are columns and
  11808. 9:07:52then generally one of those columns will
  11809. 9:07:54be the label that we're trying to
  11810. 9:07:56predict.
  11811. 9:07:58And our model is going to try to learn
  11812. 9:08:00some type of relationship between those
  11813. 9:08:02inputs and the output label. So the
  11814. 9:08:05output label could be like fraud not
  11815. 9:08:07fraud, cancer, not cancer, uh a price, a
  11816. 9:08:11temperature, those kind of things.
  11817. 9:08:15So let's talk about that. inside of um
  11818. 9:08:18supervised learning there are two
  11819. 9:08:19different types of learning that we can
  11820. 9:08:23do and they're really based on the label
  11821. 9:08:26or sometimes that label is known as the
  11822. 9:08:29target that we're trying to predict. Um
  11823. 9:08:32and depending on that type we get these
  11824. 9:08:35two different categories of learning or
  11825. 9:08:36two different types of learning. One is
  11826. 9:08:39known as regression. So that's generally
  11827. 9:08:42when we are predicting something that is
  11828. 9:08:44continuous or something that is a
  11829. 9:08:47numerical.
  11830. 9:08:49So numerical
  11831. 9:08:52numerical value. So think of price,
  11832. 9:08:55think of temperature, think of revenue.
  11833. 9:08:57We're trying to predict something like
  11834. 9:08:59that. Um versus something that is
  11835. 9:09:02categorical. So that the predicting
  11836. 9:09:05something categorical would be like
  11837. 9:09:06fraud, not fraud, spam, not spam. um
  11838. 9:09:09those are discrete categories and the
  11839. 9:09:13problem of predicting categories is is
  11840. 9:09:15known as classification because we're
  11841. 9:09:19trying to classify examples as belonging
  11842. 9:09:22to one category or another.
  11843. 9:09:26So we have these two main types of
  11844. 9:09:29supervised learning problems. we have
  11845. 9:09:31regression and we have classification
  11846. 9:09:33and they're going to be handled slightly
  11847. 9:09:36differently um for many reasons that
  11848. 9:09:39we're going to uncover. Um one of the
  11849. 9:09:42primary reasons is that of course we're
  11850. 9:09:45predicting something that's continuous
  11851. 9:09:47in the regression case versus something
  11852. 9:09:48discreet. So the models have to be
  11853. 9:09:50slightly different to account for that.
  11854. 9:09:53Um but then a step beyond that is the
  11855. 9:09:57evaluation has to be different too. Um I
  11856. 9:10:00kind of alluded to this last week, but
  11857. 9:10:02when you're predicting a regression,
  11858. 9:10:03it's very very difficult to to get the
  11859. 9:10:05exact numerical answer. So um generally
  11860. 9:10:11we don't care about that. Um generally
  11861. 9:10:15we don't care about getting it exactly
  11862. 9:10:18uh we don't care about getting it
  11863. 9:10:20exactly right.
  11864. 9:10:22um we just care about getting it um
  11865. 9:10:25we're just we care about getting it
  11866. 9:10:26nearby, getting it close enough. Um
  11867. 9:10:30whereas classification, we do care about
  11868. 9:10:33getting it exactly right because it's a
  11869. 9:10:34discrete category. So we're going to be
  11870. 9:10:36able to evaluate that a little bit
  11871. 9:10:38differently to say did we get the answer
  11872. 9:10:40right or wrong. Regression is going to
  11873. 9:10:42be did we get close? Um because it's we
  11874. 9:10:45assume it's going to be nearly
  11875. 9:10:46impossible to predict a continuous
  11876. 9:10:49number. Um, that's very hard to do.
  11877. 9:10:54Okay.
  11878. 9:10:56So, any questions on
  11879. 9:10:58uh that?
  11880. 9:11:04Any questions on those two differences?
  11881. 9:11:06Let me give you some examples. Maybe
  11882. 9:11:07it'll it'll help too.
  11883. 9:11:11So, again, the classification is going
  11884. 9:11:12to be predicting uh something that's
  11885. 9:11:14categorical. regression is going to be
  11886. 9:11:17predicting something that is continuous.
  11887. 9:11:24So think about trying to predict the
  11888. 9:11:26price of a house based on those other
  11889. 9:11:27features we talked about before like
  11890. 9:11:29square footage, bedrooms, bathrooms, all
  11891. 9:11:32of those things we predict the price.
  11892. 9:11:34That would be a regression problem
  11893. 9:11:35because the price is a continuous value.
  11894. 9:11:39Let's take a look at an example here.
  11895. 9:11:42Um, imagine we were trying to uh predict
  11896. 9:11:46the temperature tomorrow. That's going
  11897. 9:11:49to be a regression problem, a a
  11898. 9:11:51supervised learning kind of regression
  11899. 9:11:53problem because we're trying to predict
  11900. 9:11:55a numerical temperature.
  11901. 9:11:59Okay? And versus a category like a
  11902. 9:12:03discrete category would be this would be
  11903. 9:12:05a classification. So this is a
  11904. 9:12:07regression on the left. This is a
  11905. 9:12:09classification
  11906. 9:12:11on the right. Classification
  11907. 9:12:16um because we are um predicting one of
  11908. 9:12:21two categories. Is it just hot or cold?
  11909. 9:12:23Now, we're not saying exactly where that
  11910. 9:12:25threshold is on what's hot or cold. That
  11911. 9:12:27would be a decision on on what we want
  11912. 9:12:29to what our discrete categories actually
  11913. 9:12:32mean.
  11914. 9:12:33But, um we only have two choices, hot or
  11915. 9:12:37cold.
  11916. 9:12:38versus predicting the entire temperature
  11917. 9:12:41which would be um a numerical prediction
  11918. 9:12:44of some exact number. Right? So that'd
  11919. 9:12:48be a regression and then on the right
  11920. 9:12:50would be a classification.
  11921. 9:12:52Um now again why is this so different?
  11922. 9:12:55You can see the types of predictions
  11923. 9:12:56we're making are completely different.
  11924. 9:12:57One's a number, one's a category. But
  11925. 9:13:00again with the valuation it's like if
  11926. 9:13:03the if the true answer in our labels was
  11927. 9:13:0684
  11928. 9:13:07and we predicted 83 that's a pretty good
  11929. 9:13:10result. That's still pretty close.
  11930. 9:13:13That's pretty close to this. So from an
  11931. 9:13:14evaluation perspective that's pretty
  11932. 9:13:17good. Um whereas like if I predicted
  11933. 9:13:20cold and it's actually hot that's that's
  11934. 9:13:22a wrong answer. So they're evaluated
  11935. 9:13:25slightly different.
  11936. 9:13:27Um, and that's something we're going to
  11937. 9:13:30see as we talk about evaluation of our
  11938. 9:13:33models once we build them is depending
  11939. 9:13:35on if it's classification regression,
  11940. 9:13:37there's going to be different ways of
  11941. 9:13:38evaluating them.
  11942. 9:13:41You can kind of see why it's very
  11943. 9:13:43difficult to say, okay, we got exactly
  11944. 9:13:4684 when it could be any number. Our
  11945. 9:13:50model is going to be predicting a
  11946. 9:13:51number. That's really hard to pin down
  11947. 9:13:54an exact floatingoint number. So, the
  11948. 9:13:56best we can do is kind of say, how close
  11949. 9:13:58did I get? Like, this would be a worse
  11950. 9:14:00answer. If I got something all the way
  11951. 9:14:02down here, that's a really long distance
  11952. 9:14:04to here. That's bad. That's a bad
  11953. 9:14:06prediction. But if I get something
  11954. 9:14:08really close, that's better, right?
  11955. 9:14:11That's a decent prediction because it's
  11956. 9:14:13pretty close,
  11957. 9:14:15right?
  11958. 9:14:17Of course, being perfect would be
  11959. 9:14:18getting exactly right, but that would be
  11960. 9:14:20nearly impossible to do.
  11961. 9:14:33Okay.
  11962. 9:14:38All right. Any questions on this?
  11963. 9:14:41Does it make sense on regression versus
  11964. 9:14:43classification? We're going to use those
  11965. 9:14:44words quite a bit as we go along. So,
  11966. 9:14:47regression predicting that continuous
  11967. 9:14:48value. Classification predicting a
  11968. 9:14:51category.
  11969. 9:14:54And they're going to be um different
  11970. 9:14:56models that do that
  11971. 9:15:01different models being used for
  11972. 9:15:02regression versus different models being
  11973. 9:15:04used for classification.
  11974. 9:15:12All right, let's talk about supervised
  11975. 9:15:15learning uh applications here. So just
  11976. 9:15:18to name a few, we have HR operations. is
  11977. 9:15:22imaginary recruiter tasked with finding
  11978. 9:15:23the best candidates. Um so supervised
  11979. 9:15:26learning can help by um rejecting or
  11980. 9:15:29accepting candidates. Now this is
  11981. 9:15:30something that happens quite a bit even
  11982. 9:15:32today. Um and that it's kind of like uh
  11983. 9:15:37how recommendations happen like this
  11984. 9:15:40this resume should be um recommended
  11985. 9:15:42this should not um from a whole pool of
  11986. 9:15:45applications. Um so there's those kind
  11987. 9:15:49of use cases of of um predicting a
  11988. 9:15:52category that would be like a
  11989. 9:15:54classification. Should we should we
  11990. 9:15:56accept or reject the the candidate?
  11991. 9:15:59Um finance you see this all the time
  11992. 9:16:01with things like risk and loan
  11993. 9:16:04approvals.
  11994. 9:16:05Um you can uh predict the the the
  11995. 9:16:09category of like if the if the loan if
  11996. 9:16:12we should accept or reject the loan
  11997. 9:16:14application.
  11998. 9:16:15um you know that would be a
  11999. 9:16:17classification.
  12000. 9:16:19Um what's interesting about
  12001. 9:16:21classifications by the way so it says
  12002. 9:16:23here like we can predict the likelihood
  12003. 9:16:26of a of a loan being repaid
  12004. 9:16:28um is a lot of classifications um we we
  12005. 9:16:33say that they predict a category but
  12006. 9:16:36under the hood they can actually predict
  12007. 9:16:38a probability and we turn that
  12008. 9:16:40probability into a category. So, um, you
  12009. 9:16:45know, like we could say what's we could
  12010. 9:16:47say the likelihood of her loan being
  12011. 9:16:49repaid is very low. Let's say it's less
  12012. 9:16:51than 50% probability. Um, then we could
  12013. 9:16:54label this as reject,
  12014. 9:16:57right? We could label that as a
  12015. 9:16:59rejection. Um, if it's greater than 50%.
  12016. 9:17:03Then we could label this as accept. So
  12017. 9:17:06we can set a threshold there
  12018. 9:17:09and say okay truly we're predicting a
  12019. 9:17:12prob like our model spits out a
  12020. 9:17:14probability but we turn that into a
  12021. 9:17:16category by saying should we accept if
  12022. 9:17:19it's less than 50% we should reject if
  12023. 9:17:22it's greater than we should accept.
  12024. 9:17:25Okay, so that's something we will see
  12025. 9:17:26with some of our classification models
  12026. 9:17:28is that they actually produce a
  12027. 9:17:30probability and we turn that probability
  12028. 9:17:32into a category label
  12029. 9:17:36um by by doing something simple like
  12030. 9:17:38this putting a threshold on it um for
  12031. 9:17:41the for the category.
  12032. 9:17:45So finances is used all over the place.
  12033. 9:17:47Not only just loans like fraud, we
  12034. 9:17:49talked about fraud, not fraud. That
  12035. 9:17:50would be a classification.
  12036. 9:17:52Um predicting sales revenue, that would
  12037. 9:17:56be a regression, right? What is the
  12038. 9:17:58revenue going to be in the next two
  12039. 9:18:00quarters? That's going to be a
  12040. 9:18:02regression problem.
  12041. 9:18:05Uh emails like spam, not spam, that's
  12042. 9:18:07going to be a classification.
  12043. 9:18:10um that's going to operate on the that's
  12044. 9:18:12going to take the text input and predict
  12045. 9:18:14if this email is a spam or not spam.
  12046. 9:18:18That's going to be a uh supervised
  12047. 9:18:20learning problem, but it's going to be a
  12048. 9:18:22classification problem,
  12049. 9:18:25right? Uh manufacturing supervised
  12050. 9:18:28learning is used to inspect and uh
  12051. 9:18:31quality and classify products in
  12052. 9:18:32different grades. For example, a factory
  12053. 9:18:34might use a model to check for defects.
  12054. 9:18:36So this actually something that happens
  12055. 9:18:37is you look at images of products as
  12056. 9:18:39they go through the assembly line and
  12057. 9:18:42you can take a look at those images and
  12058. 9:18:43predict if it's a high quality, low
  12059. 9:18:45quality, medium quality. Um so they can
  12060. 9:18:48be this is a classification, right?
  12061. 9:18:50They're going into different categories
  12062. 9:18:52of quality. Um so it's much much like a
  12063. 9:18:55manual kind of intervention by some uh
  12064. 9:18:58QA or quality control uh specialist.
  12065. 9:19:04Okay. But that's a classification.
  12066. 9:19:10So in the maritime industry, supervised
  12067. 9:19:13learning can be used to predict current,
  12068. 9:19:15so current level
  12069. 9:19:17um and that can be used to forecast uh
  12070. 9:19:20supply and demand. Um so those would be
  12071. 9:19:23like regression models that are used to
  12072. 9:19:26predict um kind of like temperature, but
  12073. 9:19:28in this case like title levels.
  12074. 9:19:34We talked about fraud already, so that's
  12075. 9:19:36there. Um, that would be a
  12076. 9:19:38classification.
  12077. 9:19:43Okay,
  12078. 9:19:46any questions on these uh examples?
  12079. 9:19:50Of course, there's many more. Um
  12080. 9:19:53recommendation is kind of like a
  12081. 9:19:56supervised learning problem uh where you
  12082. 9:19:59are
  12083. 9:20:01taking examples of things that people
  12084. 9:20:03have viewed in the past or or reviewed
  12085. 9:20:06in the past and using that to predict
  12086. 9:20:08what they would want to watch in the
  12087. 9:20:10future. Um so recommendation is
  12088. 9:20:14supervised learning. Um and it's like a
  12089. 9:20:18classification, you know, trying to
  12090. 9:20:20predict um uh certain number of
  12091. 9:20:23categories of of uh shows or movies that
  12092. 9:20:27you would want to watch. Um
  12093. 9:20:31and that's something that we will study
  12094. 9:20:33in the future. Recommend we'll we'll
  12095. 9:20:35have a whole lesson dedicated to
  12096. 9:20:36recommendation as well.
  12097. 9:20:41All right.
  12098. 9:20:43So when it comes down to the uh actual
  12099. 9:20:48models themselves, so there's going to
  12100. 9:20:50be lots of different models that we are
  12101. 9:20:51going to cover. Um and they are um going
  12102. 9:20:55to be different in their purpose and
  12103. 9:20:58kind of their uh what kinds of problems
  12104. 9:21:01they're used for. Um and uh their their
  12105. 9:21:06how they actually train is going to be
  12106. 9:21:08different. Um, but at a high level,
  12107. 9:21:11they're all trying to do the same thing,
  12108. 9:21:13which is learn some sort of relationship
  12109. 9:21:15between the input data and the and the
  12110. 9:21:17label, right? That's really what they're
  12111. 9:21:19trying to do because they're all
  12112. 9:21:20supervised. They're they have those
  12113. 9:21:22labels, trying to build some
  12114. 9:21:24relationship there. Um, they just do it
  12115. 9:21:27differently.
  12116. 9:21:29And what we're going to study is the
  12117. 9:21:31pros and cons of a lot of these models,
  12118. 9:21:33like when would I use one of them, when
  12119. 9:21:34would I use another. Um, so we'll try to
  12120. 9:21:37talk about that as we go along. Um, but
  12121. 9:21:40they're all trying to learn some
  12122. 9:21:42relationship between the input features
  12123. 9:21:44and the output, right? So we have to
  12124. 9:21:47keep that in mind. They're trying to
  12125. 9:21:49model that relationship. They just do it
  12126. 9:21:51in different ways. Okay? So as we go
  12127. 9:21:53along and learn about new models, um, we
  12128. 9:21:57will learn the details. will learn the
  12129. 9:21:58ins and outs um and those pros and cons,
  12130. 9:22:02but they're no matter what, they're all
  12131. 9:22:04trying to
  12132. 9:22:06uh learn that relationship, right? And
  12133. 9:22:08be able to make predictions on new data.
  12134. 9:22:13Okay,
  12135. 9:22:15so here's a list of models that we will
  12136. 9:22:18cover and work on throughout the uh the
  12137. 9:22:22sessions that we have.
  12138. 9:22:24um we're not going to do them all in one
  12139. 9:22:26one sitting, but um the first one that
  12140. 9:22:28we're going to start with and that we'll
  12141. 9:22:30cover today is going to be linear
  12142. 9:22:32regression.
  12143. 9:22:33So we will cover linear regression and
  12144. 9:22:35then we'll cover the rest of these guys
  12145. 9:22:38mostly in the context of uh
  12146. 9:22:41classification.
  12147. 9:22:42So, um, what's interesting is some of
  12148. 9:22:46these guys can actually be used for both
  12149. 9:22:48regression and classification as long as
  12150. 9:22:50you make, um, certain adjustments to
  12151. 9:22:53them. They have variations that can be
  12152. 9:22:56used to do classification and regression
  12153. 9:22:58is very interesting. Um, but we're going
  12154. 9:23:02to start with linear regression today
  12155. 9:23:05and then work our way through the rest
  12156. 9:23:07of these models when we do um, we're
  12157. 9:23:09going to do a separate lesson four on
  12158. 9:23:11classification. And so these all these
  12159. 9:23:13guys will come from lesson four.
  12160. 9:23:18Um and then uh we will do this guy in
  12161. 9:23:23lesson three in the 3.2 notebook. We'll
  12162. 9:23:26do all about linear regression.
  12163. 9:23:30Yeah. I so logistic regression is a
  12164. 9:23:32classification. Um which is kind of
  12165. 9:23:35strange that its name is regression but
  12166. 9:23:37it's doing a classification. But the the
  12167. 9:23:40reason is that the logistic regression
  12168. 9:23:42um computes a probability. So it does a
  12169. 9:23:46regression to predict a number but that
  12170. 9:23:49number is actually a probability. So it
  12171. 9:23:51it produces a result that's between it
  12172. 9:23:54produces a probability that's between um
  12173. 9:23:58obviously uh zero and one.
  12174. 9:24:03So it uh and then we take that
  12175. 9:24:05probability and we turn it into a
  12176. 9:24:07category
  12177. 9:24:09like a spam not spam fraud not fraud.
  12178. 9:24:12Um but so so logistic regression is kind
  12179. 9:24:14of special. It's sort of like a
  12180. 9:24:16regression but it's predicting a very
  12181. 9:24:18specific type of value which is a
  12182. 9:24:20probability. So for for that reason it's
  12183. 9:24:23a classification uh algorithm primarily.
  12184. 9:24:32So we'll study that one in lesson four.
  12185. 9:24:35Uh but yeah, that's that's why it's
  12186. 9:24:37under that kind of umbrella of
  12187. 9:24:39classification is because it's it's
  12188. 9:24:40producing a probability as its main
  12189. 9:24:42output which we can then turn into a
  12190. 9:24:46category as long as we interpret that
  12191. 9:24:48probability as um in the right way. Uh
  12192. 9:24:52like the probability of spam,
  12193. 9:24:54probability of not spam.
  12194. 9:25:00Okay.
  12195. 9:25:03Okay. So, let me focus on um
  12196. 9:25:07let me focus on linear regression. I'm
  12197. 9:25:09not going to go through all of these
  12198. 9:25:10other use cases because we haven't
  12199. 9:25:12learned these models yet. Um so, I don't
  12200. 9:25:16think they're good. Uh I don't think
  12201. 9:25:19it's good to read about them yet until
  12202. 9:25:21we've covered them. So, once we cover
  12203. 9:25:23them in lesson four, I'll come back and
  12204. 9:25:25describe these examples to you guys and
  12205. 9:25:28we'll see why it makes sense. But I
  12206. 9:25:30think for linear regression um which is
  12207. 9:25:32what we'll cover next, let me talk about
  12208. 9:25:34that example. So a prototypical example
  12209. 9:25:36would be like predicting the house
  12210. 9:25:38prices that we've seen in that house
  12211. 9:25:40price data set.
  12212. 9:25:42So um if we wanted to uh if we wanted to
  12213. 9:25:47predict um if we wanted to estimate the
  12214. 9:25:50market value of a house so the price
  12215. 9:25:55um we could do that by using the
  12216. 9:25:57features such as number of bedrooms,
  12217. 9:25:59square footage, location, age of the
  12218. 9:26:02property. Um and you know then when a
  12219. 9:26:06new when a new house comes on the market
  12220. 9:26:08we could estimate what the price should
  12221. 9:26:10be based on those features. So linear
  12222. 9:26:14regression is a good one to predict the
  12223. 9:26:16price like a housing price. Um and we'll
  12224. 9:26:19actually practice that in the next uh
  12225. 9:26:22notebook.
  12226. 9:26:25So we'll we'll uh and then all these
  12227. 9:26:28other now there's descriptions of these
  12228. 9:26:30other models but again we haven't
  12229. 9:26:31covered these guys yet. So I don't want
  12230. 9:26:32to really go through those until we get
  12231. 9:26:35to those models. So we get to those I'll
  12232. 9:26:37come back and mention the example.
  12233. 9:26:40Uh can K andN be used for clustering?
  12234. 9:26:43No. So um the clustering model is going
  12235. 9:26:46to be different. It's going to be uh K
  12236. 9:26:49means
  12237. 9:26:51K means that's the primary clustering
  12238. 9:26:53model. Not K nearest neighbors. K
  12239. 9:26:56nearest neighbors is used for uh it can
  12240. 9:26:59be used for regression. It can be used
  12241. 9:27:00for classification.
  12242. 9:27:04So we'll we'll talk about K andN which
  12243. 9:27:06is the K nearest neighbors in lesson
  12244. 9:27:08four.
  12245. 9:27:11It sounds really similar. Yeah, it
  12246. 9:27:13sounds really similar but K means is a
  12247. 9:27:16clustering algorithm that's that's
  12248. 9:27:17slightly different
  12249. 9:27:19different uh there's no labels used at
  12250. 9:27:22all. This K nearest neighbors is a is a
  12251. 9:27:25supervised learning algorithm. It uses
  12252. 9:27:28uh labels.
  12253. 9:27:38Good. Any any other questions so far?
  12254. 9:27:56Okay.
  12255. 9:27:58So that being said, let's move on to the
  12256. 9:28:013.2 notebook.
  12257. 9:28:04Let's move on to that which will be our
  12258. 9:28:07um first discussion around uh
  12259. 9:28:11regression. So going into supervised
  12260. 9:28:14learning and regression. Give you guys a
  12261. 9:28:16moment to pull up this notebook.
  12262. 9:28:19But yeah, you want to pull up the 3.2.
  12263. 9:28:21We'll do this one next. So we'll focus
  12264. 9:28:23in. So our plan is to do regression
  12265. 9:28:26first and then we'll talk about
  12266. 9:28:28classification in lesson four
  12267. 9:28:35which we will cover all those other
  12268. 9:28:37models which you you could use for
  12269. 9:28:39classification uh on that list. But then
  12270. 9:28:42we're going to talk about linear
  12271. 9:28:43regression uh first.
  12272. 9:28:53All right. So we have a a big agenda.
  12273. 9:28:56This is a big notebook um to go through
  12274. 9:28:58a lot of material here surrounding
  12275. 9:29:02regression. So we're we're going to
  12276. 9:29:04start with linear regression and see um
  12277. 9:29:07how we actually perform it, what that
  12278. 9:29:10model is doing. Um which we've kind of
  12279. 9:29:13seen the idea of it a little bit
  12280. 9:29:14already, so it should be somewhat
  12281. 9:29:16familiar. Um and then we'll talk about
  12282. 9:29:19how to adapt that linear regression idea
  12283. 9:29:22to um nonlinear what's called nonlinear
  12284. 9:29:26regression which is going to be using
  12285. 9:29:27like polomial uh features. We'll talk
  12286. 9:29:31about how to do that. Um and then a big
  12287. 9:29:34big big topic for us is going to be
  12288. 9:29:36evaluating the model. So it'll be it'll
  12289. 9:29:39be quite easy to actually build it.
  12290. 9:29:41building the model will be really easy
  12291. 9:29:43but evaluating and interpreting that
  12292. 9:29:46will be uh a lot of interesting work
  12293. 9:29:50there um because we want to know what
  12294. 9:29:53the performance of that model is once we
  12295. 9:29:55have it built right we want to know how
  12296. 9:29:57good of a model is it is it worth using
  12297. 9:30:00or do we need to retrain it or get new
  12298. 9:30:02data or change the model up to talk
  12299. 9:30:05about that um how do you determine what
  12300. 9:30:07to do based on that performance
  12301. 9:30:10um and then we'll We'll talk about here
  12302. 9:30:13um a couple things. We may not get to
  12303. 9:30:14this today, but regularization
  12304. 9:30:17which is used to boost the performance
  12305. 9:30:19uh in certain situations um whenever the
  12306. 9:30:23model is kind of performing um poorly
  12307. 9:30:27against test data even though it
  12308. 9:30:28performs pretty well on training data.
  12309. 9:30:31In that scenario, you can use offshoots
  12310. 9:30:34of linear regression that do some uh
  12311. 9:30:36what's called regularization. We'll talk
  12312. 9:30:38about that.
  12313. 9:30:39Um, and then we'll talk about
  12314. 9:30:41hyperparameter tuning, uh, generally as
  12315. 9:30:44a strategy, which is something you
  12316. 9:30:46generally do want to do when you're
  12317. 9:30:47training machine learning models. Um, so
  12318. 9:30:51again, these two we may not get to
  12319. 9:30:53today, but um quite a quite a lot to get
  12320. 9:30:57to be prior to that mainly centered
  12321. 9:31:00around evaluation and building linear
  12322. 9:31:03regression.
  12323. 9:31:05Okay, so pretty cool. we'll get to our
  12324. 9:31:07first kind of model here. This linear
  12325. 9:31:09regression
  12326. 9:31:10to start with.
  12327. 9:31:14Okay,
  12328. 9:31:16so let's start with uh linear regression
  12329. 9:31:19here. Um, and really what linear
  12330. 9:31:24regression is attempting to do and I
  12331. 9:31:28want to show you this in this picture is
  12332. 9:31:31draw this line sometimes what is known
  12333. 9:31:34as the line of best fit. So this is our
  12334. 9:31:37model that kind of goes through the data
  12335. 9:31:40and it's generally a good predictor
  12336. 9:31:44um because if you give me um features uh
  12337. 9:31:50if you give me new features and let's
  12338. 9:31:53say they are let's say you give me a
  12339. 9:31:55feature that's right here.
  12340. 9:31:57So you say, okay, I have a feature
  12341. 9:31:59that's this value on the x- axis. Then I
  12342. 9:32:02know all I have to do is plug that into
  12343. 9:32:05my line equation, and I will generate a
  12344. 9:32:09a value that's like right here.
  12345. 9:32:12Okay, that's pretty that's on that line
  12346. 9:32:14at that input. And that's going to be my
  12347. 9:32:17prediction for what the output variable
  12348. 9:32:19should be. It's just going to be
  12349. 9:32:20something on that line. And what you can
  12350. 9:32:23see is this line is a decent estimate
  12351. 9:32:27for this data because it slices through
  12352. 9:32:30this pretty evenly. So it's a good guess
  12353. 9:32:32as to what the output should be given
  12354. 9:32:36any one of these inputs. It's a it's a
  12355. 9:32:38good estimator this line. And so our
  12356. 9:32:41goal building a linear regression is to
  12357. 9:32:43kind of build the equation of this line.
  12358. 9:32:46So we want this equation.
  12359. 9:32:50Equation of this line
  12360. 9:32:54is going to be our model.
  12361. 9:33:01Yes, it's going to look just like that.
  12362. 9:33:03MX plus B or yeah, MX plus C. It's going
  12363. 9:33:05to look exactly like that. uh except
  12364. 9:33:08that it's going to be more than just MX
  12365. 9:33:12because we have um generally more than
  12366. 9:33:16one feature. So you think of X as a
  12367. 9:33:17feature um it will be more than just MX.
  12368. 9:33:20It will generally be like uh it'll
  12369. 9:33:23generally look like this
  12370. 9:33:30and then plus maybe some bias here plus
  12371. 9:33:34an intercept. Yeah, it'll generally look
  12372. 9:33:36like that. So, yeah, you're exactly
  12373. 9:33:38right. MX plus B is the right idea.
  12374. 9:33:41Exactly right.
  12375. 9:33:44It'll generally look like that.
  12376. 9:33:48Nonlinear, it can be adapted to
  12377. 9:33:50nonlinear. Yeah. If we transform, we're
  12378. 9:33:52going to talk about that. If we
  12379. 9:33:54transform all of our features in a
  12380. 9:33:56nonlinear way, um we can apply linear
  12381. 9:33:59regression to it. Yes. And and that
  12382. 9:34:01would be a nonlinear regression. So yes,
  12383. 9:34:05we can do nonlinear things too.
  12384. 9:34:08We'll talk about that.
  12385. 9:34:18Okay. So linear regression again is the
  12386. 9:34:21art or science I should say not really
  12387. 9:34:24art but it is an exact science of
  12388. 9:34:26finding the equation of this line that
  12389. 9:34:29fits through this data. Um now why one
  12390. 9:34:32thing you should be thinking about is
  12391. 9:34:34why is this line a good predictor and
  12392. 9:34:38the argument is that if you take a look
  12393. 9:34:40at this distance from these blue points
  12394. 9:34:42so let's say these blue points are
  12395. 9:34:44actual data points this line is going to
  12396. 9:34:48be found such that it minimizes this
  12397. 9:34:52distance
  12398. 9:34:54from the points to actually I should
  12399. 9:34:57draw it this way from the points to the
  12400. 9:34:59line. So, we want this distance to be um
  12401. 9:35:04actually I should draw it that way. This
  12402. 9:35:06way. We want this distance to be kind of
  12403. 9:35:09at a minimum. So, it would be bad to
  12404. 9:35:12draw a line all the way out here because
  12405. 9:35:14then that's a lot of distance, right?
  12406. 9:35:16So, and that would be a lot of error um
  12407. 9:35:18contributed from not being able to
  12408. 9:35:20predict those points in our data set
  12409. 9:35:22very well. Um which is our training
  12410. 9:35:25data. That's why we have labels, right?
  12411. 9:35:27that that guide us in building this
  12412. 9:35:29line. Um so our goal is to build that
  12413. 9:35:32line especially so that this error or
  12414. 9:35:36this distance can be as minimum as
  12415. 9:35:39possible. Right? Which are all these
  12416. 9:35:42distances from these points to the line.
  12417. 9:35:45We want those to be as minimum as
  12418. 9:35:48possible. So our goal is to find this
  12419. 9:35:50equation.
  12420. 9:35:52So we're going to build a model that's
  12421. 9:35:54going to find this equation.
  12422. 9:35:58of the line
  12423. 9:36:01um such that our error
  12424. 9:36:06is minimal.
  12425. 9:36:09And what is the error? The error is the
  12426. 9:36:12distance
  12427. 9:36:16of our data points
  12428. 9:36:22to to
  12429. 9:36:26the line that we build. So essentially
  12430. 9:36:28what we'll do in order to train this
  12431. 9:36:30will be to adjust the parameters or the
  12432. 9:36:33or in that like I think is really good
  12433. 9:36:36you brought up the MX plus C. Basically
  12434. 9:36:38the M and the C will adjust. So we
  12435. 9:36:41adjust those accordingly to make this
  12436. 9:36:44distance as small as possible.
  12437. 9:36:48Okay? To minimize that distance as much
  12438. 9:36:50as possible.
  12439. 9:36:59Okay.
  12440. 9:37:01So um where is regression used? We've
  12441. 9:37:05already seen some examples. Here's some
  12442. 9:37:07more uh advertising like predicting
  12443. 9:37:10sales, predicting um oil and uh oil
  12444. 9:37:15production and demand. Those are like
  12445. 9:37:17forecast those are regression problems.
  12446. 9:37:20Um retail like demand forecasting for
  12447. 9:37:22inventory. Um healthcare predicting um
  12448. 9:37:27uh the levels of certain um uh blood
  12449. 9:37:32markers or you know something like that.
  12450. 9:37:34um real estate predicting prices based
  12451. 9:37:37on those uh talked about like square
  12452. 9:37:40footage, bedrooms, bathrooms, those
  12453. 9:37:42things. So regression is used again
  12454. 9:37:44whenever we want to predict a number a
  12455. 9:37:46numerical output um that's a regression
  12456. 9:37:49problem.
  12457. 9:37:55So this kind of regression we're talking
  12458. 9:37:57about here is generally
  12459. 9:38:01um known as uh a when that equation is
  12460. 9:38:05linear that is known as a linear
  12461. 9:38:08regression. So go back to that picture
  12462. 9:38:10when we have a when that equation of the
  12463. 9:38:13line that we find is a linear equation
  12464. 9:38:16meaning that it is exactly the form I've
  12465. 9:38:20been telling you. So it's it's something
  12466. 9:38:22like um weight time feature
  12467. 9:38:26plus weight time feature
  12468. 9:38:29plus weight time feature
  12469. 9:38:33and then maybe some intercept um term
  12470. 9:38:37like some some bias term there.
  12471. 9:38:40Um this is a linear equation because all
  12472. 9:38:44of the features are to the single power.
  12473. 9:38:47So it's a linear power and this is a
  12474. 9:38:49linear combination of features with with
  12475. 9:38:52those different weights. So this is a
  12476. 9:38:54linear model
  12477. 9:38:56because it is uh it's what in math we
  12478. 9:39:00would call this a linear equation right
  12479. 9:39:03everything is to the first power. It
  12480. 9:39:05resembles mx plus b. It is a linear
  12481. 9:39:07equation or linear model. Um so when we
  12482. 9:39:13talk about linear regression that is a
  12483. 9:39:16regression model so we're predicting
  12484. 9:39:17some continuous target that assumes we
  12485. 9:39:21are model our model is formed from this
  12486. 9:39:25kind of equation a linear equation.
  12487. 9:39:28So this is going to be our our model for
  12488. 9:39:31a linear
  12489. 9:39:33uh regression.
  12490. 9:39:37Okay.
  12491. 9:39:40And so when you when you train a linear
  12492. 9:39:42regression, your goal is to learn these
  12493. 9:39:45weights so that you can plug in um you
  12494. 9:39:49can plug in any one of your uh input
  12495. 9:39:52features and you um can generate a
  12496. 9:39:56prediction. You can which is going to be
  12497. 9:39:58something on that line, right? It's
  12498. 9:40:00going to be a value that's sitting here
  12499. 9:40:02on this line.
  12500. 9:40:04We put in all of our features and we end
  12501. 9:40:06up there somewhere on that line.
  12502. 9:40:10This output.
  12503. 9:40:14Okay.
  12504. 9:40:25Okay. Let me pause there. Any questions
  12505. 9:40:27on the linear model here or why it's
  12506. 9:40:32called linear regression?
  12507. 9:40:47Okay. And by the way in these notes um
  12508. 9:40:51this bullet point here where it says it
  12509. 9:40:52uses the least squares criterion to
  12510. 9:40:54estimate the coefficients that is
  12511. 9:40:56exactly what I said earlier with the
  12512. 9:40:58distance. So the distance is based on
  12513. 9:41:00the square
  12514. 9:41:02of this this quantity like how far away
  12515. 9:41:05you are from the line is based on this
  12516. 9:41:07square distance here and here and here
  12517. 9:41:11and here. So what we're trying to do is
  12518. 9:41:14find the least distance or least squares
  12519. 9:41:18which is that minimum distance. So
  12520. 9:41:20that's how we find all of these weights
  12521. 9:41:23is from minimize. We basically tune them
  12522. 9:41:25enough using our labels. So here's our
  12523. 9:41:29label which is the y. We basically plug
  12524. 9:41:32in our data and tune those enough to
  12525. 9:41:34minimize the error. It's it's a it's an
  12526. 9:41:36optimization problem, right? We we're
  12527. 9:41:39trying to find the minimum of this
  12528. 9:41:44quantity which is that best fit line.
  12529. 9:41:59Okay.
  12530. 9:42:01So we have linear regression
  12531. 9:42:04um and we can do a simple linear
  12532. 9:42:08regression that only has one feature. So
  12533. 9:42:11if it only has one feature that's
  12534. 9:42:13exactly the so if there's only one input
  12535. 9:42:16feature sometimes that is known as um
  12536. 9:42:19simple regression or simple linear
  12537. 9:42:21regression and there's only basically
  12538. 9:42:23there's only one feature. So one
  12539. 9:42:24independent variable is the feature.
  12540. 9:42:28There's only one feature. And so this
  12541. 9:42:31equation resembles the
  12542. 9:42:36exact equation that you guys just put in
  12543. 9:42:38there, which is um mx plus b,
  12544. 9:42:42right? It resembles exactly that. um
  12545. 9:42:45we're just using different symbols for
  12546. 9:42:46those like beta beta 0 and beta 1 but um
  12547. 9:42:50basically exactly that simple line
  12548. 9:42:54there's only one feature. So and that's
  12549. 9:42:57because that line is going to um that
  12550. 9:43:01line is going to be generated uh
  12551. 9:43:04according to that equation. So here's
  12552. 9:43:05kind of what it looks like.
  12553. 9:43:08This is the best fit line through all of
  12554. 9:43:10these blue dots. This is something we're
  12555. 9:43:12going to be able to build. we're going
  12556. 9:43:14to be able to build that equation um
  12557. 9:43:16pretty easily in scikitlearn.
  12558. 9:43:20So we'll be able to find that um and it
  12559. 9:43:22won't be too hard. So this line will be
  12560. 9:43:26um y = beta 0 plus beta 1. So some
  12561. 9:43:31weight beta 1 times the only feature we
  12562. 9:43:35have x1.
  12563. 9:43:37Okay. So in this case um we would be
  12564. 9:43:41predicting sales. So sales would be the
  12565. 9:43:43value basically the label that we're
  12566. 9:43:45trying to predict and the feature that
  12567. 9:43:48we're putting in is uh I think it's the
  12568. 9:43:53number of TV expenses. Yep. TV expenses
  12569. 9:43:58which is on the x- axis. So there's one
  12570. 9:43:59feature which is um TV expense.
  12571. 9:44:08So um on this graph this would be this
  12572. 9:44:12would be our model.
  12573. 9:44:14Okay that would be our model. We only
  12574. 9:44:16have one feature and we have um these
  12575. 9:44:20two weights. We have an intercept B 0
  12576. 9:44:22and or beta 0 and then a one weight
  12577. 9:44:25which gets applied to that one feature
  12578. 9:44:28beta 1. And so our model would have
  12579. 9:44:31certain value for beta 0 and a certain
  12580. 9:44:33value for beta 1. That's what get that's
  12581. 9:44:36these guys get learned
  12582. 9:44:40learned during
  12583. 9:44:44model
  12584. 9:44:47training.
  12585. 9:44:53Okay. So those are what get learned
  12586. 9:44:56during our model training and they get
  12587. 9:44:58learned by a a a least what's called a
  12588. 9:45:01lease squares algorithm that is trying
  12589. 9:45:03to minimize that distance. It tries to
  12590. 9:45:05tweak beta 0 beta 1 to minimize this
  12591. 9:45:08distance of this line
  12592. 9:45:11um this line
  12593. 9:45:13to all of these points
  12594. 9:45:18trying to minimize this.
  12595. 9:45:22So imagine taking a line and kind of
  12596. 9:45:24moving it around and turning its its
  12597. 9:45:28slope, its angle um to try to find that
  12598. 9:45:31best fit,
  12599. 9:45:33which reduces that error the most.
  12600. 9:45:35Right? That's kind of what we're doing.
  12601. 9:46:00Uh can I explain? Yeah. So uh sales is
  12602. 9:46:04in dollars and and TV expense
  12603. 9:46:08um
  12604. 9:46:10uh
  12605. 9:46:12TV actually I think it's the other way
  12606. 9:46:13around. I think the sales is actually a
  12607. 9:46:15quantity. So I this is number of sales
  12608. 9:46:18that we have and TV expense is um I
  12609. 9:46:22think I think it's in dollars. So how
  12610. 9:46:25much money how much expense um did we
  12611. 9:46:28put into the into the product and then
  12612. 9:46:31this is how many sales did we have of
  12613. 9:46:33that product.
  12614. 9:46:37So I think it's the other way around.
  12615. 9:46:44But what this what this graph is showing
  12616. 9:46:46is the blue points are our actual data
  12617. 9:46:51points. Okay. So so we have a collection
  12618. 9:46:54like we have a data frame that has so
  12619. 9:46:57imagine we had a data frame that has the
  12620. 9:46:59uh true values.
  12621. 9:47:02So it has the um TV expenses. Um, it has
  12622. 9:47:06points that are like one. So, it has
  12623. 9:47:10points that are like 120 and then the
  12624. 9:47:13sale sales could be like 700
  12625. 9:47:16700 units, let's say. And then it has um
  12626. 9:47:20so this is just our data set, right?
  12627. 9:47:21This would be like in a data frame that
  12628. 9:47:23we have. And then we had ones that were
  12629. 9:47:26um 50 and then this could be um this
  12630. 9:47:30could be 400, let's say. And on and on
  12631. 9:47:33and on, right? So this is our data and
  12632. 9:47:36this data is plotted in the blue. So
  12633. 9:47:38these are these blue points here,
  12634. 9:47:41right? So these are the blue points here
  12635. 9:47:43and the red points are is our model. So
  12636. 9:47:47we built a linear regression model
  12637. 9:47:50um where we are putting in some values.
  12638. 9:47:53We're putting in some e fake x values
  12639. 9:47:56here and generating some predictions
  12640. 9:47:58which is this line
  12641. 9:48:01this linear uh regression line.
  12642. 9:48:04Right? And that line is derived from
  12643. 9:48:07this data. Right? It gets learned from
  12644. 9:48:10this supervised uh examples.
  12645. 9:48:15Does that make sense?
  12646. 9:48:17That line is derived from the data. It's
  12647. 9:48:20actually um learned from like the line
  12648. 9:48:23of best fit is learned from that data
  12649. 9:48:30and the actual data is in the blue.
  12650. 9:48:34So you can see we're trying to build
  12651. 9:48:35this such that this distance is kind of
  12652. 9:48:37a minimum.
  12653. 9:48:39So it's an optimal fit
  12654. 9:48:44to balance out these distances.
  12655. 9:48:54So it's just plotting. So it's just
  12656. 9:48:55building that relationship between the
  12657. 9:48:57input and output. like when the when the
  12658. 9:48:59expenses are higher, um we seem to have
  12659. 9:49:01more sales.
  12660. 9:49:30Uh what's perpendicular like the
  12661. 9:49:32distance? This should be this should be
  12662. 9:49:34perpendicular because it's a distance
  12663. 9:49:35here.
  12664. 9:49:39Is that what you mean? Like the distance
  12665. 9:49:40from the real points to the line. Yeah,
  12666. 9:49:42that should be perpendicular
  12667. 9:49:44because it's it's a it's a distance
  12668. 9:49:46formula.
  12669. 9:50:06Okay.
  12670. 9:50:10All right. So, more generally now do do
  12671. 9:50:14we usually have one feature? No. So
  12672. 9:50:17generally we expand this to the more
  12673. 9:50:21general case where we have more than one
  12674. 9:50:24feature like what we see in the housing
  12675. 9:50:26data right where we could predict a
  12676. 9:50:28price but we have many different inputs
  12677. 9:50:30like bedrooms, bathrooms, square footage
  12678. 9:50:33etc.
  12679. 9:50:35So more broadly
  12680. 9:50:39instead of simple linear regression we
  12681. 9:50:42have what's known as multiple linear
  12682. 9:50:44linear regression which means we have
  12683. 9:50:46multiple variables or multiple features.
  12684. 9:50:49Um so this is exactly the equation I've
  12685. 9:50:52been talking about. Um so we just extend
  12686. 9:50:56that that one into many features. So
  12687. 9:50:59which is this case and then a intercept
  12688. 9:51:02term which is uh um there as sometimes
  12689. 9:51:06known as the bias. Um
  12690. 9:51:10but this is the intercept term to kind
  12691. 9:51:12of orient the line to start out in the
  12692. 9:51:14right place. Um and uh but this is the
  12693. 9:51:20um this is the equation that we would be
  12694. 9:51:23building the model. This is our model
  12695. 9:51:25essentially, right? This is the equation
  12696. 9:51:26we would be learning.
  12697. 9:51:29Intercept is like a constant. Yeah. So
  12698. 9:51:31if if all of the features were zero, um
  12699. 9:51:33this is what our our data would be. This
  12700. 9:51:36is what our result would be. If
  12701. 9:51:38basically if this was zero, this was
  12702. 9:51:39zero, this was zero, it would reduce to
  12703. 9:51:42this as the prediction. Yeah. It's like
  12704. 9:51:44a constant. Yes.
  12705. 9:51:53So in in geometry, the intercept is
  12706. 9:51:56actually really important because it it
  12707. 9:51:57orients where your line should start. So
  12708. 9:51:59it orients like so so these values are
  12709. 9:52:02kind of like the slope. They orient the
  12710. 9:52:04tilt of it. Like should it be tilted
  12711. 9:52:07like this or should it be more sloped?
  12712. 9:52:09But the intercept orients where it
  12713. 9:52:12should start like vertically like should
  12714. 9:52:14it start all the way up here? Should it
  12715. 9:52:16start more down here?
  12716. 9:52:18Um, that's what the intercept kind of
  12717. 9:52:21tells us.
  12718. 9:52:31Okay, so this is the situation. This is
  12719. 9:52:34going to be our linear regression model
  12720. 9:52:35that we will be building most of the
  12721. 9:52:37time because we will have again these
  12722. 9:52:39are all going to be features.
  12723. 9:52:41So this is some feature the X this is
  12724. 9:52:45some feature this is some feature
  12725. 9:52:49X1 etc. These are all features and what
  12726. 9:52:53gets learned during the training are
  12727. 9:52:55these coefficients. So all of these
  12728. 9:52:56coefficients including the beta 0ero um
  12729. 9:53:01will get learned. So these will get
  12730. 9:53:03learned
  12731. 9:53:05um from our data right they get learned
  12732. 9:53:07they will be trained from our data um in
  12733. 9:53:11order and and how do they get trained
  12734. 9:53:13it's from reducing that distance we try
  12735. 9:53:16to get that line of best fit by tweaking
  12736. 9:53:18those betas enough to uh until we reach
  12737. 9:53:22a minimum distance but there's there's
  12738. 9:53:24an algorithm behind that um that that
  12739. 9:53:26scikitlearn will run for us to find that
  12740. 9:53:30best fit Um, so we don't need to do that
  12741. 9:53:33manually, but that's that's the process
  12742. 9:53:35is basically tweaking those weights to
  12743. 9:53:38end up with that line of best fit. So in
  12744. 9:53:41higher dimensions, instead of a line,
  12745. 9:53:43you get more of what's called a plane
  12746. 9:53:46here. Um, which kind of looks like this.
  12747. 9:53:48So the best fit is actually this plane
  12748. 9:53:51where all um, it kind of dissects all
  12749. 9:53:54these points just like that um, in
  12750. 9:53:57higher dimensions. So this is uh instead
  12751. 9:53:59of a line you get this in in three
  12752. 9:54:01dimensions you get this plane like this
  12753. 9:54:04but it's still it's like a line of best
  12754. 9:54:07it's just a more general line of best
  12755. 9:54:09fit. It's still the same idea. Um we're
  12756. 9:54:12still trying to um come up with the best
  12757. 9:54:15coefficients to minimize that distance
  12758. 9:54:17from our from our points to the line.
  12759. 9:54:21Although in higher dimensions it's no
  12760. 9:54:23longer a line. It's more like a plane
  12761. 9:54:24like this. So you're trying to minimize
  12762. 9:54:26this distance from here down to the
  12763. 9:54:28plane
  12764. 9:54:29here up to the plane
  12765. 9:54:32in higher dimensions. So I want you to
  12766. 9:54:34keep in mind what we're trying to do
  12767. 9:54:37before we go into the code because the
  12768. 9:54:39code's going to make it seem really
  12769. 9:54:40really simple and that's because
  12770. 9:54:42scikitlearn is great and that's what it
  12771. 9:54:44does.
  12772. 9:54:46But we should realize that there's
  12773. 9:54:47something really complex going on which
  12774. 9:54:49is again finding the best value of these
  12775. 9:54:53weights
  12776. 9:54:55that minimizes the distance of this line
  12777. 9:54:58to the data points that we have. So
  12778. 9:55:02there's an algorithm there that will
  12779. 9:55:04keep trying to make adjustments to this
  12780. 9:55:07based on those distances. So it's going
  12781. 9:55:10to use those distances as a guide to
  12782. 9:55:12kind of tweak them to find the one that
  12783. 9:55:16results in the lowest amount of
  12784. 9:55:18distance. So we keep making tweaks, keep
  12785. 9:55:19making tweaks, keep making tweaks and
  12786. 9:55:21eventually we try to find we converge to
  12787. 9:55:24the set of weights that gives us that
  12788. 9:55:26best fitting line. Um and and there's an
  12789. 9:55:30algorithm there that occurs. Now luckily
  12790. 9:55:33that gets abstracted for us a bit behind
  12791. 9:55:36um scikitlearn
  12792. 9:55:38um finding that best fit. So there'll be
  12793. 9:55:41a function that we use in scikitlearn
  12794. 9:55:44when we build the model that will go
  12795. 9:55:46ahead and find the best weights for us
  12796. 9:55:50and that's then we now have our optimal
  12797. 9:55:53model right that then we can just plug
  12798. 9:55:55in different values of these features
  12799. 9:55:58and generate a prediction which is going
  12800. 9:56:01to be this uh result right so so that's
  12801. 9:56:04what we're ultimately trying to do is uh
  12802. 9:56:07train the model which will uh find all
  12803. 9:56:10those optimal weights and then uh we can
  12804. 9:56:13predict with it which would be plugging
  12805. 9:56:15in different feature values to to
  12806. 9:56:18generate a prediction.
  12807. 9:56:21Okay,
  12808. 9:56:23so let's see how that happens. It's
  12809. 9:56:24actually going to be super easy um with
  12810. 9:56:27scikitlearn.
  12811. 9:56:29So uh in this scenario we have um we're
  12812. 9:56:33going to import our pandas because we're
  12813. 9:56:35going to load our data from that. Um, so
  12814. 9:56:38of course we need some data to work
  12815. 9:56:40with. So we're going to load this uh
  12816. 9:56:42CSV.
  12817. 9:56:43Um, I
  12818. 9:56:46uh so I was not actually able to find
  12819. 9:56:49this CSV for this example, but I mean
  12820. 9:56:52that's okay because we'll do some we'll
  12821. 9:56:53do other examples where we'll work with
  12822. 9:56:55the data. If you happen to have it, um,
  12823. 9:56:58great. I didn't see it in in my files.
  12824. 9:57:02So just have to take the word for it
  12825. 9:57:03that these are the this is that TV and
  12826. 9:57:06sales columns here um from this data
  12827. 9:57:10set.
  12828. 9:57:11Okay. Um as an example. So um just to
  12829. 9:57:17see how it's fit um what we're going to
  12830. 9:57:21do and this is going to be a very
  12831. 9:57:24standard process for us for building a
  12832. 9:57:27model. These steps are going to be very
  12833. 9:57:29very standard for us which is going to
  12834. 9:57:31be first of all splitting the features
  12835. 9:57:35away from the label. That's the first
  12836. 9:57:38step that we always will take. So if you
  12837. 9:57:40take a look at this code, it's taking
  12838. 9:57:43all rows but only the first column.
  12839. 9:57:48Okay, so it's extracting all the
  12840. 9:57:50features from the dataf frame um which
  12841. 9:57:52happen to be which is just the first the
  12842. 9:57:55first column uh which is the TV uh
  12843. 9:57:59column right just that column there and
  12844. 9:58:02our target variable which is our label.
  12845. 9:58:06So our target variable aka the label um
  12846. 9:58:10is the second column, right? It's that
  12847. 9:58:14that sales column.
  12848. 9:58:17Um and so our first step here, let me
  12849. 9:58:21call that out here. First step is to
  12850. 9:58:24always split apart
  12851. 9:58:28features from labels.
  12852. 9:58:31Okay, so we put all those features into
  12853. 9:58:34a data frame called X and we have all of
  12854. 9:58:37our labels into technically a series but
  12855. 9:58:40uh sort of like a data frame, right? Um
  12856. 9:58:44called Y, which is just the um which is
  12857. 9:58:48just the uh uh labels. So that's just
  12858. 9:58:52the TV values. Um now you're going to
  12859. 9:58:55see why we do that. It's because we need
  12860. 9:58:59um our our features and labels split
  12861. 9:59:02apart to put them into the model
  12862. 9:59:04building function. It expects our
  12863. 9:59:08independent variables or our features to
  12864. 9:59:10be separated from our answers or our
  12865. 9:59:14labels that guide the model building.
  12866. 9:59:17That's the first thing you got to do is
  12867. 9:59:18separate those.
  12868. 9:59:20Okay, so this code will separate those
  12869. 9:59:22out into a capital X and a lowercase Y.
  12870. 9:59:26And that's actually pretty industry
  12871. 9:59:27standard notation. Whenever you split
  12872. 9:59:30apart all your features, usually you put
  12873. 9:59:33them into a data frame called capital X
  12874. 9:59:35and then you have a lowercase Y to
  12875. 9:59:38represent your labels. That's actually
  12876. 9:59:40pretty standard.
  12877. 9:59:42So it's pretty standard that um X
  12878. 9:59:46represents
  12879. 9:59:49features
  12880. 9:59:50and
  12881. 9:59:52Y represents labels
  12882. 9:59:57label column
  12883. 9:59:59whatever our label column is in this
  12884. 10:00:01case it is the sales because we're going
  12885. 10:00:04to be predicting sales
  12886. 10:00:09using the TV column the TV quant uh
  12887. 10:00:12expense quantity.
  12888. 10:00:19Yeah. So what it so the assignment is
  12889. 10:00:23that we are um the assignment is that we
  12890. 10:00:27are
  12891. 10:00:29uh we are um splitting apart our data.
  12892. 10:00:34So that when we first read in the data
  12893. 10:00:37um it is a data frame right that has two
  12894. 10:00:40columns TV and sales.
  12895. 10:00:44Oh perfect thank you Tim. I will I will
  12896. 10:00:48go ahead and so if we look at this data
  12897. 10:00:52it only has those two columns right it
  12898. 10:00:55only has those two columns. Okay. So
  12899. 10:00:59what we're doing with this is we are
  12900. 10:01:02splitting apart
  12901. 10:01:04our our independent variable our
  12902. 10:01:07features. So this this X will contain
  12903. 10:01:11our features
  12904. 10:01:15and Y will contain
  12905. 10:01:19our label.
  12906. 10:01:23Does that make sense? We're splitting
  12907. 10:01:24this data apart. So, we're only grabbing
  12908. 10:01:26that first column here to be our
  12909. 10:01:29features. And then we're we're grabbing
  12910. 10:01:32the second column, which is the sales,
  12911. 10:01:33because we're going to predict the
  12912. 10:01:34sales. This is our label. We're going to
  12913. 10:01:37we're going to build a model to predict
  12914. 10:01:38the sales given the TV input, TV expense
  12915. 10:01:43input. So, the first thing we have to do
  12916. 10:01:46is split apart the features and the
  12917. 10:01:48label.
  12918. 10:01:51Okay, that's the first step we usually
  12919. 10:01:52will take. And the reason we have to do
  12920. 10:01:55that um just to reiterate the reason we
  12921. 10:01:58have to do that is because our model
  12922. 10:02:02will expect our our data features to be
  12923. 10:02:05separate from the label. We will pass
  12924. 10:02:07those in separately.
  12925. 10:02:10X is TV. It's the first column
  12926. 10:02:14because we're using right. It's the it's
  12927. 10:02:16all rows but the first column
  12928. 10:02:19which is TV.
  12929. 10:02:24Why is sales? This is what we're
  12930. 10:02:26predicting.
  12931. 10:02:28We are predicting the sales given the TV
  12932. 10:02:32expense value.
  12933. 10:02:37Yeah. Which is why we split it into So
  12934. 10:02:40this is the second column, right? The
  12935. 10:02:42index one column.
  12936. 10:02:47Uh you just put in read CSV and pass in
  12937. 10:02:50the URL.
  12938. 10:02:52So you could So exactly the code that
  12939. 10:02:54was up earlier from temp
  12940. 10:02:57um you just do this
  12941. 10:03:01and then data equals ddread CSV URL.
  12942. 10:03:09So we split our data into X and Y here.
  12943. 10:03:13All right. Now, one other step that
  12944. 10:03:15we're going to take that's a very very
  12945. 10:03:18critical step and you're gonna we're
  12946. 10:03:19going to see this step over and over and
  12947. 10:03:23over and over again. So, splitting apart
  12948. 10:03:25into X and Y will become we'll do that
  12949. 10:03:27over and over and over and over again.
  12950. 10:03:30Not only that, but doing this next step
  12951. 10:03:34which is what's called a train test
  12952. 10:03:37split. Now, let me show you what the
  12953. 10:03:39train test split does. It takes our data
  12954. 10:03:44And it's going to split apart our data
  12955. 10:03:47that we have, our X and our Y data. It's
  12956. 10:03:50going to split it apart into a
  12957. 10:03:52percentage that will be used to train
  12958. 10:03:54the data
  12959. 10:03:57and then a percentage that will be used
  12960. 10:03:59to test. Now, why would we want to do
  12961. 10:04:02that? It's mainly so we can do
  12962. 10:04:05evaluation. So we build the model over
  12963. 10:04:08here and then we test it on data that
  12964. 10:04:11has not seen before. So we reserve a
  12965. 10:04:15percentage of the data to be used for
  12966. 10:04:16test. Usually this this data is um
  12967. 10:04:21somewhere between uh 20 to 30%.
  12968. 10:04:26So somewhere between 20 to 30% of the
  12969. 10:04:29original data. So that means the
  12970. 10:04:31majority of it is used for training. So
  12971. 10:04:34the majority of the of that X and Y over
  12972. 10:04:37here is going to be between 70 to 80%.
  12973. 10:04:42Will generally be used for for uh for
  12974. 10:04:47training. Okay. So somewhere between 20
  12975. 10:04:50to 30 the industry standard is some
  12976. 10:04:52anywhere in between there. Um a lot of
  12977. 10:04:54people like to use 30%, some people like
  12978. 10:04:56to use 20%. Um anything in that range is
  12979. 10:05:00acceptable. um we will I think we
  12980. 10:05:03generally will favor like 30%.
  12981. 10:05:06Um to be used for testing but um the the
  12982. 10:05:11point is we don't we don't want to mix
  12983. 10:05:13those together. We want those to be
  12984. 10:05:15separated out so that we can have a fair
  12985. 10:05:20evaluation, right? We want to train our
  12986. 10:05:22data on this train our model on this
  12987. 10:05:24data and then see how well it performs
  12988. 10:05:28on this data that it has never seen
  12989. 10:05:30before.
  12990. 10:05:32Right? So in order to have data it's
  12991. 10:05:34never seen before, we're going to take
  12992. 10:05:35our X and our Y and we're going to split
  12993. 10:05:37it using this function called train test
  12994. 10:05:41split that will do this kind of
  12995. 10:05:44splitting for us. Okay. So scikitlearn
  12996. 10:05:47has a function called train test split
  12997. 10:05:50that will go ahead and we're going to
  12998. 10:05:51pass our x and our y and we'll pass in a
  12999. 10:05:54percentage like 30% that we want to
  13000. 10:05:58split out into a test set and then the
  13001. 10:06:00remainder of that the 70% will be used
  13002. 10:06:03for training the model.
  13003. 10:06:07Okay.
  13004. 10:06:10So what we're going to get let me redraw
  13005. 10:06:13that. So, what we're going to get out of
  13006. 10:06:15this for the train test split is we're
  13007. 10:06:17going to we're going to have an X and a
  13008. 10:06:19Y per
  13009. 10:06:22training and test. So, we're going to
  13010. 10:06:24get now we're going to get an X train
  13011. 10:06:31and a Y train.
  13012. 10:06:35So, we're going to get training features
  13013. 10:06:36and training labels. And then we're
  13014. 10:06:39going to get test features
  13015. 10:06:44to plug into our model and and test
  13016. 10:06:47answers or test labels
  13017. 10:06:52to do evaluation because what we should
  13018. 10:06:54be able to do is build the model over
  13019. 10:06:56here and then apply the model on this
  13020. 10:06:58data. Meaning we can take these features
  13021. 10:07:01and plug it into our model and then see
  13022. 10:07:04what answers we get and compare those
  13023. 10:07:07answers to this testing data. Right? We
  13024. 10:07:10should be able to do that to generate an
  13025. 10:07:12evaluation.
  13026. 10:07:16Okay? Now you may be wondering why do we
  13027. 10:07:19do any of that? What's the purpose of
  13028. 10:07:21that?
  13029. 10:07:22Evaluating it on this test data gives us
  13030. 10:07:26a good sense of will our model
  13031. 10:07:30generalize to new examples. Right? If it
  13032. 10:07:35performs pretty well on this data,
  13033. 10:07:38that's a good signal like when it's
  13034. 10:07:39performing pretty well on data it's
  13035. 10:07:41never seen before, that's a good
  13036. 10:07:44indicator that it's going to perform
  13037. 10:07:45pretty well when we use it on brand new
  13038. 10:07:48examples
  13039. 10:07:50um in the future.
  13040. 10:07:52Right. So that's a that's why we do this
  13041. 10:07:57evaluation on this data that it has not
  13042. 10:08:00seen before. It's going to see this
  13043. 10:08:02training data, right? We're going to
  13044. 10:08:04train the model on that data. But that
  13045. 10:08:07model will never be exposed to this test
  13046. 10:08:09data until we do the evaluation
  13047. 10:08:12and and generate some metrics to see how
  13048. 10:08:15good is this performing
  13049. 10:08:18and does it have a good chance of
  13050. 10:08:19generalizing to never before seen
  13051. 10:08:22examples which is what we want right
  13052. 10:08:24because we're going to use this model in
  13053. 10:08:25the real world. It's going to be being
  13054. 10:08:28used on new examples that it hasn't seen
  13055. 10:08:30before. We want it to perform well. So,
  13056. 10:08:33this is kind of our test, our
  13057. 10:08:35evaluation.
  13058. 10:08:39Okay. Any questions on the We're going
  13059. 10:08:42to do this in a moment. I'll show you
  13060. 10:08:44what it looks like in the code, but any
  13061. 10:08:47conceptually, any questions on the train
  13062. 10:08:49test split idea? It's a very very
  13063. 10:08:52important idea that we um basically use
  13064. 10:08:57part of the data to train it and then
  13065. 10:08:58another part of it to evaluate. It's
  13066. 10:09:01very important we do that. By the way,
  13067. 10:09:03this has a term um this in machine
  13068. 10:09:06learning this is called cross
  13069. 10:09:10validation
  13070. 10:09:14because we are using one data set to
  13071. 10:09:18train the model and then we're cross
  13072. 10:09:20over we're crossing that over into
  13073. 10:09:23another data set to validate it which is
  13074. 10:09:26the uh the the testing set.
  13075. 10:09:33So this is called cross validation. Um
  13076. 10:09:35there's actually many ways to do cross
  13077. 10:09:37validation. That's something we'll
  13078. 10:09:38study. This is a very simple way of
  13079. 10:09:40doing cross validation. There's more
  13080. 10:09:41complex ways. You can take your data and
  13081. 10:09:44you can actually divide it into many
  13082. 10:09:46sections
  13083. 10:09:47and basically train it against most of
  13084. 10:09:49these and evaluate it against one at a
  13085. 10:09:51time and then rotate. So that's another
  13086. 10:09:55way to do cross validation. We're going
  13087. 10:09:56to study that. Um but this is the this
  13088. 10:09:59is the simplest way to do it here.
  13089. 10:10:08Okay.
  13090. 10:10:11So let me show you what you get when you
  13091. 10:10:12use train test split. So uh we're going
  13092. 10:10:15to import from sklearn.
  13093. 10:10:18We're uh from the model selection
  13094. 10:10:20module. Now we haven't used this before.
  13095. 10:10:23This is our first time using it. But
  13096. 10:10:24here's our model selection. We're going
  13097. 10:10:26to import this train test split function
  13098. 10:10:30and we're going to use it on our X and Y
  13099. 10:10:33and we're going to set a test size of
  13100. 10:10:3730% which is which is.3. So our test
  13101. 10:10:40size
  13102. 10:10:42is 30%.
  13103. 10:10:45Converted to decimal
  13104. 10:10:48right converted to.3. So that means
  13105. 10:10:51we're reserving 30% for that test set.
  13106. 10:10:55Um you can set a random state. Now
  13107. 10:10:57that's completely optional. Um the
  13108. 10:11:00random state
  13109. 10:11:02is for reproducibility
  13110. 10:11:10because what the train test split is
  13111. 10:11:11going to do is it's actually going to
  13112. 10:11:13shuffle the data and then split it apart
  13113. 10:11:16into the 7030.
  13114. 10:11:18So um yes, the seed. Exactly. It's like
  13115. 10:11:21a seed. So it's it's saying like when
  13116. 10:11:24you do that shuffling every time I run
  13117. 10:11:26this notebook I'm going to get the same
  13118. 10:11:28result but it's going to be random the
  13119. 10:11:30first it's going to be random but I'm
  13120. 10:11:31gonna be able to reproduce that
  13121. 10:11:33randomness with that random state. Yes,
  13122. 10:11:38it is like a seed.
  13123. 10:11:43Uh it's you can choose any number to be
  13124. 10:11:46your your um your random state. It 42
  13125. 10:11:49isn't important. You could choose zero.
  13126. 10:11:51You could choose one. Um, you could
  13127. 10:11:53choose any positive integer. Um, 42 is
  13128. 10:11:57kind of like the uh industry standard.
  13129. 10:12:01It's it's you'd have to look it up why
  13130. 10:12:03it is. Um, apparently 42 is a special
  13131. 10:12:07number. Um,
  13132. 10:12:10in in kind of the history of development
  13133. 10:12:12of this stuff, there's nothing really
  13134. 10:12:14special about 42. You could choose a
  13135. 10:12:16random you could choose a random seed to
  13136. 10:12:18be uh zero. That's fine. It it doesn't
  13137. 10:12:21really it doesn't really matter.
  13138. 10:12:25Um you just want you could choose it to
  13139. 10:12:27be uh one, two, three. Um you could
  13140. 10:12:30choose it to be 15. You can choose it to
  13141. 10:12:32be anything you want it to be. It's
  13142. 10:12:34really so that your your shuffling is
  13143. 10:12:37consistent. Every time you run this
  13144. 10:12:39notebook, you get the same shuffle
  13145. 10:12:41result. So I'm always going to get the
  13146. 10:12:43same rows in these splits.
  13147. 10:12:49Hitch. There it is. I knew it was from
  13148. 10:12:51something.
  13149. 10:12:57Yeah. So 42 is kind of like a
  13150. 10:13:02it's it's just used ubiquitously
  13151. 10:13:06uh you know as kind of a um paying
  13152. 10:13:10tribute to the Hitchhiker's Guide to the
  13153. 10:13:11Galaxy, but it's no it's there's nothing
  13154. 10:13:14that special about 42. It doesn't it's
  13155. 10:13:16not going to change our result or
  13156. 10:13:18anything.
  13157. 10:13:19It's just so that this train set split
  13158. 10:13:22is going to shuffle our data and split
  13159. 10:13:24it apart into 7030.
  13160. 10:13:27You just want to set this to something
  13161. 10:13:28so that you get a cons every time we run
  13162. 10:13:30this notebook, we get a consistent
  13163. 10:13:33shuffle.
  13164. 10:13:34And so the data in these sets
  13165. 10:13:38are uh consistent. That's all.
  13166. 10:13:47Okay. But do you guys see how we pass in
  13167. 10:13:50our X and our Y and we generate four we
  13168. 10:13:53generate four different data uh
  13169. 10:13:56quantities here which is we generate
  13170. 10:13:58training features, test features,
  13171. 10:14:01training labels and test labels because
  13172. 10:14:03again we are generating these four
  13173. 10:14:07different we're generating data on these
  13174. 10:14:09two different sets. A training set and a
  13175. 10:14:13test set. So we have training features,
  13176. 10:14:17training label,
  13177. 10:14:19and then test features, test label.
  13178. 10:14:24Okay, that's why it's so important to
  13179. 10:14:27split apart our data into the X and the
  13180. 10:14:29Y. We need those split apart in order
  13181. 10:14:32for this part to work.
  13182. 10:14:36So by the way, these two steps we will
  13183. 10:14:38always do for any model we build. We'll
  13184. 10:14:41generally do X and Y and then train test
  13185. 10:14:44split in order to generate the data that
  13186. 10:14:48we will use for building our model.
  13187. 10:14:55Okay. So this this data here is going to
  13188. 10:14:58be what we actually use to guide the
  13189. 10:15:00training of our model. So it's
  13190. 10:15:01definitely supervised, right? Linear
  13191. 10:15:03regression
  13192. 10:15:05um we we will use that
  13193. 10:15:15Okay, so we haven't built the model yet.
  13194. 10:15:17We're just getting our data split apart
  13195. 10:15:19and ready for the training. We haven't
  13196. 10:15:21actually built our model yet, right?
  13197. 10:15:24That'll be coming up uh in a moment.
  13198. 10:15:27But this is getting our data ready. We
  13199. 10:15:29started with our data frame. We split it
  13200. 10:15:31apart into uh an x and a y. And we split
  13201. 10:15:36that into a train test split. And um you
  13202. 10:15:42know then we can uh then we can go ahead
  13203. 10:15:45and um pass in to our model training
  13204. 10:15:48which we'll do in a moment.
  13205. 10:15:55Um you that's a good question. You could
  13206. 10:15:57run so what you could do is you could
  13207. 10:16:00run
  13208. 10:16:01um should we import numpy? Let's see.
  13209. 10:16:05We did. Okay. You could run the average
  13210. 10:16:09on the um you could check the MP mean on
  13211. 10:16:14the X train and see how it compares to
  13212. 10:16:18um
  13213. 10:16:20see how it compares to X.
  13214. 10:16:25So you could you could do that and see
  13215. 10:16:26what the average of this feature is um
  13216. 10:16:29compared to the average of the original.
  13217. 10:16:31They may not be perfect because we are
  13218. 10:16:33taking a reduced data set size. So I
  13219. 10:16:36don't think there's really any good
  13220. 10:16:38there's not like a one-sizefits-all
  13221. 10:16:40validation we can do because we're
  13222. 10:16:41taking a random shuffle and taking a
  13223. 10:16:44percent. We're taking 70% of the data
  13224. 10:16:46out. So we're not guaranteed to maintain
  13225. 10:16:48the same statistics. We can see if
  13226. 10:16:50they're close.
  13227. 10:16:52Um but does that make sense? Like we're
  13228. 10:16:54not guaranteed to get the same stats
  13229. 10:16:56because we're taking a slice of it.
  13230. 10:16:57We're taking 70%.
  13231. 10:17:00So it's not guaranteed to to to
  13232. 10:17:03be the same distribution really.
  13233. 10:17:10Delete that.
  13234. 10:17:13Uh is it a good practice? Yes, it is.
  13235. 10:17:17It is. Uh 30% is the industry standard.
  13236. 10:17:20Anything between 20 to 30. So 0.2.25.3
  13237. 10:17:25any of those are acceptable. It's really
  13238. 10:17:27up to you. Um I mostly see 30%.
  13239. 10:17:31Mo I think.3 is is a good good practice
  13240. 10:17:34to use for sure.
  13241. 10:17:37Um I did explain random state. Uh random
  13242. 10:17:40state is so that you get consistent
  13243. 10:17:43shuffling. Um you can set this to any
  13244. 10:17:46integer that you want it to be. It it
  13245. 10:17:48doesn't really matter. Um you can set it
  13246. 10:17:51to uh 100, you can set it to 10, you can
  13247. 10:17:54set it to 15. Um it just ensures because
  13248. 10:17:57what this split will do is it will
  13249. 10:18:00shuffle the data first. It'll shuffle
  13250. 10:18:02the rows and then um split it apart into
  13251. 10:18:05the into the train and test sets. So you
  13252. 10:18:09set the random state so that the next
  13253. 10:18:11time you run this you get the same
  13254. 10:18:13consistent shuffling. That's the only
  13255. 10:18:15that's the only thing it it helps you
  13256. 10:18:17with because it is randomized but when
  13257. 10:18:20you set a random state um it's so that
  13258. 10:18:23like if you run it again you'll get the
  13259. 10:18:25same shuffling.
  13260. 10:18:27You'll get the same the shuffling
  13261. 10:18:28matters because it it it uh dictates
  13262. 10:18:31what ends up in in these sets.
  13263. 10:18:39Okay.
  13264. 10:18:41All right. So let's see let's do let's
  13265. 10:18:44build the model
  13266. 10:18:46um and let me show you how easy this is
  13267. 10:18:49going to be to build the model and this
  13268. 10:18:50is really how it's going to be for every
  13269. 10:18:53single scikitlearn model will basically
  13270. 10:18:55look the exact same for training it
  13271. 10:18:58which is what's going to make it really
  13272. 10:19:00really nice. So the first thing we have
  13273. 10:19:02to do is import our model. So from
  13274. 10:19:05scikitlearn we're going to be using a
  13275. 10:19:07linear from the linear model package or
  13276. 10:19:10the linear model module I should say
  13277. 10:19:13within sklearn we're going to be
  13278. 10:19:15importing the linear regression
  13279. 10:19:18and we're going to create an instance of
  13280. 10:19:20the linear regression here.
  13281. 10:19:23Okay, so linear regression and look how
  13282. 10:19:27easy this is going to be. Nearly all
  13283. 10:19:31nearly all sklearn models use
  13284. 10:19:37ffit function to train.
  13285. 10:19:42So every one of them, no matter which
  13286. 10:19:44one we use, like the decision tree, like
  13287. 10:19:48the um logistic regression, any of those
  13288. 10:19:52like we use for classification that are
  13289. 10:19:53going to be coming up in lesson four,
  13290. 10:19:55they're all going to look the same in
  13291. 10:19:57terms of it's going to run.
  13292. 10:20:00Which is um scikitlearn's
  13293. 10:20:04uh generic function for training your
  13294. 10:20:06model. So this will execute the training
  13295. 10:20:10once we run this code. And what that
  13296. 10:20:13again the linear regression training is
  13297. 10:20:15going to do that least squares distance
  13298. 10:20:18procedure or algorithm to try to find
  13299. 10:20:22the right weights. It's trying to find
  13300. 10:20:25those weights that minimize that squared
  13301. 10:20:27distance uh from our line that it's
  13302. 10:20:30trying to build to the data.
  13303. 10:20:33And what I want you to notice is what we
  13304. 10:20:36put into the ffit. See how we put in the
  13305. 10:20:39training data where we put in the
  13306. 10:20:40training features and we put in the
  13307. 10:20:43training labels. Now this is supervised.
  13308. 10:20:46So of course we put in the labels,
  13309. 10:20:50right? Of course we put in these labels
  13310. 10:20:53here and of course we put in our
  13311. 10:20:55features here. So we're putting in all
  13312. 10:20:58of our examples from our training split
  13313. 10:21:03into this ffit which is going to train
  13314. 10:21:06the model uh so that we can we can use
  13315. 10:21:10it for prediction.
  13316. 10:21:12Okay, it's really fast. If I run this,
  13317. 10:21:15it's going to be pretty much instant.
  13318. 10:21:18Pretty much instantly it gets trained.
  13319. 10:21:20And you can see here we now have a
  13320. 10:21:21linear regression. you can see in this
  13321. 10:21:23little box. Um, and it and this
  13322. 10:21:26information says that it has been
  13323. 10:21:28fitted. So, it's now ready to be used,
  13324. 10:21:30right? So, we now that's it. We've
  13325. 10:21:33trained our model. We tr That's how easy
  13326. 10:21:36that was. We did fit. Now, what we
  13327. 10:21:38should realize is there's a lot of work
  13328. 10:21:41going on behind the scenes of this ffit.
  13329. 10:21:44Okay, there's a lot of work being done
  13330. 10:21:46there to do the least squares algorithm
  13331. 10:21:50and find those weights and and create
  13332. 10:21:53that line of best fit. Right? So there
  13333. 10:21:56there's a lot of work being going on
  13334. 10:21:57there that's going on there behind the
  13335. 10:21:59scenes, but scikitlearn is abstracting
  13336. 10:22:02it away for us. Right? And all we have
  13337. 10:22:04to do is fit when we're using this code.
  13338. 10:22:08Really easy. Really easy. Fit. And there
  13339. 10:22:12we go. We've trained our linear
  13340. 10:22:14regression model.
  13341. 10:22:18And by the way, if you want to see what
  13342. 10:22:20the coefficients are, you can actually
  13343. 10:22:23extract them if you do so if you take
  13344. 10:22:25your lin regression and you do um
  13345. 10:22:28coefficients like this
  13346. 10:22:32coeff with a with an underscore. So this
  13347. 10:22:36gives us the trained
  13348. 10:22:39weights coefficients
  13349. 10:22:43also known as the coefficients right.
  13350. 10:22:46Um so if you run this you can see uh
  13351. 10:22:48right now we have this coefficient here
  13352. 10:22:53um which is the only coefficient we had
  13353. 10:22:55on our feature. So we only had one
  13354. 10:22:58feature coefficient there.
  13355. 10:23:11And we can take a look at our intercept
  13356. 10:23:17which is this.
  13357. 10:23:20So this gives us the train weights
  13358. 10:23:23and so we can look at the intercept we
  13359. 10:23:25can look at the the the coefficient. Um
  13360. 10:23:30so obviously if we have multiple
  13361. 10:23:31features our model has many features
  13362. 10:23:34it's going to have more values in that
  13363. 10:23:36coefficient but the intercept is just
  13364. 10:23:38the single value 7.23
  13365. 10:23:41and then the coefficient
  13366. 10:23:45is 0.046. So that's the weight that gets
  13367. 10:23:48learned.
  13368. 10:23:52Is there a size limit? No, not really.
  13369. 10:23:54There's no size limit. Um,
  13370. 10:23:57no, you can use as much data as you
  13371. 10:23:59want.
  13372. 10:24:01There's really no size limit other than
  13373. 10:24:03what like what you can fit in memory.
  13374. 10:24:08I'd say that's the only limit is
  13375. 10:24:09basically what the amount of data that
  13376. 10:24:11can fit in memory.
  13377. 10:24:19Okay.
  13378. 10:24:22All right. Were you guys able to run
  13379. 10:24:23this? Were you guys able to run the
  13380. 10:24:24linear regression ffit?
  13381. 10:24:28Okay, perfect.
  13382. 10:24:36Perfect. You are Okay, great. Great.
  13383. 10:24:41So, we have a model and we can use it to
  13384. 10:24:44predict. Um, and so that's actually what
  13385. 10:24:46we're going to do next. If we go down
  13386. 10:24:49here, um we're going to have a function
  13387. 10:24:51that's going to um build a scatter plot
  13388. 10:24:55of our original test data.
  13389. 10:24:58Um so we're going to have our test data
  13390. 10:25:01here.
  13391. 10:25:03Um,
  13392. 10:25:05and we're going to then take our uh
  13393. 10:25:08we're going to take our training data
  13394. 10:25:10and plot we're going to use the uh this
  13395. 10:25:14data versus our sales predictions. So
  13396. 10:25:18you can see we're going to you this is
  13397. 10:25:20how by the way this is how you use the
  13398. 10:25:22scikitlearn model to predict. You have a
  13399. 10:25:25fit to train it and look at the function
  13400. 10:25:28you use to predict. It's literally just
  13401. 10:25:29called predict. That's how easy it is.
  13402. 10:25:33and you pass in your data, all your
  13403. 10:25:35features into this predict and it
  13404. 10:25:37generates a prediction for every row. So
  13405. 10:25:40every row in these features in this data
  13406. 10:25:43frame um will end up with a prediction
  13407. 10:25:47using our model. So what we're going to
  13408. 10:25:49do is plot our training date uh features
  13409. 10:25:54against the predicted sales to see how
  13410. 10:25:58good of a fit that really was.
  13411. 10:26:01Okay. to see to see the regression fit.
  13412. 10:26:08Okay. And so there's the regression fit.
  13413. 10:26:12We have all of our test data here
  13414. 10:26:14plotted in the green. We have our blue,
  13415. 10:26:16which is our um we have our our blue,
  13416. 10:26:20which is our uh um training data line
  13417. 10:26:24that we built our model on. So that's a
  13418. 10:26:26pretty decent fit. Um, and then our test
  13419. 10:26:29data is here. We just plotted in the
  13420. 10:26:32green scatter. But the thing I want you
  13421. 10:26:34to see is this prediction, right? We we
  13422. 10:26:37were able to generate some predictions
  13423. 10:26:39on that training um by running our
  13424. 10:26:43predict function with our model. Now,
  13425. 10:26:44this model has been trained. So, we've
  13426. 10:26:47already fit it and now we're using it to
  13427. 10:26:49predict, right? And so, we're predicting
  13428. 10:26:51the sales and plotting that on the
  13429. 10:26:54y-axis.
  13430. 10:26:56So the sales are we're using the
  13431. 10:26:57predicted sales there which is our blue
  13432. 10:26:59line. So this is our line of best fit.
  13433. 10:27:04So this is our model prediction.
  13434. 10:27:12This is our model predictions. Right?
  13435. 10:27:16You can see it's a pretty decent uh
  13436. 10:27:17line, right? Pretty decent line of best
  13437. 10:27:19fit. Of course, there's some error here
  13438. 10:27:23like there, you know, it's not perfect,
  13439. 10:27:25but it it does a decent job of being a
  13440. 10:27:28best fit line.
  13441. 10:27:43Okay, so look how easy that was to
  13442. 10:27:47just to recap this to fit our model was
  13443. 10:27:50a linear regression.fit. And of course,
  13444. 10:27:52we're going to do more examples. So no
  13445. 10:27:55worries uh on that. We're going to see
  13446. 10:27:57this many many many times throughout
  13447. 10:27:59this notebook. But we have linear
  13448. 10:28:02regression.fit to train it. And then we
  13449. 10:28:05have linear regression.predict
  13450. 10:28:08to and we pass in our features and that
  13451. 10:28:10generates a predicted output.
  13452. 10:28:13Right. So what this is actually doing is
  13453. 10:28:17is computing this quantity.
  13454. 10:28:32We could do either.
  13455. 10:28:34We could do either. Um, so we could do,
  13456. 10:28:39so one thing we could do is plot uh, so
  13457. 10:28:42we could swap it out. We, we could do
  13458. 10:28:44either one. It doesn't, it's not a big
  13459. 10:28:46deal to do the training set. We could
  13460. 10:28:48do, so we could plot X test and then we
  13461. 10:28:51could plot linear regression X test.
  13462. 10:29:02So it's it's a similar line. Um it's
  13463. 10:29:06just different input features, but the
  13464. 10:29:08line is going to be the same. Just
  13465. 10:29:11different inputs,
  13466. 10:29:13but the coefficients are the same,
  13467. 10:29:14right? It's the same line. It's just we
  13468. 10:29:16generate different outputs.
  13469. 10:29:22So yeah, you could do either one.
  13470. 10:29:26This is this is honestly this is
  13471. 10:29:28probably better. I see what you're
  13472. 10:29:30saying. This is probably better because
  13473. 10:29:31this is the line of best fit through
  13474. 10:29:34this data. So that probably makes sense
  13475. 10:29:36to do to do predict on the test set.
  13476. 10:29:40Agreed on that. Probably makes about
  13477. 10:29:43most sense.
  13478. 10:29:50But you could do either one.
  13479. 10:30:02Yeah, I think that would be the most I
  13480. 10:30:04think that makes the most sense is for
  13481. 10:30:05it to be on the same one just to
  13482. 10:30:07validate. So like we could do we could
  13483. 10:30:09do training here and then train and
  13484. 10:30:12train just to see how that data lines
  13485. 10:30:15up.
  13486. 10:30:16Really, what we're trying to do is have
  13487. 10:30:18our scattered data and then our line of
  13488. 10:30:20best fit on the same plot. That's all
  13489. 10:30:23we're trying to do, right? So, yeah, I
  13490. 10:30:25think I think they should be the same.
  13491. 10:30:30I think that makes sense.
  13492. 10:30:34These values
  13493. 10:30:37or which values do you want to see?
  13494. 10:30:44Yeah, we could uh we could generate
  13495. 10:30:46those if we just do um let's go down
  13496. 10:30:49here. So the the line values
  13497. 10:30:53um are going to be uh the prediction.
  13498. 10:30:57So, um the the uh test
  13499. 10:31:03predictions
  13500. 10:31:05equals um
  13501. 10:31:10test predictions equals linear
  13502. 10:31:12regression.predict x test and then we
  13503. 10:31:14could uh we could print out our test
  13504. 10:31:17predictions.
  13505. 10:31:21Yeah. So, we can see what those actual
  13506. 10:31:23values are on our uh on the test set.
  13507. 10:31:27Yeah.
  13508. 10:31:37Um, we will do that. Yeah. So, you
  13509. 10:31:39thought we were checking how well our
  13510. 10:31:41data was trained. We will do that. Yes.
  13511. 10:31:43We haven't learned how to evaluate this
  13512. 10:31:44yet. We're going to talk about that
  13513. 10:31:46coming up next. Yeah. We will do that.
  13514. 10:31:49We just haven't learned how to do proper
  13515. 10:31:51evaluation
  13516. 10:31:53of a regression model.
  13517. 10:31:55But yeah, it's something we're going to
  13518. 10:31:57talk about for sure
  13519. 10:31:59and see how to do in our code.
  13520. 10:32:06Okay.
  13521. 10:32:09All right. Any other uh questions on
  13522. 10:32:12this example?
  13523. 10:32:18Again big takeaways
  13524. 10:32:21fit to train it and then predict to use
  13525. 10:32:26it
  13526. 10:32:28predict on the features to use the model
  13527. 10:32:31and make predictions with it.
  13528. 10:32:36So here is an example we we made all the
  13529. 10:32:38predictions. This these are all the
  13530. 10:32:39values that are on that line.
  13531. 10:32:42These are all our predictions and notice
  13532. 10:32:44they this is a truly regression right?
  13533. 10:32:45These are all floatingoint values. Um,
  13534. 10:32:48so this is definitely a regression,
  13535. 10:32:50right?
  13536. 10:33:02Okay.
  13537. 10:33:10Uh, that's a good question. Um,
  13538. 10:33:14I'm not sure if there is
  13539. 10:33:18If there's like a verbose
  13540. 10:33:22there's not really no there's not really
  13541. 10:33:24a verbose you can I mean you can look at
  13542. 10:33:26the source code if you really want to
  13543. 10:33:28see you can view the source code to see
  13544. 10:33:31um how it's done I can tell you I mean
  13545. 10:33:34so generally linear regression is done
  13546. 10:33:37in two ways either you use a formula um
  13547. 10:33:40to to solve the optimization problem of
  13548. 10:33:44minimizing like this this distance from
  13549. 10:33:47the points to to the line. Um,
  13550. 10:33:51or you use something called gradient
  13551. 10:33:53descent, which is how a lot of these
  13552. 10:33:55things do it is they iterate through a
  13553. 10:33:59bunch of different iterations where they
  13554. 10:34:00update these weights according to um a
  13555. 10:34:04certain uh basically a gradient of the
  13556. 10:34:08the error function. The error function
  13557. 10:34:10in this case is the is the squared
  13558. 10:34:13distance from the line to the uh to to
  13559. 10:34:19the points.
  13560. 10:34:20So uh we can compute the gradient of
  13561. 10:34:23that and do um gradient descent. So if
  13562. 10:34:26you really want to look into it, I would
  13563. 10:34:28do some research on like linear
  13564. 10:34:30regression gradient descent.
  13565. 10:34:32Okay, linear regression gradient descent
  13566. 10:34:35to see how that's uh how that's being
  13567. 10:34:37done. Yeah, it it's it's a pretty simple
  13568. 10:34:41procedure. Um, again, you have the the
  13569. 10:34:45notion is that you want to minimize
  13570. 10:34:48minimize the loss or the error. Uh, in
  13571. 10:34:52this case, the loss is the square
  13572. 10:34:54distance. So, it's like um there's like
  13573. 10:34:58a it's a formula. It's like a sum of a
  13574. 10:35:01square distance from your prediction
  13575. 10:35:04um or your label sorry to your model
  13576. 10:35:07which is the beta 0 um plus beta 1 x1
  13577. 10:35:13plus beta 2 x2
  13578. 10:35:16etc like your model and then squared. So
  13579. 10:35:19this squared this is the squared
  13580. 10:35:21distance here and you're minimizing this
  13581. 10:35:24guy which is like a calculus problem.
  13582. 10:35:27You you find you basically find the this
  13583. 10:35:30is this is I'm getting so far into the
  13584. 10:35:32weeds of this, but this is like a
  13585. 10:35:34parabola and you work your way No, no,
  13586. 10:35:37you're good. It's it's it's a good
  13587. 10:35:39question. Um you work your way down to
  13588. 10:35:42the minimum of it. Does that make sense?
  13589. 10:35:44Like you're working your way down here
  13590. 10:35:46and you do that through a descent
  13591. 10:35:48process, like a descent iteration.
  13592. 10:35:51Um
  13593. 10:35:53so
  13594. 10:35:54that's how these are found.
  13595. 10:35:57Um, but you don't see that happening in
  13596. 10:36:01the background. But if you look at the
  13597. 10:36:02source code, it I guarantee you it would
  13598. 10:36:04be it's either going to be this or
  13599. 10:36:06they're going to use the they're going
  13600. 10:36:07to use a a a matrix formula to basically
  13601. 10:36:11solve an equation um that involves this
  13602. 10:36:17basically the derivative of this set
  13603. 10:36:19equal to zero and you find the minimum.
  13604. 10:36:22Either way, you're finding the minimum
  13605. 10:36:23of this.
  13606. 10:36:28Okay. But yeah, I don't think Psycharn
  13607. 10:36:31has like a uh maybe there's some type of
  13608. 10:36:34verbose flag you can look for.
  13609. 10:36:38I don't think they have that though. Not
  13610. 10:36:40that I've seen.
  13611. 10:36:51All right.
  13612. 10:36:53So I have uh an important um concept to
  13613. 10:36:57talk about next which is going to be uh
  13614. 10:37:00called overfitting and underfitting
  13615. 10:37:03um which is a really important concept
  13616. 10:37:05that's related to the training and test
  13617. 10:37:08data we just split apart to do
  13618. 10:37:11evaluation.
  13619. 10:37:13And um essentially the the issue with
  13620. 10:37:16machine learning is that it's not
  13621. 10:37:18perfect and it can struggle in different
  13622. 10:37:20ways. And the two ways that it primarily
  13623. 10:37:23struggles is going to be overfitting and
  13624. 10:37:24underfitting. So overfitting is a
  13625. 10:37:28situation where the model basically
  13626. 10:37:33memorizes the training data so well that
  13627. 10:37:36it's it fails to generalize to new
  13628. 10:37:40examples. So what we see with
  13629. 10:37:41overfitting is this exact sign here
  13630. 10:37:45where we have really good performance on
  13631. 10:37:47the training data. So when so when we do
  13632. 10:37:49that train test split we see a really
  13633. 10:37:51good accuracy or really low error on the
  13634. 10:37:56training data but it does not perform
  13635. 10:38:00anywhere near that on that test data
  13636. 10:38:02split. So what that means is that the
  13637. 10:38:05model is overfitting to the training
  13638. 10:38:08data. it's basically memorizing it and
  13639. 10:38:11it's not able to generalize very well.
  13640. 10:38:15Now, why does that happen? It's usually
  13641. 10:38:18because the model is way too complex.
  13642. 10:38:21And that means generally you need to do
  13643. 10:38:24something to reduce the complexity.
  13644. 10:38:27Either you need to use a simpler model
  13645. 10:38:30or you need to use some type of
  13646. 10:38:32technique to mitigate overfitting. And
  13647. 10:38:35we're going to we're going to study some
  13648. 10:38:37of those techniques coming up in this
  13649. 10:38:38notebook. Uh we might not get to it
  13650. 10:38:40today, but we're going to study
  13651. 10:38:42particularly what can we do to prevent
  13652. 10:38:44overfitting because overfitting is the
  13653. 10:38:46more common issue with machine learning
  13654. 10:38:48models. They tend to do so well at
  13655. 10:38:52learning from data that they pick up on
  13656. 10:38:54small details and patterns in the
  13657. 10:38:57training examples that they're exposed
  13658. 10:38:58to. They don't do a great job at
  13659. 10:39:01generalizing to new examples. they can
  13660. 10:39:03struggle with that. So that's
  13661. 10:39:06overfitting is struggling to generalize
  13662. 10:39:09to new examples, but you do really well
  13663. 10:39:11on your training data. So it appears
  13664. 10:39:13like you have a good model, but it it's
  13665. 10:39:16not able to go and make predictions on
  13666. 10:39:17test data very well, which means we
  13667. 10:39:20would not want to use that model in the
  13668. 10:39:21real world, right? Because it's not able
  13669. 10:39:24to generalize outside of what it's
  13670. 10:39:26already seen. And that's not a good
  13671. 10:39:28thing if we're trying to use it for real
  13672. 10:39:29world examples, right?
  13673. 10:39:32So overfitting is a real issue. Um you
  13674. 10:39:35see it all the time. I've seen it many
  13675. 10:39:37many times in the real world, real
  13676. 10:39:39industry uh work that I've done.
  13677. 10:39:42Overfitting is a is a challenge for a
  13678. 10:39:44lot of machine learning models. And so
  13679. 10:39:46we need some techniques to overcome
  13680. 10:39:49overfitting and we're going to study
  13681. 10:39:51some of those uh coming up shortly.
  13682. 10:39:55Um, one of the things that we can do,
  13683. 10:39:58one of the one of the things that we can
  13684. 10:40:00do to detect overfitting is exactly what
  13685. 10:40:03we just did, which is you split apart
  13686. 10:40:06your data into training and testing so
  13687. 10:40:08that you have a chance to do an
  13688. 10:40:10evaluation to see if you're even
  13689. 10:40:12overfitting in the first place. You want
  13690. 10:40:14to see that performance be consistent
  13691. 10:40:17from train to test, right? You want to
  13692. 10:40:20see consistency. What you don't want to
  13693. 10:40:22see is performance that drops off on the
  13694. 10:40:25test data. It's much worse. You don't
  13695. 10:40:28want to see that. That means that your
  13696. 10:40:30model is overfit uh to your training
  13697. 10:40:32data and it's not going to perform well
  13698. 10:40:34in the real world.
  13699. 10:40:37Okay. So, we're going to have a couple
  13700. 10:40:38ways to uh overcome that. Talk about
  13701. 10:40:42that. Um now, the opposite can actually
  13702. 10:40:45happen as well, which is called
  13703. 10:40:47underfitting.
  13704. 10:40:48And underfitting
  13705. 10:40:50refers to the fact that a model is too
  13706. 10:40:53simple and it actually just performs
  13707. 10:40:57poorly across the board. So if we see
  13708. 10:40:59poor performance on the training and
  13709. 10:41:04testing data, that's a good signal that
  13710. 10:41:06the model's underfit and that means it's
  13711. 10:41:10too simple usually and you should try
  13712. 10:41:12using something more complex. Um, so the
  13713. 10:41:15best way to combat underfitting is to
  13714. 10:41:17use a more complex model. And as we go
  13715. 10:41:21through and learn about the models,
  13716. 10:41:23we're going to learn about which ones
  13717. 10:41:24are simple and which ones are complex.
  13718. 10:41:26So we're going to have a scale of kind
  13719. 10:41:29of complexity. And if you're
  13720. 10:41:31underfitting, you want to bump up to the
  13721. 10:41:33to a more complex model. If you're if
  13722. 10:41:36you're overfitting, one way of combating
  13723. 10:41:38that is to actually go down to something
  13724. 10:41:40more simple. Go the opposite way to
  13725. 10:41:42something simpler. So we need to learn
  13726. 10:41:44right now we've only learned linear
  13727. 10:41:46regression
  13728. 10:41:47but we will learn other models you know
  13729. 10:41:49in the future and we'll we'll talk about
  13730. 10:41:52uh their complexity and how they're
  13731. 10:41:53related to each other.
  13732. 10:41:56Okay, but these are two issues we see
  13733. 10:41:58just to draw that out again is if we
  13734. 10:42:01have a train test split where we have
  13735. 10:42:037030 split let's say and we perform
  13736. 10:42:06really well over here but we go to apply
  13737. 10:42:08that model over here and it fails it's
  13738. 10:42:11accuracy drops off significantly more
  13739. 10:42:14error that's that's definitely
  13740. 10:42:15overfitting which is not good
  13741. 10:42:20right and then underfitting is just not
  13742. 10:42:22performing well in either case so even
  13743. 10:42:24on the training data itself self your
  13744. 10:42:26your accuracy is not very good. So
  13745. 10:42:28you're not really learning effectively.
  13746. 10:42:31You're underfitting your model. So
  13747. 10:42:34that's that's um underfitting case.
  13748. 10:42:41Okay.
  13749. 10:42:46All right. Now the issue is that it can
  13750. 10:42:49be very difficult to balance these two
  13751. 10:42:52and get it correct. That's what makes
  13752. 10:42:53machine learning a little bit
  13753. 10:42:54challenging is getting this balance
  13754. 10:42:57correct of simplicity and complexity. So
  13755. 10:43:01you don't want to be overly complex that
  13756. 10:43:03you overfit, but you don't want to be
  13757. 10:43:05overly simple that you underfit and
  13758. 10:43:08you're not able to learn effectively. So
  13759. 10:43:11there's a bit of a tradeoff there. And
  13760. 10:43:12this trade-off is typically known in the
  13761. 10:43:14community as bias variance trade-off. Um
  13762. 10:43:18in which case, uh it's basically like a
  13763. 10:43:20complexity simplicity trade-off. It's
  13764. 10:43:22another word for that. Um,
  13765. 10:43:25and so, uh, it's it's thought that, um,
  13766. 10:43:30if you, uh, if you have very, um, if you
  13767. 10:43:35have a situation where you're able to
  13768. 10:43:36fit the training data very well, you
  13769. 10:43:39risk not being able to generalize. In
  13770. 10:43:42other words, you risk overfitting, and
  13771. 10:43:44it's hard to um, it's hard to combat
  13772. 10:43:48that in a way. Um, and um, on the
  13773. 10:43:53reverse side, if you have something
  13774. 10:43:54really simple, um, you risk not learning
  13775. 10:43:58enough. Even if you're trying to combat
  13776. 10:44:00that overfitting, you risk not learning
  13777. 10:44:03enough and your model just doesn't
  13778. 10:44:05perform as well as it could. So, there's
  13779. 10:44:07a bit of a trade-off there of trying to
  13780. 10:44:09find the right balance between something
  13781. 10:44:11complex enough to learn, but something
  13782. 10:44:14not overly complex that it's going to
  13783. 10:44:17not generalize to new data. That's the
  13784. 10:44:21challenge. Um, like I said, we are going
  13785. 10:44:24to have techniques to overcome this. So
  13786. 10:44:28luckily there are things to basically
  13787. 10:44:30overcome this trade-off and um and help
  13788. 10:44:34us along the way so that we don't
  13789. 10:44:36overfit. They basically prevent
  13790. 10:44:38overfitting
  13791. 10:44:40um and allow us to use complex enough
  13792. 10:44:42models um that that won't be overfit.
  13793. 10:44:47This is in the um this was in our uh
  13794. 10:44:51lesson 3.2 notebook. So you want to pull
  13795. 10:44:54that one back up. We were working on
  13796. 10:44:55Monday.
  13797. 10:44:56Um, and just to recap this a little bit,
  13798. 10:44:59remember we were building a linear
  13799. 10:45:02regression, I wanted to recap some of
  13800. 10:45:04the steps we took there, um, that we
  13801. 10:45:08will be doing over and over again. And
  13802. 10:45:10really the same kind of steps, uh, that
  13803. 10:45:12we do here, we'll do in a lot of our
  13804. 10:45:15model building. Pretty much all of our
  13805. 10:45:17model building um, that we do, whether
  13806. 10:45:19it's regression or classification,
  13807. 10:45:21doesn't really matter. um we'll still be
  13808. 10:45:23doing a lot of these steps which are um
  13809. 10:45:27remember first we split apart our data
  13810. 10:45:29into kind of a features and a label
  13811. 10:45:33uh x and y and the reason that's
  13812. 10:45:36important is because um the model
  13813. 10:45:39training uses the features and the label
  13814. 10:45:43um to help train the model right they
  13815. 10:45:46use those separately um so we want to
  13816. 10:45:48split those apart whenever we can and so
  13817. 10:45:50we have usually Uh it's a good practice
  13818. 10:45:53to call your features capital X and your
  13819. 10:45:56labels lowercase Y. And what we do with
  13820. 10:45:59that is remember we immediately split
  13821. 10:46:02that into what we called a training and
  13822. 10:46:05a test set. And the picture we had for
  13823. 10:46:07that was something like this
  13824. 10:46:11where we had about 70% of the data
  13825. 10:46:15we used to train the model against and
  13826. 10:46:18then the other 30% of the data we use to
  13827. 10:46:21test the model against. Meaning that we
  13828. 10:46:24build a model over here and we apply it
  13829. 10:46:27to this set over here um to make
  13830. 10:46:30predictions. And then the that's where
  13831. 10:46:32the supervised learning really comes
  13832. 10:46:33into play, right? is on this test set.
  13833. 10:46:37We already have the answers. We already
  13834. 10:46:39have the label. And so we can apply our
  13835. 10:46:41model to this to the features over here.
  13836. 10:46:44Predict uh what the the label should be
  13837. 10:46:47and compare that. We can get a a metric,
  13838. 10:46:50right, that compares how close we are in
  13839. 10:46:53our prediction to the actual values. Um
  13840. 10:46:56and that was some of our performance
  13841. 10:46:58metrics. I'll recap some of those that
  13842. 10:47:00kind of measure that distance away from
  13843. 10:47:02our predictions to what the actual label
  13844. 10:47:05is. Um, but remember we had this train
  13845. 10:47:09test split function which helps us split
  13846. 10:47:12apart our features and our labels into
  13847. 10:47:15these uh four sets of data. So we have
  13848. 10:47:18our training features, our testing
  13849. 10:47:20features and then our training labels
  13850. 10:47:22and our testing labels. So we have all
  13851. 10:47:24of those and um really these two guys
  13852. 10:47:27are going to be used to train the model.
  13853. 10:47:30That's why they're called underscore
  13854. 10:47:32train. They're going to be used to train
  13855. 10:47:33that model and then the then we're going
  13856. 10:47:35to predict on these set of features and
  13857. 10:47:39then com use those predictions to
  13858. 10:47:41compare to this set of labels right
  13859. 10:47:44that's on the test test set. Um and you
  13860. 10:47:47notice here our test size is set to 30%.
  13861. 10:47:50Um, that's a pretty standard number.
  13862. 10:47:52Anywhere between like 20 to 30% is
  13863. 10:47:54pretty standard. Um, we'll typically
  13864. 10:47:57use.3, but it could be 02. Anywhere in
  13865. 10:48:00between is fine.
  13866. 10:48:04Okay, so we had that. Hopefully that uh
  13867. 10:48:06we remember that from Monday.
  13868. 10:48:09So we had a train and a test set. And
  13869. 10:48:11then building the model was actually
  13870. 10:48:13really really easy. Once you have those
  13871. 10:48:15train and test sets, um, we just import
  13872. 10:48:17our model object. So from uh scikitlearn
  13873. 10:48:20sklearn
  13874. 10:48:22um linear model uh module from that
  13875. 10:48:25package we import the linear regression
  13876. 10:48:27model and then we do um linear
  13877. 10:48:31regression.fit
  13878. 10:48:32and we pass in our features and our
  13879. 10:48:34labels and this is again this is where
  13880. 10:48:37that supervised learning is really
  13881. 10:48:38coming into play because we're passing
  13882. 10:48:41in these labels.
  13883. 10:48:43That's really what makes this work,
  13884. 10:48:44right? We need those labels to help
  13885. 10:48:46guide the model to make those updates.
  13886. 10:48:48If you guys remember, the model is
  13887. 10:48:51something that looks like this.
  13888. 10:48:57So, this was a bunch of different
  13889. 10:49:00coefficients
  13890. 10:49:01um times the features,
  13891. 10:49:04however many we have. Um, and so these
  13892. 10:49:09labels are really taking the place of
  13893. 10:49:11this and they're helping us um make the
  13894. 10:49:15correct updates to these to these
  13895. 10:49:17coefficients or sometimes we call them
  13896. 10:49:20weights. Um, these B 0, B1, B2. Um, we
  13897. 10:49:25find out what the optimal one is to get
  13898. 10:49:27the best fit, right? To get the line of
  13899. 10:49:29best fit. Um, that's what the model
  13900. 10:49:33training when we call this fit. That's
  13901. 10:49:35really what it's doing in the background
  13902. 10:49:36is finding all those coefficients,
  13903. 10:49:38right, to end up with the line of best
  13904. 10:49:40fit that has the lowest amount of error.
  13905. 10:49:46Okay, so hopefully that makes sense.
  13906. 10:49:48That's just a fit um to train our
  13907. 10:49:51models. And that's really going to be um
  13908. 10:49:53the case for
  13909. 10:49:56uh pretty much every single model that
  13910. 10:49:59we uh train with scikitlearn. It's
  13911. 10:50:01pretty much going to be a fit. we pass
  13912. 10:50:03in our training uh features and our
  13913. 10:50:06training labels.
  13914. 10:50:09Okay, so we had that and this was the
  13915. 10:50:12visualization of that where we had our
  13916. 10:50:14test points kind of scattered and we see
  13917. 10:50:17our line of best fit is the one that
  13918. 10:50:19goes through there with that minimal
  13919. 10:50:21error. That's that's the whole goal.
  13920. 10:50:25Pretty decent predictor.
  13921. 10:50:30Okay. And then we talked about
  13922. 10:50:32overfitting, underfitting. So just to
  13923. 10:50:34recap this, overfitting is the concept
  13924. 10:50:37of our model basically memorizing our
  13925. 10:50:39training data. It performs really well
  13926. 10:50:41on that training set, but it is not able
  13927. 10:50:44to generalize outside of that. So it
  13928. 10:50:47performs poorly on the test set or data
  13929. 10:50:49that it's never seen before. Um, and
  13930. 10:50:52that's overfitting. So the reason that
  13931. 10:50:56it overfits is generally the model is
  13932. 10:50:58too complex and it needs to be um it
  13933. 10:51:01needs to be simplified a bit. And one of
  13934. 10:51:04the things we're going to do today is
  13935. 10:51:06see a couple of ways we can alter the
  13936. 10:51:08linear regression model um if we are
  13937. 10:51:11overfitting to prevent overfitting. Um
  13938. 10:51:15so there's going to be ways to handle
  13939. 10:51:17this. Um and so we're going to explore
  13940. 10:51:19some of those today.
  13941. 10:51:22Uh underfitting is kind of the reverse
  13942. 10:51:24of that. Remember it's where the model
  13943. 10:51:26is not learning enough. So the
  13944. 10:51:27performance is poor even on the training
  13945. 10:51:29data. It's not good on the test data
  13946. 10:51:32either. Um that is a sign that the model
  13947. 10:51:35is probably too simple and maybe we
  13948. 10:51:38should use something more complex like
  13949. 10:51:40go from a linear regression maybe use a
  13950. 10:51:42polomial regression. Um or maybe use an
  13951. 10:51:45entirely different model altogether. Um,
  13952. 10:51:48if we're underfitting, our performance
  13953. 10:51:49is poor, it's a good signal we should
  13954. 10:51:52try something else. Um,
  13955. 10:51:55okay.
  13956. 10:51:57So, we talked about those
  13957. 10:52:01and one of the things we also talked
  13958. 10:52:02about was evaluations. If you guys
  13959. 10:52:05remember, we had different metrics that
  13960. 10:52:07we could compute to get a gauge of how
  13961. 10:52:10good our model is actually performing.
  13962. 10:52:12Um, one of those was MSE, which is this
  13963. 10:52:15mean squared error function. Um so we
  13964. 10:52:17did this example during class last time
  13965. 10:52:19on Monday um where we uh were able to
  13966. 10:52:24generate the mean squared error. That's
  13967. 10:52:27one of our metrics. And we can see what
  13968. 10:52:29the mean squared error is on the
  13969. 10:52:32training set and see what it is on the
  13970. 10:52:33test set by um just passing in our um
  13971. 10:52:37training predictions and our training
  13972. 10:52:38labels, our test predictions and our
  13973. 10:52:41test labels. pass those into this mean
  13974. 10:52:43squared error function and it computes
  13975. 10:52:45the MSE and that's that's a helpful
  13976. 10:52:47function from the scikitlearn metrics
  13977. 10:52:50um package um or module I should say and
  13978. 10:52:55we'll be using that quite a bit to do
  13979. 10:52:58you know evaluation of of especially of
  13980. 10:53:00regression right mean squared error is
  13981. 10:53:02pretty is probably the most common uh
  13982. 10:53:06performance metric we can have and if
  13983. 10:53:08you guys remember what it's really doing
  13984. 10:53:10is measuring these distances So mean
  13985. 10:53:12squared error is kind of like the
  13986. 10:53:13average distance away from our our
  13987. 10:53:16points to the actual um to the
  13988. 10:53:19predictions which the predictions are
  13989. 10:53:21all on this line. Um so it's like
  13990. 10:53:25measuring on average how how much error
  13991. 10:53:27do we have on average right? Um, and the
  13992. 10:53:30idea is the closer to zero the better.
  13993. 10:53:33Generally means that the distance away
  13994. 10:53:35from our prediction to our points is
  13995. 10:53:37pretty low. The closer to zero it is.
  13996. 10:53:40Um, which is pretty desirable.
  13997. 10:53:43So a low MSE is kind of what we're
  13998. 10:53:45looking for. Um, closer to zero the
  13999. 10:53:47better. And so um if one model has if
  14000. 10:53:51one model has um a low lower MSE than
  14001. 10:53:56another, it's it's a better performing
  14002. 10:53:58model, right? It has less error.
  14003. 10:54:02Okay. And then we also looked at the R R
  14004. 10:54:05squared or sometimes known as R2 um
  14005. 10:54:08score. Um this is another metric that we
  14006. 10:54:12could use that measures the the
  14007. 10:54:15variability
  14008. 10:54:16um of uh the predictions and if our
  14009. 10:54:21model is capturing that variability um
  14010. 10:54:23well um and so R squ is has a range of 0
  14011. 10:54:27to one one is better that means the
  14012. 10:54:29model is capturing the the changes in in
  14013. 10:54:32the um output it um our predictions
  14014. 10:54:36follow along with those same changes um
  14015. 10:54:38so they're pretty close um so closer to
  14016. 10:54:41one would be a better score. So we have
  14017. 10:54:45those kind of metrics. So like on this
  14018. 10:54:47data um this would this would show that
  14019. 10:54:50this model was underfitting remember
  14020. 10:54:52because this
  14021. 10:54:54mean this MSE was bad and this MSE was
  14022. 10:54:58bad.
  14023. 10:54:59Um and what we should think of these in
  14024. 10:55:02the units of what our labels are. um
  14025. 10:55:06especially if we take the square root of
  14026. 10:55:08this the RMSSE that was another metric
  14027. 10:55:11we had um the square root of this is
  14028. 10:55:14actually in the exact units that we um
  14029. 10:55:17have for our labels. So uh in this
  14030. 10:55:20example this was the um this was the the
  14031. 10:55:24units or the sales versus the TV
  14032. 10:55:27products, right? Um and so this would
  14033. 10:55:31indicate that on average if we take the
  14034. 10:55:33square root of this um
  14035. 10:55:35in the square root of this um we have uh
  14036. 10:55:40um we're on average about 11 sales units
  14037. 10:55:44off squared. So if we take the square
  14038. 10:55:45roo of that um it's somewhere around 3
  14039. 10:55:47to four um somewhere in between three
  14040. 10:55:51and four units off. And this is as well.
  14041. 10:55:54Um, and because both of these are still
  14042. 10:55:57not close to zero, um, this would be
  14043. 10:55:59under fit. And this shows that as well.
  14044. 10:56:02This isn't that close to one. It's
  14045. 10:56:04decent, but it's not, um, not that close
  14046. 10:56:06to one. So, we would say, and
  14047. 10:56:08performance is poor on both training and
  14048. 10:56:11test sets. That's the key indicator of
  14049. 10:56:13underfitting. It's poor on both.
  14050. 10:56:20Yeah, exactly. High MSE correlates to
  14051. 10:56:23underfitting. Yes. Yes. And it what's
  14052. 10:56:25key is it's high MSE on both on both the
  14053. 10:56:29training and the test sets.
  14054. 10:56:33If you have a high MSE on your test set
  14055. 10:56:35but a low MSE on your training set,
  14056. 10:56:37that's overfitting, right? Where it's
  14057. 10:56:40not generalizing from the training set
  14058. 10:56:42to the test data that it hasn't seen
  14059. 10:56:44before. That's overfitting. So the key
  14060. 10:56:47is high MSE on both sets.
  14061. 10:56:52All right. So we talked about that. Um
  14062. 10:56:55we did polomial regression last time. So
  14063. 10:56:58that was um doing
  14064. 10:57:02that was uh making a curved graph um by
  14065. 10:57:06transforming the features into polomial
  14066. 10:57:08features and then doing linear
  14067. 10:57:09regression with that. So you guys
  14068. 10:57:11remember from Monday we did this where
  14069. 10:57:14um we took our features and uh
  14070. 10:57:17transformed them according to this
  14071. 10:57:19polomial features from scikitlearn. So
  14072. 10:57:22we can go all the way up to degree
  14073. 10:57:23whatever degree we want. So we put in
  14074. 10:57:25four here but there's nothing special
  14075. 10:57:26about four really. This is just testing
  14076. 10:57:28it out. um and we generate the the
  14077. 10:57:31polomial features and we can fit a
  14078. 10:57:34linear regression on those polomial
  14079. 10:57:36features and we get a slightly better
  14080. 10:57:39model, right? Um it fits the data a
  14081. 10:57:42little bit better than just a straight
  14082. 10:57:43line. This curved line with the polomial
  14083. 10:57:48features um performs a little bit better
  14084. 10:57:49and we could see that with the MSE,
  14085. 10:57:51right? or we could evaluate the MSE of
  14086. 10:57:53this um and it would be lower.
  14087. 10:57:57It would be lower than the curve line.
  14088. 10:57:58And so that's something we could do. Um
  14089. 10:58:01we would just have to pass in these test
  14090. 10:58:03predictions, the training predictions
  14091. 10:58:05and then the the test labels and
  14092. 10:58:07training labels and pass those into the
  14093. 10:58:09mean squared error function and we could
  14094. 10:58:10compute that, right? Wouldn't be hard to
  14095. 10:58:12do.
  14096. 10:58:16All right. And then finally where we
  14097. 10:58:18left off um you know is on our
  14098. 10:58:21performance metrics. So we talked about
  14099. 10:58:23mean squared error. That's that average
  14100. 10:58:25distance away from the labels to our
  14101. 10:58:28predictions. Um and we take the square
  14102. 10:58:31root of that. It's it's basically
  14103. 10:58:33measuring the same thing but it's the
  14104. 10:58:35square root of it is um more
  14105. 10:58:37interpretable because it's in the same
  14106. 10:58:38units as our label.
  14107. 10:58:41um mean absolute error is is the average
  14108. 10:58:45distance of the absolute value. So it's
  14109. 10:58:47not the squared distance formula like a
  14110. 10:58:49uklidian distance but it is a absolute
  14111. 10:58:52value. So it's a little bit um less
  14112. 10:58:54sensitive to outliers. They don't get
  14113. 10:58:56magnified as much. Um but it's not
  14114. 10:59:00typically used as much as a mean squared
  14115. 10:59:03error would be with regression. um we
  14116. 10:59:05talked about the last time because um
  14117. 10:59:08the distance formula or that distance is
  14118. 10:59:11actually what's used to train the model.
  14119. 10:59:13So it's a more natural um fit for a
  14120. 10:59:17performance metric for it.
  14121. 10:59:22All right. And then we had R square. We
  14122. 10:59:23just talked about that closer to zero
  14123. 10:59:25would be um worse. Closer to one would
  14124. 10:59:28be better. That means that the model
  14125. 10:59:30explains um all the variability in the
  14126. 10:59:33in the predictions. Uh it captures those
  14127. 10:59:36predictions um closely to the labels
  14128. 10:59:41um very well. So uh one would be better.
  14129. 10:59:46Closer to one would be better.
  14130. 10:59:49All right. So that's where we left off.
  14131. 10:59:51Um we're gonna pick up from there with
  14132. 10:59:53cross validation. um we've actually
  14133. 10:59:56already seen one method of cross
  14134. 10:59:58validation. So we're going to study um
  14135. 11:00:00we're going to kind of recap that and
  14136. 11:00:01and then um talk about cross validation
  14137. 11:00:04in general um and look at some more
  14138. 11:00:08sophisticated techniques of it um coming
  14139. 11:00:10up next. But before I do that, any
  14140. 11:00:13questions about anything we've covered
  14141. 11:00:16um to this point in in the recap or
  14142. 11:00:19anything from Monday? Any questions on
  14143. 11:00:22that?
  14144. 11:00:24All right. So let's talk about uh cross
  14145. 11:00:27validation. Um now this term cross
  14146. 11:00:32validation refers to a technique that
  14147. 11:00:36evaluates performance. And what it does
  14148. 11:00:39is it divides our data into essentially
  14149. 11:00:43um training and test sets which we've
  14150. 11:00:45kind of already seen. And then we are
  14151. 11:00:47able to train a model on on the training
  14152. 11:00:50set, evaluate it on the test set. And
  14153. 11:00:52that's where that's where we get the
  14154. 11:00:53name cross validation because we're
  14155. 11:00:56crossing over our model from one batch
  14156. 11:00:58of data used to train it over to another
  14157. 11:01:01set of data used to validate those
  14158. 11:01:03predictions. Um, and there's actually
  14159. 11:01:06different ways to do cross validation.
  14160. 11:01:08So cross validation is a bit of an
  14161. 11:01:10umbrella term for multiple ways to do
  14162. 11:01:12that. We've already seen one way of
  14163. 11:01:14doing that um which I'm going to scroll
  14164. 11:01:16down to is um known as a hold out cross
  14165. 11:01:21validation. So that's um what we've been
  14166. 11:01:23doing so far. So this is just um
  14167. 11:01:26generating a train and a test set
  14168. 11:01:30train um split.
  14169. 11:01:33Um that's the that's what's known as the
  14170. 11:01:36hold out cross validation method. Um and
  14171. 11:01:39and this is exactly what we've been
  14172. 11:01:41doing so far, which is you split your
  14173. 11:01:43data into some type of split, usually
  14174. 11:01:467030,
  14175. 11:01:48um of a train and test
  14176. 11:01:51and then you um train your model on this
  14177. 11:01:54section of data and then apply it to
  14178. 11:01:56this to evaluate performance. Right? So
  14179. 11:01:59that's that's what's known as the hold
  14180. 11:02:01out method. Um it is uh you know
  14181. 11:02:06relatively simple. It's pretty fast to
  14182. 11:02:08do. Um, but there are more robust ways
  14183. 11:02:13to try to divide up our data a little
  14184. 11:02:16bit uh more evenly. Instead of just
  14185. 11:02:19having one split, we can actually do
  14186. 11:02:21many splits, which is the idea of um the
  14187. 11:02:24next kind of cross validation I'll
  14188. 11:02:26cover. But hold out method is one that
  14189. 11:02:29we've already studied. It's the most
  14190. 11:02:30basic type of cross validation you can
  14191. 11:02:33have. Um so hold out this is the most
  14192. 11:02:36basic
  14193. 11:02:38and we we've already been we've already
  14194. 11:02:41been uh working with this type. Okay.
  14195. 11:02:46So we've we've already seen hold out
  14196. 11:02:48method. Let me uh explain to you a more
  14197. 11:02:51sophisticated method a little bit more
  14198. 11:02:53advanced of a cross validation um which
  14199. 11:02:56is known as Kfold cross validation. So
  14200. 11:02:59this is um going to be a little bit more
  14201. 11:03:02advanced of a technique but this is the
  14202. 11:03:04idea of kfold is that you take your data
  14203. 11:03:07set
  14204. 11:03:09and you split it into k number of what
  14205. 11:03:14are called splits or folds. So you take
  14206. 11:03:17your data and you let's say it was let's
  14207. 11:03:19say k equals 5. So we have five splits
  14208. 11:03:22here.
  14209. 11:03:25Okay. So let's say k equals 5. We have
  14210. 11:03:28five splits. So what we're going to do
  14211. 11:03:33is we're going to we're going to train
  14212. 11:03:35our model on K minus one of those folds.
  14213. 11:03:39So if K was five, we had five splits.
  14214. 11:03:42We're going to take our model and train
  14215. 11:03:44it on four out of five of those uh
  14216. 11:03:48splits. So let's say it's these four.
  14217. 11:03:52We'll train it on these four.
  14218. 11:03:56Okay. And then what we do is the one
  14219. 11:03:59split that's left over, we will we will
  14220. 11:04:03test our model against that split. So
  14221. 11:04:05we'll test here.
  14222. 11:04:10Okay. Now, this sounds very similar to
  14223. 11:04:12the hold out method where we're doing a
  14224. 11:04:14train test split, but it's a little bit
  14225. 11:04:16this kful cross validation a little bit
  14226. 11:04:18more sophisticated because we repeat
  14227. 11:04:20this process that I just mentioned over
  14228. 11:04:23and over for all combinations of the
  14229. 11:04:25splits. So then what we'll do, this is
  14230. 11:04:28just one trial that we'll do it again,
  14231. 11:04:33but this time we will pick um four
  14232. 11:04:36different splits. So, this time we might
  14233. 11:04:39pick,
  14234. 11:04:41let me do blue. This time we might pick
  14235. 11:04:44this one, this one,
  14236. 11:04:47um,
  14237. 11:04:49this one,
  14238. 11:04:51and this one.
  14239. 11:04:55And then those four we will train our
  14240. 11:04:57data on. And then we will test against
  14241. 11:04:59this one. Okay. And we'll do we'll
  14242. 11:05:02repeat this
  14243. 11:05:05repeat for all combos of the folds.
  14244. 11:05:15Okay. So we'll repeat that. So
  14245. 11:05:17essentially what we're doing is rotating
  14246. 11:05:19through. Every time we rotate through
  14247. 11:05:22one of the folds is going to be left out
  14248. 11:05:23as a test set. Now this is a little bit
  14249. 11:05:27more robust than just a train test
  14250. 11:05:29split, right? because we are exposing
  14251. 11:05:32our model to more of the data in in
  14252. 11:05:35doing this, right? Because we're going
  14253. 11:05:37to split it evenly into five or 10
  14254. 11:05:39splits. Those are pretty common um
  14255. 11:05:42number of folds to use. 10 or five. Um
  14256. 11:05:45those are the ones I've most commonly
  14257. 11:05:47seen. Um but we're going to by rotating
  14258. 11:05:52through which folds are being used for
  14259. 11:05:53training, which ones being left out. um
  14260. 11:05:56we are exposing our our model to more of
  14261. 11:05:59the data this way than just doing a
  14262. 11:06:00single train test split. Right? So now
  14263. 11:06:04what do we do with with the results is
  14264. 11:06:07every time we do this we we generate um
  14265. 11:06:10an MSE let's say or some type of
  14266. 11:06:12performance metric. So let's say we
  14267. 11:06:14generate an MSE from this guy
  14268. 11:06:17we generate an MSE from this version and
  14269. 11:06:20we generate an MSE for all combos.
  14270. 11:06:24each combo we generate MSE and then what
  14271. 11:06:27we do is we average
  14272. 11:06:30the metrics
  14273. 11:06:33or the in this case uh if we use MSE we
  14274. 11:06:36would average those together. So every
  14275. 11:06:39time we do a fold combination and we
  14276. 11:06:41keep four of them for training, one for
  14277. 11:06:42test and we rotate through all those
  14278. 11:06:45combinations, we are going to generate
  14279. 11:06:47an MSE for every combination
  14280. 11:06:50then we're just going to average those
  14281. 11:06:52MSSE's to get a final. So the final MSE
  14282. 11:06:56of cross val of this K-fold.
  14283. 11:07:00So the final metric
  14284. 11:07:03is just the average of the uh
  14285. 11:07:06performance on all of the fold
  14286. 11:07:08combinations. Okay. So our final MSE, we
  14287. 11:07:12just average all those MSE from all of
  14288. 11:07:14our combinations.
  14289. 11:07:16Okay.
  14290. 11:07:18Now, what's the advantage to doing this?
  14291. 11:07:21It's way more robust of a estimate of
  14292. 11:07:24the of the performance of the model
  14293. 11:07:26because we're exposing it to all
  14294. 11:07:29basically all of our data, right? We're
  14295. 11:07:31getting a sense of how it performs
  14296. 11:07:32across all those different folds. Um
  14297. 11:07:35rather than just doing a single train
  14298. 11:07:37test split, which is a bit it's basic,
  14299. 11:07:40it works, but it's a bit basic. Um so
  14300. 11:07:42this is more robust estimate of the
  14301. 11:07:46performance.
  14302. 11:07:47Now, what's the drawback to doing this
  14303. 11:07:50is that it's more intensive. So, if you
  14304. 11:07:52have a lot of data, this is going to be
  14305. 11:07:54pretty expensive to do because you're
  14306. 11:07:55going to have to especially you have a
  14307. 11:07:57high number of folds, right? You're
  14308. 11:07:58going to have to divide your data into k
  14309. 11:08:01number of folds and you're going to have
  14310. 11:08:03to do this over and over again. Um, and
  14311. 11:08:05if it's a large data set, it might take
  14312. 11:08:07your model a long time to train. It's
  14313. 11:08:09going to be a little bit more uh
  14314. 11:08:12computationally intense than if we just
  14315. 11:08:15did a train test split.
  14316. 11:08:17Okay, we just did a single like 7030
  14317. 11:08:19split. We only do that once. We only
  14318. 11:08:22train the model once, right? We train it
  14319. 11:08:24on the 70, apply it to the 30% test data
  14320. 11:08:28and evaluate performance that way. Um,
  14321. 11:08:31so we're only really using the model and
  14322. 11:08:33training the model once, but in this
  14323. 11:08:36kfold, we're going to do it um, you
  14324. 11:08:38know, k number of times essentially
  14325. 11:08:42or I should say one for every
  14326. 11:08:44combination that we have to work through
  14327. 11:08:46of of all the folds.
  14328. 11:08:52Okay.
  14329. 11:08:55All right. Does that make sense? Any any
  14330. 11:08:58questions on kf fold cross validation?
  14331. 11:09:01So k K is an important uh number here.
  14332. 11:09:05It it's how many folds how many splits
  14333. 11:09:08do you have? A typical value for K is
  14334. 11:09:10going to be somewhere like five or 10.
  14335. 11:09:14So 10 folds or five folds. Those are
  14336. 11:09:17pretty pretty standard
  14337. 11:09:21from what from what I've seen.
  14338. 11:09:25But does the does the concept make sense
  14339. 11:09:27or is there any questions on it on in
  14340. 11:09:29terms of um you're always going to leave
  14341. 11:09:31one fold out. You're going to split it
  14342. 11:09:33up into K number of folds. Always leave
  14343. 11:09:35one out. Train on the rest of it.
  14344. 11:09:38Evaluate on that one that gets left out
  14345. 11:09:39and then rotate those through. And
  14346. 11:09:41you're going to do that for every
  14347. 11:09:42combination and average all those
  14348. 11:09:44metrics.
  14349. 11:09:52And by the way, there's going to be an
  14350. 11:09:53easy function in scikitlearn that will
  14351. 11:09:56do this for us. So managing all these
  14352. 11:09:58combinations will be really easy. It's
  14353. 11:10:01actually just built into scikitlearn. So
  14354. 11:10:03we don't have to um we don't have to do
  14355. 11:10:06this all by hand. Okay, this will be in
  14356. 11:10:08scikitlearn. It'll handle doing all
  14357. 11:10:10these combinations of folds for us and
  14358. 11:10:13computing the average metric will be
  14359. 11:10:15really easy. So um
  14360. 11:10:19we don't have to worry about that. We're
  14361. 11:10:20going to see an example of this coming
  14362. 11:10:21up shortly.
  14363. 11:10:23All right, of kfold cross validation,
  14364. 11:10:28but this is a this is a really widely
  14365. 11:10:30used technique. And again, like the
  14366. 11:10:32purpose, you may be wondering like
  14367. 11:10:33what's the purpose ultimately of doing
  14368. 11:10:35this? It's to get a sense of if our
  14369. 11:10:37model is going to perform well on new
  14370. 11:10:39data. That's really what we want to
  14371. 11:10:41know. Like is the model going to perform
  14372. 11:10:43well when I start to use it on new data
  14373. 11:10:45that it's never seen before? And this
  14374. 11:10:48kffold is a decent indicator of that
  14375. 11:10:52because we are varying which data it
  14376. 11:10:55sees across many different folds. Right?
  14377. 11:10:58So it's a it's kind of a good um proxy
  14378. 11:11:02to exposing it to different kinds of
  14379. 11:11:05data each time and seeing how it
  14380. 11:11:07performs.
  14381. 11:11:09All right? Because we're working our way
  14382. 11:11:10through each one of the folds. There's
  14383. 11:11:11always going to be one fold left out.
  14384. 11:11:13We're going to change which fold gets
  14385. 11:11:15left out each time. And um that's sort
  14386. 11:11:18of mimicking the idea of we're going to
  14387. 11:11:20apply our model to new data and see how
  14388. 11:11:22it performs. And it's it's new data
  14389. 11:11:26every fold.
  14390. 11:11:45um how we know which model is best suits
  14391. 11:11:49for which scenario because we have Yeah,
  14392. 11:11:52that's a good question. Um,
  14393. 11:11:54so my we're going to learn this as we go
  14394. 11:11:57along because we haven't covered all the
  14395. 11:11:59models yet, but generally the best
  14396. 11:12:02advice I can give on that is
  14397. 11:12:05you you generally want to start as
  14398. 11:12:09simple as you can get and then if it's
  14399. 11:12:11not performing well then work your way
  14400. 11:12:13up to something more complex.
  14401. 11:12:16So we are going to have models that are
  14402. 11:12:18simpler. We're going to have models that
  14403. 11:12:19are more complex. The rule of thumb is
  14404. 11:12:22to start with the most simple model that
  14405. 11:12:24works.
  14406. 11:12:26So you're usually going to have the same
  14407. 11:12:30ones that you're going to try in the
  14408. 11:12:31beginning. And linear regression is a
  14409. 11:12:34very simple model. It's usually the
  14410. 11:12:36first one you want to try for regression
  14411. 11:12:38because it's the simplest.
  14412. 11:12:40Um, and for classification, we're going
  14413. 11:12:42to have a similar like logistic
  14414. 11:12:44regression is the simplest kind of
  14415. 11:12:46classification model we could have. So
  14416. 11:12:49usually want to start with that and then
  14417. 11:12:51if it underfits like if we see it's
  14418. 11:12:54producing a lot of error then we work
  14419. 11:12:56our way up to a more sophisticated
  14420. 11:12:59model.
  14421. 11:13:01So um that's the way we that's the way
  14422. 11:13:05it should usually go is simple to
  14423. 11:13:07complex it based on their performance.
  14424. 11:13:09So we evaluate it and then we can repeat
  14425. 11:13:11the process. If it's not performing well
  14426. 11:13:13we can try something different that's
  14427. 11:13:15more complex if it's underfitting.
  14428. 11:13:24Uh this is a good question. Does a model
  14429. 11:13:26reset after training each K minus one
  14430. 11:13:28fold? Um yeah, it's essentially like a
  14431. 11:13:31blank model every time uh every fold. So
  14432. 11:13:35um we imagine like you have a brand you
  14433. 11:13:38have a fresh model every um k minus one
  14434. 11:13:42combination. Yes.
  14435. 11:13:50And the reason the reason it has to be
  14436. 11:13:52that way is because you don't want the
  14437. 11:13:55other folds influencing the model that
  14438. 11:13:59like on on the next combination. You
  14439. 11:14:02don't want the previous combination to
  14440. 11:14:03influence the results on the next one,
  14441. 11:14:05right? Um you want it to be a fresh
  14442. 11:14:08evaluation on every combination of
  14443. 11:14:11folds.
  14444. 11:14:30Okay.
  14445. 11:14:32All right. So, let me describe to you a
  14446. 11:14:35variation on what we just um talked
  14447. 11:14:38about with the K-fold. So, there's
  14448. 11:14:40another cross validation known as
  14449. 11:14:42stratified K-fold. And um this is the
  14450. 11:14:46same exact procedure as k-fold except
  14451. 11:14:49that when we this is used for
  14452. 11:14:51classification.
  14453. 11:14:52Um so when we do classification
  14454. 11:14:55uh we want to make sure that the
  14455. 11:14:57different categories are going to be um
  14456. 11:15:00split amongst those folds in a
  14457. 11:15:03proportional way. So we don't what we
  14458. 11:15:05don't want to happen is um when we split
  14459. 11:15:08apart the data. So, let's say we have
  14460. 11:15:10let's say we're predicting um spam not
  14461. 11:15:13spam. What we don't want to have happen
  14462. 11:15:15when we do our splits is we don't want
  14463. 11:15:18to have all of the spams end up in one
  14464. 11:15:21fold and then every other fold has no
  14465. 11:15:24spam, no spam, no spam, no spam, right?
  14466. 11:15:28That's not very good. Um because if we
  14467. 11:15:31if we train against all these guys, we
  14468. 11:15:33have no shot at predicting spam when
  14469. 11:15:35they've never seen spam before. So
  14470. 11:15:38stratify kffold is is used in
  14471. 11:15:40classification
  14472. 11:15:42and it's to um it's to make our splits
  14473. 11:15:46ensure that they have basically a
  14474. 11:15:49balanced number of categories for each
  14475. 11:15:51split. Um so that we don't end up with
  14476. 11:15:54certain splits with way more spams than
  14477. 11:15:57not spams. Um so we we do what's called
  14478. 11:16:00stratifying where we make sure the
  14479. 11:16:02proportions are balanced across each uh
  14480. 11:16:05split. So this is only really useful in
  14481. 11:16:07classification, not really necessary in
  14482. 11:16:10regression because we're predicting a
  14483. 11:16:11value. But if we were predicting a
  14484. 11:16:13category,
  14485. 11:16:15like in classification like fraud, not
  14486. 11:16:18fraud, we don't want to do the split and
  14487. 11:16:20have every single fraud example um by
  14488. 11:16:23bad luck in our shuffling and split end
  14489. 11:16:25up in one split and every other um every
  14490. 11:16:29other split has no examples of fraud.
  14491. 11:16:31Right? Right. So we want to stratify
  14492. 11:16:32this to spread out those um frauds
  14493. 11:16:35against all the other splits. Um so uh
  14494. 11:16:40again um scikitlearn will take care of
  14495. 11:16:42that for you. Um but if you're doing
  14496. 11:16:44classification and you have an
  14497. 11:16:46imbalanced data set um you you really
  14498. 11:16:49want to make sure you stratify kfold. um
  14499. 11:16:52imbalanced meaning that you have a a um
  14500. 11:16:56different number. Like if you're doing
  14501. 11:16:58fraud, not fraud, you have way more not
  14502. 11:17:00frauds than frauds. Um when where that
  14503. 11:17:02category is imbalanced,
  14504. 11:17:05you want to make sure it's balanced
  14505. 11:17:06across all your splits.
  14506. 11:17:09Um so this is this is useful in
  14507. 11:17:12classification only, not really
  14508. 11:17:13regression, which is what we're talking
  14509. 11:17:14about right now. Um but it's just a
  14510. 11:17:17variation on this that ensures when we
  14511. 11:17:19do those folds um the data is
  14512. 11:17:22distributed evenly amongst those folds
  14513. 11:17:24as much as we can. The labels are I
  14514. 11:17:26should say.
  14515. 11:17:28Okay. So that's stratified kfold. It's
  14516. 11:17:32the same same procedure once we have our
  14517. 11:17:34splits. It's the same where we do k
  14518. 11:17:35minus one of them. We train test on that
  14519. 11:17:38last fold um and then rotate through all
  14520. 11:17:42the folds and and average all the
  14521. 11:17:44metrics. the same exact procedure. It's
  14522. 11:17:46just the splitting itself um is going to
  14523. 11:17:49be balanced in a stratified kfold.
  14524. 11:17:55Okay, so hold out we've already talked
  14525. 11:17:56about um is just doing a single train
  14526. 11:17:59test split. We've talked about that. One
  14527. 11:18:02more variation that is a bit of an
  14528. 11:18:03extreme version of Kfold. So it's
  14529. 11:18:06actually the same process as Kfold, but
  14530. 11:18:08it's an extreme version is if you set K
  14531. 11:18:12equal to the number of data points. So
  14532. 11:18:14you basically are um this is a really
  14533. 11:18:17really extreme kfold where you um
  14534. 11:18:20basically are training on all the data.
  14535. 11:18:23Um so you're training on all the data
  14536. 11:18:26except one point and then you test
  14537. 11:18:29against that one point. Um now why would
  14538. 11:18:33you ever do this? Um it's mainly so for
  14539. 11:18:36this reason here. it's to um maximize
  14540. 11:18:40the amount of training data that your
  14541. 11:18:42model gets exposed to because instead of
  14542. 11:18:44just doing instead of just doing five
  14543. 11:18:46splits
  14544. 11:18:48um which would be like
  14545. 11:18:51you know these four folds are going to
  14546. 11:18:53be used and then we um test against one
  14547. 11:18:55fold um we're essentially going to use
  14548. 11:18:5999% of the data right one point is going
  14549. 11:19:02to be left out 99% of the data gets used
  14550. 11:19:05to train um and then we're always going
  14551. 11:19:08to leave out one point and and the issue
  14552. 11:19:10is we're actually going to do that over
  14553. 11:19:11and over and over again and rotate that
  14554. 11:19:13one point to cover the whole data set.
  14555. 11:19:16So we're going to train on 99% leave one
  14556. 11:19:19that one point out
  14557. 11:19:21and then rotate through every
  14558. 11:19:23combination of points until we've left
  14559. 11:19:25out every single point and then average
  14560. 11:19:28all those together. Um so this is a this
  14561. 11:19:30is an extreme kffold. Again the number
  14562. 11:19:33of folds is actually equal to the number
  14563. 11:19:35of data points in this case. So we have
  14564. 11:19:37every point is its own fold and we train
  14565. 11:19:40on everything but one test on that one.
  14566. 11:19:43This gets you the maximum size of your
  14567. 11:19:46training data because you're basically
  14568. 11:19:48going to have every point but one used
  14569. 11:19:50in the training.
  14570. 11:19:52This gets you the maximum size. However,
  14571. 11:19:54it gets you the maximum uh expense
  14572. 11:19:58especially for large data sets. This is
  14573. 11:20:00going to be usually you're not going to
  14574. 11:20:01use this um especially for large data
  14575. 11:20:04sets because it's just too extreme. It's
  14576. 11:20:08going to take you a really long time to
  14577. 11:20:09work through every single point being
  14578. 11:20:12left out. Um it's just going to take a
  14579. 11:20:15while to do.
  14580. 11:20:17So for that reason, the leave one out um
  14581. 11:20:20that that's why it's called leave one
  14582. 11:20:21out because it's you're leaving one out
  14583. 11:20:24every single time. Um is rarely used. I
  14584. 11:20:27I don't really see it used that often,
  14585. 11:20:30but it is an extreme version of K-fold
  14586. 11:20:33cross validation.
  14587. 11:20:36Okay. But rarely ever actually used. I
  14588. 11:20:38think the the ones that get used the
  14589. 11:20:40most are definitely the hold out method
  14590. 11:20:42with just a regular train test split. Um
  14591. 11:20:45and then uh the other one that gets used
  14592. 11:20:47quite a bit is is Kfold
  14593. 11:20:50or stratified K-fold if you're if you're
  14594. 11:20:52doing classification, but certainly
  14595. 11:20:54K-fold in the in a regression case.
  14596. 11:20:58Okay.
  14597. 11:21:02All right. Um, we're going to do an
  14598. 11:21:05example with these guys. So, we'll do
  14599. 11:21:07that next. Um, with with the different
  14600. 11:21:09cross validation techniques. Um, but any
  14601. 11:21:14questions on what they are doing
  14602. 11:21:17conceptually before we actually do the
  14603. 11:21:19code example?
  14604. 11:21:36Okay,
  14605. 11:21:38very good.
  14606. 11:21:44All right, so let's see some examples.
  14607. 11:21:47Um let's go into our code and build a
  14608. 11:21:51model and do the different cross
  14609. 11:21:54validation techniques on it. Um you're
  14610. 11:21:57going to see it's actually going to be
  14611. 11:21:58really easy to do and we it sounds
  14612. 11:22:00complex like doing the kfold and leaving
  14613. 11:22:03one out and testing it sounds kind of
  14614. 11:22:05complex but I promise you scikitlearn
  14615. 11:22:07makes it really easy to do. Um
  14616. 11:22:11and so uh we won't need to do too much
  14617. 11:22:14besides just use the right uh tools from
  14618. 11:22:17scikitlearn. Uh so we're going to we're
  14619. 11:22:19going to see that. Um so here we have
  14620. 11:22:21some imports. The um primary uh thing
  14621. 11:22:25that's a little bit new for us is going
  14622. 11:22:27to be these um different kinds of cross
  14623. 11:22:30validation techniques. So we have our
  14624. 11:22:31kfold, we have our stratified kfold,
  14625. 11:22:33leave one out um which are those
  14626. 11:22:35different cross validation techniques.
  14627. 11:22:38Um these are going to be used in
  14628. 11:22:40combination with this cross val score
  14629. 11:22:45which is going to keep track of the
  14630. 11:22:47different um metrics and then average
  14631. 11:22:49them
  14632. 11:22:51uh while we do one of these um cross
  14633. 11:22:55validation techniques. So this guy gets
  14634. 11:22:58used in combination with one of these to
  14635. 11:23:01um as as we're going to see in the code
  14636. 11:23:04uh to average those metrics. um doing
  14637. 11:23:07the different folds, right? Perform
  14638. 11:23:09doing performance against the different
  14639. 11:23:10folds. Okay. And then of course we need
  14640. 11:23:13a model
  14641. 11:23:15using linear regression. That's that's
  14642. 11:23:16the one we've studied so far. Um and
  14643. 11:23:20then we have just a regular metrics if
  14644. 11:23:22we want to compute those. Um using maybe
  14645. 11:23:25just hold out, right? And hold out um
  14646. 11:23:28which which is just a regular train test
  14647. 11:23:30split. Um we could use these guys to
  14648. 11:23:32evaluate performance.
  14649. 11:23:34But in a more sophisticated K-fold style
  14650. 11:23:37of cross audition, we're going to use
  14651. 11:23:39this to evaluate the the performance.
  14652. 11:23:45Okay, let's see.
  14653. 11:23:48So, we're going to be working with this
  14654. 11:23:50housing with ocean proximity data. Um,
  14655. 11:23:53you guys should have this one. Uh, so
  14656. 11:23:58you guys should have this one. So, if
  14657. 11:24:00you want to follow along and run it
  14658. 11:24:01yourself, um, you can load that one in.
  14659. 11:24:06Um, I want to make sure that I have it.
  14660. 11:24:11Let me pull that one in. So, it should
  14661. 11:24:12be this guy.
  14662. 11:24:23I'm going to load that in so I can make
  14663. 11:24:25sure I run it with you guys.
  14664. 11:24:28Um,
  14665. 11:24:37so let me run this.
  14666. 11:24:41Do you guys have that data?
  14667. 11:24:44The housing with ocean proximity?
  14668. 11:24:48It's another it's another housing data
  14669. 11:24:50set. Um,
  14670. 11:24:53but it it's a little bit different than
  14671. 11:24:55the ones we've seen before. It has a a
  14672. 11:24:58special feature for how close it is to
  14673. 11:25:00the ocean, the different locations.
  14674. 11:25:10So, it looks kind of like this. If we
  14675. 11:25:12load it in and do our head, which is
  14676. 11:25:14usually what we do, right, we can see um
  14677. 11:25:17we can see that it's got these features.
  14678. 11:25:19So, it's got uh uh bedrooms, total
  14679. 11:25:23rooms,
  14680. 11:25:25um it's got uh median age. Now, this is
  14681. 11:25:28this is looks a little strange for total
  14682. 11:25:30rooms and um uh bedrooms and population,
  14683. 11:25:35etc., but it's um
  14684. 11:25:39it's it's got those uh it's got those
  14685. 11:25:42because it's representing an entire
  14686. 11:25:44neighborhood. So, it's an entire
  14687. 11:25:46neighborhood and we're looking at this
  14688. 11:25:49um this is actually going to be our
  14689. 11:25:50label is this median house value for the
  14690. 11:25:52entire neighborhood. So, what's that
  14691. 11:25:54median value uh in the neighborhood? And
  14692. 11:25:58this is the total number of bedrooms,
  14693. 11:26:00total number of rooms, um population,
  14694. 11:26:03households. So, how many houses are
  14695. 11:26:06there? Um median income. And of course,
  14696. 11:26:08these are scaled. So these are um likely
  14697. 11:26:11times you know uh thousands
  14698. 11:26:16um
  14699. 11:26:18but um that's our data. We could
  14700. 11:26:22describe it.
  14701. 11:26:27So we can see the average age median age
  14702. 11:26:31um which sounds a little um weird but
  14703. 11:26:33that's it's because again this is the
  14704. 11:26:36median of data within a neighborhood. Um
  14705. 11:26:40so the average of those is about 28 or
  14706. 11:26:4329. Um we have
  14707. 11:26:47u
  14708. 11:26:51total bedrooms. The we can look at the
  14709. 11:26:53min. There's some data that only has
  14710. 11:26:56one. So it's likely only one house in
  14711. 11:26:58there. Um which is what this represents.
  14712. 11:27:01There's only one house. So there there
  14713. 11:27:03is some neighborhood that only has one
  14714. 11:27:05house. Um, and we see the median, um, we
  14715. 11:27:10see the minimum, uh, median house values
  14716. 11:27:13there. And then the maximum down here,
  14717. 11:27:16um, is a pretty big number.
  14718. 11:27:216,000 households is the largest that we
  14719. 11:27:23have in any any one of these
  14720. 11:27:25neighborhoods.
  14721. 11:27:27Okay. So, just a little bit of
  14722. 11:27:29description of the data.
  14723. 11:27:45Okay. So then we can run.info. So this
  14724. 11:27:47is um let me ask you guys, were you able
  14725. 11:27:50to load this? Were you able to run this?
  14726. 11:27:56If you're following along, were you able
  14727. 11:27:57to
  14728. 11:27:59load it and take a look at dot head.
  14729. 11:28:09Okay, great. Great.
  14730. 11:28:12Okay, so we're able to load that and
  14731. 11:28:15then look at dot head. Perfect. Um
  14732. 11:28:20okay.
  14733. 11:28:22Um and then we run describe which gives
  14734. 11:28:25us that uh usual kind of statistical
  14735. 11:28:27description. Uh so we can see some
  14736. 11:28:30interesting stats about those.
  14737. 11:28:37What do you guys notice about the info?
  14738. 11:28:40Anything interesting that we see from
  14739. 11:28:41there?
  14740. 11:28:56Is there any missing data?
  14741. 11:29:01Any features that have missing data? Can
  14742. 11:29:04we see
  14743. 11:29:07object? Yeah, object type usually is
  14744. 11:29:09string. If it's an object type, that
  14745. 11:29:11usually means string. Python when we
  14746. 11:29:14read it into pandas it usually is just a
  14747. 11:29:16string.
  14748. 11:29:19So that that makes sense like we have
  14749. 11:29:20mostly numerical features but then we
  14750. 11:29:22have a this ocean proximity which is a
  14751. 11:29:25string.
  14752. 11:29:34Yeah. Total bedrooms has nles. That's
  14753. 11:29:36right. Because you can see here this
  14754. 11:29:38does not equal the number of uh rows
  14755. 11:29:41that we have. So, this is the number of
  14756. 11:29:42rows. There's about 20,000 rows. That's
  14757. 11:29:44a good size data set, right? 20,000
  14758. 11:29:47rows. That's decent. Um, we're
  14759. 11:29:49definitely missing some data here for
  14760. 11:29:52sure. Um, we could count how much we're
  14761. 11:29:54missing exactly by running this is NATO
  14762. 11:29:58sum. Um,
  14763. 11:30:00and so we see that total bedrooms is
  14764. 11:30:02missing about 200 uh 200 rows are
  14765. 11:30:06missing total bedroom uh value.
  14766. 11:30:12Okay. And then one thing I wanted to
  14767. 11:30:14look at is yes, this is a string. So
  14768. 11:30:16what remember what we can do with those?
  14769. 11:30:18That's a categorical.
  14770. 11:30:20So ocean proximity
  14771. 11:30:26is a categorical
  14772. 11:30:28string
  14773. 11:30:30feature.
  14774. 11:30:33So we can take a look at its value
  14775. 11:30:35counts, which is usually a good idea to
  14776. 11:30:37take a look and see what possible values
  14777. 11:30:40that feature could be. So if we look at
  14778. 11:30:43our
  14779. 11:30:45um what are we calling this? Housing
  14780. 11:30:47data.
  14781. 11:30:51Housing data
  14782. 11:30:54ocean
  14783. 11:30:56proximity
  14784. 11:31:01dot value counts.
  14785. 11:31:09So here's the different types that that
  14786. 11:31:11one can be. So there's some
  14787. 11:31:12neighborhoods that are less than 1 hour
  14788. 11:31:14from the ocean. There's some that are
  14789. 11:31:16inland. There's some that are near the
  14790. 11:31:18ocean. There's some that are near a bay.
  14791. 11:31:21There's even five of them that are on an
  14792. 11:31:22island. So, these are the different
  14793. 11:31:25values of the ocean proximity. So,
  14794. 11:31:28remember, you can always do that. If you
  14795. 11:31:29see a string feature, you can always
  14796. 11:31:32take a look at what its um categories
  14797. 11:31:34are. And it looks like most things are
  14798. 11:31:37less than 1 hour from the ocean, but
  14799. 11:31:38it's kind of evenly distributed here. Um
  14800. 11:31:42otherwise
  14801. 11:31:44very few islands.
  14802. 11:31:52But as you can imagine like this feature
  14803. 11:31:54is probably going to be important for
  14804. 11:31:56determining um what the value is, right?
  14805. 11:32:00Probably going to be important.
  14806. 11:32:07Okay. So, um, we need to deal with these
  14807. 11:32:10NLES. If we're going to build a model,
  14808. 11:32:12right? So, um, this is all of our
  14809. 11:32:14typical data prep. If we want to build a
  14810. 11:32:17model, we're going to have to deal with
  14811. 11:32:18these NLES. What do you guys think we
  14812. 11:32:20should do with the NLES? What would you
  14813. 11:32:22what do you think for total bedrooms?
  14814. 11:32:24What do you think is a good strategy to
  14815. 11:32:26do? Keep in mind, we have 20,000 points,
  14816. 11:32:3120,000 rows I should say, and about 200
  14817. 11:32:35of them are null.
  14818. 11:32:42Right. So about 200 are null. Um so what
  14819. 11:32:47do you what do you guys think would be
  14820. 11:32:48like a good strategy to deal with those
  14821. 11:32:49nles in that case?
  14822. 11:33:10average. We can't ignore it because we
  14823. 11:33:14can't ignore that column.
  14824. 11:33:17We can't ignore the whole column. So,
  14825. 11:33:20something needs to go there.
  14826. 11:33:29Probably don't want to make it zero.
  14827. 11:33:31I think average is a decent average is a
  14828. 11:33:34decent idea. Probably don't want to make
  14829. 11:33:36it zero because um that would indicate
  14830. 11:33:39that there's no bedrooms and yet we
  14831. 11:33:41still have a bunch of total rooms. So,
  14832. 11:33:43it probably doesn't make sense to do
  14833. 11:33:45zero.
  14834. 11:33:50Average, I think average could be a
  14835. 11:33:52decent one.
  14836. 11:33:54Now, in this example, what we're
  14837. 11:33:56actually going to do is we're
  14838. 11:34:01rows.
  14839. 11:34:03We're actually going to drop the rows al
  14840. 11:34:04together. Now, why are we doing that?
  14841. 11:34:06It's because we have so much data and
  14842. 11:34:10only 200 of them are null.
  14843. 11:34:14Okay, only 200 of them are null. So,
  14844. 11:34:16we're actually just going to drop the
  14845. 11:34:17rows. Now, that's a choice.
  14846. 11:34:22Um, that's a choice, right? Is that we
  14847. 11:34:25could fill in with the average like you
  14848. 11:34:26guys are suggesting. What we're actually
  14849. 11:34:29going to do is just drop the rows. It It
  14850. 11:34:31makes up less. It makes up about 1% of
  14851. 11:34:35the whole data. So it's not that much of
  14852. 11:34:38it is missing. We can drop those rows.
  14853. 11:34:41So that's actually what we're going to
  14854. 11:34:42do here is we remove all the roles with
  14855. 11:34:46the NLES by doing drop NA. So this just
  14856. 11:34:48drops them. So those rows are cut out.
  14857. 11:34:51Um, it's arguable that we could replace
  14858. 11:34:57it's arguable that we could just replace
  14859. 11:34:58it with something and I think you guys
  14860. 11:35:00have good thoughts which is the average
  14861. 11:35:02a default
  14862. 11:35:04um assume total bedrooms. We could we
  14863. 11:35:07could try that. Yeah.
  14864. 11:35:10Assign a value based on comparable home
  14865. 11:35:13value. Yes, you could do that too.
  14866. 11:35:14That's a good strategy is to look at the
  14867. 11:35:17other rows that are similar to it and
  14868. 11:35:19fill in a value. That's absolutely fair.
  14869. 11:35:22Um, in this example, we're actually just
  14870. 11:35:24going to drop those rows,
  14871. 11:35:28but I think that's totally um totally
  14872. 11:35:31valid.
  14873. 11:35:33This is a choice.
  14874. 11:35:36We could fill NA with different values
  14875. 11:35:42such as the average,
  14876. 11:35:45total bedrooms,
  14877. 11:35:49um, derive a value, etc. So, we could
  14878. 11:35:53derive something, which I think Brent,
  14879. 11:35:55you have a good suggestion. That's a
  14880. 11:35:57good suggestion. Um, we could derive
  14881. 11:35:59something like that, uh, and fill in the
  14882. 11:36:01blank, and that's I think that's totally
  14883. 11:36:03valid. Um, we could take the average of
  14884. 11:36:06the um bedrooms. Uh, I meant total rooms
  14885. 11:36:10here. Sorry, total rooms. Um, we could
  14886. 11:36:13fill in we could fill it in with the
  14887. 11:36:15total rooms for that category um or for
  14888. 11:36:18that row. Um, many options. In this
  14889. 11:36:23case, we're actually just going to drop
  14890. 11:36:24those rows because they make up such a
  14891. 11:36:27small percentage relative to the 20,000
  14892. 11:36:30rows that we have. It's about 1%, right?
  14893. 11:36:33200 rows is about 1% of 20,000.
  14894. 11:36:37So, we're just going to drop them. But
  14895. 11:36:39that's a choice. We don't have to drop
  14896. 11:36:41them. We could fill in with something.
  14897. 11:36:43Um, and if we did that, we would use
  14898. 11:36:46fill na rather than drop NA, right?
  14899. 11:37:02Uh after dropping the rows, how many? So
  14900. 11:37:05it's just so after we drop the rows, um
  14901. 11:37:08after we drop the rows, it's just going
  14902. 11:37:09to be we still have all our other rows
  14903. 11:37:12are intact, right? So if we look at this
  14904. 11:37:14now,
  14905. 11:37:22we now have um slightly uh slightly less
  14906. 11:37:26entries.
  14907. 11:37:28So now we have this this many um rather
  14908. 11:37:32than rather than this many,
  14909. 11:37:36right? We dropped those 200
  14910. 11:37:45But they're all filled in. Yeah, they're
  14911. 11:37:47So all the other columns are still
  14912. 11:37:48filled in. We're just we're we're
  14913. 11:37:50cutting out the whole row. So if you
  14914. 11:37:52think about our data set, um we have all
  14915. 11:37:54these rows and all these columns. What
  14916. 11:37:58we're doing is like if there's a null
  14917. 11:38:00here, we're just we're just getting rid
  14918. 11:38:02of that whole row, right? And so we
  14919. 11:38:05still have all the other rows intact.
  14920. 11:38:14Uh, we can drop them because we have a
  14921. 11:38:16good sample size. Yes,
  14922. 11:38:22that's exactly right, Ronald. Yep, we
  14923. 11:38:24can drop them because we have we have
  14924. 11:38:2520,000 rows and only 200 are missing
  14925. 11:38:28values. So, that's totally fine.
  14926. 11:38:35Uh, drop a removes all rows that has any
  14927. 11:38:37null. Yes, that's true. It it will go
  14928. 11:38:40ahead and just drop any row where
  14929. 11:38:42there's any null, no matter what column
  14930. 11:38:44it's in. Yes,
  14931. 11:38:55index. Yeah, the index is not getting
  14932. 11:38:57reset. Um that's true. So um what we
  14933. 11:39:03what you can always do is you can reset
  14934. 11:39:05the index. So, um, if you want to, it's
  14935. 11:39:08optional. We we're not really going to
  14936. 11:39:10use the index for anything that
  14937. 11:39:12important, right? But what we could do
  14938. 11:39:14is, uh, reset index.
  14939. 11:39:19Uh,
  14940. 11:39:21we could do that, right? Which will
  14941. 11:39:22reset it.
  14942. 11:39:36So now now it gets reset.
  14943. 11:39:56But um let me actually I don't I don't
  14944. 11:39:59really want to do that. I'm going to
  14945. 11:40:01reset this.
  14946. 11:40:08Um,
  14947. 11:40:16yeah, we could do that.
  14948. 11:40:38Okay.
  14949. 11:40:40So now importantly there should be uh no
  14950. 11:40:43missing data of this of this new one
  14951. 11:40:45where we've dropped NAS right. So now
  14952. 11:40:46this is good. If you now the reason we
  14953. 11:40:49had to do this is because if we try to
  14954. 11:40:51build a linear regression and we have
  14955. 11:40:53nles in there um the the issue is like
  14956. 11:40:57how do you build a model where you have
  14957. 11:40:59something like this
  14958. 11:41:08and these are null like what do how do
  14959. 11:41:10you multiply a number by a null?
  14960. 11:41:14Um, we can't really do that, right?
  14961. 11:41:17We can't really do that. So, um,
  14962. 11:41:27so therefore, uh, we need to get rid of
  14963. 11:41:30NLES like the the null's not really
  14964. 11:41:32going to work in there. So, uh, we need
  14965. 11:41:35to get rid of them for linear regression
  14966. 11:41:37to to really have a chance to work,
  14967. 11:41:38right? To train it and be able to use
  14968. 11:41:40it.
  14969. 11:41:42You got to get rid of those nles.
  14970. 11:41:50All right,
  14971. 11:41:52any questions so far? So, we haven't
  14972. 11:41:54done any modeling yet. We're doing some
  14973. 11:41:55We're doing some data preparation before
  14974. 11:41:57we get to the modeling. And we haven't
  14975. 11:41:59done any cross validation yet. We
  14976. 11:42:00haven't set that up. We're just doing
  14977. 11:42:02our data preparation before we get to
  14978. 11:42:04the modeling. Right. So, we've dropped
  14979. 11:42:06some NAS. We've checked it. Um, we're
  14980. 11:42:10going to do one more prep step, which is
  14981. 11:42:12to um change that ocean proximity
  14982. 11:42:16feature into something numerical because
  14983. 11:42:18again, how do you build a model where
  14984. 11:42:20you're inserting a string into those
  14985. 11:42:23like beta 1, beta 2, beta 3 times of
  14986. 11:42:25features? You can't really do that when
  14987. 11:42:27it's a string. Um, so what we're going
  14988. 11:42:30to do, and I'm going to get rid of this
  14989. 11:42:32because I don't think we really need
  14990. 11:42:33that. um is we are going to uh run this
  14991. 11:42:38get dummies function which is our um our
  14992. 11:42:42get dummies function is our usual one to
  14993. 11:42:46uh our git dummies one is our usual one
  14994. 11:42:48to um
  14995. 11:42:51uh get our one hot encoding. So this is
  14996. 11:42:55our uh one hot encoding here.
  14997. 11:43:02We now are going to have data that's
  14998. 11:43:05like this, right? So we have ocean. So
  14999. 11:43:08So by the way, this prefix
  15000. 11:43:10um this prefix is OP, which which is
  15001. 11:43:14short for ocean proximity, right? So we
  15002. 11:43:16have ocean proximity uh less than 1 hour
  15003. 11:43:19from the ocean, ocean proximity inland,
  15004. 11:43:22ocean proximity island, near bay, near
  15005. 11:43:24ocean. So these first five rows are near
  15006. 11:43:27the bay. Um so they have a one there and
  15007. 11:43:30a zero in the other spots. So this is
  15008. 11:43:32good. This one hot encodes that feature
  15009. 11:43:35into these numerical uh values,
  15010. 11:43:39right?
  15011. 11:43:41Were you guys able to run that one? The
  15012. 11:43:43get dummies
  15013. 11:43:51So the reason that Yeah, that's a great
  15014. 11:43:53question. How did it go ocean proximity?
  15015. 11:43:54It's because um that is the only uh
  15016. 11:43:58string feature we have. That's the only
  15017. 11:44:01one we have. So it it's going to look
  15018. 11:44:03for any non-numericals and one hot
  15019. 11:44:05encode those however many however many
  15020. 11:44:07there are. So whatever objects we have
  15021. 11:44:10which are strings, it's going to
  15022. 11:44:12automatically oneh hot encode those.
  15023. 11:44:23Yeah, we could have right we could have
  15024. 11:44:24went here and did Right. We could have
  15025. 11:44:27done ocean
  15026. 11:44:30proximity,
  15027. 11:44:33but we only have one of those features.
  15028. 11:44:36So, it's just going to do that to the
  15029. 11:44:38whole data frame
  15030. 11:44:40uh on that one feature. So, what we're
  15031. 11:44:43going to do is um go ahead and split it
  15032. 11:44:46into an X and a Y um which the X is
  15033. 11:44:50always what includes our features. The Y
  15034. 11:44:54is what we are trying to predict, which
  15035. 11:44:55is the label. Now, um, in order to
  15036. 11:44:59separate those out, what we're going to
  15037. 11:45:01do is assign X to be the variable that
  15038. 11:45:04is, um, our data frame minus this median
  15039. 11:45:08house value column. So what this is
  15040. 11:45:10doing is um uh it's not permanently
  15041. 11:45:14dropping because we're not uh dropping
  15042. 11:45:17it in place but it is returning us a
  15043. 11:45:21copy of the data frame with the median
  15044. 11:45:24house value column left out right it's
  15045. 11:45:26dropped. So this is this is uh something
  15046. 11:45:29we want to do because that will the rest
  15047. 11:45:31of it will contain our features, right?
  15048. 11:45:33So um this will temporarily or I should
  15049. 11:45:38say return a copy of the DF with um
  15050. 11:45:44median house value
  15051. 11:45:48dropped,
  15052. 11:45:50right? Median house value dropped. Um so
  15053. 11:45:53we go ahead and drop that one. Uh now
  15054. 11:45:57remember it's not permanent. It's just
  15055. 11:45:58giving us uh the remainder of it which
  15056. 11:46:01is this housing data dropping this and
  15057. 11:46:04it's assigning that to x and then we're
  15058. 11:46:06taking the actual median house value
  15059. 11:46:08column from the original data and
  15060. 11:46:11assigning that to y. So this is going to
  15061. 11:46:13be our labels
  15062. 11:46:17right. So this is what we are trying to
  15063. 11:46:22predict.
  15064. 11:46:25Okay, so that is our Y and that's always
  15065. 11:46:27how it is. X is our features, Y is our
  15066. 11:46:29labels. Um, hopefully that makes sense.
  15067. 11:46:32What this is doing is this is going to
  15068. 11:46:35get rid of that label column and
  15069. 11:46:37everything else will be our features.
  15070. 11:46:39And then this will get rid of this will
  15071. 11:46:41just assign the label column to Y.
  15072. 11:46:46All right. And then what we can do is
  15073. 11:46:48pass X and Y into our train test split
  15074. 11:46:51function. And this will generate the
  15075. 11:46:54hold out set. So if we want to do the
  15076. 11:46:56hold out cross validation, this is how
  15077. 11:46:58we would do it is we would split the
  15078. 11:47:01data into X-ray, X test, Y train, Y test
  15079. 11:47:04um using train test split. So this is
  15080. 11:47:08what we did last time. This would be
  15081. 11:47:10this would be for hold out cross
  15082. 11:47:15validation,
  15083. 11:47:18right? where we are uh uh just have that
  15084. 11:47:22one one set for testing and one set for
  15085. 11:47:25uh one set for training one test one set
  15086. 11:47:28for testing I should say right so this
  15087. 11:47:31is pretty standard train test split um
  15088. 11:47:33we pass in that x we pass in the y we
  15089. 11:47:36use a 30% test size is pretty standard
  15090. 11:47:39and random state so that we get the
  15091. 11:47:41consistent shuffling if we were to run
  15092. 11:47:43this multiple times um we we get that uh
  15093. 11:47:46consistent randomization
  15094. 11:47:50Okay,
  15095. 11:47:51so we have that and so now our X train
  15096. 11:47:55is a percentage um of the data frame of
  15097. 11:47:59the 20,000 uh rows and the X test is uh
  15098. 11:48:0430% of that. So it's only about 6,000
  15099. 11:48:06rows which is what um the shape of that
  15100. 11:48:08is.
  15101. 11:48:12Yeah. X so X is our features. So we're
  15102. 11:48:15we're putting all of our data in that is
  15103. 11:48:17our features into X. And so the the um
  15104. 11:48:21most efficient way of doing that is um
  15105. 11:48:24the most efficient way of doing that is
  15106. 11:48:26to
  15107. 11:48:28uh just take our data and drop the
  15108. 11:48:31median house value column because that's
  15109. 11:48:33our label column. So we just remove
  15110. 11:48:36that. The rest of the data is our
  15111. 11:48:38features. So that's what that's what
  15112. 11:48:40this X is, right? It's all of our
  15113. 11:48:42feature data. All of our columns that is
  15114. 11:48:44not the label column essentially is what
  15115. 11:48:46that's doing. And then Y is our label
  15116. 11:48:49column from our original data.
  15117. 11:48:54Right? Y is our label column. And so
  15118. 11:48:58this this will um contain all of our
  15119. 11:49:00labels which is the median house value.
  15120. 11:49:04X X X contains every column but the one
  15121. 11:49:07we're gonna so we we ultimately decide
  15122. 11:49:10that but X contains um X is everything
  15123. 11:49:15that is not our dependent variable which
  15124. 11:49:18is what we're predicting. So we're
  15125. 11:49:20removing what we are trying to predict
  15126. 11:49:22from X. X should be everything else.
  15127. 11:49:25That's always how it's going to be. X is
  15128. 11:49:27X is always going to be all of those
  15129. 11:49:29independent variables that we're using
  15130. 11:49:32to predict the median house value. So we
  15131. 11:49:36are going to predict the median house
  15132. 11:49:38value. We need to remove it from X.
  15133. 11:49:42So we're we're taking everything but
  15134. 11:49:44that column.
  15135. 11:49:49So it's the whole data frame. It's the
  15136. 11:49:52whole data frame minus this one column
  15137. 11:49:54with just the dependent variable. Right.
  15138. 11:49:59Exactly right. Removing the dependent
  15139. 11:50:00variable and keeping all the
  15140. 11:50:01independence. That's exactly right.
  15141. 11:50:04Exactly right. So think about it in
  15142. 11:50:07terms of the model. Let's go back to the
  15143. 11:50:08features. Right. Think about it in terms
  15144. 11:50:10of the model. We are trying to predict
  15145. 11:50:13this this value. We're building a model
  15146. 11:50:16to try to predict this. So we are going
  15147. 11:50:19to make sure X is everything but this
  15148. 11:50:24right. So this is actually just Y.
  15149. 11:50:27That's our label. That's our dependent
  15150. 11:50:29variable. Right? That's Y. Everything
  15151. 11:50:32else is belongs to X. Everything else
  15152. 11:50:35belongs to X including all of these.
  15153. 11:50:42Right? We choose this one to be Y
  15154. 11:50:44because we're building a model to
  15155. 11:50:45predict that. That's our label.
  15156. 11:50:49All right. So, we have our we use X and
  15157. 11:50:52Y to do our train test split. So, we
  15158. 11:50:54have our our training features and our
  15159. 11:50:57test features and then our training
  15160. 11:50:58label and test labels here. Um, pretty
  15161. 11:51:01standard there.
  15162. 11:51:04Um, okay. So, this is what's new is if
  15163. 11:51:06we want to do K-fold uh validation, what
  15164. 11:51:10we're going to do is create a kfold
  15165. 11:51:12object. So, we have this kffold from
  15166. 11:51:14scikitlearn that we already imported. we
  15167. 11:51:17are going to create a kfold um where we
  15168. 11:51:22are going to specify how many folds we
  15169. 11:51:24want. So that is the in uh inslits
  15170. 11:51:27parameter as this says um this is going
  15171. 11:51:30to be uh uh in this case we're going to
  15172. 11:51:34do 10 folds. That's pretty standard. So
  15173. 11:51:36I think the typical number of folds that
  15174. 11:51:38I've seen and I've worked with in my in
  15175. 11:51:40my career is usually five or 10.
  15176. 11:51:44Five or 10 folds is the standard.
  15177. 11:51:49Okay. So, we're doing 10 folds in this
  15178. 11:51:51case and we're setting a random state
  15179. 11:51:53because we're going to do shuffling. So,
  15180. 11:51:55in order to produce those folds, we're
  15181. 11:51:57going to shuffle the data first and then
  15182. 11:51:58split it into five folds, right? So,
  15183. 11:52:01this this kfold object is going to
  15184. 11:52:04manage creating these splits for us,
  15185. 11:52:08right? These even splits. I know I I
  15186. 11:52:10didn't draw it even, but um it's going
  15187. 11:52:13to manage these five folds for us and
  15188. 11:52:15it's going to shuffle the data and
  15189. 11:52:17assign them to these different folds and
  15190. 11:52:19we're and then what we're going to do is
  15191. 11:52:21use those to do our training. We're
  15192. 11:52:24going to execute the cross validation
  15193. 11:52:26using this k-fold object.
  15194. 11:52:29Okay, so we create the kffold
  15195. 11:52:32um we initialize our model as well. So,
  15196. 11:52:35of course, in order to train something
  15197. 11:52:38uh in the Kfolds, we're going to need a
  15198. 11:52:39model. In this case, we're using linear
  15199. 11:52:42regression, right? Which is which is the
  15200. 11:52:44model we've been studying so far. So,
  15201. 11:52:46you have a linear regression. Um now,
  15202. 11:52:49look how easy it's going to be in order
  15203. 11:52:51to execute cross validation. All we need
  15204. 11:52:54to do is um all we need to do is create
  15205. 11:52:58a crossfile score function.
  15206. 11:53:02um or I should say use the cross val
  15207. 11:53:04score function from scikitlearn. So we
  15208. 11:53:06use that with the model we want to
  15209. 11:53:08train. So our model goes first. So
  15210. 11:53:11that's the linear regression object.
  15211. 11:53:14Then our data. So our extra our features
  15212. 11:53:17and our label for our training.
  15213. 11:53:20And then um let me skip over this for a
  15214. 11:53:23second. I'll explain what this is in a
  15215. 11:53:24second. Um but then we are using uh the
  15216. 11:53:30cross validation technique is our kfold.
  15217. 11:53:33So this is where our k-fold object goes
  15218. 11:53:35in the CV parameter which is cross
  15219. 11:53:37validation. So what cross validation
  15220. 11:53:40strategy are you using? We're using
  15221. 11:53:42kfold and the kfold we're using is this
  15222. 11:53:44one we defined up here KF. So we're
  15223. 11:53:47putting that right here for this. And
  15224. 11:53:49then um in jobs um allows us to
  15225. 11:53:54parallelize this. So if we set it to
  15226. 11:53:55negative one, that's the that that's the
  15227. 11:53:58default. Um it will do it will actually
  15228. 11:54:01train across the different combinations
  15229. 11:54:03in parallel. Um which speeds it up. So
  15230. 11:54:06you want to you want to keep this to
  15231. 11:54:07negative one if you can. So um now let
  15232. 11:54:12me describe the scoring. So what this
  15233. 11:54:15means is we put in our metric here. Um
  15234. 11:54:20and so you can put mean absolute error,
  15235. 11:54:23you can put in mean squared error. Um
  15236. 11:54:25those are the two that we can use. And
  15237. 11:54:29um the reason we it has a negative in
  15238. 11:54:31front of it is because we want to find
  15239. 11:54:34the one that has the lowest score.
  15240. 11:54:38That's going to be our best model is the
  15241. 11:54:41one that has the lowest score. So we
  15242. 11:54:43take the absolute value.
  15243. 11:54:46I'm sorry. we take the abs the the the
  15244. 11:54:48metric and we take the negative of it um
  15245. 11:54:52because the highest scoring one is going
  15246. 11:54:56to be the closest to zero. Um so it's
  15247. 11:54:59just a we use the we use the negative of
  15248. 11:55:02the of the metric um because on the
  15249. 11:55:05number line like the the highest um
  15250. 11:55:08scoring one should be the least um or I
  15251. 11:55:12should say the maximum negative that we
  15252. 11:55:14can get. that's going to be closest to
  15253. 11:55:16this to zero. So if here's zero, this
  15254. 11:55:18would be like -1 is better than -10,
  15255. 11:55:22right? So something that scores um the
  15256. 11:55:25maximum negative uh absolute error would
  15257. 11:55:29be closest to zero
  15258. 11:55:32and something that has more is going to
  15259. 11:55:34be on this side.
  15260. 11:55:37So this is only the reason we need this
  15261. 11:55:39is only just to keep track of the scores
  15262. 11:55:42of each individual um fold. Okay.
  15263. 11:55:47So the one so the reason we can do that
  15264. 11:55:49is at the end we can kind of see which
  15265. 11:55:51which combination performed the best. Um
  15266. 11:55:55it's going to be the one that has the
  15267. 11:55:57highest uh highest value of the negative
  15268. 11:56:00which is closest to zero.
  15269. 11:56:08That's just a convention.
  15270. 11:56:11Yeah, it's just because um it's because
  15271. 11:56:13the cross validation is looking to
  15272. 11:56:15maximize the metric. So, whatever has
  15273. 11:56:18the best score
  15274. 11:56:20um whatever has the best score is
  15275. 11:56:23considered the best uh performance. Um
  15276. 11:56:25but we are using uh something where
  15277. 11:56:28lower is better. So we we take the
  15278. 11:56:31negative and like the the highest
  15279. 11:56:33negative would be closest to zero,
  15280. 11:56:37right? The highest negative is going to
  15281. 11:56:38be closest to zero.
  15282. 11:56:42So that so it's it's just because like
  15283. 11:56:45we want the lower score to be the best.
  15284. 11:56:49The lowest score should be the best.
  15285. 11:56:52So we take the negative of it. Um and so
  15286. 11:56:56something that is more negative is going
  15287. 11:56:58to be worse. Yeah, that's the reason.
  15288. 11:57:03So something that's down this way is
  15289. 11:57:05going to be worse. Okay, so it runs this
  15290. 11:57:11and what you can see is if we actually
  15291. 11:57:13print this out, if we print out our
  15292. 11:57:15k-fold scores, what we should get is 10
  15293. 11:57:17different scores.
  15294. 11:57:21And you can see um we have 10 different
  15295. 11:57:24uh scores here which are all negative
  15296. 11:57:27because we're taking the negative of the
  15297. 11:57:28absolute of the mean absolute error. Um
  15298. 11:57:32so what we would be looking for here is
  15299. 11:57:37um we want to take the average of these
  15300. 11:57:39scores but take the absolute value of
  15301. 11:57:42them to get the best performance. So
  15302. 11:57:44this is capturing like this is the score
  15303. 11:57:46on the first fold combination. This is
  15304. 11:57:49the score on the second fold
  15305. 11:57:50combination. This is the score on the
  15306. 11:57:52third fold combination and on and on and
  15307. 11:57:55on. And these are the absolute errors.
  15308. 11:57:58Okay, these are the absolute errors. Um
  15309. 11:58:02so if we take a look at computing the uh
  15310. 11:58:05average, which by the way, we don't need
  15311. 11:58:07this import because we're using the
  15312. 11:58:08numpy average. So that's fine. um we can
  15313. 11:58:12take the absolute value of those um and
  15314. 11:58:15take a look at the average MSE
  15315. 11:58:20or sorry MAE. Now I want you to think
  15316. 11:58:24about this this uh average performance.
  15317. 11:58:27So this is our performance right here on
  15318. 11:58:29the cross validation.
  15319. 11:58:31This is our average
  15320. 11:58:33MAE
  15321. 11:58:35across all of our fold combinations. So
  15322. 11:58:37that's a that's an indicator of our
  15323. 11:58:38performance, right? um for the cross
  15324. 11:58:41validation.
  15325. 11:58:44Now, what are the units of our original
  15326. 11:58:49uh the original median value? They're
  15327. 11:58:53already in the thousands, right? So, if
  15328. 11:58:56we go to that feature, they're already
  15329. 11:58:59in these hundreds of thousands. So, this
  15330. 11:59:03is not a very good error. It's it's kind
  15331. 11:59:07of high, right? because it's in this is
  15332. 11:59:1049,000.
  15333. 11:59:12Um that's that's how far away we are in
  15334. 11:59:15absolute value on average is $49,000 um
  15335. 11:59:19dollars on the median value. That's not
  15336. 11:59:22very good. So this score
  15337. 11:59:26this score is
  15338. 11:59:28um not very good. So this model is not
  15339. 11:59:32performing that well and we can see that
  15340. 11:59:34by comparing this error to our actual uh
  15341. 11:59:37data. So this is right around 50,000
  15342. 11:59:42and our median uh house values are in
  15343. 11:59:45the hundreds of thousands. So on average
  15344. 11:59:48we're 50,000 off when we make a
  15345. 11:59:51prediction. That's a significant amount,
  15346. 11:59:54right? It's a significant amount on
  15347. 11:59:56average um when our when our data is in
  15348. 11:59:59about the hundreds of thousands here.
  15349. 12:00:03So we are um we have a significant
  15350. 12:00:06amount of error 50,000 relative to the h
  15351. 12:00:09to our units that our our data is in.
  15352. 12:00:11Right? Um so this score is not very
  15353. 12:00:14good. Um
  15354. 12:00:17and so we see that from the cross
  15355. 12:00:18validation. So look how easy the cross
  15356. 12:00:20val is. Again we just do cross val
  15357. 12:00:22score. We put in our model. We put in
  15358. 12:00:24our data. We put in our cross validation
  15359. 12:00:27uh strategy here which is k-fold and we
  15360. 12:00:30can generate these metrics across all
  15361. 12:00:32the fold combinations. So it's this
  15362. 12:00:35function is taking care of rotating
  15363. 12:00:37those and doing every combo with just
  15364. 12:00:39the 10 different combinations here of
  15365. 12:00:42the of the folds.
  15366. 12:00:4410 different instances where you have
  15367. 12:00:46you know 10 different folds are the ones
  15368. 12:00:48that are left out for evaluation.
  15369. 12:00:50Um so it's managing that for us using
  15370. 12:00:53this data right using this training data
  15371. 12:00:55here. Um and we uh we generate these um
  15372. 12:01:02generate these scores.
  15373. 12:01:06Okay. So that's kfold. It's not hard to
  15374. 12:01:08do. All you have to do is um just use a
  15375. 12:01:12cross file score. And we could change
  15376. 12:01:14this to mean squared error. That's you
  15377. 12:01:17know we could do that too. That'd be
  15378. 12:01:18pretty easy. Um, so that'd be no issue.
  15379. 12:01:23We just happen to be using the absolute
  15380. 12:01:24error here. Of course, we could use
  15381. 12:01:27squared error.
  15382. 12:01:29Were you guys able to get this to run
  15383. 12:01:31kfold scores?
  15384. 12:01:34It produces an array of 10 10 different
  15385. 12:01:36scores, which should make sense because
  15386. 12:01:38those are these are the um we're
  15387. 12:01:41splitting our data into 10 different
  15388. 12:01:43folds,
  15389. 12:01:45right?
  15390. 12:01:4710 different folds and leaving one out
  15391. 12:01:49to do our evaluation on. So the one that
  15392. 12:01:51gets left out every time is what's
  15393. 12:01:53producing these scores. So it's 10
  15394. 12:01:55different ones get left out when we
  15395. 12:01:57rotate through all the combinations.
  15396. 12:02:02And so we average these scores
  15397. 12:02:11and we get this amount of we get about
  15398. 12:02:1350,000 in error on average.
  15399. 12:02:22Um, what do you think would be what do
  15400. 12:02:24you think would be acceptable? So, if
  15401. 12:02:25our if we're predicting the price, like
  15402. 12:02:27if we're a real estate agent and we're
  15403. 12:02:29predicting these prices and they
  15404. 12:02:32typically are
  15405. 12:02:34Yeah, close to zero would be great.
  15406. 12:02:36That'd be fantastic. Closer to zero
  15407. 12:02:38would be better. The average is um
  15408. 12:02:40206,000.
  15409. 12:02:43So 50,000 is a decent percentage of
  15410. 12:02:46that. Um so you know you can compute it
  15411. 12:02:49as a percentage right? So 50,000 is a
  15412. 12:02:52decent percentage of that. Um probably
  15413. 12:02:55you want this to be less than 20,000
  15414. 12:02:58would be about 10% error. 20,000
  15415. 12:03:03right? So maybe like 30,000 somewhere in
  15416. 12:03:07there.
  15417. 12:03:12Yeah. 10% would be 5% error. 10,000
  15418. 12:03:15would be 5% error. That's true. That's
  15419. 12:03:16true. So that would be that would be
  15420. 12:03:18much better. So being closer to zero,
  15421. 12:03:20like the smaller the better, of course.
  15422. 12:03:23Of course. Um but yeah, I would say an
  15423. 12:03:26acceptable percentage of error is
  15424. 12:03:28probably 20%.
  15425. 12:03:31Probably 20%, which would be um like
  15426. 12:03:3440,000 or less would probably be
  15427. 12:03:36acceptable.
  15428. 12:03:39Usually when we usually when you build
  15429. 12:03:41models um 80% accuracy is usually uh
  15430. 12:03:46considered decent.
  15431. 12:03:48Usually considered decent
  15432. 12:03:5180%. So I'd say 40,000 or less would be
  15433. 12:03:55kind of ideal.
  15434. 12:04:00Does that make sense to answer the
  15435. 12:04:03question?
  15436. 12:04:05That's a good question. What value is
  15437. 12:04:06acceptable? I think probably less than
  15438. 12:04:0940,000 would be ideal. That's right
  15439. 12:04:12around 20% error.
  15440. 12:04:16All right, so that's kfold. Um let's do
  15441. 12:04:19just a regular hold out now. So this is
  15442. 12:04:21just using our training and test data.
  15443. 12:04:23Um doing model.fit and calculating an
  15444. 12:04:26MSE on the test data. So this is this is
  15445. 12:04:29just the um hold out strategy here where
  15446. 12:04:32we just have um this is less robust but
  15447. 12:04:35it's a lot quicker to do and easier to
  15448. 12:04:37set up. Right? So um this is using the
  15449. 12:04:41hold out strategy. So just a regular
  15450. 12:04:46um train test split.
  15451. 12:04:50Are we going to rebuild the model? No,
  15452. 12:04:52not necessarily. There's some things we
  15453. 12:04:54could do most likely. And like one thing
  15454. 12:04:58we did not do was scale our features.
  15455. 12:05:01Remember I said that's a pretty
  15456. 12:05:02important thing to do is to scale our
  15457. 12:05:05features. We did not do that. So that
  15458. 12:05:07would be an enhancement to this that
  15459. 12:05:08we're going to So I I actually do think
  15460. 12:05:10we'll do that later. Yes. So I think we
  15461. 12:05:13will actually do that now that I'm
  15462. 12:05:14thinking about it. Yes. One of the
  15463. 12:05:17things we can do is scale these features
  15464. 12:05:19using like a minmax scaler, a standard
  15465. 12:05:21scaler. that's actually going to help us
  15466. 12:05:23um that's going to help us do better
  15467. 12:05:26predictions.
  15468. 12:05:30So that that's one thing we could do. Um
  15469. 12:05:33but yeah, we will we'll try to see if we
  15470. 12:05:36can get better.
  15471. 12:05:38It should help it. Yeah, usually you
  15472. 12:05:40want to scale you want to scale the
  15473. 12:05:42data. That's something we didn't do in
  15474. 12:05:43our preparation step. We did a lot of
  15475. 12:05:45the things we should do. We removed nles
  15476. 12:05:47and we did one hot encoding to the
  15477. 12:05:49proximity feature like this one. Um
  15478. 12:05:53those are good to do but we didn't scale
  15479. 12:05:56any of these other we didn't scale any
  15480. 12:05:58of the features right we didn't scale
  15481. 12:06:00any of them. Um it you it will have an
  15482. 12:06:04effect. It usually when we scale it
  15483. 12:06:06it'll be a better model.
  15484. 12:06:09It'll it'll learn a little bit better if
  15485. 12:06:11we can scale the data. Um so that way
  15486. 12:06:15like these
  15487. 12:06:18um like ages aren't you know drastically
  15488. 12:06:21different than like in scale then total
  15489. 12:06:23bedrooms or income
  15490. 12:06:26uh those kind of things. So we usually
  15491. 12:06:28want these to be in a similar scale
  15492. 12:06:30range.
  15493. 12:06:33So we'll we will I think we'll scale
  15494. 12:06:35them coming up in a bit and it should
  15495. 12:06:38help the model.
  15496. 12:06:41We've talked about that before, right?
  15497. 12:06:42scaling usually is a good idea to do
  15498. 12:06:44when you're prepping your data for
  15499. 12:06:46modeling.
  15500. 12:06:49No, you want to you want to scale your
  15501. 12:06:51test data as well. You're going to do
  15502. 12:06:53both. You're going to scale your
  15503. 12:06:54training data, you're going to scale it.
  15504. 12:06:56So, that's actually a good point you
  15505. 12:06:58bring up is any transformations you do
  15506. 12:07:00on your training to build your model,
  15507. 12:07:02you should also do on your test set so
  15508. 12:07:05you get an applesto apples comparison.
  15509. 12:07:07You should always do the same
  15510. 12:07:09transformations.
  15511. 12:07:11Yes.
  15512. 12:07:12Would scaling data impact K? Yeah, it
  15513. 12:07:14could it could make it better. It could
  15514. 12:07:16uh yeah, it should impact it. We should
  15515. 12:07:18get a better model. So when we do the
  15516. 12:07:20different folds, we'll get different
  15517. 12:07:22we'll get better scores. Yeah, it it
  15518. 12:07:24will impact
  15519. 12:07:28uh yeah, if they're so that's a good
  15520. 12:07:31point. If they're going to use our
  15521. 12:07:32model, then yes, they have to scale the
  15522. 12:07:34data as well. If they're gonna if we
  15523. 12:07:36build the model on the assumption that
  15524. 12:07:37the input is scaled, then yes, they have
  15525. 12:07:40to also scale their data when they're
  15526. 12:07:42using it with our model. That's true.
  15527. 12:07:49I mean, not really. I'll show you why.
  15528. 12:07:52There's something that's actually going
  15529. 12:07:53to make it easier um that that will
  15530. 12:07:56automate doing the scaling for them. So,
  15531. 12:07:58they don't they don't have to do the
  15532. 12:07:59scaling manually. it'll just it'll
  15533. 12:08:02happen automatically when they use the
  15534. 12:08:04model. I'm going to show you something
  15535. 12:08:05that's going to automate that which is
  15536. 12:08:07going to be called a pipeline.
  15537. 12:08:09So that part will be automated and they
  15538. 12:08:11won't have to do that. So it won't be
  15539. 12:08:13heavy on the user. No, in theory it is,
  15540. 12:08:17but
  15541. 12:08:18has a really helpful tool to make it
  15542. 12:08:20easy to do that. So I'm going to I'm
  15543. 12:08:22going to show us that um later on in the
  15544. 12:08:24notebook.
  15545. 12:08:27No, the data data is not for a single
  15546. 12:08:29house. It's for like a neighborhood. So
  15547. 12:08:31there's a certain number of households
  15548. 12:08:34in the neighborhood and this is the
  15549. 12:08:36we're predicting the median house value
  15550. 12:08:38of that neighborhood.
  15551. 12:08:41Yeah. So there's a there's certain
  15552. 12:08:42number of households. There's there's
  15553. 12:08:43like an a median income, a population,
  15554. 12:08:46certain number of people that live
  15555. 12:08:47there. Um proximity generally of where
  15556. 12:08:52that location is. It also has a latitude
  15557. 12:08:54and longitude.
  15558. 12:08:58So,
  15559. 12:09:01and a median age in that neighborhood.
  15560. 12:09:03So, yeah, it's not just a single house.
  15561. 12:09:14Okay, let's go back to this was the hold
  15562. 12:09:18out strategy. So, this is a lot simpler.
  15563. 12:09:20This is just model.fit, right? This is
  15564. 12:09:22just model.fit on the training uh data.
  15565. 12:09:25And then we um can predict on the test
  15566. 12:09:27features and generate test predictions.
  15567. 12:09:30And then we can compute our error on
  15568. 12:09:32those um we can compute our error
  15569. 12:09:34amongst the test predictions and our
  15570. 12:09:36test uh label. So that's our useful mean
  15571. 12:09:40squared error function, right? To to
  15572. 12:09:43compute the MSE. Um let's see what the
  15573. 12:09:47MSE is. So MSE is right here.
  15574. 12:09:53Um now what we can do is we can take the
  15575. 12:09:56MSE
  15576. 12:09:58and we can take the square root of it.
  15577. 12:09:59So let's actually do that. Let's um do
  15578. 12:10:02MP. square root of the
  15579. 12:10:05um test
  15580. 12:10:07MSE
  15581. 12:10:11and we get um 67 we get 67,000.
  15582. 12:10:16So that's pretty high on this. So when
  15583. 12:10:19we just now look at the difference of
  15584. 12:10:21that, right? When we just do a train
  15585. 12:10:23test split, um
  15586. 12:10:28when we just do a train test split, we
  15587. 12:10:30get a worse score because it's not as
  15588. 12:10:32it's not as robust, right? We're not
  15589. 12:10:34showing that to many of the other uh
  15590. 12:10:37folds. So we get a lot more error this
  15591. 12:10:40way on the test data.
  15592. 12:10:43So this is um actually worse performance
  15593. 12:10:46just doing the train test split.
  15594. 12:10:50This is a really higher.
  15595. 12:11:02Yeah, we can. We can. I'm going to I'm
  15596. 12:11:04going to show us how to how the scaling
  15597. 12:11:06will be done automatically. Yes, we can.
  15598. 12:11:10Um there's there's a really easy tool to
  15599. 12:11:13do that will scale it automatically.
  15600. 12:11:20It's going to be later in this notebook.
  15601. 12:11:21I'll show us it.
  15602. 12:11:30All right. So, just to recap this, this
  15603. 12:11:32is fitting the model.
  15604. 12:11:34This is fitting the model. This is
  15605. 12:11:36making the predictions, right?
  15606. 12:11:37Model.predict.
  15607. 12:11:39So, this is making the predictions. And
  15608. 12:11:41then this is calculating the error, the
  15609. 12:11:43mean squared error, which is looking at
  15610. 12:11:45our test labels versus our test
  15611. 12:11:47predictions, right? And this is
  15612. 12:11:49computing the distance, the average
  15613. 12:11:51distance away from these values to these
  15614. 12:11:54values,
  15615. 12:11:56right?
  15616. 12:11:57And then we can also compute the R squar
  15617. 12:12:00R R 2 and we see that it's not a very
  15618. 12:12:03good R squared. 65 uh is not a very
  15619. 12:12:06great model
  15620. 12:12:08um because it closer to one would be
  15621. 12:12:10better. So this is still this is not
  15622. 12:12:12very good.
  15623. 12:12:14We know that we knew that from the cross
  15624. 12:12:16file score. But this is just doing um
  15625. 12:12:18this is just doing a hold out uh where
  15626. 12:12:21we do a train and test split, right? So
  15627. 12:12:23it's a little bit simpler, but it's not
  15628. 12:12:25quite as robust. Um
  15629. 12:12:28it's not quite as robust as the cross
  15630. 12:12:30valve, but it works. Um it's, you know,
  15631. 12:12:34we can do hold out. Um,
  15632. 12:12:37we can do hold out uh to to quickly
  15633. 12:12:39evaluate a model and see if we need to
  15634. 12:12:42make any adjustments.
  15635. 12:12:44It's a little bit quicker to run.
  15636. 12:12:50Okay. And any questions on it? Does it
  15637. 12:12:52make sense what we're doing here?
  15638. 12:12:53Model.fit to train it predict to get our
  15639. 12:12:57predictions. Um, this is pretty
  15640. 12:13:00standard, right? To train is the
  15641. 12:13:01model.fit it and then to use the model
  15642. 12:13:03to predict. We predict on the test
  15643. 12:13:05features.
  15644. 12:13:07Um, so this is passing on on all of our
  15645. 12:13:09features into this model to generate
  15646. 12:13:13predictions for every row. That's
  15647. 12:13:15something I also want to point out that
  15648. 12:13:16may be a little bit confusing is this is
  15649. 12:13:18a data frame. So we're passing in a
  15650. 12:13:22bunch of rows of features with columns,
  15651. 12:13:24right? So um, we're passing in a bunch
  15652. 12:13:27of data that looks like this. And what
  15653. 12:13:30we're doing is essentially making a
  15654. 12:13:32prediction for every row. So this will
  15655. 12:13:34generate a prediction. This row will
  15656. 12:13:37generate a prediction. This row will
  15657. 12:13:39generate a prediction. And on and on and
  15658. 12:13:41on. So this this predict will predict
  15659. 12:13:44for every row. And so we end up with
  15660. 12:13:47this collection of predictions here for
  15661. 12:13:49each row. And we're comparing those to
  15662. 12:13:52the labels that we have for those rows
  15663. 12:13:54from our from our supervised learning,
  15664. 12:13:57right? From our data set.
  15665. 12:13:59So that's truly supervised learning,
  15666. 12:14:01right? We have the examples and we're
  15667. 12:14:04comparing those to what our model is
  15668. 12:14:06predicting to to get our performance.
  15669. 12:14:16All right.
  15670. 12:14:18So let's uh let's try the other just so
  15671. 12:14:21you can see it. The leave one out. Now
  15672. 12:14:23the leave one out cross validation is
  15673. 12:14:25going to actually work the same way
  15674. 12:14:26where we put in the leave one out um
  15675. 12:14:30strategy inside of the cross file score.
  15676. 12:14:32Now here we don't need to specify how
  15677. 12:14:34many folds there are because we know how
  15678. 12:14:37many there going to be. It's going to be
  15679. 12:14:38the number of data points, right? So
  15680. 12:14:41which is actually going to be quite
  15681. 12:14:42large because there's 20,000 rows. So
  15682. 12:14:45this is going to be extremely
  15683. 12:14:48uh extremely um intensive because we are
  15684. 12:14:53doing um you know 20,000 examples and
  15685. 12:14:57leaving one example out to be our
  15686. 12:15:00validation and then um doing that across
  15687. 12:15:03every 20,000 uh examples.
  15688. 12:15:06So we could do it though just to see how
  15689. 12:15:08it works. Um we have this again leave
  15690. 12:15:10one out. We generate our cross file
  15691. 12:15:12score from our model our data and then
  15692. 12:15:16same scoring that we had before and but
  15693. 12:15:18this time we change our cross file to be
  15694. 12:15:20instead of our k-fold object we have our
  15695. 12:15:23leave one out object which is this
  15696. 12:15:26um and then we could run this. We can
  15697. 12:15:29compute our average uh across the all
  15698. 12:15:32the folds. Now this is going to be a lot
  15699. 12:15:34bigger of an array. It's going to be a
  15700. 12:15:3620,000 size array and we're going to
  15701. 12:15:39compute the average across it.
  15702. 12:15:43So, let's do that. It's going to take a
  15703. 12:15:45moment because there's lots. So, if you
  15704. 12:15:47notice it when you run, it's going to
  15705. 12:15:48take a little bit of time to run because
  15706. 12:15:50it's running across all 20,000 examples
  15707. 12:15:53and leaving one out. So, you have 20,000
  15708. 12:15:57and then one left out to uh test
  15709. 12:16:00against. So, it's quite intensive. You
  15710. 12:16:03can see it's taking a lot more time.
  15711. 12:16:11It's still running. It's taking a while.
  15712. 12:16:23Okay, just let that run. Still running.
  15713. 12:16:28So, if you guys try running this, it's
  15714. 12:16:30going to take a little bit of time.
  15715. 12:16:31Hopefully that makes sense why it's
  15716. 12:16:33taking so long, right? It's because it's
  15717. 12:16:35instead of doing 10 folds, it's it's
  15718. 12:16:39putting every data point but one is the
  15719. 12:16:41training set and then iterating through
  15720. 12:16:43all 20,000 points.
  15721. 12:16:47This takes a while to do.
  15722. 12:17:04Let's see what our
  15723. 12:17:08RAM our memory is a little increased.
  15724. 12:17:20Okay,
  15725. 12:17:22still running. That's okay. I'll let it
  15726. 12:17:24run.
  15727. 12:17:27Come back when it's finished.
  15728. 12:17:34Yeah, exactly. This is a This is for
  15729. 12:17:37This is giving us a performance
  15730. 12:17:38evaluation. This is like the average
  15731. 12:17:41error across all of our uh different
  15732. 12:17:43folds. Um now this is the extreme case
  15733. 12:17:46where we have the number of folds equals
  15734. 12:17:48the number of points.
  15735. 12:17:51Right? So it's an extreme case but yes
  15736. 12:17:53it's just like kfold. It's giving us
  15737. 12:17:55that performance estimate.
  15738. 12:18:04Okay. It's about the same right. This is
  15739. 12:18:07still around 50,000.
  15740. 12:18:10Not much difference, right? Still right
  15741. 12:18:12around there. But look how much longer
  15742. 12:18:15it took. That took 2 minutes to run. The
  15743. 12:18:17other one was pretty instant, right? So
  15744. 12:18:19this this took about 2 minutes to run.
  15745. 12:18:22So um definitely uh
  15746. 12:18:28yeah, definitely don't want to run this
  15747. 12:18:31uh too often. I think that it's
  15748. 12:18:34generally preferred to do k-fold if
  15749. 12:18:36you're going to do cross validation.
  15750. 12:18:37Generally want to do k-fold or just the
  15751. 12:18:39regular hold out train test split. Uh
  15752. 12:18:42generally better than doing leave one
  15753. 12:18:44out. It's just going to take too long
  15754. 12:18:46and um it results in about the same kind
  15755. 12:18:49of score as the kfold.
  15756. 12:19:01Okay,
  15757. 12:19:04any questions about um the cross
  15758. 12:19:07validation that we just did.
  15759. 12:19:28Okay,
  15760. 12:19:29good. And as it says here that the
  15761. 12:19:32stratified kfold is usually used for
  15762. 12:19:34classification. Again, we're not doing
  15763. 12:19:35classification yet. That's in going to
  15764. 12:19:36be in lesson four. So, we don't need to
  15765. 12:19:39worry too much about that. Just for
  15766. 12:19:40regression, um regular kfold is
  15767. 12:19:43preferred, right? Because we don't need
  15768. 12:19:45to um worry about distributing
  15769. 12:19:48categories amongst our folds uh in any
  15770. 12:19:51regression problems.
  15771. 12:19:53And as we see the error is kind of high.
  15772. 12:19:56Um there's going to be some things we
  15773. 12:19:57can do to improve that which will be uh
  15774. 12:20:00later on we'll learn about some more
  15775. 12:20:02advanced models. This signals that the
  15776. 12:20:05performance is bad. We probably need a
  15777. 12:20:07more complex model. Um one thing we
  15778. 12:20:10could try before we try a complex model
  15779. 12:20:12is to do scaling. We will try to do
  15780. 12:20:15scaling. I'm going to show us how we can
  15781. 12:20:17do that coming up um in a in a nice
  15782. 12:20:20streamlined fashion. Um, but uh outside
  15783. 12:20:24of that, if we still had bad
  15784. 12:20:26performance, we would likely need to use
  15785. 12:20:27a more advanced model. And we'll learn
  15786. 12:20:30about more advanced models uh in the
  15787. 12:20:33next lesson. And what's great is some of
  15788. 12:20:35those advanced models can actually be
  15789. 12:20:37used for regression. So they have
  15790. 12:20:39variations that can be used for both
  15791. 12:20:41classification and regression, which is
  15792. 12:20:43pretty cool. So I'll point those out
  15793. 12:20:45when we get to them. Um, okay.
  15794. 12:20:50So what I want to talk about now is a
  15795. 12:20:53way we can combat overfitting. So if we
  15796. 12:20:56have overfitting which remember that is
  15797. 12:20:58the case where the uh the we see good
  15798. 12:21:03performance on the training data but
  15799. 12:21:05then um it doesn't generalize over to
  15800. 12:21:07the test data. We get poor performance
  15801. 12:21:09on the test data. Um there's there's a
  15802. 12:21:12drop off there. Um that would signal
  15803. 12:21:15overfitting.
  15804. 12:21:18overfitting
  15805. 12:21:19and one way of um combating overfitting
  15806. 12:21:23is to do something called regularization
  15807. 12:21:26which we're going to talk about next. So
  15808. 12:21:29the key idea in regularization
  15809. 12:21:33is to
  15810. 12:21:35change our uh the change the way we
  15811. 12:21:39train. Essentially, what we're going to
  15812. 12:21:41do is modify our training
  15813. 12:21:46uh error function or sometimes called
  15814. 12:21:50the objective function or loss function.
  15815. 12:21:53We're going to change that to add a
  15816. 12:21:56penalty to penalize excessive complex
  15817. 12:22:00complexity. Essentially the the way that
  15818. 12:22:03we're going to penalize is by making
  15819. 12:22:05sure the size of the coefficients
  15820. 12:22:08doesn't grow too much which should
  15821. 12:22:11mitigate overfitting because remember in
  15822. 12:22:14linear regression what we are learning
  15823. 12:22:16are the coefficients right we're
  15824. 12:22:18learning the beta 0 the beta 1 the beta
  15825. 12:22:212 and on and on however many betas there
  15826. 12:22:24are beta n we're learning all of those
  15827. 12:22:26guys um through the regression error
  15828. 12:22:29function we're trying to minimize that
  15829. 12:22:31error function. That's how it trains. We
  15830. 12:22:33talked about that on Monday.
  15831. 12:22:36Um so what we're going to do is um
  15832. 12:22:41basically penalize the these guys
  15833. 12:22:44growing too big and making sure we kind
  15834. 12:22:47of keep them small so that no one
  15835. 12:22:51coefficient has a dominant uh effect on
  15836. 12:22:54the model. And this should help with
  15837. 12:22:56overfitting and complexity. should make
  15838. 12:22:58the model simpler because all the
  15839. 12:23:00coefficients are going to be encouraged
  15840. 12:23:02to be smaller. They're not going to grow
  15841. 12:23:04too big. Um and this this has the effect
  15842. 12:23:07of making the model so basically make
  15843. 12:23:10the model simpler.
  15844. 12:23:14Make the model simpler is what these
  15845. 12:23:17regularization techniques are
  15846. 12:23:18essentially trying to achieve is is
  15847. 12:23:20remove complexity, make them a little
  15848. 12:23:22bit simpler, make these coefficients
  15849. 12:23:24smaller so that you can generalize a bit
  15850. 12:23:26better and and prevent overfitting. So
  15851. 12:23:30we want to prevent
  15852. 12:23:33uh overfitting,
  15853. 12:23:35right, is what we want to do. Um so
  15854. 12:23:38there's going to be a penalty and I'll
  15855. 12:23:40show you where that penalty gets added
  15856. 12:23:42and kind of what it looks like.
  15857. 12:23:44Um but uh to control the level of that
  15858. 12:23:48penalty we are actually going to
  15859. 12:23:49introduce another parameter to our model
  15860. 12:23:52um called alpha.
  15861. 12:23:55Alpha is going to scale the penalty. So
  15862. 12:23:58if alpha is really high that imposes a
  15863. 12:24:02stronger penalty on the coefficients um
  15864. 12:24:05which will make the model a lot simpler.
  15865. 12:24:08So the higher the alpha the simpler the
  15866. 12:24:10model we will get and we the the risk
  15867. 12:24:14with that is we actually underfit. So if
  15868. 12:24:17alpha is too big we may underfit the
  15869. 12:24:20training data
  15870. 12:24:22um a bit too much because it will make
  15871. 12:24:24the model way too simple. Um and I again
  15872. 12:24:27I'll show you what this means
  15873. 12:24:28mathematically in a moment. Um but on
  15874. 12:24:31the other hand if we have a lower alpha
  15875. 12:24:34this will have a lower penalty. it's a
  15876. 12:24:37weaker penalty term and that'll lead to
  15877. 12:24:40a model that is um a bit more complex.
  15878. 12:24:44Um which could um risk some level of
  15879. 12:24:47overfitting. Um so there's so there's
  15880. 12:24:51still the risk of overfitting if you
  15881. 12:24:53have a low alpha. And of course if alpha
  15882. 12:24:55goes all the way to zero, there's no
  15883. 12:24:56penalty at all. So you're back to your
  15884. 12:24:59original linear regression. Um which
  15885. 12:25:02could risk a lot of overfitting,
  15886. 12:25:04right? So you you generally want to pick
  15887. 12:25:07an alpha um effectively and actually
  15888. 12:25:10we're going to see h what's the best way
  15889. 12:25:12to pick alpha. Um we're actually going
  15890. 12:25:14to learn how to do that. I'm going to
  15891. 12:25:15show us how doing some tuning techniques
  15892. 12:25:18to pick what alpha should be. Um but um
  15893. 12:25:23a a pretty industry standard alpha that
  15894. 12:25:25most people default to is alpha equals
  15895. 12:25:28to one. So just just one which signals
  15896. 12:25:32that there should be some penalty. we
  15897. 12:25:34just have alpha equal to one is a
  15898. 12:25:35standard penalty. We don't want it to be
  15899. 12:25:37too high. We don't want it to be too
  15900. 12:25:39low. Like we don't want it to be a
  15901. 12:25:40fraction. Um but a penalty of one is
  15902. 12:25:43usually uh good enough.
  15903. 12:25:47Okay, I'm going to show you where that
  15904. 12:25:49comes into play in a moment.
  15905. 12:25:52Um but the whole purpose of doing this
  15906. 12:25:54is to mitigate overfitting, right? Um
  15907. 12:25:57that's what and and doing this penalty
  15908. 12:26:00is is called regularization. So adding
  15909. 12:26:03so going beyond just regular linear
  15910. 12:26:05regression adding this extra penalty to
  15911. 12:26:07to the training process um to penalize
  15912. 12:26:11large weights large coefficients
  15913. 12:26:14um is known as regularization.
  15914. 12:26:18Okay. Um and there's two common
  15915. 12:26:21penalties that are added. Um so there's
  15916. 12:26:23actually two different variations on the
  15917. 12:26:25penalty. Um we're going to study both of
  15918. 12:26:27them and um they're they're known as
  15919. 12:26:30lasso. So if you take linear regression
  15920. 12:26:32and add a particular type of penalty,
  15921. 12:26:34it's known as lasso. If you add another
  15922. 12:26:37type of penalty, it's known as ridge
  15923. 12:26:39regression. We're going to study both of
  15924. 12:26:41those and what their differences are.
  15925. 12:26:43But these are the primary two
  15926. 12:26:46uh regularization tech uh models that
  15927. 12:26:49are used um to take a regular both of
  15928. 12:26:52these take regular linear regression and
  15929. 12:26:54just modify the training process a
  15930. 12:26:57little bit in different ways. Two
  15931. 12:26:59different ways. um using that alpha
  15932. 12:27:03um to penalize the terms in slightly
  15933. 12:27:06different mathematical ways. So we're
  15934. 12:27:08going to learn about these two guys.
  15935. 12:27:09Lasso regression there. Both of these
  15936. 12:27:12are just offshoots of linear regression.
  15937. 12:27:14So underlying model is still linear
  15938. 12:27:16regression. It just adds different types
  15939. 12:27:19of penalties to the training process.
  15940. 12:27:22So both of these are still in the family
  15941. 12:27:24of linear regression. In fact, in um in
  15942. 12:27:29scikitlearn, they both come from they
  15943. 12:27:31both are still from the linear model
  15944. 12:27:33family in inside of the linear model
  15945. 12:27:36module, which is where linear regression
  15946. 12:27:37comes from. So there's still linear
  15947. 12:27:39regression. They just have different
  15948. 12:27:42styles of penalties added to them. Um
  15949. 12:27:45which we're going to see.
  15950. 12:27:48Okay, so just to recap that
  15951. 12:27:51regularization is the process of adding
  15952. 12:27:54a penalty to the training to discourage
  15953. 12:27:58complexity. In this case, we're going to
  15954. 12:28:00discourage large coefficients.
  15955. 12:28:04And um this should help prevent
  15956. 12:28:07overfitting.
  15957. 12:28:09And so uh these are going to lead us to
  15958. 12:28:12two different offshoots of linear
  15959. 12:28:13regression that have two different
  15960. 12:28:15penalties.
  15961. 12:28:16lasso and ridge regression, which we're
  15962. 12:28:18going to uh study next,
  15963. 12:28:21but they they function the same way as
  15964. 12:28:23linear regression. They will just have
  15965. 12:28:26different penalty terms added onto their
  15966. 12:28:28training process um to discourage
  15967. 12:28:32uh discourage um again those large
  15968. 12:28:35weights.
  15969. 12:28:40Okay, any questions about regularization
  15970. 12:28:42before we first look at our we're going
  15971. 12:28:43to look at our first uh variation on on
  15972. 12:28:47our first regularization technique which
  15973. 12:28:48is going to be called lasso regression.
  15974. 12:29:08Okay, let's look at lasso regression. So
  15975. 12:29:10what is lasso regression? It's actually
  15976. 12:29:13lasso is short for least absolute
  15977. 12:29:16shrinkage and selection operator
  15978. 12:29:18regression. Um and this will function by
  15979. 12:29:23adding a particular penalty to the
  15980. 12:29:27linear regression model. So again it's
  15981. 12:29:29based on linear regression. That's the
  15982. 12:29:30underlying model. It's just that during
  15983. 12:29:33the training process we are going to um
  15984. 12:29:36add a penalty which has the effect of
  15985. 12:29:41shrinkage of the weights. That's why
  15986. 12:29:43it's called shrinkage. It encourages
  15987. 12:29:45smaller weights through that penalty and
  15988. 12:29:48it also will shrink some of them so much
  15989. 12:29:51that they'll become zero and so it has
  15990. 12:29:53has an effect of kind of selection which
  15991. 12:29:56means that some of them get wiped out to
  15992. 12:29:58zero
  15993. 12:30:00and this means that whatever is left
  15994. 12:30:02over is kind of what's selected as our
  15995. 12:30:05features because the other ones will
  15996. 12:30:08have zero weight applied to them. So
  15997. 12:30:10this penalty will really favor small
  15998. 12:30:14weights um and penalize really large
  15999. 12:30:18weights. In fact, it will favor small
  16000. 12:30:20weights so much that some of them will
  16001. 12:30:22actually um be shrunk to zero um during
  16002. 12:30:26the training process. And the ones that
  16003. 12:30:28are left over are the ones that um are
  16004. 12:30:32the ones that are what we call selected
  16005. 12:30:35because they are the ones that remain in
  16006. 12:30:37in the training um after the other ones
  16007. 12:30:40get uh coefficients of zero. Um now when
  16008. 12:30:44you make some of the coefficients zero,
  16009. 12:30:46you are inherently making the model
  16010. 12:30:48simpler, right? There's less features
  16011. 12:30:51involved in the prediction that or less
  16012. 12:30:53features that have an effect on the
  16013. 12:30:55prediction. So this definitely makes the
  16014. 12:30:57model simpler. This lasso, this
  16015. 12:31:00shrinkage and selection uh process makes
  16016. 12:31:04makes the model simpler for sure. Um
  16017. 12:31:08and this is supposed to reduce
  16018. 12:31:10overfitting, right? If you make the
  16019. 12:31:11model simpler, it's not as complex. It
  16020. 12:31:14has less of a chance of memorizing
  16021. 12:31:16training data and not generalizing over
  16022. 12:31:19to test data. So our whole goal with uh
  16023. 12:31:23regularization is to make our model
  16024. 12:31:25better at generalization right over to
  16025. 12:31:28test data from the original training
  16026. 12:31:30data.
  16027. 12:31:32Um so how does this happen? We have to
  16028. 12:31:35go back to the
  16029. 12:31:38uh training process. If you guys
  16030. 12:31:40remember I I wrote out this equation a
  16031. 12:31:42little bit earlier which is the
  16032. 12:31:44distance. This is the sum of squared
  16033. 12:31:46distance between our labels and our
  16034. 12:31:48prediction.
  16035. 12:31:50This is basically the mean squared error
  16036. 12:31:52uh calculation that we're trying to
  16037. 12:31:54reduce when we build our model using the
  16038. 12:31:56training data. Um so this is just in
  16039. 12:31:58standard linear regression. This is the
  16040. 12:32:01um uh sum of squares uh distance right
  16041. 12:32:05so this is this is what the model is
  16042. 12:32:07trying to minimize when it learns these
  16043. 12:32:09coefficients.
  16044. 12:32:11So when it learns these coefficients,
  16045. 12:32:13it's trying to minimize this guy.
  16046. 12:32:18Minimize. It's trying to find the betas
  16047. 12:32:21that minimize this quantity.
  16048. 12:32:23Mathematically, that's what it's doing.
  16049. 12:32:25Um, and there's there's a algorithm that
  16050. 12:32:28will discover what the best betas are
  16051. 12:32:30that actually minimize this. That gives
  16052. 12:32:32us the line of best fit, right? That's
  16053. 12:32:34what we've been talking about for
  16054. 12:32:35regression.
  16055. 12:32:37Now in regularization
  16056. 12:32:41here's by the way here is that same
  16057. 12:32:42thing but we've just inserted our model
  16058. 12:32:44for the predictions. This is our model
  16059. 12:32:47just a fancy way of writing down our
  16060. 12:32:49model right it's the beta 0 plus all of
  16061. 12:32:52these betas. So beta 1 x1 plus beta 2 x2
  16062. 12:32:58plus on and on and on. Right? That's
  16063. 12:33:01that's what this uh means if you're
  16064. 12:33:03unfamiliar with the sigma notation. It
  16065. 12:33:05just means sum. So it's the sum of all
  16066. 12:33:07these guys or this term. Um so this is
  16067. 12:33:12this here is just a regular linear
  16068. 12:33:14regression
  16069. 12:33:18uh training regular linear regression
  16070. 12:33:21training. So we the training process
  16071. 12:33:24solves for these parameters right it
  16072. 12:33:27solves for these weights. We discover
  16073. 12:33:29what those are by minimizing this
  16074. 12:33:31quantity. That's the whole training
  16075. 12:33:33process. Um but when we do lasso
  16076. 12:33:37we add a penalty which is this
  16077. 12:33:43here is our penalty.
  16078. 12:33:47So basically um we take our linear
  16079. 12:33:51regression training which is this and we
  16080. 12:33:54add on a penalty which is this and you
  16081. 12:33:57can see exactly what this penalty when
  16082. 12:34:00when you minimize this penalty it's when
  16083. 12:34:03these weights are small. So this
  16084. 12:34:05encourages
  16085. 12:34:06So minimizing this quantity encourages
  16086. 12:34:10small weights
  16087. 12:34:13encourages small betas
  16088. 12:34:17beta I
  16089. 12:34:19right you or in this case beta j sorry
  16090. 12:34:23this encourages small beta js uh because
  16091. 12:34:26we want this thing to be minimized
  16092. 12:34:31minimize
  16093. 12:34:33So um what's going to make this minimal
  16094. 12:34:36is of course the line of best fit and
  16095. 12:34:38small weights right are going to make
  16096. 12:34:40are going to bring this error down the
  16097. 12:34:43most.
  16098. 12:34:46So um and here's our alpha right here's
  16099. 12:34:48our alpha. So you can encourage a higher
  16100. 12:34:51penalty with a larger alpha or a lower
  16101. 12:34:53penalty. If alpha equals zero
  16102. 12:34:56what happens to that term? It just goes
  16103. 12:34:59away. So if alpha equals zero, there's
  16104. 12:35:01no penalty and we're back to uh we're
  16105. 12:35:05back to regular linear regression.
  16106. 12:35:09We just have regular linear regression
  16107. 12:35:11because we have no penalty at that point
  16108. 12:35:12when alpha equals zero. So the smaller
  16109. 12:35:15alpha is, the less penalty we're
  16110. 12:35:19enforcing in in the regularization.
  16111. 12:35:22Okay.
  16112. 12:35:24Now what happens is in reality when you
  16113. 12:35:27train with lasso. So this is lasso is
  16114. 12:35:30this particular penalty. This is called
  16115. 12:35:31the lasso penalty
  16116. 12:35:34or sometimes um people call this the L1
  16117. 12:35:37penalty.
  16118. 12:35:39Um L1 just comes from the fact that this
  16119. 12:35:42is the first power or absolute value. Um
  16120. 12:35:46so it's not a squared penalty. It's a
  16121. 12:35:48single uh single power penalty
  16122. 12:35:51um there. But when you add this lasso
  16123. 12:35:55penalty, what can happen is it c it does
  16124. 12:35:58because the because you're minimizing
  16125. 12:36:00this, it does encourage some of these
  16126. 12:36:02weights to become zero.
  16127. 12:36:05So some if you're really trying to get
  16128. 12:36:08the lowest quantity of this,
  16129. 12:36:11the lower the better.
  16130. 12:36:14What makes this thing lower is of course
  16131. 12:36:17if some of these go away. If some of
  16132. 12:36:18these go to zero then that of course
  16133. 12:36:21will lower this as much as we as much as
  16134. 12:36:23possible. Right? So what happens during
  16135. 12:36:25the training is some of these
  16136. 12:36:27coefficients actually they're encouraged
  16137. 12:36:29to be small because of this penalty. But
  16138. 12:36:32some of them will actually becomes will
  16139. 12:36:35actually become zero um in order to get
  16140. 12:36:38the best model the best fit. Some of
  16141. 12:36:40these will actually get so small that
  16142. 12:36:42they'll basically become zero. And that
  16143. 12:36:44means that that that feature basically
  16144. 12:36:47has no effect anymore. It's it's been
  16145. 12:36:51the model has been simplified, right?
  16146. 12:36:53That feature no longer really has an
  16147. 12:36:54effect.
  16148. 12:37:00So just to call out the alpha again um
  16149. 12:37:02if alpha zero some co uh basically you
  16150. 12:37:06have your linear regression you're back
  16151. 12:37:08to linear regression because alpha 0 is
  16152. 12:37:10just wiping this out and you're back to
  16153. 12:37:12linear regression. Um if alpha is
  16154. 12:37:14infinity now if alpha is infinity that's
  16155. 12:37:16an extreme. So if alpha is infinity the
  16156. 12:37:19only way to make this minimize is if all
  16157. 12:37:21your coefficients are zero. If every
  16158. 12:37:23beta is zero, then this will lower the
  16159. 12:37:25the error as as much as possible. So you
  16160. 12:37:28basically have no model. So if all
  16161. 12:37:31coefficients are zero, you have no model
  16162. 12:37:32and that's useless. So you don't want
  16163. 12:37:35your penalty, you don't want your alpha
  16164. 12:37:37to be huge is what this is saying. You
  16165. 12:37:40also don't want your alpha to be small.
  16166. 12:37:41You're basically back to linear
  16167. 12:37:42regression. So you want something in
  16168. 12:37:44between. Um and the typical typical
  16169. 12:37:48value is alpha equals 1.
  16170. 12:37:51typical is alpha equals 1
  16171. 12:37:55to have some level of penalty there. So
  16172. 12:37:58just a regular kind of regular penalty
  16173. 12:38:00term.
  16174. 12:38:08But we are actually going to have a way
  16175. 12:38:10to test and evaluate which alphas are
  16176. 12:38:12the best.
  16177. 12:38:27Um,
  16178. 12:38:28basically you can, yeah, you can have a
  16179. 12:38:31you can have a penalty that's close to
  16180. 12:38:33zero. You can get rid of this if just a
  16181. 12:38:36regular linear regression performs
  16182. 12:38:38pretty well. You can basically have no
  16183. 12:38:40penalty in that case.
  16184. 12:38:42Yeah. So near zero or like it could be
  16185. 12:38:46that adding a little bit of penalty
  16186. 12:38:48actually helps the overfitting and it
  16187. 12:38:50could be really small. One thing that
  16188. 12:38:52we're basically going to do is have a
  16189. 12:38:54strategy to try out different alphas,
  16190. 12:39:00try different alphas
  16191. 12:39:03and evaluate performance
  16192. 12:39:07and then we can decide which. So that's
  16193. 12:39:09what we're going to do is have a
  16194. 12:39:11strategy to just plug in different
  16195. 12:39:12alphas, generate the like train the
  16196. 12:39:15model, and then see what its performance
  16197. 12:39:17is and see if those alphas are good.
  16198. 12:39:19What what which alpha is the best? We
  16199. 12:39:22can evaluate that
  16200. 12:39:24because we can train the model and see
  16201. 12:39:26what it performance is,
  16202. 12:39:31right?
  16203. 12:39:34Yeah. Yeah. So we'll do that. We'll
  16204. 12:39:35practice that.
  16205. 12:39:44Okay, great. Any other questions about
  16206. 12:39:46this lasso regression? So, remember this
  16207. 12:39:48is linear regression here. This is the
  16208. 12:39:50this is how you're training to find the
  16209. 12:39:52betas in linear regression. So, this is
  16210. 12:39:55just linear regression uh um training
  16211. 12:40:00function there.
  16212. 12:40:02We're adding a penalty which is this is
  16213. 12:40:04the lasso penalty
  16214. 12:40:07lasso penalty there right we're adding
  16215. 12:40:09that this is known as regularization
  16216. 12:40:12and the goal of regularization is to
  16217. 12:40:15prevent overfitting so you add a penalty
  16218. 12:40:18here this makes the model simpler which
  16219. 12:40:21prevents overfitting it helps you
  16220. 12:40:23generalize better when it's simpler Any
  16221. 12:40:36questions conceptually on this? We're
  16222. 12:40:38going to do a code example with it
  16223. 12:40:39coming up, but any questions on this?
  16224. 12:40:57Uh yeah, you you so that's the thing,
  16225. 12:40:59Ronald, is you may be willing to
  16226. 12:41:01sacrifice some accuracy in order to
  16227. 12:41:04generalize to unseen data because
  16228. 12:41:06remember that's what we're really trying
  16229. 12:41:08to get after is we may be willing to
  16230. 12:41:10sacrifice some accuracy on this training
  16231. 12:41:12data in order to have it perform better
  16232. 12:41:14on the test data, right? We may be
  16233. 12:41:17willing to do that. That's a willing
  16234. 12:41:19that's an okay sacrifice
  16235. 12:41:22as long like if if it generalizes
  16236. 12:41:24better. That's what we want. That's what
  16237. 12:41:27we're trying to do here is add a
  16238. 12:41:29penalty, make the model simpler and help
  16239. 12:41:33it generalize better to new and unseen
  16240. 12:41:36data. Right?
  16241. 12:41:39That's that picture I've been using with
  16242. 12:41:41the with the um train and test split.
  16243. 12:41:45Where is the square?
  16244. 12:41:47So in the model there's no square. So
  16245. 12:41:50remember the model is the model is this
  16246. 12:41:56um equation uh that has no squares in
  16247. 12:41:58it. Right? It's beta 0 plus beta 1 x1
  16248. 12:42:03plus beta 2 x2 plus beta n xn.
  16249. 12:42:09That's the that's the linear regression
  16250. 12:42:11model. This is the now this this is the
  16251. 12:42:15model but this is the equation that
  16252. 12:42:18helps us train and find the betas. This
  16253. 12:42:21is how this is what we find the betas
  16254. 12:42:23with. So we'll continue. Um we were
  16255. 12:42:27talking about the lasso regression which
  16256. 12:42:30uh adds it takes linear regression right
  16257. 12:42:33which is this optimization and adds in a
  16258. 12:42:36penalty um scaled by the alpha. Um, and
  16259. 12:42:40what that does in order to minimize this
  16260. 12:42:43whole thing, it encourages these to be
  16261. 12:42:46small uh as possible. Um, which makes
  16262. 12:42:49the model simpler, right? The weights
  16263. 12:42:52don't get overly big and complex. Um,
  16264. 12:42:55they they tend to stay small. In fact,
  16265. 12:42:57some of them can even go all the way to
  16266. 12:42:58zero. Um, which makes the model even
  16267. 12:43:01more simpler,
  16268. 12:43:03right? Um, so let's practice uh using it
  16269. 12:43:07in code. It's actually really easy to
  16270. 12:43:08use. It's going to be essentially the
  16271. 12:43:10same uh style and and code as linear
  16272. 12:43:14regression except we are um just going
  16273. 12:43:17to have to uh put in our alpha parameter
  16274. 12:43:21um when we use the lasso. So here we are
  16275. 12:43:26um from the linear model family right
  16276. 12:43:30which makes sense. It's a linear
  16277. 12:43:31regression offshoot that has this
  16278. 12:43:33penalty in it during the training. um we
  16279. 12:43:35are grabbing our lasso regression. Um it
  16280. 12:43:38also has a version of the lasso that
  16281. 12:43:42we're going to take a look at that is
  16282. 12:43:43used for cross validation which is
  16283. 12:43:46really um convenient as well. So it has
  16284. 12:43:49a cross validation lasso which is a
  16285. 12:43:51really convenient um combination of
  16286. 12:43:54basically cross val score and lasso um
  16287. 12:43:57all in one. So it actually is really
  16288. 12:43:59nice to use that way. Um so we'll take a
  16289. 12:44:02look at that example. Um, but we are
  16290. 12:44:05importing it. The main thing is going to
  16291. 12:44:07be the lasso model here. Um, we're going
  16292. 12:44:09to be using a different data set for
  16293. 12:44:11this one. So, not the ocean uh data, but
  16294. 12:44:14this hitters data, which is a baseball
  16295. 12:44:16data set. Um, so it has 322 rows um with
  16296. 12:44:2120 different columns and it looks like
  16297. 12:44:23this. So, you want to download that one.
  16298. 12:44:26Um, hopefully you guys have access to
  16299. 12:44:28that one.
  16300. 12:44:31Um,
  16301. 12:44:33so I will upload it into
  16302. 12:44:36this.
  16303. 12:44:39So give me a moment.
  16304. 12:44:45There's that. And then we can run this.
  16305. 12:44:49Okay. So we are displaying the data and
  16306. 12:44:52so it has um the the hitters names and
  16307. 12:44:56then it has a bunch of different
  16308. 12:44:57statistics. These are all baseball
  16309. 12:44:58statistics.
  16310. 12:45:00Um, if you're unfamiliar with with them,
  16311. 12:45:01that's okay. It's not a big deal. Um,
  16312. 12:45:04but just different baseball stats here.
  16313. 12:45:09Okay. Were you guys able to load that?
  16314. 12:45:11Um, if you're following along, were you
  16315. 12:45:12able to load that? You should have
  16316. 12:45:14access to this data. The hitters CSV.
  16317. 12:45:18This is the one we're going to use for
  16318. 12:45:19the lasso model
  16319. 12:45:21to build a lasso model.
  16320. 12:45:32Yeah.
  16321. 12:45:39Okay. Able to load that one. Perfect.
  16322. 12:45:42Okay. So, able to load that one. Um, and
  16323. 12:45:45we take a look at the the head. Um, so
  16324. 12:45:48we're actually going to uh drop this
  16325. 12:45:51unnamed column because we don't care
  16326. 12:45:53about their name. it's actually just the
  16327. 12:45:55batter's name, which is not going to be
  16328. 12:45:57useful in modeling. Um, so and remember
  16329. 12:46:00that's generally true like an ID, a user
  16330. 12:46:03ID, like a customer ID, a name, that's
  16331. 12:46:06usually not going to be useful in any
  16332. 12:46:08kind of modeling. So we're actually just
  16333. 12:46:09going to drop that uh column and we're
  16334. 12:46:12going to do it in place.
  16335. 12:46:14And access equals 1 means we're dropping
  16336. 12:46:16that column. Um, so we're going to drop
  16337. 12:46:19that and we should no longer have that
  16338. 12:46:21column. And we have all of these guys
  16339. 12:46:23now. So you want to run that. This will
  16340. 12:46:25drop that. Um this will drop drops the
  16341. 12:46:30column in place.
  16342. 12:46:33Um and now we can see we have uh all we
  16343. 12:46:37have this data where um we have this
  16344. 12:46:41data where it's now removed. So this
  16345. 12:46:44that column is now gone and now we have
  16346. 12:46:46these guys. Um, do you notice anything
  16347. 12:46:50about this
  16348. 12:46:52from the info?
  16349. 12:46:58Looks like we have a couple categorical
  16350. 12:46:59features, a few of them, league and
  16351. 12:47:02division
  16352. 12:47:04and new league. What do you notice about
  16353. 12:47:07this?
  16354. 12:47:16Nolles. Yep. So, there's definitely some
  16355. 12:47:17missing data there um that we're going
  16356. 12:47:20to have to deal with.
  16357. 12:47:25So, it looks like there are uh there are
  16358. 12:47:2959.
  16359. 12:47:31Um there are 59. Now, we could we the
  16360. 12:47:35alternative to doing that is we could uh
  16361. 12:47:37we could just use our usual code where
  16362. 12:47:39we do df.is is uh is null.
  16363. 12:47:44Um and then we do uh dot sum to total
  16364. 12:47:49those up across our different columns.
  16365. 12:47:51And we can see that uh we have 59 of
  16366. 12:47:54those in this salary column. That's this
  16367. 12:47:57is the standard way of doing that,
  16368. 12:47:58right?
  16369. 12:48:02Standard way of doing that. And we have
  16370. 12:48:03so we have 59 of those
  16371. 12:48:1259 of those. So, we have to deal with
  16372. 12:48:15it. Any ideas on how to deal with it?
  16373. 12:48:22Any ideas on how to deal with it? This
  16374. 12:48:24is now This is 59 out of 300.
  16375. 12:48:28So,
  16376. 12:48:29what do you guys think about that? It's
  16377. 12:48:31a little bit different than 200 out of
  16378. 12:48:3220,000. A little bit different. We have
  16379. 12:48:36We have about 60 out of 300.
  16380. 12:48:40There's a decent amount.
  16381. 12:48:45Any ideas on how to handle this one?
  16382. 12:48:50Replace. Yep, we should replace. What do
  16383. 12:48:52you think we should replace with?
  16384. 12:48:55It's a float. It's a floating point uh
  16385. 12:48:58value.
  16386. 12:49:09By the way, something unique about this
  16387. 12:49:11that's a little different than usual,
  16388. 12:49:12too, is that the uh this is actually the
  16389. 12:49:16column we're going to use as our label.
  16390. 12:49:18So, we're actually going to predict the
  16391. 12:49:19salary based on the uh based on the um
  16392. 12:49:24rest of the features. So, we definitely
  16393. 12:49:27need to fill in these nles, right?
  16394. 12:49:29because they're actually going to be the
  16395. 12:49:30labels
  16396. 12:49:32and we're missing some labels uh in our
  16397. 12:49:34data. We we definitely need to fill them
  16398. 12:49:37in. Yeah. So, we're going to replace
  16399. 12:49:39them.
  16400. 12:49:49All right. So, we'll we will replace
  16401. 12:49:51them down below. That's going to be
  16402. 12:49:52coming up. Uh we'll come back and
  16403. 12:49:54replace them. um before we replace them,
  16404. 12:49:56we're actually going to get our uh one
  16405. 12:50:00hot encodings for those three different
  16406. 12:50:02um features we have. Um so we do uh get
  16407. 12:50:09dummies with this. Now um of course we
  16408. 12:50:12don't need to do this if we just so this
  16409. 12:50:15code we don't need to do if we just pass
  16410. 12:50:17in the dtype here
  16411. 12:50:21um which is uh then we don't need to do
  16412. 12:50:24this. So we can comment this out.
  16413. 12:50:28Um so now what I want you guys to notice
  16414. 12:50:31is this is the alternative to what we
  16415. 12:50:32did before where we are purposely just
  16416. 12:50:35doing these columns not the whole data
  16417. 12:50:37frame but just doing these columns and
  16418. 12:50:40then we can um concatenate those these
  16419. 12:50:46one hot encodings. We're going to
  16420. 12:50:47concatenate back to the data frame.
  16421. 12:50:50Right? So if we do our dummies and then
  16422. 12:50:53do dummies.info info. Um, we can see
  16423. 12:50:56that we end up with six new columns. And
  16424. 12:50:59in fact, we can do dummies.head
  16425. 12:51:03and take a look at what those are.
  16426. 12:51:05Right? So, these are league A, league,
  16427. 12:51:08uh, N, division E, W, division W, new
  16428. 12:51:11league A, new league N.
  16429. 12:51:14Okay.
  16430. 12:51:16So, um these are uh these are our one
  16431. 12:51:20hot encodings for these three different
  16432. 12:51:22features which are strings, right? So,
  16433. 12:51:24those those features were strings. If
  16434. 12:51:25you go back up, those were our only
  16435. 12:51:27string features we had. So, we've one
  16436. 12:51:30hot encoded those so we can use them in
  16437. 12:51:31our model. What we need to do is just
  16438. 12:51:34concatenate this back to our data frame.
  16439. 12:51:37Right? So, we just need to concatenate
  16440. 12:51:39it back into our data.
  16441. 12:51:48Okay. So, what we're going to do then is
  16442. 12:51:51we're going to grab um we're going to
  16443. 12:51:55grab Y as our salary. And of course,
  16444. 12:51:57we're going to fill nles on that Y
  16445. 12:51:59coming up shortly. But we're going to
  16446. 12:52:01grab Y as our salary and X new. Now
  16447. 12:52:05before building a full X, we're going to
  16448. 12:52:08take a look at X numerical as our data
  16449. 12:52:11frame minus these columns. The reason
  16450. 12:52:14we're doing minus those is because we
  16451. 12:52:17are going to concatenate our dummy
  16452. 12:52:19variables back into this that are going
  16453. 12:52:22to replace these guys. So we're going to
  16454. 12:52:24replace these anyways with our one hot
  16455. 12:52:26encodings. We don't want the strings. So
  16456. 12:52:29we're going to get rid of those. And
  16457. 12:52:32we're also going to get rid of the
  16458. 12:52:33salary because that's going to be part
  16459. 12:52:34of our that's just the label. So we
  16460. 12:52:36don't want that in the X, the eventual
  16461. 12:52:38X.
  16462. 12:52:44Are you guys able to run this one?
  16463. 12:52:48Hope I'm not going too fast. You guys
  16464. 12:52:50able to run this? And does it make
  16465. 12:52:52sense? What we're doing is we're putting
  16466. 12:52:54our labels in Y, which is what we
  16467. 12:52:56usually do. So we're going to predict
  16468. 12:52:58the salary
  16469. 12:53:00and we're getting ready to build the X.
  16470. 12:53:02But before we first want to get rid of
  16471. 12:53:04those one hot the the strings. This is
  16472. 12:53:06getting rid of the strings
  16473. 12:53:09and this is getting rid of the label and
  16474. 12:53:11that's going to be part of our features.
  16475. 12:53:12What we need to do is build our final X
  16476. 12:53:14by concatenating our dummies with this.
  16477. 12:53:17Do you guys see that? We're going to
  16478. 12:53:19concatenate our dummies with this to
  16479. 12:53:21build our final X.
  16480. 12:53:23But but prior to doing that, we need to
  16481. 12:53:26get rid of these string columns here. So
  16482. 12:53:28we're dropping those
  16483. 12:53:31dropping those from the uh data frame uh
  16484. 12:53:35and getting a numerical uh x numerical
  16485. 12:53:38here.
  16486. 12:53:40You can see the columns of that are just
  16487. 12:53:42these guys here. So the the results we
  16488. 12:53:45need to concatenate our we need to
  16489. 12:53:46concatenate this guy um into this and
  16490. 12:53:50then that'll be our full x all of our
  16491. 12:53:52features.
  16492. 12:53:57Okay. So you can see x is going to be
  16493. 12:54:00pd.con
  16494. 12:54:02of this with our dummies.
  16495. 12:54:06This with our dummies. And um
  16496. 12:54:11uh instead of doing this, I'm actually
  16497. 12:54:13going to do the full dummies. We don't
  16498. 12:54:14need to um pick just a few columns.
  16499. 12:54:18We're actually going to do our full
  16500. 12:54:20dummies here and um do x equals 1. Now,
  16501. 12:54:24the reason that's the case is because um
  16502. 12:54:28this will get rid of one column per
  16503. 12:54:32feature and basically assume that if you
  16504. 12:54:35have a if you have a zero, the other one
  16505. 12:54:37should be a one. If you have a one, the
  16506. 12:54:39other one should be a zero. Um so it
  16507. 12:54:42basically makes that assumption because
  16508. 12:54:44we only have two of them. Um so whenever
  16509. 12:54:47there's a one, the other should be zero.
  16510. 12:54:50Um, so you can get away with just having
  16511. 12:54:52these three, but um I think it makes
  16512. 12:54:55more sense to just have to have the full
  16513. 12:54:58dummies,
  16514. 12:55:00but by process of elimination, you can
  16515. 12:55:03get away with just using two of them
  16516. 12:55:04because anytime you have a zero, the
  16517. 12:55:06other one should be the other feature
  16518. 12:55:08would have been would have been a one,
  16519. 12:55:11right? And vice versa, when there's a
  16520. 12:55:13one, the other feature would have been a
  16521. 12:55:14zero.
  16522. 12:55:20So we do that one.
  16523. 12:55:31And you can see all of our uh all of our
  16524. 12:55:34one hot encoding features end up back in
  16525. 12:55:36there.
  16526. 12:55:40So this is the code that I want you guys
  16527. 12:55:42to run. I think it makes more sense. It
  16528. 12:55:45follows along what we've been doing.
  16529. 12:55:47um which will concatenate our dummies
  16530. 12:55:49back to our features here to build out
  16531. 12:55:52our full X. So now X is all of our
  16532. 12:55:54features. Um remember X
  16533. 12:55:58X X contains all of our features
  16534. 12:56:04now.
  16535. 12:56:05So X contains all of our features and so
  16536. 12:56:09we have all of this now.
  16537. 12:56:15Okay. Were you guys able to run this
  16538. 12:56:17one?
  16539. 12:56:19D. We have y, we have x. We still need
  16540. 12:56:23to deal with the nles in y. So that
  16541. 12:56:26something we still need to deal with.
  16542. 12:56:33But hopefully you have this. Now
  16543. 12:56:38all of these are numerical.
  16544. 12:56:41So that should be good with the model.
  16545. 12:56:44That's one thing about X is you should
  16546. 12:56:47you our X should have all numerical
  16547. 12:56:51features, right? Because it's going to
  16548. 12:56:53go into a model to to learn those betas.
  16549. 12:56:56So it needs to have all numerical
  16550. 12:56:58features,
  16551. 12:57:00right? These are going to be all
  16552. 12:57:01numerical, which makes sense. We change
  16553. 12:57:03we did one hot encoding to change all
  16554. 12:57:05those guys to numerical.
  16555. 12:57:15Sorry, I'm scrolling down.
  16556. 12:57:20Okay, we do fill in the nator. Okay.
  16557. 12:57:33Okay.
  16558. 12:57:35Any questions so far? So, we're just
  16559. 12:57:36getting our data ready. We haven't
  16560. 12:57:37applied the lasso yet, but we're just
  16561. 12:57:39doing some prep. Now, hopefully you guys
  16562. 12:57:42recognize th these are some standard
  16563. 12:57:45steps that we're taking when we do our
  16564. 12:57:48modeling. We have to do these data prep
  16565. 12:57:50steps. They're necessary. And so, if it
  16566. 12:57:53seems like it's a lot of work, that's
  16567. 12:57:55because it is. It is work that you do to
  16568. 12:57:59prepare your data to get ready for
  16569. 12:58:01modeling. You have to do that. Okay.
  16570. 12:58:07So, we're doing that here. Um, now we're
  16571. 12:58:10going to do our train test split because
  16572. 12:58:12we're just going to do uh we're going to
  16573. 12:58:14do hold out here. So, we're doing a
  16574. 12:58:16train test split with about with a test
  16575. 12:58:18size of about 0.25. So, again, anywhere
  16576. 12:58:20between 02 to.3 would be okay.
  16577. 12:58:24Um, so uh it's our choice. We could do
  16578. 12:58:2902. We could do 3. We could do anywhere
  16579. 12:58:31in between there. We're doing 0.25.
  16580. 12:58:33That's fine. Um, that's okay. So, we we
  16581. 12:58:37build our train test split right there.
  16582. 12:58:40Um, so pretty pretty simple and we've
  16583. 12:58:42seen that a bunch of times with our X
  16584. 12:58:45and our Y data frames. There we have our
  16585. 12:58:50train test split.
  16586. 12:58:53Okay.
  16587. 12:58:55Um, now what we're going to do is do our
  16588. 12:58:58our scaling. So, we're we didn't do this
  16589. 12:59:01last time, but we're going to do this
  16590. 12:59:02now as uh because we should get in the
  16591. 12:59:05habit of doing that. Um is um we're
  16592. 12:59:10going to um go ahead and scale our
  16593. 12:59:13features and we're going to use the
  16594. 12:59:15standard scaler here uh to do that
  16595. 12:59:19scaling. Okay. Now, we could use minmax
  16596. 12:59:22scaler that's fine, too. We're just
  16597. 12:59:24going to use the standard scaler here.
  16598. 12:59:26Um and remember we are going to uh um
  16599. 12:59:31use the standard scaler from sklearn and
  16600. 12:59:34we're going to transform our features uh
  16601. 12:59:38uh according to our um according to our
  16602. 12:59:43training data. So we have our
  16603. 12:59:47pre-processing standard scaler here. So
  16604. 12:59:49we import that guy and then we um build
  16605. 12:59:53our standard scaler and fit it on the
  16606. 12:59:56training data only on the numerical
  16607. 12:59:59features. Um so that's which is going to
  16608. 13:00:04be uh all of these guys. So we're doing
  16609. 13:00:08the scaling on all of these guys. Now,
  16610. 13:00:10something to note is that we are not
  16611. 13:00:13scaling all of these one hot encodings
  16612. 13:00:16mainly because it doesn't make sense to
  16613. 13:00:18scale those really. They're zero or one.
  16614. 13:00:20They don't need to be scaled, right?
  16615. 13:00:23They're already zero and one. So,
  16616. 13:00:25they're they don't need to be even if we
  16617. 13:00:27were doing minmax scaling, it's going to
  16618. 13:00:29put them between zero and one. It
  16619. 13:00:30wouldn't affect it really, right? So,
  16620. 13:00:33these one hot encoding features, we're
  16621. 13:00:35not going to scale because they're
  16622. 13:00:36they're always going to be zero or one.
  16623. 13:00:39There's no need to scale them really.
  16624. 13:00:42Um, but we're going to scale all the
  16625. 13:00:44other features here that are floats.
  16626. 13:00:46So that's these guys here. These
  16627. 13:00:50numerical features we're going to scale.
  16628. 13:00:53Okay.
  16629. 13:00:54Don't need to we don't really need to
  16630. 13:00:56scale the one hot encoding. Uh, it's
  16631. 13:00:58pretty much already scaled.
  16632. 13:01:05Oh, you should change that. Um, go back
  16633. 13:01:08and rerun go back and rerun this. But
  16634. 13:01:10make sure you have your data type as int
  16635. 13:01:13here.
  16636. 13:01:15Make sure you add that in there to
  16637. 13:01:17change that over to integer and rerun
  16638. 13:01:19that and then rerun the rerun the
  16639. 13:01:23concatenation.
  16640. 13:01:24So make sure you run this
  16641. 13:01:27and then u make sure you rerun this and
  16642. 13:01:29rerun the concatenation part which is uh
  16643. 13:01:34this
  16644. 13:01:47Okay. So, we go ahead and fit the um
  16645. 13:01:52scaler to this data and then we're going
  16646. 13:01:54to transform our training features,
  16647. 13:01:56those numerical features. Um we're and
  16648. 13:02:00then we're going to uh transform these
  16649. 13:02:02features uh uh the test features in the
  16650. 13:02:05same way. So we're going to perform the
  16651. 13:02:08same transformation from the scaler on
  16652. 13:02:10the test data. So that's something
  16653. 13:02:12really important I want to note here is
  16654. 13:02:13that we always scale both the training
  16655. 13:02:19and test data. We always scale both. Of
  16656. 13:02:22course, we're going to train the model
  16657. 13:02:23on the training data. Um, but we are
  16658. 13:02:27going to also test it on the testing
  16659. 13:02:30data and it also needs to be scaled
  16660. 13:02:32because our model that we build is going
  16661. 13:02:34to assume scaled features. The
  16662. 13:02:37coefficients that it learns are going to
  16663. 13:02:39be assuming scaled features.
  16664. 13:02:42So we need to also scale our test data
  16665. 13:02:46in the same way. So we're doing that as
  16666. 13:02:49well.
  16667. 13:02:53So, we scale that. And now we have our
  16668. 13:02:56uh training and testing features have
  16669. 13:02:58been scaled.
  16670. 13:03:01No, we haven't replaced. We're going to
  16671. 13:03:02do that. We have not yet. We're going to
  16672. 13:03:05do that coming up in a minute. Yeah, we
  16673. 13:03:08haven't done that. Um, it is it is the
  16674. 13:03:11label. We definitely need to replace
  16675. 13:03:13NLES. We just haven't done it yet
  16676. 13:03:15because it's not in the features and
  16677. 13:03:16we're doing all of our uh uh
  16678. 13:03:18pre-processing to our pre-processing to
  16679. 13:03:20our features.
  16680. 13:03:30Yeah. So, we're definitely we need to
  16681. 13:03:32we're going to in a minute.
  16682. 13:03:36Okay. So, if you look at the data now,
  16683. 13:03:38it's all been scaled. So, these are all
  16684. 13:03:40um zcores. These are all on a much
  16685. 13:03:43better scale now. Um, and these are we
  16686. 13:03:48still have our one hot encoding features
  16687. 13:03:49which are zero or one. So this scaling
  16688. 13:03:53should lead to a better model than if we
  16689. 13:03:56didn't scale. So scaling is really
  16690. 13:03:59important. We can see that here.
  16691. 13:04:06Okay.
  16692. 13:04:08Now, um, let me ask you guys, were you
  16693. 13:04:10able to run the scaling? Are you caught
  16694. 13:04:12up to here? If you're following along,
  16695. 13:04:15were you able to run the scaling?
  16696. 13:04:28Okay, great. Great.
  16697. 13:04:36Awesome.
  16698. 13:04:39Okay. So, uh what we're going to do now
  16699. 13:04:42is replace nulls in the uh replace nles
  16700. 13:04:48by calculating the median of the data.
  16701. 13:04:52So, what I want you to notice is that we
  16702. 13:04:55are taking the NLES now this is um this
  16703. 13:04:59is on purpose is we are purposely taking
  16704. 13:05:02the NLES um out of the median
  16705. 13:05:05calculation. So we're skipping the NLES
  16706. 13:05:07when we compute the median because we
  16707. 13:05:09don't want those NLES to affect the
  16708. 13:05:11median calculation.
  16709. 13:05:13Um so we compute a median salary here
  16710. 13:05:16and then we fill our NLES with the
  16711. 13:05:19median salary um from the training data.
  16712. 13:05:23So this is our choice. This is a choice
  16713. 13:05:26um to use the median and it's also a
  16714. 13:05:30choice to use the training set median
  16715. 13:05:34for both train and test. What we could
  16716. 13:05:38have done, this is an alternative that
  16717. 13:05:40we could have done is use the entire
  16718. 13:05:43column and then um use the median of all
  16719. 13:05:48of the data to replace. That's really up
  16720. 13:05:50to us. Um this is one way of doing it.
  16721. 13:05:53We could have done before we did the
  16722. 13:05:56split. We could have um filled in with
  16723. 13:06:00the median earlier. We chose to do it
  16724. 13:06:03here mainly because it doesn't affect
  16725. 13:06:05the features. So, we could have done
  16726. 13:06:07this earlier and did it before we did
  16727. 13:06:09the split and filled the NAS. Um really
  16728. 13:06:12doesn't it's doesn't matter that much
  16729. 13:06:14which way we do it. Um but we do need to
  16730. 13:06:17fill in NLES. We cannot have those be
  16731. 13:06:19null when we when we put it into our
  16732. 13:06:21model. So some way we need to fill in
  16733. 13:06:23NLES. Um and so in this strategy we're
  16734. 13:06:27filling in our Y train um with the
  16735. 13:06:30median salary from our training data.
  16736. 13:06:32And same with this we're filling in with
  16737. 13:06:34the median salary of the training data
  16738. 13:06:36as well. But that's a choice f we could
  16739. 13:06:39fill in with the mean with the average.
  16740. 13:06:43Um we could fill in with the we could do
  16741. 13:06:46it with all the data together before we
  16742. 13:06:48split it. we could have filled in with
  16743. 13:06:50all of the the median across the whole
  16744. 13:06:52data set. Um either one works. You can
  16745. 13:06:55do it either way, but we we did it um
  16746. 13:06:59later here to show that it doesn't
  16747. 13:07:01really affect the features. So, we can
  16748. 13:07:03choose when we do it, right? It doesn't
  16749. 13:07:05affect the features at all. So, we can
  16750. 13:07:08do all of our pre-processing on the
  16751. 13:07:09features and then do our label uh
  16752. 13:07:12filling and nulls um if if we have them.
  16753. 13:07:20uh x numerical. Um make sure you're
  16754. 13:07:23running uh this
  16755. 13:07:27uh x numerical was defined here
  16756. 13:07:32when we split it apart um from
  16757. 13:07:37uh when we dropped these columns here.
  16758. 13:07:39So make sure you're running this. This
  16759. 13:07:41is x numerical
  16760. 13:07:43gets defined there.
  16761. 13:07:45So, go back up to uh this cell
  16762. 13:07:49where we split apart the y and we and we
  16763. 13:07:51have the x here x numerical.
  16764. 13:07:55Make sure you run this.
  16765. 13:08:04Make sure you run this. And then you can
  16766. 13:08:06run these. Then you run this to build x.
  16767. 13:08:19All right.
  16768. 13:08:21Are we up to here with this filling in
  16769. 13:08:24the labels?
  16770. 13:08:26Uh because then we can build our model
  16771. 13:08:29once we're up to here. We've scaled
  16772. 13:08:30everything. We filled in our NLES.
  16773. 13:08:34We've gotten one hot encoding.
  16774. 13:08:41Yeah, it is. That's why you know that's
  16775. 13:08:42why we spend a lot of uh time on model
  16776. 13:08:45on data preparation with pandas, right?
  16777. 13:08:47That's why we did all that pandas work
  16778. 13:08:49for sure. Yes, there is a lot of work
  16779. 13:08:51before we can build a model.
  16780. 13:08:54Yes, the mo do you guys notice that like
  16781. 13:08:57the modeling is relatively easy. It's
  16782. 13:08:58just a fit and predict. The modeling is
  16783. 13:09:01actually really easy. It's all the other
  16784. 13:09:03work that's that's more involved, right?
  16785. 13:09:07more code.
  16786. 13:09:10The modeling itself is really easy.
  16787. 13:09:13It's just it's just one line of like
  16788. 13:09:15ffit.
  16789. 13:09:18Yeah, pretty easy to do.
  16790. 13:09:23And then you do evaluation which is a
  16791. 13:09:25couple lines.
  16792. 13:09:47Yep. There's these are all the these are
  16793. 13:09:50the common steps. All these steps we're
  16794. 13:09:52doing are very very prototypical in
  16795. 13:09:54model building is you let's just go back
  16796. 13:09:57through this to see what we did right we
  16797. 13:09:59imported our data
  16798. 13:10:01um we analy we dropped this name column
  16799. 13:10:04because it's not useful to us so we
  16800. 13:10:06dropped that um we filled in the nles
  16801. 13:10:10eventually um but you know if there were
  16802. 13:10:13any nles in our features we would have
  16803. 13:10:14to deal with those as well by replacing
  16804. 13:10:16them or dropping the rows like we did
  16805. 13:10:18earlier Um
  16806. 13:10:21and then we do one hot encoding because
  16807. 13:10:24of course we can't have any string
  16808. 13:10:25columns going in our models. We got a
  16809. 13:10:27one hot encode.
  16810. 13:10:29Um we uh then build our X and Y by
  16811. 13:10:33concatenating the one hot encoded back
  16812. 13:10:36to the numerical features.
  16813. 13:10:39Then we train test split. Right? That's
  16814. 13:10:41pretty common. Or we could do cross
  16815. 13:10:43validation either way. Um the K full
  16816. 13:10:46cross validation. Then we scale. So, we
  16817. 13:10:49didn't do this last time, but this is
  16818. 13:10:50something we should get in the habit of
  16819. 13:10:52is scaling um our features. So, we do
  16820. 13:10:55that and now we're ready to model. So,
  16821. 13:10:58now we're ready to model. Um so, that's
  16822. 13:11:01this part.
  16823. 13:11:04Okay. So, let's do the model. Um the
  16824. 13:11:07model's actually uh pretty easy to do.
  16825. 13:11:10So, we're going to use a lasso. So, we
  16826. 13:11:12have a lasso model here. Notice what
  16827. 13:11:14we're setting our alpha to. So the big
  16828. 13:11:16parameter we really need ignore this
  16829. 13:11:19iterations. We actually don't really
  16830. 13:11:20need the we don't really need that
  16831. 13:11:21parameter. Um so just ignore it for the
  16832. 13:11:24moment. But the big one that we're
  16833. 13:11:26setting here is the alpha. So when we
  16834. 13:11:29did linear regression, we didn't need
  16835. 13:11:31any parameters to go inside the linear
  16836. 13:11:33regression object. We didn't need any
  16837. 13:11:35parameters, right? Because there are
  16838. 13:11:37really no parameters of it. But for
  16839. 13:11:39lasso, the important one is the alpha.
  16840. 13:11:42And so we need to know what to set alpha
  16841. 13:11:44to. Um let's start with alpha equals 1.
  16842. 13:11:49That's a good starting place. So a
  16843. 13:11:51typical um starting point
  16844. 13:11:55for alpha
  16845. 13:11:58um is uh is one. So that's a typical
  16846. 13:12:03starting point. And so we can set alpha
  16847. 13:12:05equals to one. This max iterations is
  16848. 13:12:08the the parameter that governs the
  16849. 13:12:11training process because it is
  16850. 13:12:13iterative. So if for some reason we we
  16851. 13:12:16can't converge to the right betas and
  16852. 13:12:18we've run it for 10,000 steps once we
  16853. 13:12:21pass 10,000 steps, uh it will stop and
  16854. 13:12:24just give us the betas at that point.
  16855. 13:12:26But it will likely never hit this
  16856. 13:12:29number. It'll converge before then. So
  16857. 13:12:31um we don't really need to um specify
  16858. 13:12:34it. So, I'm actually just going to get
  16859. 13:12:36rid of it. Um, it's not really a big
  16860. 13:12:38deal. It should converge before then.
  16861. 13:12:41Um, but if if we want to set like a
  16862. 13:12:44maximum step size in the optimization,
  16863. 13:12:46we definitely could there. Uh, but not
  16864. 13:12:49concerned about that too much. But
  16865. 13:12:51here's our lasso. And then we're just
  16866. 13:12:53going to do a fit on our data. So, look
  16867. 13:12:56how easy that is. Just like a linear
  16868. 13:12:58regression. Lasso.fit,
  16869. 13:13:02right? So, we do fit. Um,
  16870. 13:13:10oh, I didn't run this. I'm sorry. I got
  16871. 13:13:12to run this. Okay. Actually, that's a
  16872. 13:13:16good example of what happens when you
  16873. 13:13:17don't when you have nulls, right? So, it
  16874. 13:13:19says our our null contains nan. That's
  16875. 13:13:21because I didn't run this. But now, that
  16876. 13:13:24should be filled in. Now, we should be
  16877. 13:13:26able to run this. Okay, perfect. So it
  16878. 13:13:28runs.
  16879. 13:13:36Okay. So you can see what the intercept
  16880. 13:13:38is. Um this is one of our coefficients,
  16881. 13:13:40right? The intercept is 457. And look
  16882. 13:13:43now what's really interesting about the
  16883. 13:13:44coefficients is look at what some of the
  16884. 13:13:47coefficients are.
  16885. 13:13:49Some of them are actually zero, which is
  16886. 13:13:53really So some of them ended up being at
  16887. 13:13:55zero, which is very very interesting.
  16888. 13:13:58that means that those features get
  16889. 13:14:02cancelceled out and they're basically
  16890. 13:14:03not part of the model which is really
  16891. 13:14:06interesting. Um so we have all these
  16892. 13:14:09coefficients and some of them are zero.
  16893. 13:14:16Yeah, negative0 is just because of the
  16894. 13:14:19convergence like they started out
  16895. 13:14:21negative and worked their way up to
  16896. 13:14:23zero. it. Negative zero really just
  16897. 13:14:26means zero, but they just were coming
  16898. 13:14:29from they were like small negatives and
  16899. 13:14:31ended up at zero
  16900. 13:14:34during the training process. They were
  16901. 13:14:36negative at one point and ended up zero.
  16902. 13:14:39Um
  16903. 13:14:40so yeah, negative 0 just obviously means
  16904. 13:14:43zero. Um it's still still zero there.
  16905. 13:14:54So what's interesting is some of these
  16906. 13:14:56features ended up uh being zero which
  16907. 13:14:59you don't usually see in a linear
  16908. 13:15:01regression. So if we were to train this
  16909. 13:15:03using a linear regression we typically
  16910. 13:15:05wouldn't see that but some of these turn
  16911. 13:15:07out to be zero because again we're
  16912. 13:15:10encouraging those betas to be small.
  16913. 13:15:13we're encouraging them to be uh small
  16914. 13:15:16and so um you know what happens is some
  16915. 13:15:20of them can be shrunk all the way down
  16916. 13:15:22to zero meaning those features don't
  16917. 13:15:24contribute that's a really simple model
  16918. 13:15:26at that point right so we've taken
  16919. 13:15:29something complex that includes all of
  16920. 13:15:32these features and actually reduced it
  16921. 13:15:34into something simple that only includes
  16922. 13:15:36these features
  16923. 13:15:38right
  16924. 13:15:40so that's what it does um now we need to
  16925. 13:15:44evaluate this to see how good of a model
  16926. 13:15:46it is. But that's what this is saying
  16927. 13:15:49here in this text is that um a positive
  16928. 13:15:52uh coefficient indicates that as the
  16929. 13:15:54independent variable increases the
  16930. 13:15:56dependent variable also increases.
  16931. 13:15:58Negative coefficient means as the
  16932. 13:16:00independent variable increases dependent
  16933. 13:16:02decreases because it's reducing the
  16934. 13:16:04value. Um and lasso is known for feature
  16935. 13:16:09selection by shrinking some of them to
  16936. 13:16:10zero effectively removing those
  16937. 13:16:13variables from the model from the
  16938. 13:16:15equation right
  16939. 13:16:17um
  16940. 13:16:19so that's what happens
  16941. 13:16:23some of them end up being zero
  16942. 13:16:27were you guys able to run this this
  16943. 13:16:30lasso uh fit which is the training of
  16944. 13:16:33the lasso
  16945. 13:16:52No, it doesn't ensure there's no
  16946. 13:16:53overfit, but it helps with overfitting.
  16947. 13:16:56It's supposed to help by making the
  16948. 13:16:58model simpler. And this is definitely a
  16949. 13:16:59simpler model because it's removing some
  16950. 13:17:02of the features from the model
  16951. 13:17:03essentially, right? Because some of the
  16952. 13:17:05features aren't going to contribute.
  16953. 13:17:07It's a simpler model.
  16954. 13:17:09It doesn't it doesn't mean there's not
  16955. 13:17:11going to be any overfitting, but it
  16956. 13:17:13helps prevent it. That's what it's
  16957. 13:17:15designed to do, help prevent it.
  16958. 13:17:19Yeah. So, higher coefficient. Yes. The
  16959. 13:17:22higher coefficient means it's a more
  16960. 13:17:24important feature towards the
  16961. 13:17:26prediction.
  16962. 13:17:27Yes. That's what it means for sure. The
  16963. 13:17:30higher the magnitude, the more of a
  16964. 13:17:32contributor towards that prediction. Uh
  16965. 13:17:35it is. Yes.
  16966. 13:17:43And it's not just it's it could be
  16967. 13:17:45higher positive or negative there. Like
  16968. 13:17:47a higher negative is also a pretty big
  16969. 13:17:50factor,
  16970. 13:17:52right? So So you want to think about it
  16971. 13:17:53in terms of absolute value.
  16972. 13:18:01does not guarantee but helps. Yes, it
  16973. 13:18:03doesn't guarantee it but it's designed
  16974. 13:18:05to help overfitting, help prevent it.
  16975. 13:18:07Yes, absolutely.
  16976. 13:18:32Okay.
  16977. 13:18:35So let's do some evaluation. Um so let's
  16978. 13:18:38do in this case we are going to do our
  16979. 13:18:42predict
  16980. 13:19:02Oh, yeah. I'm not sure why that's the
  16981. 13:19:05case.
  16982. 13:19:07Interesting.
  16983. 13:19:27We could try increasing the um max
  16984. 13:19:31iterations.
  16985. 13:19:43Okay, that's why. Yeah. So then you get
  16986. 13:19:45that result with the with the higher max
  16987. 13:19:47iterations.
  16988. 13:19:49It doesn't get cut off there.
  16989. 13:19:53I think that's why you probably left
  16990. 13:19:54this in there,
  16991. 13:19:57which is fine. You get about the same
  16992. 13:19:58numbers.
  16993. 13:20:06Yeah.
  16994. 13:20:16All right. Let's evaluate this. So,
  16995. 13:20:17we're going to to to do evaluation. I
  16996. 13:20:19want you guys to see again, we should
  16997. 13:20:21get in the habit of doing evaluation,
  16998. 13:20:24which is taking our model and predicting
  16999. 13:20:26on the training and predicting on the
  17000. 13:20:29test sets, right? So we predict on the
  17001. 13:20:31train set and calculate our MSE
  17002. 13:20:35and we um calculate our R2 score um or R
  17003. 13:20:41squar score I should say. Uh but again
  17004. 13:20:44the MSE is the one we're really going to
  17005. 13:20:46use mostly. Um but we calculate so we do
  17006. 13:20:50our predictions and then we compare that
  17007. 13:20:52into our mean squared error with our
  17008. 13:20:54labels
  17009. 13:20:56and we uh go ahead and do the same thing
  17010. 13:21:00with the test. Right? So we do uh
  17011. 13:21:02lasso.predict
  17012. 13:21:04on our test features and we go ahead and
  17013. 13:21:07compare that with the test labels. And
  17014. 13:21:10so what we're doing there is generating
  17015. 13:21:12our MSE.
  17016. 13:21:15So we we take a look at our MSE and we
  17017. 13:21:18get uh 84,000
  17018. 13:21:21MSE. Um and so of course we could take
  17019. 13:21:25the um what we could do with that is
  17020. 13:21:29take a look at the um MSE on the uh we
  17021. 13:21:34could do um MP. Square root
  17022. 13:21:39and do the square root of the MSE test.
  17023. 13:21:45and we get um 340. So this would be in
  17024. 13:21:49the units of our label. So, if we go
  17025. 13:21:51back and look at our label um for some
  17026. 13:21:54of those um
  17027. 13:22:08so uh we are in 300s and our data is
  17028. 13:22:12like right around the 500. So, of
  17029. 13:22:14course, if we describe this um we could
  17030. 13:22:16see what the statistics are of it. So,
  17031. 13:22:19we could do df.escribe describe and
  17032. 13:22:21generate that. But that doesn't look
  17033. 13:22:22like a very good error, right? If these
  17034. 13:22:24are in the 400s, um that's that's not a
  17035. 13:22:27very good error.
  17036. 13:22:29So again, it's not a very great model.
  17037. 13:22:32But one thing I want you to see is that
  17038. 13:22:33it's it's not overfitting.
  17039. 13:22:36Um if anything, it's actually
  17040. 13:22:38underfitting, which is what this kind of
  17041. 13:22:41um MSE suggests, right? because our
  17042. 13:22:43error here is 84 uh excuse me 84,000.
  17043. 13:22:48Um
  17044. 13:22:54our our area here is 84,000,
  17045. 13:22:58excuse me. And on the test set it's
  17046. 13:23:01116,000.
  17047. 13:23:02Um so these two errors are both bad. So
  17048. 13:23:08it's not overfitting. This is actually
  17049. 13:23:10underfitting. So it's not overfitting,
  17050. 13:23:13it's actually underfitting. Um, and so
  17051. 13:23:16that's the risk with something like
  17052. 13:23:17lasso is that it's making the model a
  17053. 13:23:20bit too simple and we actually risk
  17054. 13:23:23underfitting, which is what happens. We
  17055. 13:23:26have too much error across both the
  17056. 13:23:29training and the test set. Overfitting
  17057. 13:23:31is when we do we have really good
  17058. 13:23:34performance on the training set, but bad
  17059. 13:23:36performance on the test set. We're not
  17060. 13:23:38overfitting.
  17061. 13:23:40um we are uh underfitting because our
  17062. 13:23:43performance is not good either way. Even
  17063. 13:23:45this R squar is pretty low. It's not
  17064. 13:23:48even at 50%.
  17065. 13:23:56Okay, so that's so we we do the
  17066. 13:23:59evaluation and again the evaluation just
  17067. 13:24:00comes down to making predictions and
  17068. 13:24:03computing our error amongst those
  17069. 13:24:05predictions to our labels. That's always
  17070. 13:24:07what the uh evaluation is going to be
  17071. 13:24:14for MSE.
  17072. 13:24:17What's the ideal MSE? What do you think
  17073. 13:24:20it should be? What is So, think about it
  17074. 13:24:23like this. The MSE represents the
  17075. 13:24:25average distance between our predictions
  17076. 13:24:30and the labels.
  17077. 13:24:33So, if we're getting it right all the
  17078. 13:24:36time, what's that distance going to be
  17079. 13:24:38if we're always right? What's our
  17080. 13:24:40distance from what's our distance from
  17081. 13:24:43our predictions to our labels going to
  17082. 13:24:46be if we're always getting it right?
  17083. 13:24:49Zero. Yeah, there's not going to be any
  17084. 13:24:51distance. It's going to be right. It's
  17085. 13:24:53going to be perfectly aligned, right?
  17086. 13:24:55There's going to be no distance there.
  17087. 13:24:57So, yeah, an ideal MSE is zero.
  17088. 13:25:00That's an ideal MSE.
  17089. 13:25:03So, anything close to like the smaller
  17090. 13:25:06the better for MSE. The smaller the
  17091. 13:25:09better. Um, for this R squared, uh, it's
  17092. 13:25:13it's a scale between 0 to one where one
  17093. 13:25:16is the best. So, one would be perfectly
  17094. 13:25:18aligned predictions. Um, so, and again,
  17095. 13:25:22this this is we actually multiply by 100
  17096. 13:25:25to get uh because it's it's a number
  17097. 13:25:26between 0 and one. So we get about 47%
  17098. 13:25:30which is not good.
  17099. 13:25:41Okay.
  17100. 13:25:56All right. Any questions on this
  17101. 13:25:59evaluation?
  17102. 13:26:13All right. I want to show you something
  17103. 13:26:15which is
  17104. 13:26:18Yeah, this that's true. the scale of it
  17105. 13:26:21matters on the data because we should be
  17106. 13:26:22you should always interpret your MSE in
  17107. 13:26:25the scale of
  17108. 13:26:27um your your labels because your labels
  17109. 13:26:32like in this case our labels um you know
  17110. 13:26:34we could take uh for example we could
  17111. 13:26:37easily let's actually do that let's take
  17112. 13:26:39the average
  17113. 13:26:42let's take the average of our labels on
  17114. 13:26:45the training data
  17115. 13:26:49and and we can see what those are. Um,
  17116. 13:26:52so the average is 500,
  17117. 13:26:55right? The average is 500. And look at
  17118. 13:26:58what our uh square root of our MSE is,
  17119. 13:27:01which is in the same units as our
  17120. 13:27:03original. Um, so we have uh quite a bit
  17121. 13:27:07of error. 340 when our units are right
  17122. 13:27:10around 500.
  17123. 13:27:13So that's quite a bit of error.
  17124. 13:27:23Yeah, MSSE of zero means our our uh our
  17125. 13:27:26predictions are nearly identical to the
  17126. 13:27:29test labels. Yes, that's what MSSE of
  17127. 13:27:33zero means. There's zero distance.
  17128. 13:27:38So closer to zero, the better.
  17129. 13:27:47But we talked about it as you you really
  17130. 13:27:49so the rule of thumb should be what is
  17131. 13:27:53your RMSSE as a percentage of your
  17132. 13:27:57typical value. So your typical value is
  17133. 13:28:00in the 500s. Our our RMSSE is 340.
  17134. 13:28:05That's just really high. That's over
  17135. 13:28:07like 60% of that value.
  17136. 13:28:11So that's just a lot. That's too much
  17137. 13:28:14error. What we would love this RMSSE to
  17138. 13:28:16be is under 20% of the typical value. So
  17139. 13:28:20that means on average we are 20% or less
  17140. 13:28:25off in our prediction. That would be
  17141. 13:28:28good. That would be pretty good. That
  17142. 13:28:30means we're like 80% accurate,
  17143. 13:28:34right? That'd be pretty ideal. So you
  17144. 13:28:36got to think about it in terms of this
  17145. 13:28:37RMSSE which is in the same units as your
  17146. 13:28:40labels.
  17147. 13:28:43This is the
  17148. 13:28:47RMSSE
  17149. 13:28:50which is in the same units as the
  17150. 13:28:54labels.
  17151. 13:28:57So and then to interpret this we have
  17152. 13:29:00340
  17153. 13:29:02is compared to
  17154. 13:29:05typical
  17155. 13:29:07um salary unit of 500
  17156. 13:29:11right so this is uh quite a bit when the
  17157. 13:29:15typical value is 500 and we are off on
  17158. 13:29:18average by 340 units
  17159. 13:29:21that's so much relative to the typical
  17160. 13:29:24value
  17161. 13:29:26that's just too. That's a lot of error.
  17162. 13:29:28That's not a very good model, right?
  17163. 13:29:31It's underfitting. It's definitely
  17164. 13:29:33underfitting.
  17165. 13:29:45Yeah. So, that's a great question. What
  17166. 13:29:46should we do from here? So, um because
  17167. 13:29:49we're underfitting
  17168. 13:29:57um we should use a more complex model.
  17169. 13:30:02So uh we're going to learn about those
  17170. 13:30:04in lesson four, but we should use
  17171. 13:30:06something different. This linear
  17172. 13:30:08regression is still too basic. Even with
  17173. 13:30:10lasso, it's still too basic.
  17174. 13:30:16Yeah, we're underfitting because we But
  17175. 13:30:18it could also be we're underfitting with
  17176. 13:30:20a regular linear regression. We should
  17177. 13:30:22test that out. Um, and maybe it would be
  17178. 13:30:24an exercise for you guys um to test that
  17179. 13:30:28out yourself. It shouldn't be hard to
  17180. 13:30:29do. Um, you already have all the data
  17181. 13:30:32scaled. You So, do you see how you would
  17182. 13:30:35do that? You would just come in here and
  17183. 13:30:36build a linear regression rather than a
  17184. 13:30:38lasso and dofit. And then you would
  17185. 13:30:41evaluate it the same way with a predict.
  17186. 13:30:44It's really easy to do that. And then we
  17187. 13:30:46can compare that um to to this. It
  17188. 13:30:51shouldn't be that hard to do that,
  17189. 13:30:52right?
  17190. 13:30:54And something you guys could do for
  17191. 13:30:55sure. Um,
  17192. 13:30:58is build the linear regression and
  17193. 13:31:00actually compare it and see what kind of
  17194. 13:31:04difference it makes. I mean, we honestly
  17195. 13:31:06we could do it ourselves. We could do it
  17196. 13:31:07right now. Maybe it's worth trying that.
  17197. 13:31:12So, let's build a linear regression
  17198. 13:31:16for comparison.
  17199. 13:31:20So we have our linear regression
  17200. 13:31:24uh is linear regression and then we do
  17201. 13:31:28ffit linear regression.fit fit
  17202. 13:31:32right so so this will train it um and
  17203. 13:31:35then we can evaluate it so lin mse is
  17204. 13:31:41mean squared error
  17205. 13:31:44and then we can do our um let's do our
  17206. 13:31:47training let's do the training and then
  17207. 13:31:52um let's predict
  17208. 13:31:54actually let me do that here
  17209. 13:31:57uh y prediction
  17210. 13:32:00train
  17211. 13:32:03linear
  17212. 13:32:05equals um linear regression.predict
  17213. 13:32:10and then we're going to predict on our
  17214. 13:32:12training features.
  17215. 13:32:17Okay, do you guys see what I'm doing?
  17216. 13:32:18I'm building a linear regression for
  17217. 13:32:20comparison.
  17218. 13:32:22I'm doing fit here to train it and then
  17219. 13:32:25I'm making some predictions on the
  17220. 13:32:26training set and we're going to evaluate
  17221. 13:32:29those. I'm going to replace that here
  17222. 13:32:30with y prred
  17223. 13:32:34uh train
  17224. 13:32:37linear. So these predictions
  17225. 13:33:02Okay. So, if you guys want this code, I
  17226. 13:33:03can paste it in.
  17227. 13:33:15So, let's see what the RMSSE for just a
  17228. 13:33:18linear model is.
  17229. 13:33:21It's a little bit better. It's better
  17230. 13:33:23for sure.
  17231. 13:33:26So 289 is better than this 340. It's
  17232. 13:33:30better. It's getting closer to zero.
  17233. 13:33:33It's still underfitting though,
  17234. 13:33:36right? And that's just on the training
  17235. 13:33:38set. Let's look at the Let's do the same
  17236. 13:33:41thing, but on
  17237. 13:33:45Let's change this. Let's swap this out
  17238. 13:33:47for um test
  17239. 13:33:51And then let's do test.
  17240. 13:33:55And then let's do test
  17241. 13:34:01test.
  17242. 13:34:04And then
  17243. 13:34:06test test.
  17244. 13:34:17Okay. Okay, so this is producing test
  17245. 13:34:19predictions on the test set.
  17246. 13:34:22We are generating an MSE test
  17247. 13:34:27and then we're doing MSE test
  17248. 13:34:30which is using the test labels and our
  17249. 13:34:32test predictions and then we take the
  17250. 13:34:35square root of that for RMSSE and then
  17251. 13:34:36we're going to generate that. So it's
  17252. 13:34:39still under fit. I mean this is still
  17253. 13:34:41high. This is still high um on the test
  17254. 13:34:44set and versus on the training set. So,
  17255. 13:34:46it's still pretty high. Um, even the
  17256. 13:34:49basic linear regression is under is
  17257. 13:34:50still underfitting. Still underfitting,
  17258. 13:34:53right? Even without the lasso,
  17259. 13:34:56which is lasso is supposed to help with
  17260. 13:34:58overfitting. It's definitely not
  17261. 13:34:59overfitting. Um, it's definitely
  17262. 13:35:02underfitting, but this is a signal that
  17263. 13:35:05it's kind of overfitting because this is
  17264. 13:35:07performing better on the training data
  17265. 13:35:09and then it gets worse on the test data.
  17266. 13:35:13Definitely gets worse, right?
  17267. 13:35:23Did you guys follow?
  17268. 13:35:27I'm just running this above I'm running
  17269. 13:35:29this above this. It doesn't matter where
  17270. 13:35:31you put it. We could uh we could move it
  17271. 13:35:33down.
  17272. 13:35:40We could move it down to I just ran I
  17273. 13:35:43just picked a new cell right here and
  17274. 13:35:46ran it. But we could move it actually
  17275. 13:35:48let's do that. Let's move it down
  17276. 13:35:53to
  17277. 13:35:56after the lasso evaluation.
  17278. 13:36:01Okay. So I just moved it there.
  17279. 13:36:04And then let's move
  17280. 13:36:07this down.
  17281. 13:36:09So I just put it here after the um after
  17282. 13:36:12this. So this is the um this is
  17283. 13:36:16basically the objective function right
  17284. 13:36:19of the training process. So during the
  17285. 13:36:22algorithm that runs when we call ffit in
  17286. 13:36:25scikitlearn it's going to find these
  17287. 13:36:28betas right it's actually going to learn
  17288. 13:36:30what these best betas are for our model.
  17289. 13:36:34Um this is our model here, right? It's
  17290. 13:36:36the combination of betas times our
  17291. 13:36:37features um plus an intercept beta. Uh
  17292. 13:36:42so that's our model. But um we penalize
  17293. 13:36:45those large uh weights in absolute value
  17294. 13:36:49by um adding a penalty term like this um
  17295. 13:36:53where alpha is some level of penalty
  17296. 13:36:56that we want to provide. Usually alpha
  17297. 13:36:58equals 1 is okay. But um actually what
  17298. 13:37:01we're going to learn uh to finish out
  17299. 13:37:02this section is there's going to be a
  17300. 13:37:04systematic way we can test out different
  17301. 13:37:06alphas um that represent the level of
  17302. 13:37:09penalty we want to uh apply to lasso or
  17303. 13:37:12even ridge
  17304. 13:37:14uh regression. So that was the lasso and
  17305. 13:37:17um if you guys remember using it was
  17306. 13:37:19super easy. Uh we worked through this
  17307. 13:37:22problem with this um baseball data um
  17308. 13:37:26and we had uh
  17309. 13:37:29let's see scrolling down we um split out
  17310. 13:37:31our numerical data and we did uh we one
  17311. 13:37:36hot encoded our our categorical data
  17312. 13:37:39combined it back together. Hopefully
  17313. 13:37:41that um rings a bell there. Um and we
  17314. 13:37:44actually scaled our data which is pretty
  17315. 13:37:46standard to do is we do some type of
  17316. 13:37:48scaling to our features especially our
  17317. 13:37:50numerical features right want to scale
  17318. 13:37:52those in some way whether it's minmax
  17319. 13:37:54scale or standard scaler um want to do
  17320. 13:37:57that and so we did that for this example
  17321. 13:37:59and then we um ran the lasso regression
  17322. 13:38:04which is pretty easy to use. You just
  17323. 13:38:05use the lasso object and you pick an
  17324. 13:38:07alpha here. Um, again, we are going to
  17325. 13:38:10have a way to test out different alphas
  17326. 13:38:14that could be candidates and we can see
  17327. 13:38:16which one's the best. Um, so I'm going
  17328. 13:38:19to show us that today coming up shortly.
  17329. 13:38:23But that was that was the lasso. If you
  17330. 13:38:25guys remember, we did that. Um, this it
  17331. 13:38:28we compared that to a basic linear
  17332. 13:38:30regression which is just this pretty
  17333. 13:38:32straightforward just a fit and then
  17334. 13:38:34predict and then we can generate mean
  17335. 13:38:35squed error. Um, still not a very good
  17336. 13:38:39mean squared error on this data, it's
  17337. 13:38:41still fairly large. Um, so it's still
  17338. 13:38:44not, no matter which model we use, it's
  17339. 13:38:46still not very good, but at least we can
  17340. 13:38:49practice doing that comparison. That's
  17341. 13:38:50what we did last time. We did this on
  17342. 13:38:53Wednesday.
  17343. 13:38:54Um
  17344. 13:38:56and then
  17345. 13:38:58we saw that the effect of different
  17346. 13:38:59alphas we had a lasso um
  17347. 13:39:03we had a lasso uh cross validation
  17348. 13:39:06example here. So beyond just using a
  17349. 13:39:08regular lasso model that um scikitlearn
  17350. 13:39:10has a lasso cv which allows you to try
  17351. 13:39:13out different alphas uh with cross
  17352. 13:39:16validation and um figure out what the
  17353. 13:39:18best alpha is. Um, now we're actually
  17354. 13:39:21going to have a different strategy
  17355. 13:39:22that'll instead of just picking random
  17356. 13:39:24ones, we can actually um supply multiple
  17357. 13:39:28parameters that we may want to test. Um,
  17358. 13:39:31as many as the models may support. And
  17359. 13:39:33in some more complex models, we'll have
  17360. 13:39:36more than one parameter like lasso only
  17361. 13:39:38has the alpha. Um, technically it also
  17362. 13:39:41has this max iterations, but really the
  17363. 13:39:43only one that matters is this alpha.
  17364. 13:39:45Other models have many more
  17365. 13:39:47hyperparameters that we can um uh change
  17366. 13:39:52and so we want a way to systematically
  17367. 13:39:54test out those different combinations
  17368. 13:39:57and to see which one leads to the best
  17369. 13:39:59uh version of that model. Let's say the
  17370. 13:40:01best results. So um we're going to
  17371. 13:40:04explore that coming up. So we had lasso.
  17372. 13:40:08Um now this is where we ended last time.
  17373. 13:40:10We had ridge regression. If you guys
  17374. 13:40:12remember, this one is just a slightly
  17375. 13:40:15different penalty. Um,
  17376. 13:40:18it takes the it I drew it out for us. It
  17377. 13:40:21takes the same penalty we had before.
  17378. 13:40:23So, it has that um residual sum of
  17379. 13:40:26squares error, which is the main one we
  17380. 13:40:29use for linear regression, but it has a
  17381. 13:40:31penalty with an alpha. And then it has
  17382. 13:40:33the sum of the beta squares
  17383. 13:40:37beta i squares. So it penalizes it has a
  17384. 13:40:42penalty but it penalizes slightly
  17385. 13:40:44differently where it uses the square not
  17386. 13:40:46the absolute value. That's the ridge
  17387. 13:40:48regression. And this has the similar
  17388. 13:40:50effect of you don't in order to minimize
  17389. 13:40:52this right because our goal in training
  17390. 13:40:54a model was to minimize this thing
  17391. 13:40:57minimize this um quantity and find the
  17392. 13:41:01best betas that minimize this. Um so
  17393. 13:41:04generally yes you want to encourage
  17394. 13:41:06lower values but the um once you get
  17395. 13:41:10values that are a fraction if you square
  17396. 13:41:12them they actually get smaller. Um, so,
  17397. 13:41:16uh, it's it's not, um, it's not
  17398. 13:41:20necessary to shrink them all the way to
  17399. 13:41:22zero. They will get smaller as soon as
  17400. 13:41:24they're kind of below one. Um, so they
  17401. 13:41:27don't encourage it to completely go
  17402. 13:41:29away, uh, like the absolute value does.
  17403. 13:41:32It's just slightly different
  17404. 13:41:33minimization. Um, so what we see with
  17405. 13:41:36the ridge is we don't see the features
  17406. 13:41:38kind of get wiped out completely like we
  17407. 13:41:40do with a lasso. and lasso they get
  17408. 13:41:42encouraged to be um to become zero
  17409. 13:41:45because that's kind of the only way to
  17410. 13:41:46minimize an absolute value. But with
  17411. 13:41:48squares they can keep getting smaller
  17412. 13:41:50and smaller and smaller um fractions and
  17413. 13:41:54they don't have to become zero. It's not
  17414. 13:41:56as harsh of a of a penalty.
  17415. 13:41:59Um so uh the ridge was easy to use as
  17416. 13:42:05well. Um and it also has an alpha that
  17417. 13:42:09we can set. So, it's literally the same
  17418. 13:42:12exact code, just a different model, just
  17419. 13:42:15slightly different penalty, and it
  17420. 13:42:17results in different coefficients. You
  17421. 13:42:19notice that none of them are exactly
  17422. 13:42:20zero. Like with the lasso, you can get
  17423. 13:42:22ones that are exactly zero. We don't see
  17424. 13:42:25that with the ridge. You remember that.
  17425. 13:42:28Um, so we we s pointed out that last
  17426. 13:42:30time. Notice the coefficients aren't
  17427. 13:42:31zero. Um, and then we can evaluate it.
  17428. 13:42:34So we did our MSE calculation which is a
  17429. 13:42:37pretty standard thing where we use our
  17430. 13:42:39model to predict on a training set,
  17431. 13:42:41predict on a test set, evaluate those um
  17432. 13:42:45by computing the metric like the mean
  17433. 13:42:47squared error and we can see if we're
  17434. 13:42:49overfitting underfitting. This is
  17435. 13:42:51definitely the same kind of story we've
  17436. 13:42:53seen with all these models is
  17437. 13:42:54underfitting because the error is so big
  17438. 13:42:56across both sets
  17439. 13:42:58across training and test. So it's it's
  17440. 13:43:00definitely underfitting.
  17441. 13:43:02Um
  17442. 13:43:04and same thing as lasso, it has a cross
  17443. 13:43:06validation uh variation on it that
  17444. 13:43:09allows you to try out different alphas
  17445. 13:43:12and um do different folds. So 10 folds,
  17446. 13:43:15five folds, whatever, and compute the um
  17447. 13:43:19try to find the best alpha that way.
  17448. 13:43:23Okay.
  17449. 13:43:27All right.
  17450. 13:43:29Any questions on this so far from last
  17451. 13:43:32time from reviewing that a little bit?
  17452. 13:43:35Hopefully that uh hopefully that is
  17453. 13:43:38jogging your memory a little bit on
  17454. 13:43:40ridge and lasso. Um you know where we're
  17455. 13:43:43going to pick it up today is to finish
  17456. 13:43:45out this lesson with one more model
  17457. 13:43:49which is going to be a combination of
  17458. 13:43:51ridge and lasso. So you can actually
  17459. 13:43:54combine them together
  17460. 13:43:56um in a linear fashion those penalties.
  17461. 13:44:00So you can actually have both penalties,
  17462. 13:44:02the absolute value and the square. And
  17463. 13:44:04when you have both penalties um that's a
  17464. 13:44:08special model called the elastic net uh
  17465. 13:44:11regression or elastic net model. Um so
  17466. 13:44:14this is a combination of lasso and ridge
  17467. 13:44:18together. So you have lasso, you have
  17468. 13:44:20ridge and then you have elastic net
  17469. 13:44:21which combines both of those penalties.
  17470. 13:44:24Um let me show you the equation.
  17471. 13:44:28So here is the uh so here is the the
  17472. 13:44:33model. This is the same that we've
  17473. 13:44:35always had. This is our usual u model
  17474. 13:44:39fitting for linear. This is a basic
  17475. 13:44:42linear regression um loss function or
  17476. 13:44:44objective function that we're trying to
  17477. 13:44:46minimize to find the betas. Notice how
  17478. 13:44:48we have both of our penalties though
  17479. 13:44:50this time. So instead of just having one
  17480. 13:44:52of the penalties, we actually have both.
  17481. 13:44:54So we have the lasso penalty
  17482. 13:44:57and then we have the ridge penalty here.
  17483. 13:45:00So we actually use both of them and um
  17484. 13:45:04try to find a balance of minimizing
  17485. 13:45:06those two uh those two penalties.
  17486. 13:45:11Okay. And notice how they instead of
  17487. 13:45:13just a single alpha, we kind of have a
  17488. 13:45:14balance on both of them.
  17489. 13:45:18So, we can actually weight the lasso one
  17490. 13:45:20more. We can weight the ridge one more.
  17491. 13:45:23We can weight them the same. Uh we can
  17492. 13:45:27um change that around as much as we
  17493. 13:45:28want. So, they have two different
  17494. 13:45:29weights there um that they could be.
  17495. 13:45:34Um now what happens in reality is uh
  17496. 13:45:39we're going to see this in the model is
  17497. 13:45:41that um usually what happens is these
  17498. 13:45:44get combined into a fraction. So there's
  17499. 13:45:47usually a ratio of lambda 1 to lambda 2
  17500. 13:45:51and this is known as the um this is
  17501. 13:45:54sometimes known as the L1 ratio
  17502. 13:45:58and this is a this is a a parameter
  17503. 13:46:00inside the model that we'll be able to
  17504. 13:46:02set um along with alpha. So we'll be
  17505. 13:46:05able to set an alpha and then this
  17506. 13:46:07ratio. Um the idea is is that um the
  17507. 13:46:12ratio will uh allow us to control which
  17508. 13:46:16one is more dominant. So if this number
  17509. 13:46:19is bigger the um this lasso penalty will
  17510. 13:46:23will be weighted more. If this ratio is
  17511. 13:46:26smaller if it's less than one for
  17512. 13:46:28example that means that the um ridge
  17513. 13:46:31regression is more uh dominant. Um but
  17514. 13:46:36the so we'll have this we'll have really
  17515. 13:46:38this and this at our disposal and alpha
  17516. 13:46:43is um
  17517. 13:46:46alpha is kind of like a a you can think
  17518. 13:46:48of it as a scale that is um so lambda 1
  17519. 13:46:53kind of like lambda 1 plus lambda 2 um
  17520. 13:46:57combined to equal alpha.
  17521. 13:47:00So it's like our total level of penalty
  17522. 13:47:03um our total level of penalty and we can
  17523. 13:47:06set that equal to one. We can set it
  17524. 13:47:08equal to whatever we want. Um and so
  17525. 13:47:11these will be in this ratio and there'll
  17526. 13:47:13be a total level of penalty that we can
  17527. 13:47:15apply. So the model will actually use
  17528. 13:47:18these two parameters when we when we do
  17529. 13:47:20it. But that's how they're that's how
  17530. 13:47:22they're all related.
  17531. 13:47:25Okay. So ridge uses both penalties.
  17532. 13:47:28That's the only difference between lasso
  17533. 13:47:30or sorry elastic net uses both
  17534. 13:47:32penalties. Um so one thing I want you to
  17535. 13:47:35notice is that uh if we um if we want we
  17536. 13:47:41could set this L1 ratio all the way to
  17537. 13:47:43zero
  17538. 13:47:45um which uh if we do that um the only
  17539. 13:47:50way this L1 ratio could be zero would be
  17540. 13:47:52if lambda 1 is zero. So it would just
  17541. 13:47:54revert back to ridge regression. So it
  17542. 13:47:56complet if if this is zero this will
  17543. 13:47:59wipe out this term and we'll be back to
  17544. 13:48:00ridge if the L1 ratio is zero.
  17545. 13:48:05Okay.
  17546. 13:48:10All right. So we have a elastic net
  17547. 13:48:13model. Um now it's used the exact same
  17548. 13:48:17way as we did the other models in the
  17549. 13:48:19code. So we have elastic net um uh from
  17550. 13:48:23the scikitlearn linear model family just
  17551. 13:48:26exactly where we had linear regression
  17552. 13:48:29lasso ridge all of those came from this
  17553. 13:48:32linear model um elastic net also comes
  17554. 13:48:35from there and then the cross validation
  17555. 13:48:36version also comes from there um so
  17556. 13:48:41let's see so when we build our model
  17557. 13:48:43it's going to be um very very simple
  17558. 13:48:46easy stuff because it's the same code
  17559. 13:48:48that we always have um we just use the
  17560. 13:48:51elastic net. We set an alpha alpha
  17561. 13:48:54equals 1 is pretty standard um just like
  17562. 13:48:56it is in in the last one ridge that's
  17563. 13:48:58industry standard is one and then an L
  17564. 13:49:02L1 ratio of.5
  17565. 13:49:04that's pretty standard as well. What the
  17566. 13:49:05L1 ratio.5 is is kind of a um
  17567. 13:49:11uh kind of a that means that the lambda
  17568. 13:49:141 to lambda 2 ratio is 1/2. Um, so
  17569. 13:49:18that's that's a pretty standard uh ratio
  17570. 13:49:20as well, but again, we could set this
  17571. 13:49:23equal to one and they'd be kind of
  17572. 13:49:25equally weighted. Um, L1 ratio of a half
  17573. 13:49:28means that the uh ridge regard the the
  17574. 13:49:32ridge penalty is a little bit more
  17575. 13:49:34weighted uh in that in that situation.
  17576. 13:49:39Okay.
  17577. 13:49:40So uh once we have this model um we can
  17578. 13:49:44do ffit and we can run that on our
  17579. 13:49:46training data and we can um get we can
  17580. 13:49:50figure out what our parameters are like
  17581. 13:49:51our coefficients and our intercepts. Our
  17582. 13:49:53model will have that but more
  17583. 13:49:55importantly we can use our model to
  17584. 13:49:56predict right so we can predict on the
  17585. 13:49:58test set. Um let me go back and load our
  17586. 13:50:02data and actually run this.
  17587. 13:50:06So, we're going to be using the same
  17588. 13:50:08data that we did for
  17589. 13:50:10uh lasso,
  17590. 13:50:14which is the I'm scrolling back up so I
  17591. 13:50:16can load it. It's the baseball data
  17592. 13:50:18here.
  17593. 13:50:22Um,
  17594. 13:50:25just run it from there.
  17595. 13:50:27It's this hitters.csv. So, hopefully you
  17596. 13:50:30have that one.
  17597. 13:50:42Let me load this.
  17598. 13:50:50Okay, so we loaded that and then that
  17599. 13:50:52should load.
  17600. 13:50:54Drop that unnamed column.
  17601. 13:51:01We will get our dummies
  17602. 13:51:10and then concatenate those split
  17603. 13:51:15scale. I'm just rerunning things. I'm
  17604. 13:51:18rerunning things so we can see our model
  17605. 13:51:19one more time.
  17606. 13:51:22So rerun that. Take a look at that. That
  17607. 13:51:23looks good. and then
  17608. 13:51:27fill in the NLES on the on those.
  17609. 13:51:30Okay. So, we should be able to run our
  17610. 13:51:34uh elastic net now.
  17611. 13:51:43Okay. So, let's import that and then
  17612. 13:51:46let's build our model. So, there we go.
  17613. 13:51:47We build our model and the intercept is
  17614. 13:51:51that. Now, of course, we can look at our
  17615. 13:51:53coefficients. Let's look at that.
  17616. 13:51:59Look at our coefficients. So remember
  17617. 13:52:01the coefficients are the uh betas. These
  17618. 13:52:03are our betas that are in our model. Um
  17619. 13:52:06so we can take a look at those. Now um
  17620. 13:52:08they're it's somewhere in between. It's
  17621. 13:52:11not a full lasso where we're going to
  17622. 13:52:12see some of these be zero. It's not a
  17623. 13:52:14full ridge. Um so the coefficients we
  17624. 13:52:17get are different. They're somewhere in
  17625. 13:52:19between there. Those two models that
  17626. 13:52:21we've already built. So not quite the
  17627. 13:52:23same um somewhere in between there.
  17628. 13:52:30Um and then we can use our model to make
  17629. 13:52:32predictions and and compute the MSE
  17630. 13:52:35uh or the RMSSE I should say as well. So
  17631. 13:52:38we can take the mean squared error, pass
  17632. 13:52:39that into the square root and compute
  17633. 13:52:41the RMSSE. So still pretty bad. Um this
  17634. 13:52:44is right around that 300 range of what
  17635. 13:52:46we've gotten for our other RMSSE. So,
  17636. 13:52:48it's not like elastic net is any better
  17637. 13:52:50than those other like linear or lasso or
  17638. 13:52:53ridge. And that's not surprising because
  17639. 13:52:56it's just adding those extra penalties.
  17640. 13:52:58We don't expect it to magically get
  17641. 13:53:00better. It's actually a more complex
  17642. 13:53:02um when we add when we add those in,
  17643. 13:53:05we're actually reducing it and making it
  17644. 13:53:07simpler. And we need something more
  17645. 13:53:09complex, I should say. So, we're making
  17646. 13:53:11it simpler um by by making penalizing
  17647. 13:53:16our weights a little bit more. And so,
  17648. 13:53:18it's still not a good fit. That's not
  17649. 13:53:21really surprising, right? It's still not
  17650. 13:53:23really a great fit.
  17651. 13:53:25And we can we can even double check
  17652. 13:53:27that. We know our RMSSE is pretty bad.
  17653. 13:53:30Um but we can double check it with this
  17654. 13:53:31R2 score. And it's, you know, still not
  17655. 13:53:34good. Remember, a one would be really
  17656. 13:53:36good. Um that'd be like a perfect linear
  17657. 13:53:38model. This is um still pretty bad.
  17658. 13:53:46Okay, so as we said, the alpha controls
  17659. 13:53:48the overall strength. Um so the higher
  17660. 13:53:51the alpha, the more overall penalty
  17661. 13:53:54we're supplying, which makes the model
  17662. 13:53:57simpler. Um uh but the L1 um ratio
  17663. 13:54:02determines the mix or that ratio of the
  17664. 13:54:06lambdas, the lasso to the ridge. Um if
  17665. 13:54:09you have it be um exactly zero, you you
  17666. 13:54:14revert all the way back to um if you if
  17667. 13:54:18you put it at zero, you revert all the
  17668. 13:54:20way back to ridge. One would be all the
  17669. 13:54:22way to pure lassos. Somewhere in
  17670. 13:54:23between, like one half is is good.
  17671. 13:54:35Okay,
  17672. 13:54:37so this is another example of trying out
  17673. 13:54:40different values of alpha in the CV to
  17674. 13:54:42see which one works. Now again, I'm
  17675. 13:54:45going to show us in a minute a
  17676. 13:54:46systematic way to do this, but this is
  17677. 13:54:49just trying out um different alphas that
  17678. 13:54:51we set up in this uh in this um
  17679. 13:54:56uh range. So we have different uh values
  17680. 13:54:59between minus2 and two um
  17681. 13:55:02logarithmically.
  17682. 13:55:03Um so these are uh logarithm values that
  17683. 13:55:07are between this between minus2 and two
  17684. 13:55:09and we choose a 100 different alphas and
  17685. 13:55:12then we choose a 100 different um L1
  17686. 13:55:14ratios between 0.01 and one and we run
  17687. 13:55:18that we run this um cross validation
  17688. 13:55:20with 10folds. So this is quite a bit.
  17689. 13:55:22So, we're doing 10 folds and we're
  17690. 13:55:25trying out a hundred different um
  17691. 13:55:27options. Uh every time we do an option,
  17692. 13:55:30we're trying out 10 folds to evaluate
  17693. 13:55:31it. So, it's going to take a minute to
  17694. 13:55:34run.
  17695. 13:55:46It's still running here. But again, what
  17696. 13:55:49this is doing is trying out different
  17697. 13:55:50alphas and it's it's going to do a cross
  17698. 13:55:54validation. And you guys remember the
  17699. 13:55:56t-fold cross validation is where we take
  17700. 13:55:58our data and we divide it into 10 folds
  17701. 13:56:03and then we um train on nine of those
  17702. 13:56:06and then test on the remaining fold and
  17703. 13:56:08then we rotate all the folds 10 times.
  17704. 13:56:11and that we average those mean squared
  17705. 13:56:14error metrics together um against those
  17706. 13:56:1810 different uh fold options to generate
  17707. 13:56:22a basically like an average performance
  17708. 13:56:25for that value of alpha. And we're doing
  17709. 13:56:27that a 100 times for all these different
  17710. 13:56:29100 alphas that there are and 100
  17711. 13:56:32different L1 ratios that we're trying
  17712. 13:56:33with them.
  17713. 13:56:39So that's quite a bit of processing but
  17714. 13:56:42uh it did finish.
  17715. 13:56:46So we can see what our best alpha is and
  17716. 13:56:48our best one ratio. So we get the best
  17717. 13:56:50alpha is this best one ratio is this. Um
  17718. 13:56:54and therefore we can uh build a model
  17719. 13:56:57with those with just these two guys as
  17720. 13:56:59the alpha and the L1 and um see how that
  17721. 13:57:04performs.
  17722. 13:57:06We build that model and then we predict
  17723. 13:57:08on the test set and we generate the
  17724. 13:57:10RMSSE. It's just a little bit better.
  17725. 13:57:12It's still not It's just a little bit
  17726. 13:57:14better, but it's still not good, right?
  17727. 13:57:16It's still 338. It is just way too big.
  17728. 13:57:20Remember, this is RMSSE, so it's in the
  17729. 13:57:23units of our uh target variable. So,
  17730. 13:57:27it's in the units of if we go back to
  17731. 13:57:30our data, actually, I could just print
  17732. 13:57:32it out here.
  17733. 13:57:34um this RMSSE.
  17734. 13:57:38If I just do this, we could take a look
  17735. 13:57:40at um DF
  17736. 13:57:42or I could look at Y test
  17737. 13:57:47and you can see some of these values.
  17738. 13:57:48These are these salary values in the
  17739. 13:57:50hundreds, right? Some of them are in the
  17740. 13:57:51thousands. Um but an error of like 338
  17741. 13:57:56is just too big. That's a really big
  17742. 13:57:58error. That means we would be off by an
  17743. 13:57:59average of 300 when our our values if we
  17744. 13:58:02just do the mean
  17745. 13:58:06um
  17746. 13:58:08is only 550 as on average is 550 but we
  17747. 13:58:12have this amount of error on average um
  17748. 13:58:15so that's just a way too big of a
  17749. 13:58:17proportion of error right it's not a
  17750. 13:58:19very good model and again we can verify
  17751. 13:58:21that by looking at this R2 for.
  17752. 13:58:31So if we go down here,
  17753. 13:58:39still not very good.
  17754. 13:58:43Here's some of our coefficients. So
  17755. 13:58:45remember, you can always take your
  17756. 13:58:46coefficients and line them up to your
  17757. 13:58:48your data columns. Uh so that you can
  17758. 13:58:51get a sense of what coefficient belongs
  17759. 13:58:53with what feature. So that's all we're
  17760. 13:58:56doing here is just creating a series
  17761. 13:58:57where those coefficients instead of just
  17762. 13:58:59printing out the coefficients, we're
  17763. 13:59:00actually lining them up to the columns.
  17764. 13:59:02So this tells us um remember the larger
  17765. 13:59:05it is the more influence it kind of has
  17766. 13:59:07on the on the final result. Um either
  17767. 13:59:10way, so like this has a big negative
  17768. 13:59:12influence. um this has a large positive
  17769. 13:59:15influence.
  17770. 13:59:22Okay, let me pause there. Any questions
  17771. 13:59:25about the
  17772. 13:59:27elastic net model?
  17773. 13:59:31This is a really this model is a really
  17774. 13:59:33good one to use when you are building a
  17775. 13:59:36linear regression and it's performing
  17776. 13:59:37well but it's overfitting. This is a
  17777. 13:59:40really good one to use because you can
  17778. 13:59:41balance
  17779. 13:59:43lasso and ridge you can get the best of
  17780. 13:59:45both worlds. So the the main strategy is
  17781. 13:59:48if you are using a linear regression and
  17782. 13:59:51you see overfitting
  17783. 13:59:53um meaning that it's performing decently
  17784. 13:59:56so on the training set
  17785. 13:59:59it's performing okay but then on the
  17786. 14:00:01test set like you know it's it's not
  17787. 14:00:04underfitting. it's performing pretty
  17788. 14:00:05well on the training set, but then on
  17789. 14:00:07the test set it's um performance is much
  17790. 14:00:11worse. That's overfitting. If you're
  17791. 14:00:14overfitting, then this is a great model
  17792. 14:00:15to use because we can try basically by
  17793. 14:00:18by rotating through different alphas and
  17794. 14:00:20different L1 ratios, we can try out
  17795. 14:00:23different strengths of penalty and
  17796. 14:00:26different variations on lasso and ridge
  17797. 14:00:28together. This is a really good model to
  17798. 14:00:30to use for those overfitting cases where
  17799. 14:00:33linear regression is doing decently. Um,
  17800. 14:00:37but it's overfitting,
  17801. 14:00:39right? So far, we haven't ran that case
  17802. 14:00:42because so far, no matter what model
  17803. 14:00:44we've used, it's always underfit. So,
  17804. 14:00:48anytime we have those underfitting
  17805. 14:00:50cases, it signals that we should likely
  17806. 14:00:53just use a more complex model. And we
  17807. 14:00:56haven't learned about those yet.
  17808. 14:00:58um we will coming up in lesson four, but
  17809. 14:01:02um that's for this data. That's ultim
  17810. 14:01:05ultimately what we'd want to do is
  17811. 14:01:06probably use a more advanced model
  17812. 14:01:08because it's underfitting um just using
  17813. 14:01:10a linear regression and and then using
  17814. 14:01:12the the overfitting variations of linear
  17815. 14:01:14regression like lasso ridge and elastic
  17816. 14:01:16net.
  17817. 14:01:22Okay.
  17818. 14:01:26Any questions on this on elastic then
  17819. 14:01:35the TV? Yeah, we Yeah, I think I have
  17820. 14:01:37it. I can share it with you.
  17821. 14:01:44I said that and now I can't find it. I
  17822. 14:01:46thought I had it.
  17823. 14:01:55I don't have it. I thought I had it, but
  17824. 14:01:57I don't.
  17825. 14:01:59If anyone does have that one.
  17826. 14:02:05Yeah, I'll look one more time. I thought
  17827. 14:02:07I had that one.
  17828. 14:02:11Um,
  17829. 14:02:17yeah, it's not in there. I had it. Let
  17830. 14:02:19me see.
  17831. 14:02:32Yeah, I don't have it either. I thought
  17832. 14:02:33I had it in here.
  17833. 14:02:42Yeah, I don't have that one.
  17834. 14:02:46I don't have that one. I'll have to find
  17835. 14:02:48it. Uh I have this marketing data. I
  17836. 14:02:50don't think this is the same one.
  17837. 14:02:54I have this marketing data. I don't
  17838. 14:02:55think that's the right one, but you can
  17839. 14:02:56take a look at it.
  17840. 14:03:01No, we're using So, for this example,
  17841. 14:03:03we're using the same hitters data set
  17842. 14:03:05that we used earlier for lasso.
  17843. 14:03:08No, that's an earlier one.
  17844. 14:03:14That's from the uh very beginning of the
  17845. 14:03:18notebook. So that's the that's from this
  17846. 14:03:21one.
  17847. 14:03:25Oh, this Oh, this is where it is. Sorry.
  17848. 14:03:27This is where it is. You can find it
  17849. 14:03:29here.
  17850. 14:03:34That's right. It was from a URL.
  17851. 14:03:40It was used in the very beginning of the
  17852. 14:03:41notebook.
  17853. 14:03:43And we did we did this.
  17854. 14:03:47Okay.
  17855. 14:03:49That's right. That's why I didn't have
  17856. 14:03:50it downloaded.
  17857. 14:03:55Okay.
  17858. 14:03:59All right. Any other questions on the
  17859. 14:04:01elastic before I move I'm going to move
  17860. 14:04:03on to uh finding those a systematic way
  17861. 14:04:07to find the best hyperparameters.
  17862. 14:04:10Um, I'm going to show you a couple
  17863. 14:04:11strategies to doing that. Um, so far
  17864. 14:04:14we've just ran CV with some random
  17865. 14:04:16choices. Um, I'm going to show you a
  17866. 14:04:18better, more systematic approach. That's
  17867. 14:04:20kind of the industry standard for doing
  17868. 14:04:22tuning. Um, so I'm going to I'm going to
  17869. 14:04:25show you that next, but any questions on
  17870. 14:04:26the elastic net?
  17871. 14:04:33Okay. And again like you know
  17872. 14:04:35scikitlearn makes it really easy for you
  17873. 14:04:37guys
  17874. 14:04:39because it just behaves the same way as
  17875. 14:04:41any other model. You use the object and
  17876. 14:04:44then you do ffit and predict right? So
  17877. 14:04:46the ffit is going to train it um and the
  17878. 14:04:50predict is going to allow you to use
  17879. 14:04:52that model to predict. It's it's super
  17880. 14:04:54easy that way. Every scikitlearn model
  17881. 14:04:56is like that dofitit and predict. So it
  17882. 14:04:59provides a really simple way to use
  17883. 14:05:01basically every model.
  17884. 14:05:07Okay,
  17885. 14:05:09let's talk about let's finish up this
  17886. 14:05:11lesson with a couple things. Um, one of
  17887. 14:05:14those things is going to be
  17888. 14:05:15hyperparameter tuning. So what is this?
  17889. 14:05:19The hyperparameter tuning is a
  17890. 14:05:21systematic way to find the best
  17891. 14:05:24parameters in a machine learning model.
  17892. 14:05:28So a lot of machine learning models have
  17893. 14:05:30what are called hyperparameters.
  17894. 14:05:33These are not the betas that we learn
  17895. 14:05:35during the training that's learned from
  17896. 14:05:37the data. These are settings that we set
  17897. 14:05:40ahead of time like the alpha. That's a
  17898. 14:05:43perfect example like alpha L1 ratio in
  17899. 14:05:45in the elastic net. We set those up
  17900. 14:05:48ahead of time and depending on what we
  17901. 14:05:50pick for those we get different
  17902. 14:05:51performance, right? And so what we
  17903. 14:05:54really need is a systematic way to find
  17904. 14:05:57the best settings for those
  17905. 14:06:00hyperparameters as we are training our
  17906. 14:06:02models. Um the the the main like idea
  17907. 14:06:08behind this process though is going to
  17908. 14:06:10be to systematically try out different
  17909. 14:06:14combinations as many as we want to try.
  17910. 14:06:17And so we're we're basically going to
  17911. 14:06:19have a strategy for tuning that is going
  17912. 14:06:22to exhaust all the combinations of those
  17913. 14:06:26hyperparameters that we want to try
  17914. 14:06:28until we find the one that performs the
  17915. 14:06:31best. Um and and that strategy is known
  17916. 14:06:34as grid search. Um and essentially what
  17917. 14:06:39it does is it sets up a grid um where
  17918. 14:06:42which is basically like a matrix to say
  17919. 14:06:45okay which parameters do you want to
  17920. 14:06:47try? I want to try um alpha and I want
  17921. 14:06:50to try L1 ratio
  17922. 14:06:53um L1 ratio like let's say I want to try
  17923. 14:06:57these two. So we set these up in a grid
  17924. 14:06:59where we say, "Okay, I want to try this
  17925. 14:07:01value. I want to try this value. I want
  17926. 14:07:02to try this value. This one, this one,
  17927. 14:07:04this one, and on and as many as we want
  17928. 14:07:06to try." So we could set up set those up
  17929. 14:07:09systematically like a linear um a
  17930. 14:07:12linearly spaced like I want to try every
  17931. 14:07:14alpha between between 0 and 10 spaced by
  17932. 14:07:18one um whatever. You know, we can set up
  17933. 14:07:21different ranges of those, but that's
  17934. 14:07:23going to be in this grid. And then the
  17935. 14:07:25L1 ratio, same thing. We can try out
  17936. 14:07:27different values of these that we want
  17937. 14:07:28to try. Let's say there's many of those.
  17938. 14:07:32Um maybe every um tenth between 0 to one
  17939. 14:07:36I want to try out. Um so you set up your
  17940. 14:07:40parameters and you can set up as many as
  17941. 14:07:41you want in the grid. And then
  17942. 14:07:43essentially what you're going to do to
  17943. 14:07:45do grid search is you're going to work
  17944. 14:07:47your way through every combination of
  17945. 14:07:49those. So you're going to try out this
  17946. 14:07:50combo. You're going to try out this
  17947. 14:07:53combo. You're going to try out this
  17948. 14:07:54combo.
  17949. 14:07:56this combo. So the first value of alpha
  17950. 14:08:00with every possible L1 ratio, then go to
  17951. 14:08:03the next, try out the next value of
  17952. 14:08:05alpha with every L1 ratio, and on and on
  17953. 14:08:08and on. So we're going to try
  17954. 14:08:11all combos
  17955. 14:08:14in the grid.
  17956. 14:08:16We're going to try all combos and we're
  17957. 14:08:18going to find the lowest MSE
  17958. 14:08:22combination. find lowest
  17959. 14:08:25MSE
  17960. 14:08:27combo.
  17961. 14:08:28So whatever leads to the best model is
  17962. 14:08:31going to be the um parameters that are
  17963. 14:08:34that are deemed to be the best. And the
  17964. 14:08:36idea is once we have found those we know
  17965. 14:08:40that we can use we can go ahead and
  17966. 14:08:42train a model with those best alpha and
  17967. 14:08:44len ratio and on and on and on.
  17968. 14:08:54Yeah, when you get an So this goes back
  17969. 14:08:56to the error. Remember that for a
  17970. 14:08:59regression,
  17971. 14:09:01the error is this measurement of how far
  17972. 14:09:04off we are, right? So if we have a bunch
  17973. 14:09:06of points and we draw we fit a line
  17974. 14:09:08through there, the the MSE is measuring
  17975. 14:09:12this distance, right? So what do you
  17976. 14:09:14think is a good distance? Like if our
  17977. 14:09:16model is perfect,
  17978. 14:09:19what's the best distance from our
  17979. 14:09:21predictions to the actual points? Zero.
  17980. 14:09:25Yes. So the lower the better. The lower
  17981. 14:09:29the better. Um so for an R RMSSE, the
  17982. 14:09:32lower the closer to zero the better.
  17983. 14:09:35However, the RMSSE can be it's its units
  17984. 14:09:39are interpreted in the units of our
  17985. 14:09:41target.
  17986. 14:09:43So what is deemed to be good is relative
  17987. 14:09:46to our target. Like let's say our target
  17988. 14:09:48is in the thousands. Like it averages in
  17989. 14:09:51the thousands. If we produce an MSE of
  17990. 14:09:5450 or sorry an RMSSE of 50, that's
  17991. 14:09:59pretty good, right? Because our units
  17992. 14:10:01are in the thousands
  17993. 14:10:03and we're only on average we are off by
  17994. 14:10:0750 units,
  17995. 14:10:09right? Our distance away is about 50
  17996. 14:10:11units. That's pretty good. So the RMSSE
  17997. 14:10:15is relative to your target variable.
  17998. 14:10:18Does that make sense? Yeah. It depends
  17999. 14:10:20on the target. It depends on what you're
  18000. 14:10:22trying to predict.
  18001. 14:10:25So that's why we got RMSSE that were in
  18002. 14:10:27the 300s for those hitters, but the
  18003. 14:10:29average was the average of the target
  18004. 14:10:31was in the 500s. So that's a really bad
  18005. 14:10:35proportion of error relative to the
  18006. 14:10:37average target value. Right? If our
  18007. 14:10:41RMSSE was 300,
  18008. 14:10:43but the target was sitting in the 500s,
  18009. 14:10:47that's just too much error. Way too much
  18010. 14:10:50error, right? That's just too big of a
  18011. 14:10:52value. Um, our predictions are just off
  18012. 14:10:56way too much
  18013. 14:10:58in terms of that distance. So, this
  18014. 14:10:59would be this was a bad model. It was
  18015. 14:11:02underfit.
  18016. 14:11:04We know that from the the RMSSE. So
  18017. 14:11:07yeah, the RMSSE closer to zero, no
  18018. 14:11:09matter what is good,
  18019. 14:11:12zero is being perfect. Um, but it to
  18020. 14:11:16know what's good, you need to know what
  18021. 14:11:18your target is on average and then think
  18022. 14:11:20of this as kind of a ratio to that
  18023. 14:11:23average target. I think that's the best
  18024. 14:11:26way to think about it.
  18025. 14:11:37Okay, so going back to this grid idea is
  18026. 14:11:42so the grid is just basically laying out
  18027. 14:11:44all possible parameter combinations and
  18028. 14:11:48trying them all out by fitting and
  18029. 14:11:50predicting until and generating an a
  18030. 14:11:53metric like an MSE
  18031. 14:11:55until we find the one with the lowest
  18032. 14:11:58MSE. So find the lowest MSE combination
  18033. 14:12:02and that will be the best
  18034. 14:12:04that will be the best combo and then if
  18035. 14:12:07we once we know that best combo we can
  18036. 14:12:09use that we can use that alpha we can
  18037. 14:12:11use that L1 ratio and use that model
  18038. 14:12:14going forward we can we can use those
  18039. 14:12:16parameters in our model so this strategy
  18040. 14:12:20it has a name it's known as grid search
  18041. 14:12:24so it is a hyperparameter tuning process
  18042. 14:12:26that tries out all combinations S.
  18043. 14:12:30So what's the what's the U benefit to
  18044. 14:12:34this is that we get to test out a lot of
  18045. 14:12:37different combo combos of those
  18046. 14:12:38parameters like the alpha and L1. So we
  18047. 14:12:41can be confident what the best model is,
  18048. 14:12:43right? So we can pick the alpha and L1
  18049. 14:12:46perfectly because we're trying out a
  18050. 14:12:47bunch of different combinations on the
  18051. 14:12:49data to see which one's the best. What's
  18052. 14:12:52the downside?
  18053. 14:12:54It's expensive, right? It's an
  18054. 14:12:56exhaustive search. So if you have many
  18055. 14:13:00different parameters and you're trying
  18056. 14:13:03out many different combinations, it can
  18057. 14:13:06get exponentially
  18058. 14:13:08expensive
  18059. 14:13:09to perform this search. Okay, so grid
  18060. 14:13:12search is great except for the fact that
  18061. 14:13:15it can be expensive if you have many
  18062. 14:13:17parameters with with very wide ranges
  18063. 14:13:20that you're searching over because that
  18064. 14:13:21that's a lot of combinations you have to
  18065. 14:13:23test, right? And especially if you have
  18066. 14:13:26a lot of data, that's going to be
  18067. 14:13:28expensive
  18068. 14:13:30um to do.
  18069. 14:13:33So, uh we're going to practice doing
  18070. 14:13:35grid search, but that is that's the pro
  18071. 14:13:37and con. The pro is that we get to try
  18072. 14:13:38out all these combinations and see which
  18073. 14:13:40one's the best. The downside is it can
  18074. 14:13:43be expensive to do that if you have a
  18075. 14:13:44lot of parameters um that you want to
  18076. 14:13:47tune for your model um and you have very
  18077. 14:13:52uh many different choices that you're
  18078. 14:13:53trying to evaluate for those and it just
  18079. 14:13:56creates a really big um collection of
  18080. 14:13:59combinations that you have to try out,
  18081. 14:14:02right? Um that's the only downside to
  18082. 14:14:05grid search.
  18083. 14:14:08Now on the opposite end of the spectrum
  18084. 14:14:10of that is a randomized search or random
  18085. 14:14:13search and this will basically just um
  18086. 14:14:17do a sampling of those parameters from
  18087. 14:14:23um kind of fixed uh specified
  18088. 14:14:25distribution. So essentially what you do
  18089. 14:14:28is similarly you define your range. So
  18090. 14:14:31you say I want to look at alphas um
  18091. 14:14:34between zero sorry between let's say
  18092. 14:14:38yeah 0 to 10. I want to look at a bunch
  18093. 14:14:40of different alphas. Um, and I want to
  18094. 14:14:42look at a bunch of different L1 ratios
  18095. 14:14:45that are between 0ero to one.
  18096. 14:14:490 to one. And um, what we do is we say,
  18097. 14:14:53okay, I'm going to restrict only testing
  18098. 14:14:5720, 30, 40 times. I'm not going to do
  18099. 14:14:59all possible combinations. I'm just
  18100. 14:15:02going to randomly sample something in
  18101. 14:15:04this range and randomly sample something
  18102. 14:15:07in this range. And so and I'm going to
  18103. 14:15:10perform that experiment a fixed number
  18104. 14:15:12of times. So let's say I set the uh
  18105. 14:15:16sampling where I'm only going to do um
  18106. 14:15:1920 evaluations.
  18107. 14:15:21And so 20 times we're going to pick a
  18108. 14:15:24combo randomly. So I'm going to pick an
  18109. 14:15:28alpha and I'm going to pick an L1 ratio.
  18110. 14:15:33L1 ratio.
  18111. 14:15:36And um we are we are just going to uh
  18112. 14:15:40sample those randomly from this range.
  18113. 14:15:44Um and we're going to use those and test
  18114. 14:15:47those out and then it's but otherwise
  18115. 14:15:49it's the same as grid search. Whatever
  18116. 14:15:50is the lowest MSE
  18117. 14:15:53um so whatever is the lowest MSE is the
  18118. 14:15:56best.
  18119. 14:15:57So we evaluate those. We sample we train
  18120. 14:16:00the model. Evaluate it. Whatever is the
  18121. 14:16:03lowest MSE
  18122. 14:16:05is the best is the best combo. Now,
  18123. 14:16:09what's the benefit to this is it's a
  18124. 14:16:12much more controlled experiment in the
  18125. 14:16:16sense that we um aren't going to iterate
  18126. 14:16:18through every possible combination in
  18127. 14:16:20the grid. We're we basically set up a
  18128. 14:16:23fixed number of times we're going to try
  18129. 14:16:24out stuff.
  18130. 14:16:26The risk to doing this is that you're
  18131. 14:16:28not you're not exploring all
  18132. 14:16:31combinations, right? Because you're
  18133. 14:16:33randomly sampling, you may get unlucky
  18134. 14:16:36and you may not stumble into the best.
  18135. 14:16:39You you can make um samples and figure
  18136. 14:16:42out what's the best amongst your
  18137. 14:16:43samples, but you may not be covering all
  18138. 14:16:46the combinations. Does that make sense?
  18139. 14:16:48The grid search is going to try every
  18140. 14:16:50combo. The random search is going to
  18141. 14:16:53randomly sample those combos.
  18142. 14:16:56So, it's not going to try every single
  18143. 14:16:58one. It's going to try a limited number,
  18144. 14:17:00however many you set up. Now, if you set
  18145. 14:17:03that number really, really, really high.
  18146. 14:17:06Now, you're starting to approach a grid
  18147. 14:17:07search because now you're sampling so
  18148. 14:17:10many of those combos that you basically
  18149. 14:17:12are trying them all at that point,
  18150. 14:17:16right? Um,
  18151. 14:17:20so, so that's the way the random search.
  18152. 14:17:22So by the way, both of these use cross
  18153. 14:17:24validation in the sense that when you
  18154. 14:17:27evaluate accommodation, you're actually
  18155. 14:17:29doing it with cross validation. So when
  18156. 14:17:31you do an evaluation, you're going to do
  18157. 14:17:34probably 10 or five folds where you
  18158. 14:17:36split your data and then you test it on
  18159. 14:17:39the rest of the folds and evaluate or
  18160. 14:17:41train it on the rest of the folds,
  18161. 14:17:42evaluate it on one of them and generate
  18162. 14:17:44an average MSE to get your evaluation.
  18163. 14:17:49So every evaluation is using cross
  18164. 14:17:51validation.
  18165. 14:17:52That's why that's and hopefully you can
  18166. 14:17:54see why this would be so expensive for a
  18167. 14:17:56really big grid, right? Because you're
  18168. 14:17:58trying out many different combinations
  18169. 14:18:02and every combination is going to do a
  18170. 14:18:04cross validation procedure. So it's
  18171. 14:18:07going to train 10 times and test against
  18172. 14:18:1010 different folds and average those
  18173. 14:18:12together. it's going to be a pretty
  18174. 14:18:13expensive operation
  18175. 14:18:15for a really big grid, right, of of
  18176. 14:18:18parameters.
  18177. 14:18:19Um, but these are the two kind of
  18178. 14:18:21systematic approaches we have at trying
  18179. 14:18:24out different hyperparameters. Remember
  18180. 14:18:27those those things are called
  18181. 14:18:28hyperparameters. These are those choices
  18182. 14:18:31that we have before we train our model.
  18183. 14:18:34Um, those choices we have that affect
  18184. 14:18:36the performance of the model like the
  18185. 14:18:38alphas, the L1 ratios, those kind of
  18186. 14:18:40things. um we have control over what
  18187. 14:18:43they're going to be. This is a
  18188. 14:18:44systematic approach to find out what the
  18189. 14:18:46best
  18190. 14:18:48uh value of those parameters is going to
  18191. 14:18:51be on our data.
  18192. 14:18:53Right?
  18193. 14:18:57Okay. So before we practice this, we're
  18194. 14:18:59going to practice with grid search
  18195. 14:19:01first. Um
  18196. 14:19:04any questions?
  18197. 14:19:14Uh, I don't know if it has a built-in
  18198. 14:19:16That's a good question. By time limit, I
  18199. 14:19:17don't know if it has a built-in way of
  18200. 14:19:18doing it, but you could certainly set up
  18201. 14:19:20like a a a loop um to like to wrap
  18202. 14:19:25around. Do you know what I mean? Like
  18203. 14:19:26you could set up a loop where you check
  18204. 14:19:28the time if it's if if the time elapsed
  18205. 14:19:31as you're doing a search if the time
  18206. 14:19:32elapsed is greater than the the time
  18207. 14:19:35limit then you can kind of break early.
  18208. 14:19:38Um so it's not hard to implement that
  18209. 14:19:40but I don't know if it has that built
  18210. 14:19:42in. I don't think it does
  18211. 14:19:44because I don't think it really cares
  18212. 14:19:46how long every evaluation takes. It's
  18213. 14:19:48just going to exhaust all those
  18214. 14:19:50especially in a grid search.
  18215. 14:19:53But um yeah, I there's probably a way to
  18216. 14:19:56manually kind of set up a time time
  18217. 14:19:58loop.
  18218. 14:20:04So hyperparameters are um settings that
  18219. 14:20:08we have on the model itself. And a
  18220. 14:20:11really good example of this is like the
  18221. 14:20:13alpha and L1 ratio in the in the elastic
  18222. 14:20:15net. So they're not things that we um
  18223. 14:20:20learn from the data directly like the
  18224. 14:20:22betas in the model like those get
  18225. 14:20:24trained directly by doing the um least
  18226. 14:20:28squares process right um by doing that
  18227. 14:20:31gradient descent and all that
  18228. 14:20:32optimization.
  18229. 14:20:34Um so these are not learned from that.
  18230. 14:20:36They're actually set ahead of time. And
  18231. 14:20:39so what we're saying is the best way to
  18232. 14:20:42understand the effects of those is to
  18233. 14:20:44try out different combinations of those
  18234. 14:20:46until we land on the best one. Right? So
  18235. 14:20:49hyperparameters are those options we
  18236. 14:20:51have in the model like the alpha like
  18237. 14:20:54the alpha and l1 ratio in the uh elastic
  18238. 14:20:57net. Many models have hyperparameters.
  18239. 14:21:01Um we're actually going to see that in
  18240. 14:21:03in future models that we study. they
  18241. 14:21:05have options that you can set that
  18242. 14:21:07affect their performance.
  18243. 14:21:09And so this this is just a strategy to
  18244. 14:21:11evaluate those different options to see
  18245. 14:21:13which one's the best.
  18246. 14:21:25Yeah. So again, hyperparameters, those
  18247. 14:21:28are settings on the model itself um that
  18248. 14:21:32affect the performance of it.
  18249. 14:21:36And basically we have the two two
  18250. 14:21:38strategies here. We can set up an
  18251. 14:21:40exhaustive grid and search through all
  18252. 14:21:41of those until we find the lowest MSE uh
  18253. 14:21:44option or we can randomly sample
  18254. 14:21:48potential options, try them out and see
  18255. 14:21:50which one's the lowest as well. And do
  18256. 14:21:52that a fixed number of times. Um
  18257. 14:21:56sort of like a fixed number of trials
  18258. 14:21:58almost. um which has a risk of not
  18259. 14:22:02trying out every option but but
  18260. 14:22:04hopefully you try out enough that you've
  18261. 14:22:06explored the space a bit and you get
  18262. 14:22:09some quality choices there but no
  18263. 14:22:12guarantees right no guarantees you try
  18264. 14:22:14everything which is what a grid search
  18265. 14:22:15will do it will try everything
  18266. 14:22:20okay now luckily per usual scikitlearn
  18267. 14:22:26has something to manage this process for
  18268. 14:22:28us in terms of grid search. Um so in
  18269. 14:22:33that way we will not need to manage this
  18270. 14:22:36process ourselves. We can just rely on
  18271. 14:22:37scikitlearn. And so if you're doing
  18272. 14:22:40hyperparameter tuning um this is going
  18273. 14:22:42to come from the model selection module
  18274. 14:22:45inside of sklearn. So we're going to
  18275. 14:22:48import from from sklearn the model
  18276. 14:22:50selection module. We have our grid
  18277. 14:22:52search cross validation.
  18278. 14:22:55Okay, that's what the CV stands for.
  18279. 14:22:57grid search cross validation. So, this
  18280. 14:23:00is going to do that grid search
  18281. 14:23:01strategy. Um, we're going to set it up
  18282. 14:23:03with our dictionary essentially of
  18283. 14:23:06choices. So, we're going to say, hey,
  18284. 14:23:08here's the alphas I want to try. Here's
  18285. 14:23:09the L1 ratios I want to try. Um, and
  18286. 14:23:12here's my other settings like uh how
  18287. 14:23:15many folds I want to use, what my random
  18288. 14:23:17state is for the shuffling. So, we'll
  18289. 14:23:19set all that up. Um,
  18290. 14:23:22and then we'll just run the grid search.
  18291. 14:23:24And then what should come out of that is
  18292. 14:23:26the best options for our parameters from
  18293. 14:23:29the grid and then we can use those going
  18294. 14:23:31forward in the we can build a model with
  18295. 14:23:34those best options right so we're really
  18296. 14:23:37doing some evaluation here of what is
  18297. 14:23:39going to be those best alphas those best
  18298. 14:23:410 to1 ratios on our data set right and
  18299. 14:23:44the only way to really know that is to
  18300. 14:23:47evaluate them because they're not things
  18301. 14:23:48that are learned during the training
  18302. 14:23:51hopefully that makes sense right they're
  18303. 14:23:53not things that we learn directly from
  18304. 14:23:55training. There are things that we have
  18305. 14:23:57to set and then kind of evaluate and see
  18306. 14:23:59how they affect things.
  18307. 14:24:03Okay, so we have grid search CV. That's
  18308. 14:24:05going to be our primary um tool to do
  18309. 14:24:08the evaluations of the different
  18310. 14:24:10hyperparameter options.
  18311. 14:24:13Grid search CV. Um we're going to set up
  18312. 14:24:15our cross validation uh object here. Now
  18313. 14:24:19I want you to pay attention to this is
  18314. 14:24:21that um it's a slightly different
  18315. 14:24:24version than the kfold we had earlier.
  18316. 14:24:26So we've used k-fold before with a
  18317. 14:24:27certain number of folds. This would be
  18318. 14:24:2910 folds and we can set a random state
  18319. 14:24:32for the shuffling um that happens in the
  18320. 14:24:34folds.
  18321. 14:24:36But this is actually a slight different
  18322. 14:24:37variation on it where it is a repeated
  18323. 14:24:39kfold where we do three repeated trials.
  18324. 14:24:43Now why would we do that? It's to be
  18325. 14:24:46extra extra extra careful with the
  18326. 14:24:49shuffling.
  18327. 14:24:51So this what this means is we do three
  18328. 14:24:52different shuffles. So we do kfold, we
  18329. 14:24:56actually repeat it three times with
  18330. 14:24:58three different shufflings. That's all
  18331. 14:24:59that means. So the repeated kfold is
  18332. 14:25:03actually a bit beyond the just basic
  18333. 14:25:06kfold. What basic kfold will do will
  18334. 14:25:09we'll will shuffle and then do our
  18335. 14:25:12splits into 10 splits and then train on
  18336. 14:25:14nine of those. test on the other one and
  18337. 14:25:16rotate through all the splits.
  18338. 14:25:19We're actually going to do that process
  18339. 14:25:21three different times with three
  18340. 14:25:24different shuffles. So this and we're
  18341. 14:25:26going to average 30 results instead of
  18342. 14:25:29just 10. So repeated kfold is just going
  18343. 14:25:33above and beyond to do extra to repeat
  18344. 14:25:36the kfold three different times. In this
  18345. 14:25:38case only three. We could do more,
  18346. 14:25:41but um now is that necessary to do? You
  18347. 14:25:44could argue not necessarily. Um but it
  18348. 14:25:47just provides extra robustness
  18349. 14:25:50uh beyond just our single shuffle and
  18350. 14:25:53then split and then rotation of those
  18351. 14:25:55folds, right? We're doing it actually
  18352. 14:25:57three different shuffles. Um so we're
  18353. 14:26:00repeating our kfold three times uh for
  18354. 14:26:03every now is the thing is we're doing
  18355. 14:26:05that for every evaluation. So it is
  18356. 14:26:07going to be more expensive than just a
  18357. 14:26:09basic K-fold.
  18358. 14:26:25So we have three different K-fold trials
  18359. 14:26:28that we're doing essentially.
  18360. 14:26:31Okay, hopefully that makes sense. This
  18361. 14:26:33is the repeated K-fold. We haven't
  18362. 14:26:35really seen that before. we've only
  18363. 14:26:36worked with the Kfold, which would get
  18364. 14:26:38rid of this repeats option and only have
  18365. 14:26:41uh 10 splits in a random state for the
  18366. 14:26:44for the single shuffle that we do. So,
  18367. 14:26:46we can um recreate that same shuffle
  18368. 14:26:48every time. Um but now we're actually
  18369. 14:26:51going to do three random shuffles, uh
  18370. 14:26:54three different trials. So, one shuffle
  18371. 14:26:56creates the and then create the 10
  18372. 14:26:58splits, evaluate, then go back and do
  18373. 14:27:00another shuffle, another new 10 splits.
  18374. 14:27:03So, one thing that should be um clear is
  18375. 14:27:06that we get different splits every time
  18376. 14:27:09because we're going to shuffle once,
  18377. 14:27:12right? We're going to shuffle once and
  18378. 14:27:14generate our splits
  18379. 14:27:16and then we're going to shuffle again,
  18380. 14:27:18generate these splits which are going to
  18381. 14:27:19be different and then shuffle one more
  18382. 14:27:21time for for three different times,
  18383. 14:27:23right? And then get get these splits and
  18384. 14:27:26then we're going to get 10 metrics here,
  18385. 14:27:2810 metrics here, 10 metrics here, and
  18386. 14:27:30then average all of those together.
  18387. 14:27:34So, it's a bit more just going up extra
  18388. 14:27:37above and beyond for a K-fold. Okay.
  18389. 14:27:42All right. So, here comes the fun of
  18390. 14:27:44when you do grid search. Now, the grid
  18391. 14:27:48is actually just a dictionary. It's a
  18392. 14:27:50Python dictionary where you declare what
  18393. 14:27:54your parameters are going to be inside
  18394. 14:27:55the dictionary and you set up a range of
  18395. 14:27:58values that you're go or a list. It can
  18396. 14:28:02be a list. It can be a range
  18397. 14:28:04but some declaration of what you are
  18398. 14:28:07going to test and evaluate inside of
  18399. 14:28:09your grid search. So the grid is
  18400. 14:28:12initialized as an empty dictionary.
  18401. 14:28:15And then what we do is we say okay in my
  18402. 14:28:18grid I want to check different alphas.
  18403. 14:28:20So we're going to add a collection of
  18404. 14:28:22alphas in here that we're going to test.
  18405. 14:28:26So let me make a comment there. We add a
  18406. 14:28:30add a range of alphas to test. And this
  18407. 14:28:36range is a this is just like the Python
  18408. 14:28:40range. Um
  18409. 14:28:43this is just like a Python range um uh
  18410. 14:28:46operator here where this is going to be
  18411. 14:28:49uh every so it's going to be um every
  18412. 14:28:54uh value between
  18413. 14:28:57zero and one um uh steps with a step
  18414. 14:29:03size
  18415. 14:29:05of 0.1. So, it's going to try a bunch of
  18416. 14:29:09different alphas um between uh zero and
  18417. 14:29:130.1
  18418. 14:29:14sorry 0 and one stepping by 0.1. So,
  18419. 14:29:16it's going to try zero.1
  18420. 14:29:182.3 point 4.5 6 right all the way up to
  18421. 14:29:22one.
  18422. 14:29:24So, that's what this will do. And it's a
  18423. 14:29:25numpy range. So, it's just all those
  18424. 14:29:27decimals between 0 to one.
  18425. 14:29:31It you can use either that's valid.
  18426. 14:29:34Yeah, you can do you can do that to
  18427. 14:29:36create a dictionary or you can use the
  18428. 14:29:38keyword um dict. You can use either one.
  18429. 14:29:41Either one works.
  18430. 14:29:45Whatever whatever you want to use.
  18431. 14:29:46They're the same.
  18432. 14:29:49Yeah. The the reason people prefer
  18433. 14:29:52dictionary is because um sets are
  18434. 14:29:56created with the same braces.
  18435. 14:29:59So it it makes it clear what you're
  18436. 14:30:01creating as a dictionary. If you use if
  18437. 14:30:03you use this, that's the only advantage
  18438. 14:30:06is it's just plainly obvious what you're
  18439. 14:30:08making. Uh because technically you can
  18440. 14:30:10make a set with the curly braces as
  18441. 14:30:13well.
  18442. 14:30:18Yeah,
  18443. 14:30:21no worries. Um okay, so we have our
  18444. 14:30:24alphas here. So what I want you to
  18445. 14:30:27notice is that we are going to try out
  18446. 14:30:29different alphas and we are that's the
  18447. 14:30:31only parameter we are going to try in
  18448. 14:30:33our ridge regression. So we're going to
  18449. 14:30:36we're going to try ridge but just try
  18450. 14:30:38different alphas in the in this range um
  18451. 14:30:41in our grid search. So the grid search
  18452. 14:30:44CV takes in a model. It takes in our
  18453. 14:30:47grid dictionary which is really
  18454. 14:30:48critical. We need that dictionary to
  18455. 14:30:50declare what we're going to try.
  18456. 14:30:53um we need a scoring to say to find the
  18457. 14:30:56best. Now remember it uses the negative
  18458. 14:30:59to find the lowest which is going to be
  18459. 14:31:02the the least negative option.
  18460. 14:31:06Um otherwise it wouldn't um just based
  18461. 14:31:09on the optimization it would look for
  18462. 14:31:11the highest value. Um so the highest
  18463. 14:31:14would be closest to zero in this
  18464. 14:31:15situation. Um so we use negative and
  18465. 14:31:19again we could use squared error. It's
  18466. 14:31:21using absolute. We could use um squared
  18467. 14:31:25uh either either one works.
  18468. 14:31:29Um more typical would probably be
  18469. 14:31:31squared error, but um absolute is fine.
  18470. 14:31:35Here's where we have our repeated kfold.
  18471. 14:31:37So we pass in our um how we're doing CV.
  18472. 14:31:40That can be it can be a kfold object. It
  18473. 14:31:42can actually just be an integer, which
  18474. 14:31:44is say I just want to do 10 splits or
  18475. 14:31:46five splits um to to do every
  18476. 14:31:49evaluation. But these are the bare
  18477. 14:31:51minimum that you need. Just really the
  18478. 14:31:53model and the grid and your CV. Um what
  18479. 14:31:57metric you're using to evaluate what's
  18480. 14:31:59going to be the best. And then this end
  18481. 14:32:01jobs is to parallelize. If you have it
  18482. 14:32:03set to minus one, it's going to it's
  18483. 14:32:04going to try out all the grid options in
  18484. 14:32:06parallel. Um which is nice. It's going
  18485. 14:32:09to help speed up the overall search.
  18486. 14:32:12Okay. So let me mark that down as n
  18487. 14:32:17jobs equals minus one.
  18488. 14:32:21tries out the combos in parallel.
  18489. 14:32:26So in this situation, we actually don't
  18490. 14:32:28have more than one parameter. We only
  18491. 14:32:30have the alpha. So we're really just
  18492. 14:32:32going to be systematically working our
  18493. 14:32:34way through every alpha and evaluating
  18494. 14:32:36which one's the best right with this.
  18495. 14:32:39And notice that in order to use this
  18496. 14:32:41grid search, all we have to do is call
  18497. 14:32:43search.fit. So it works kind of like
  18498. 14:32:46every other model does, right? It's the
  18499. 14:32:49grid search.fit.
  18500. 14:32:52And we pass in our data.
  18501. 14:32:54And we um once we're once this prints
  18502. 14:32:58out the results, you get a results
  18503. 14:33:00object um which has a best score and
  18504. 14:33:04then a dictionary with your best
  18505. 14:33:06parameters. So, whatever your best grid
  18506. 14:33:08member was or grid members, um it prints
  18507. 14:33:12that out and you can So, for from that,
  18508. 14:33:14we can um grab our best alpha, which
  18509. 14:33:18which let's confirm what that ends up
  18510. 14:33:20being.
  18511. 14:33:26Oops. We need to import repeated kfold.
  18512. 14:33:37So we'll import that.
  18513. 14:33:44Oh, I didn't. Let's do from
  18514. 14:33:47sklearn.linear
  18515. 14:33:52model import ridge.
  18516. 14:33:56Okay.
  18517. 14:34:04Okay. So, it completed the search and
  18518. 14:34:06what we found is this is the best score
  18519. 14:34:08is 238 for the mean absolute error and
  18520. 14:34:12the best alpha that we got was 0.9. So,
  18521. 14:34:16the best alpha that worked here, the one
  18522. 14:34:19that gave us the best score was actually
  18523. 14:34:210.9 as the alpha. So what it did is it
  18524. 14:34:24tried out everything between this range
  18525. 14:34:27and 0.9 was the best. So it did cross
  18526. 14:34:30validation, tried out every single combo
  18527. 14:34:33in our grid.
  18528. 14:34:35So if we want we could actually print
  18529. 14:34:37out
  18530. 14:34:40print our grid so we can see
  18531. 14:34:43what our combinations were.
  18532. 14:34:49So, it tried out all of these guys and
  18533. 14:34:51the best one that we had was 0.9.
  18534. 14:35:00Okay, so pretty cool how that works. And
  18535. 14:35:03if we had other parameters, like if we
  18536. 14:35:06were doing a elastic net, we could add
  18537. 14:35:08those into our dictionary and it would
  18538. 14:35:10do all combinations of those. So if we
  18539. 14:35:13did um so for instance to add to our
  18540. 14:35:16grid we could do grid
  18541. 14:35:18um L1 ratio
  18542. 14:35:21this would be for like an elastic net
  18543. 14:35:22right now the ridge regression by itself
  18544. 14:35:24doesn't have an L1 ratio parameter but
  18545. 14:35:26just as an example um we could try out
  18546. 14:35:29different ranges um similar range
  18547. 14:35:32different one um maybe an exact list
  18548. 14:35:35whatever we want to do. So this is going
  18549. 14:35:37to try out different ones between 0ero
  18550. 14:35:38to one
  18551. 14:35:40as well. And so it's going to try out
  18552. 14:35:42every combination of these from this
  18553. 14:35:45grid.
  18554. 14:35:47Okay, if we did that. But again, this
  18555. 14:35:49the ridge regression doesn't have an L1
  18556. 14:35:52ratio. The elastic net does. So that the
  18557. 14:35:55ridge regression only has an alpha to as
  18558. 14:35:58a hyperparameter. So we're only testing
  18559. 14:36:00out that one.
  18560. 14:36:05Okay. So that's grid search CV.
  18561. 14:36:09Pretty useful. This is pretty useful in
  18562. 14:36:11doing parameter tuning again when you
  18563. 14:36:13want to try out ranges of different
  18564. 14:36:15values and you can evaluate those to see
  18565. 14:36:18which one is your best and then we can
  18566. 14:36:21use that best going forward. So we can
  18567. 14:36:23for instance this is what this code does
  18568. 14:36:26below it is it fetches the best. Um you
  18569. 14:36:29can do it this way or you can do it um
  18570. 14:36:32the alternative is to do results.b best
  18571. 14:36:34params
  18572. 14:36:38and then you can just grab it like this
  18573. 14:36:41alpha.
  18574. 14:36:42Either way you can do get or like this
  18575. 14:36:46um and it this is just a dictionary,
  18576. 14:36:48right? And you can grab your alpha. So
  18577. 14:36:50that's the 0.9 um and we can pass that
  18578. 14:36:53alpha into the ridge regression and go
  18579. 14:36:56back and refit it to our data um and
  18580. 14:36:59then use that model going forward. So
  18581. 14:37:01the grid search really just evaluates
  18582. 14:37:04those different options, tells you
  18583. 14:37:06what's the best according to this score,
  18584. 14:37:10right?
  18585. 14:37:12And you should, by the way, you should
  18586. 14:37:14interpret this score in the positive
  18587. 14:37:16sense. It's only negative because we're
  18588. 14:37:19purposely making it negative to find out
  18589. 14:37:22what the lowest option is, right?
  18590. 14:37:24Because the lower is the better. So we
  18591. 14:37:26we purposely make it negative to make it
  18592. 14:37:28whatever is the least negative is the
  18593. 14:37:30winner. Um more negative is worse.
  18594. 14:37:35So it's really positive version of it is
  18595. 14:37:39the is the true result for the error. Um
  18596. 14:37:42and they are a tool from scikitlearn to
  18597. 14:37:45put together your model with your
  18598. 14:37:47pre-processing steps. So they kind of
  18599. 14:37:49get automated together. Um and they
  18600. 14:37:52combine everything into kind of a
  18601. 14:37:54streamline process. You're going to see
  18602. 14:37:56what that looks like, but it's a really
  18603. 14:37:58nice um feature of scikitlearn. Um why
  18604. 14:38:02would we care about pipelines? They help
  18605. 14:38:05organize our code um so that we ensure
  18606. 14:38:08that we basically always run the
  18607. 14:38:10pre-processing steps before we train and
  18608. 14:38:12use a model to with the predictions. Um,
  18609. 14:38:15so it bundles those steps together,
  18610. 14:38:17minimizes the risk of forgetting a step
  18611. 14:38:20because one of the things that can
  18612. 14:38:21happen is when you do pre-processing, if
  18613. 14:38:24you're doing it on the training set, you
  18614. 14:38:25have to do it on new test data as well
  18615. 14:38:27when you put it through your model
  18616. 14:38:29because your model is training against
  18617. 14:38:30that pre-processed data.
  18618. 14:38:33So in order to make sure you never
  18619. 14:38:35forget that, you can bundle it all
  18620. 14:38:37together in a pipeline which is going to
  18621. 14:38:39make things really really easy to use
  18622. 14:38:42and and make sure that those steps
  18623. 14:38:44happen in a repeatable way. Um and it
  18624. 14:38:49makes things easier to uh deploy that
  18625. 14:38:52model as well because everything is
  18626. 14:38:54together in one pipeline. So in the in
  18627. 14:38:57the industry, I've seen this a lot. Um
  18628. 14:39:00you know, people will do their initial
  18629. 14:39:03exploration steps and initial model
  18630. 14:39:05building. They may not use pipelines
  18631. 14:39:07right away, but as they found their
  18632. 14:39:10model, um they'll generally move it into
  18633. 14:39:13a pipeline and all their steps into a
  18634. 14:39:14pipeline so that it's uh easier to work
  18635. 14:39:16with um when you're when you're
  18636. 14:39:18deploying it, actually using it uh in in
  18637. 14:39:22the real world. Um, so here's what a
  18638. 14:39:25pipeline generally looks like. It's from
  18639. 14:39:27scikitlearn. It's this pipeline object.
  18640. 14:39:30Um, and the pipeline is made up of steps
  18641. 14:39:33that we're going to see that that are
  18642. 14:39:35various um uh basically um kinds of
  18643. 14:39:40pre-processing we've seen before like a
  18644. 14:39:42scaler or um filling in missing values.
  18645. 14:39:46Those kind of things we can put here in
  18646. 14:39:48the steps which is basically a list. um
  18647. 14:39:51steps is just going to be a list of
  18648. 14:39:53scikitlearn functions that we can apply
  18649. 14:39:54to data. One of those being a model. Um
  18650. 14:39:58and then whenever we use the pipeline,
  18651. 14:40:00it's basically um you know, it's going
  18652. 14:40:02to be something like pipeline.fit
  18653. 14:40:05or pipeline.predict.
  18654. 14:40:07So the pipeline kind of behaves like a
  18655. 14:40:10model. It's just going to contain many
  18656. 14:40:12more steps than that like the
  18657. 14:40:14pre-processing steps we've worked with
  18658. 14:40:16before. Um, and it also has some
  18659. 14:40:19capabilities for caching. So you can
  18660. 14:40:21like uh cache some of the data in
  18661. 14:40:24memory. Um, so that if you're reusing
  18662. 14:40:26the predictions, it kind of goes faster.
  18663. 14:40:29Um, so there's some options for that
  18664. 14:40:31too. I'm not too concerned about that at
  18665. 14:40:33this stage, but the main thing is going
  18666. 14:40:35to be filling out our steps and then
  18667. 14:40:37using the pipeline.
  18668. 14:40:40Okay.
  18669. 14:40:42Um, so some important bits of
  18670. 14:40:45information about the pipeline is that
  18671. 14:40:46it is going to be a sequence of data
  18672. 14:40:48transformations that will have at the
  18673. 14:40:51very end of the pipeline the model
  18674. 14:40:53because of course we're going to do
  18675. 14:40:55transformations and then train a model
  18676. 14:40:58or predict with a model. So every
  18677. 14:41:03Oh, can you guys hear me? Okay,
  18678. 14:41:06not able to hear me. Thanks for letting
  18679. 14:41:08me know. Can you guys were you able to
  18680. 14:41:09hear me so far?
  18681. 14:41:14Okay. Make sure. Yeah, it might be on
  18682. 14:41:16your internet or your your uh Yeah, it
  18683. 14:41:20seems like seems like it's good. So,
  18684. 14:41:24no, you can't hear me. Check your
  18685. 14:41:26volume. Check your headphones if you're
  18686. 14:41:28wearing headphones. Oh, no issues. Okay,
  18687. 14:41:31perfect.
  18688. 14:41:32Okay. Yeah, local internet issue. Yeah.
  18689. 14:41:37Okay.
  18690. 14:41:39Always let me know. always let me know
  18691. 14:41:40cuz it could be the case that it is me.
  18692. 14:41:43So, um always always make sure to let me
  18693. 14:41:46know. Um but sounds like yeah, you may
  18694. 14:41:49want to check on that. Um
  18695. 14:41:52so, okay. What I was saying is every
  18696. 14:41:55pipeline's going to have a uh a sequence
  18697. 14:41:57of steps that go first and then the
  18698. 14:41:59model at the end. Um so, the order
  18699. 14:42:02really matters. Um
  18700. 14:42:05uh so the order matters in the sense
  18701. 14:42:08that we want our transformations to go
  18702. 14:42:09first. Things like scaling, things like
  18703. 14:42:12filling in missing values, we want those
  18704. 14:42:13to be first and then we want our uh
  18705. 14:42:17model to be last because we want those
  18706. 14:42:19transformations to happen prior to
  18707. 14:42:21training or prior to prediction. So
  18708. 14:42:24usually what you'll see in these
  18709. 14:42:26pipelines is a model at the end, right?
  18710. 14:42:28a model that's going to be at the end of
  18711. 14:42:31the pipeline because we want basically
  18712. 14:42:34our processing steps then our training
  18713. 14:42:36or our processing steps then our
  18714. 14:42:38predictions. Um so everything in the
  18715. 14:42:42pipeline though is going to be from
  18716. 14:42:43scikitlearn. Uh that's how it gets
  18717. 14:42:46automated in the sense that all of those
  18718. 14:42:48things are going to have fit and
  18719. 14:42:49transform functions built into them so
  18720. 14:42:51the pipeline can use them. Uh, and then
  18721. 14:42:54the last step is going to be a model
  18722. 14:42:56that has a fit and a predict. So it's
  18723. 14:42:59pretty standard that the last part of
  18724. 14:43:00the pipeline is just going to be a
  18725. 14:43:01model. Um,
  18726. 14:43:05uh,
  18727. 14:43:06so we can um, as we do more modeling,
  18728. 14:43:11we're going to play around with the
  18729. 14:43:12pipelines quite a bit and see how we can
  18730. 14:43:14change up some of the parameters. like
  18731. 14:43:15if we want to change a model's parameter
  18732. 14:43:18um we can actually adjust it to do
  18733. 14:43:20things like uh grid search or cross
  18734. 14:43:22validation. So um we're going to see
  18735. 14:43:26some examples of some pipelines but for
  18736. 14:43:28right now mostly what we're going to see
  18737. 14:43:30is how to build one and then how to use
  18738. 14:43:32one. And then as we get into lesson
  18739. 14:43:35four, we'll get some more practice with
  18740. 14:43:37pipelines cuz we're going to start using
  18741. 14:43:38them quite a bit uh to build our models
  18742. 14:43:41rather than do manual steps uh all the
  18743. 14:43:45manual pre-processing
  18744. 14:43:47um and then kind of building a model
  18745. 14:43:49from there. We'll just include all of it
  18746. 14:43:51together in a pipeline.
  18747. 14:43:55Okay, so the example we're going to do
  18748. 14:43:56is with this housing with ocean
  18749. 14:43:59proximity. So we've actually looked at
  18750. 14:44:00this data set before. Um so we have uh
  18751. 14:44:05this ocean proximity data set that has
  18752. 14:44:07the feature of like how close it is to
  18753. 14:44:09the ocean like the bay or the less than
  18754. 14:44:111 hour. Remember we had that and it had
  18755. 14:44:14the median house value for different
  18756. 14:44:16neighborhoods. Um so we're going to work
  18757. 14:44:18with that one again. Let me make sure I
  18758. 14:44:20have that one uploaded.
  18759. 14:44:24You guys should have this one. It should
  18760. 14:44:25be in your uh data sets.
  18761. 14:44:29Um, I'll I can upload it here in case
  18762. 14:44:31you don't have it though.
  18763. 14:44:42Does this use multi-threading? I think
  18764. 14:44:44it does. Yeah, I think in order to do it
  18765. 14:44:46can do uh um I think it can do
  18766. 14:44:49processing in parallel for some of the
  18767. 14:44:51pipeline steps. Um, now does it use that
  18768. 14:44:54all the time? Not necessarily because
  18769. 14:44:57some of it is sequential in nature where
  18770. 14:44:59you have to do one step and then you do
  18771. 14:45:01the next step and then you do the next
  18772. 14:45:03step. So it's not like you can do them
  18773. 14:45:04in parallel.
  18774. 14:45:06Um in terms of the like you need to know
  18775. 14:45:08the output of one step to compute the
  18776. 14:45:10the output of the next step. Um so it
  18777. 14:45:15can but it it doesn't always lend itself
  18778. 14:45:18well. The thing that will use
  18779. 14:45:20multi-threading is is like the training
  18780. 14:45:22process could be parallelized
  18781. 14:45:25like the fit um can be for some models
  18782. 14:45:29it can be parallelized not every model
  18783. 14:45:35it so long answer is or the short answer
  18784. 14:45:38is that it depends
  18785. 14:45:40depends on what kind of transforms
  18786. 14:45:41you're doing and what kind of model
  18787. 14:45:42you're using if you can really take
  18788. 14:45:44advantage of
  18789. 14:45:50Okay. So, we load our data here and take
  18790. 14:45:53a look at that. Um, do you guys have
  18791. 14:45:56this data set? Are you able to load it
  18792. 14:45:58in? If you're following along, are you
  18793. 14:46:00able to load it?
  18794. 14:46:08Okay.
  18795. 14:46:10And and again, we've worked with this
  18796. 14:46:11data before, so hopefully it's somewhat
  18797. 14:46:13familiar. Remember, every row represents
  18798. 14:46:16a neighborhood and it has a we're going
  18799. 14:46:17to end up trying to predict this median
  18800. 14:46:20house value as our target um variable,
  18801. 14:46:24our dependent variable. Um and we're
  18802. 14:46:26going to use the rest of these features.
  18803. 14:46:28Remember that um this feature is in
  18804. 14:46:31particular going to need to be one hot
  18805. 14:46:33encoded,
  18806. 14:46:35right? It's going to be one hot encoded
  18807. 14:46:37because it is currently a string and we
  18808. 14:46:39need to turn that into a numerical
  18809. 14:46:42feature which is the one hot encoded
  18810. 14:46:43feature. So we're going to have to do
  18811. 14:46:46that but we're going to do that as part
  18812. 14:46:48of our pipeline.
  18813. 14:46:50Okay. So we'll be able to include that
  18814. 14:46:52in our pipeline steps uh to to do one
  18815. 14:46:55hot encoding which is nice.
  18816. 14:46:59All right. So we're going to split apart
  18817. 14:47:00our data um as we normally do. So we're
  18818. 14:47:04going to uh create our feature uh data
  18819. 14:47:08frame which is everything but this
  18820. 14:47:10median house value. So we go ahead and
  18821. 14:47:12drop that column and then our target is
  18822. 14:47:14the median house value. So it is just
  18823. 14:47:16that column here. Pretty standard. Um
  18824. 14:47:20and then we're going to train test split
  18825. 14:47:23and um split it into 30%
  18826. 14:47:27uh test data. And again random state you
  18827. 14:47:30can choose whatever you want to be. that
  18828. 14:47:31just affects the shuffling. Um, so
  18829. 14:47:34whatever doesn't really matter what it
  18830. 14:47:36is. It's just so that when you rerun
  18831. 14:47:37this, you get the same result in the in
  18832. 14:47:39the shuffle.
  18833. 14:47:42Okay, so we have our train and our test.
  18834. 14:47:47So you want to make sure you run those.
  18835. 14:47:50All right, so what we're going to do is
  18836. 14:47:52take a look at our data
  18837. 14:47:54and see if we have any null values. Um
  18838. 14:47:58if you guys remember this data actually
  18839. 14:48:01did have null values. You can see it
  18840. 14:48:02here in this this guy and exactly how
  18841. 14:48:05many there are is from this the sum. So
  18842. 14:48:08we have um 162 nles in in this data. Uh
  18843. 14:48:13and this is just a training data. So of
  18844. 14:48:15course you know the test data could have
  18845. 14:48:17that in there as well. Um so that's
  18846. 14:48:19something we're going to want to make
  18847. 14:48:21sure we fill in the blanks on any data
  18848. 14:48:23set we use whether we're using the
  18849. 14:48:25training or test set. Um, like if we're
  18850. 14:48:27doing training, we want to make sure
  18851. 14:48:29that gets filled in. If we're doing
  18852. 14:48:30predictions with the test set, want to
  18853. 14:48:32make sure that gets filled in. Um, so we
  18854. 14:48:35we should be doing that. Um, now
  18855. 14:48:41what we're going to do is use this data
  18856. 14:48:44to help uh train our pipeline or or use
  18857. 14:48:49with our pipeline. We need to construct
  18858. 14:48:51our pipeline. So far, we've just split
  18859. 14:48:53apart our data. We haven't done anything
  18860. 14:48:55with our processing steps in our model
  18861. 14:48:57yet. Um so roughly
  18862. 14:49:02it this should be the flow of our
  18863. 14:49:03pipeline. What should happen is we
  18864. 14:49:05should be doing some type of feature
  18865. 14:49:07scaling
  18866. 14:49:08um some type of uh feature um
  18867. 14:49:12manipulation. So that could be
  18868. 14:49:13engineering, that could be um that could
  18869. 14:49:17be uh doing the one hot encoding. Um so
  18870. 14:49:21extracting new features like one hot
  18871. 14:49:23encoding,
  18872. 14:49:26one hot encoding. Um we are going to be
  18873. 14:49:29doing that and and by the way, this is
  18874. 14:49:31split up into this is when we use our
  18875. 14:49:33pipeline for training.
  18876. 14:49:36Um it's going to look like this where we
  18877. 14:49:37do our scaling, we do one hot encoding,
  18878. 14:49:40um we have our model here. Um, so that
  18879. 14:49:43could be a linear regression, that could
  18880. 14:49:44be a lasso, that could be a ridge, it
  18881. 14:49:46could be elastic net. Whatever model we
  18882. 14:49:48end up using is going to be last in the
  18883. 14:49:50pipeline. And we're going to run this
  18884. 14:49:53pipeline. Ultimately, we're going to run
  18885. 14:49:55pipeline.fit,
  18886. 14:50:00right? We're going to run a fit function
  18887. 14:50:02and we get a fitted model as the result
  18888. 14:50:04of this pipeline.
  18889. 14:50:07Then when we use it when we use our
  18890. 14:50:10model for prediction,
  18891. 14:50:14we use our model for prediction in this
  18892. 14:50:16lower part. It's the same pipeline, same
  18893. 14:50:19exact pipeline, but it's this model has
  18894. 14:50:22now been trained.
  18895. 14:50:24So we now have a trained model here. So
  18896. 14:50:26the great thing about the pipeline is
  18897. 14:50:28it's the same this is the same pipeline
  18898. 14:50:30that we're using here. So it's just
  18899. 14:50:33going to it's going to repeat those same
  18900. 14:50:35transformations. It's going to do our
  18901. 14:50:37scaling. It's going to do our one hot
  18902. 14:50:39encoding. It's going to use our model
  18903. 14:50:41and it's going to generate predictions
  18904. 14:50:43and generate uh we can we can do
  18905. 14:50:45predictions. We can do evaluation like
  18906. 14:50:47in a cross validation. Um we can use it
  18907. 14:50:50however we want to use it. Uh but notice
  18908. 14:50:54that the pipeline makes it consistent
  18909. 14:50:57between training and test. We're using
  18910. 14:50:58the exact same transformations
  18911. 14:51:01and the model is last. It's it's either
  18912. 14:51:04being trained or it's being used for
  18913. 14:51:05prediction, but it's last. Our
  18914. 14:51:07transformations are upfront, which are
  18915. 14:51:10things like our scaling, things like our
  18916. 14:51:11one hot encoding, right? Those happen
  18917. 14:51:14first. No matter what data we put
  18918. 14:51:17through there, we put our training data
  18919. 14:51:19through there, we put our test data
  18920. 14:51:20through there, they're going to go
  18921. 14:51:21through the same steps,
  18922. 14:51:25right?
  18923. 14:51:28So that's that's the design of the
  18924. 14:51:29pipeline. That's what it's supposed to
  18925. 14:51:31do. So our job is to create those steps.
  18926. 14:51:36So we need to create those relevant
  18927. 14:51:38steps and then put them together into
  18928. 14:51:40this pipeline. Okay. So that's going to
  18929. 14:51:43be the code we're going to see coming up
  18930. 14:51:44is we're going to build out these steps
  18931. 14:51:47and then put them together into the
  18932. 14:51:48pipeline.
  18933. 14:51:58Um any questions on this diagram? Does
  18934. 14:52:00it make sense what we're trying to do
  18935. 14:52:02with this pipeline? We want to have
  18936. 14:52:04repeatable steps during the training,
  18937. 14:52:06during a prediction process.
  18938. 14:52:09Okay.
  18939. 14:52:15All right.
  18940. 14:52:21All right. So, um, a couple of things
  18941. 14:52:24we're going to need is, uh, to first of
  18942. 14:52:27all, let's jot down what steps we're
  18943. 14:52:29actually going to do. We're going to
  18944. 14:52:30need to deal with missing values. So,
  18945. 14:52:32we're going to fill in we're going to
  18946. 14:52:33need a pre-processing pre-processing
  18947. 14:52:36step that fills in any nulls. We always
  18948. 14:52:39need that, right? So, if there's nles,
  18949. 14:52:42we're going to fill them in somehow.
  18950. 14:52:44We're going to define how we do that in
  18951. 14:52:46our in our step. Um and we also need to
  18952. 14:52:50one hot encode and we need to scale
  18953. 14:52:53right those are pretty standard steps
  18954. 14:52:56that we've dealt with whenever we're
  18955. 14:52:57building these models right so pretty
  18956. 14:52:59standard things fill in nles one hot
  18957. 14:53:02encode any categorical data whatever
  18958. 14:53:04however much we have and then go ahead
  18959. 14:53:07and um standardize which is the scaling
  18960. 14:53:10so this this just is the same word for
  18961. 14:53:13scaling our numeric features so we're
  18962. 14:53:15going to we're going to define Windows.
  18963. 14:53:18Um, so that's why we're going to go
  18964. 14:53:21ahead and import from pre-processing.
  18965. 14:53:23We're going to import our scaler. Um,
  18966. 14:53:25again, we could use minmax scaler here.
  18967. 14:53:27We're going to use standard scaler. Um,
  18968. 14:53:30but we could use minmax. Um, we have our
  18969. 14:53:33one hot encoder here. Now, usually when
  18970. 14:53:37we do oneh hot encoding, we use pd.get
  18971. 14:53:41dummies. This does the same thing as
  18972. 14:53:44that, but because we're going to be
  18973. 14:53:46building a pipeline, we actually want
  18974. 14:53:48the scikitlearn version of git dummies.
  18975. 14:53:52So this is the scikitlearn version of
  18976. 14:53:54git dummies here. And it and we have to
  18977. 14:53:56use that version in the pipeline because
  18978. 14:53:59everything in the pipeline needs to be
  18979. 14:54:00an sklearn object. It needs to be an
  18980. 14:54:03sklearn tool or object.
  18981. 14:54:06So um instead of using pandis get
  18982. 14:54:09dummies we're using one hot encoder
  18983. 14:54:11which is does the same thing. Okay. In
  18984. 14:54:15fact it just this basically just uses
  18985. 14:54:18pd.get dummies um under the hood.
  18986. 14:54:24Okay. So it just uses that uh anyways.
  18987. 14:54:26It's just code that builds on builds on
  18988. 14:54:28that.
  18989. 14:54:30Now what's really nice here is we're
  18990. 14:54:32also going to use from sklearn.impute
  18991. 14:54:35impute. We're going to use a simple
  18992. 14:54:36imper now what this is is an automated
  18993. 14:54:40way to fill in missing values. So this
  18994. 14:54:42is a fancy way of basically doing the
  18995. 14:54:45the fill na on a data frame. So simple
  18996. 14:54:48imputer um we are going to basically
  18997. 14:54:52fill in the blanks. What we're going to
  18998. 14:54:54do when we create this object is give it
  18999. 14:54:56a strategy of how to fill in blanks.
  19000. 14:54:58Should you use the average? Should you
  19001. 14:55:00use the median? Should you use the max?
  19002. 14:55:02Should you use the min? should use a
  19003. 14:55:04default value. We're going to tell it
  19004. 14:55:06what to do in this object.
  19005. 14:55:09Okay. So, we're going to we're so we're
  19006. 14:55:12going to use this as our automated tool
  19007. 14:55:14for filling in missing values. So,
  19008. 14:55:16that's really nice. It has so this is
  19009. 14:55:18going to be a critical part of our
  19010. 14:55:20pipeline an imputer that's going to fill
  19011. 14:55:23in missing values.
  19012. 14:55:26So, we have that.
  19013. 14:55:31Yeah. Coding to reduce coding. Exactly.
  19014. 14:55:34Uh we have our pipeline now. So we have
  19015. 14:55:36our pipeline. So our pipeline is going
  19016. 14:55:38to hold everything. So we need the
  19017. 14:55:39pipeline object um to hold everything
  19018. 14:55:42and that comes from sklearn.pipeline.
  19019. 14:55:45Um so everything's going to actually go
  19020. 14:55:47into a pipeline object. We're going to
  19021. 14:55:49see how that looks. Um and finally we're
  19022. 14:55:53going to from skarn.compose we're going
  19023. 14:55:56to use a column transformer. The reason
  19024. 14:55:58we're going to do this is because we are
  19025. 14:56:01going to specify for some columns like
  19026. 14:56:04the numerical features we should be
  19027. 14:56:06scaling
  19028. 14:56:07for some columns like the categorical
  19029. 14:56:10features we should be one hot encoding.
  19030. 14:56:13So the column transformer will allow us
  19031. 14:56:15to map different transformations to
  19032. 14:56:18different sections of columns which is
  19033. 14:56:20really useful. So this is actually going
  19034. 14:56:22to be a critical part of our pipeline to
  19035. 14:56:25apply to make sure we only apply this to
  19036. 14:56:27numerical features and only apply this
  19037. 14:56:30to categorical features. Right? So this
  19038. 14:56:33column transformer will help us um to to
  19039. 14:56:38apply pre-processing to particular
  19040. 14:56:40columns. Um like that ocean proximity is
  19041. 14:56:44the only one that really needs this but
  19042. 14:56:46every other column is going to need this
  19043. 14:56:48all the numerical features.
  19044. 14:56:50So, we're going to use this column
  19045. 14:56:52transformer. And again, we're going to
  19046. 14:56:53see how this looks, but just trying to
  19047. 14:56:56give you an idea of why we're importing
  19048. 14:56:57all these things.
  19049. 14:57:03Okay. So, let's import those.
  19050. 14:57:06Uh, this mentions about the column
  19051. 14:57:08transformer. We just talked about it. It
  19052. 14:57:10allows us to have a particular column or
  19053. 14:57:13group of columns get the right
  19054. 14:57:14transformation. So again, uh, looking
  19055. 14:57:17ahead to our pipeline, the numerical
  19056. 14:57:20features are the ones that are going to
  19057. 14:57:21need scaling, but the categorical
  19058. 14:57:24features are the ones that are going to
  19059. 14:57:26need one hot encoding. However many
  19060. 14:57:27categoricals there are. In this case,
  19061. 14:57:29there's really only one, which is that
  19062. 14:57:30ocean proximity. Go back to our data.
  19063. 14:57:33Um, you can even see that in the info,
  19064. 14:57:36there's just that one. Um, and we see
  19065. 14:57:38that here, right? Just this one string
  19066. 14:57:40column that should be one hot encoded.
  19067. 14:57:42All these other guys should be scaled.
  19068. 14:57:45Right? They should all be uh uh standard
  19069. 14:57:47scaled.
  19070. 14:57:49So this will allow us to specify those
  19071. 14:57:52distinctions.
  19072. 14:57:56All right. So let's get started building
  19073. 14:58:01our pipeline. So this is going to be
  19074. 14:58:02really cool. We're going to build out
  19075. 14:58:03the pipeline. Um let's extract our
  19076. 14:58:08numerical data and our categorical data.
  19077. 14:58:10Now this is a really neat way of doing
  19078. 14:58:12that that I'm not sure we've seen
  19079. 14:58:13before.
  19080. 14:58:14Um so what this does is we'll take our
  19081. 14:58:18data frame particular our training data
  19082. 14:58:21frame and select our data
  19083. 14:58:26that's what this select dtypes does is
  19084. 14:58:28select data from it um which includes
  19085. 14:58:31only the object type columns so only the
  19086. 14:58:35object types. Now what's that?
  19087. 14:58:38The object type is the string right? So
  19088. 14:58:41this should select only this column
  19089. 14:58:44because it's in the include.
  19090. 14:58:47We go here include only object types in
  19091. 14:58:50the result. And so this should only have
  19092. 14:58:53our one categorical column which is
  19093. 14:58:56ocean proximity. So, housing cat is
  19094. 14:58:59going to have a reference to our uh it's
  19095. 14:59:03going to be a list that has a a
  19096. 14:59:05basically just our ocean proximity
  19097. 14:59:08feature because this select dtypes will
  19098. 14:59:11make sure we only pick object types and
  19099. 14:59:15um
  19100. 14:59:16grab those columns. So this is a way to
  19101. 14:59:20neatly grab um our categorical features
  19102. 14:59:24here by including the object types. Now
  19103. 14:59:28on the flip side we can exclude object
  19104. 14:59:30types and get everything else. So this
  19105. 14:59:32is going to be all other columns which
  19106. 14:59:35is excluding the object. So this is
  19107. 14:59:38excluding this meaning we should get all
  19108. 14:59:41of our numerical features that way. So
  19109. 14:59:44this will be all of our numericals
  19110. 14:59:47by excluding the object type and this
  19111. 14:59:51will be our housing num which is short
  19112. 14:59:54for numerical. So this excludes
  19113. 14:59:58the uh object type meaning all numerical
  19114. 15:00:06features
  19115. 15:00:08right all numerical features there.
  19116. 15:00:12Okay.
  19117. 15:00:15So, if we were to uh let's double check
  19118. 15:00:18this. Let's sanity check this. If we
  19119. 15:00:19were to print out the housing
  19120. 15:00:23cat, um this should be just the ocean
  19121. 15:00:27proximity feature, which it is. So, just
  19122. 15:00:30that one. If we were to print out the
  19123. 15:00:32housing num, this should be all the
  19124. 15:00:34numerical features, which are all these
  19125. 15:00:37guys. So it's just a reference to those
  19126. 15:00:39columns so that we can uh use those
  19127. 15:00:43later when we're mapping uh this
  19128. 15:00:45transform needs to go to this column
  19129. 15:00:47like the one hot encoding needs to go to
  19130. 15:00:49this column and the scaling needs to go
  19131. 15:00:52to these columns right so we have those
  19132. 15:00:56uh names of those columns already at our
  19133. 15:00:58disposal. So, we're just doing that.
  19134. 15:01:04And this is just a
  19135. 15:01:07simple check.
  19136. 15:01:11Uh, are you guys able to run this?
  19137. 15:01:16If you're following along, let me pause
  19138. 15:01:18there. Make sure I'm not going too fast.
  19139. 15:01:26Uh it so the the issue with a specific
  19140. 15:01:29data type like that is none of these are
  19141. 15:01:31ants. They're actually all floats. So we
  19142. 15:01:34did float. I think that should work. But
  19143. 15:01:36yes, that's the idea.
  19144. 15:01:41Great. I'm glad to hear that right there
  19145. 15:01:43with me. Great. Glad to hear that.
  19146. 15:01:54Okay. So, we have our columns picked out
  19147. 15:01:56here, which we're going to use later.
  19148. 15:01:59Okay.
  19149. 15:02:02All right. So, let's go ahead and build
  19150. 15:02:06out our steps for each of these types.
  19151. 15:02:10So, um for our numerical features, let's
  19152. 15:02:15build out our pipeline steps. So what
  19153. 15:02:17we're going to do is build out a
  19154. 15:02:18numerical pipeline. And it's going to be
  19155. 15:02:21a pipeline with a list
  19156. 15:02:25of tupils. And the reason these are
  19157. 15:02:28tupils is because every tupil has a
  19158. 15:02:30name. So here this is a name that we can
  19159. 15:02:33it can be whatever we want it to be. So
  19160. 15:02:35we're calling it imputer. We could call
  19161. 15:02:38it anything we want. We could call it
  19162. 15:02:39fill in the blanks. We could call it
  19163. 15:02:41null filling. Call it whatever you want.
  19164. 15:02:45We're calling it imputer because that's
  19165. 15:02:46that's a pretty um easy name for it. An
  19166. 15:02:50accurate name to what it's doing. Um but
  19167. 15:02:53the important thing is after the name
  19168. 15:02:55you give it, you put in the scikitlearn
  19169. 15:02:59object that you are going to use to
  19170. 15:03:01operate on your data. So in this case,
  19171. 15:03:04we're using a simple impery
  19172. 15:03:08of median. Now that's a choice. We could
  19173. 15:03:11use a strategy of mean, max. Um, we
  19174. 15:03:16could provide it a constant default
  19175. 15:03:18value. But what this means is we are
  19176. 15:03:22going to fill any blanks we find in
  19177. 15:03:24those columns with the median value of
  19178. 15:03:27that column. That's the strategy for the
  19179. 15:03:29computer. So that's pretty cool. This is
  19180. 15:03:31kind of an automated way to fill in the
  19181. 15:03:32blanks using for any column using its
  19182. 15:03:37median,
  19183. 15:03:39right? And so we could change that. We
  19184. 15:03:40could put mean here or max or min or
  19185. 15:03:43whatever. Um
  19186. 15:03:46but we are filling in the blank on any
  19187. 15:03:48column with its median. And the reason
  19188. 15:03:51this works is because we are going to
  19189. 15:03:53apply this pipeline only to these
  19190. 15:03:55numerical features. So that is fine.
  19191. 15:03:59We're we're not going to apply it to the
  19192. 15:04:01categorical features. We're going to
  19193. 15:04:02apply it to only those numerical. So it
  19194. 15:04:05should have a median value, right? So
  19195. 15:04:08that that's totally fine. So we're going
  19196. 15:04:11to now look at how we're constructing
  19197. 15:04:13the steps. We have a list of tupils.
  19198. 15:04:16Here's one tupole
  19199. 15:04:19which is the imputer with a simple imper
  19200. 15:04:22of strategy median. And then we can have
  19201. 15:04:25as many tupils as we want which
  19202. 15:04:27represent processing steps. So every let
  19203. 15:04:30me write that down. Every tupil
  19204. 15:04:34represents
  19205. 15:04:36a pre-processing
  19206. 15:04:38step on our data.
  19207. 15:04:42Okay, so we have an imputer step named
  19208. 15:04:46imputer and the reason it has a name is
  19209. 15:04:49just so you can reference it in the
  19210. 15:04:51pipeline if you need to. So you so it
  19211. 15:04:53has like a a reference name um that you
  19212. 15:04:56give it. Um but this is the more
  19213. 15:04:59important part is the actual scikitlearn
  19214. 15:05:01object that's doing the processing. So
  19215. 15:05:04in this case a simple computer but
  19216. 15:05:06notice that we have a secondary step
  19217. 15:05:08which is our scaling. Now this makes
  19218. 15:05:09sense. This is something we should be
  19219. 15:05:11doing to our features is we should be
  19220. 15:05:14scaling them. So here we we say okay
  19221. 15:05:17let's fill in any blanks first.
  19222. 15:05:20By the way order
  19223. 15:05:23matters.
  19224. 15:05:26So, and what I mean by that is the
  19225. 15:05:30simple imputer
  19226. 15:05:33is before the scaler. Now, that's
  19227. 15:05:37important because what that means is we
  19228. 15:05:40should be filling in any blanks before
  19229. 15:05:42we attempt scaling.
  19230. 15:05:45So, that order actually matters. We're
  19231. 15:05:47going to fill in blanks first in this
  19232. 15:05:50list. That's first. We're going to fill
  19233. 15:05:52in blanks. Then we are going to scale
  19234. 15:05:58right then we scale which makes sense
  19235. 15:06:01right so we we fill in blanks first then
  19236. 15:06:03we apply the scaler to scale our
  19237. 15:06:05features so those are our two steps
  19238. 15:06:10so so pretty simple um we are building
  19239. 15:06:14out our two steps now this is just one
  19240. 15:06:17piece of the puzzle we are going to put
  19241. 15:06:18this pipeline together with our one hot
  19242. 15:06:21encoding that's going to be coming up
  19243. 15:06:23next and build out our final pipeline.
  19244. 15:06:27But this is um a a pipeline that has two
  19245. 15:06:30steps that will actually be used with a
  19246. 15:06:32larger pipeline coming up where we we do
  19247. 15:06:35one hot encoding to our categoricals and
  19248. 15:06:37then we put a model in there at the end
  19249. 15:06:40to train and and use for prediction. So
  19250. 15:06:44um pipelines can actually be composed is
  19251. 15:06:48is uh something to realize there is that
  19252. 15:06:50we can have a pipeline that contains a
  19253. 15:06:53few steps. We can have another pipeline
  19254. 15:06:54over here that contains a few steps and
  19255. 15:06:56we can actually um kind of put them
  19256. 15:06:58together into a final pipeline that has
  19257. 15:07:00both pipelines uh kind of merged
  19258. 15:07:02together. Okay. So we're going to see
  19259. 15:07:05that coming up when we construct our
  19260. 15:07:07final one. Our final one, as you can
  19261. 15:07:09imagine, needs to handle this mapping of
  19262. 15:07:12basically saying, let's do one hot
  19263. 15:07:13encoding to these guys and then do this
  19264. 15:07:17pipeline here to these numerical
  19265. 15:07:20features. That's what our final pipeline
  19266. 15:07:23needs to handle. And it will. We're
  19267. 15:07:25going to build that out.
  19268. 15:07:28But let me pause here. Um, were you guys
  19269. 15:07:32able to run this? Are you with me on
  19270. 15:07:35this this pipeline here?
  19271. 15:07:38Does that make sense? Those two steps
  19272. 15:07:40one is filling in blanks with a median
  19273. 15:07:44whatever column. So where so this is
  19274. 15:07:47this is what's so amazing about this is
  19275. 15:07:50this is going to automatically search
  19276. 15:07:52for nulls and if you come across a
  19277. 15:07:56column with a null, it's going to use
  19278. 15:07:58the median of that column
  19279. 15:08:02to fill in the blank, right? To fill in
  19280. 15:08:04those nles.
  19281. 15:08:18Okay,
  19282. 15:08:21great. Glad to hear. Glad to hear.
  19283. 15:08:25Okay.
  19284. 15:08:27All right. So we are going to now um put
  19285. 15:08:32this together with a column transformer
  19286. 15:08:37to basically say what steps are going to
  19287. 15:08:40be mapped to what columns.
  19288. 15:08:44Um so now you can see what we're doing
  19289. 15:08:47here is using the column transformer
  19290. 15:08:49which is going to be a list of tupils
  19291. 15:08:51again. So this is another um list of
  19292. 15:08:55tupils.
  19293. 15:08:57But the important thing is um
  19294. 15:09:01each tupil
  19295. 15:09:04has a name
  19296. 15:09:06followed by so it has a name uh which
  19297. 15:09:10again is is generic. You can say
  19298. 15:09:12whatever you want it to be. So here
  19299. 15:09:14we're kind of shortening this to
  19300. 15:09:15numerical. This is short for
  19301. 15:09:16categorical. But the important thing is
  19302. 15:09:18it's followed by a pipeline
  19303. 15:09:23slashstep
  19304. 15:09:26followed by a pipeline slashstep
  19305. 15:09:29um followed by a uh followed by a list
  19306. 15:09:34of columns that it applies to. So you
  19307. 15:09:39can see that pattern here. What we're
  19308. 15:09:41saying is we're going to apply that
  19309. 15:09:44numerical pipeline we just defined. So
  19310. 15:09:46this is saved in a numerical pipeline
  19311. 15:09:48object here. We're going to apply that
  19312. 15:09:51to those numerical features. So this is
  19313. 15:09:54that list
  19314. 15:09:56of numerical features here. So that's
  19315. 15:09:59how we do the mapping. We have a tupil
  19316. 15:10:01here that says okay apply these steps to
  19317. 15:10:04these columns.
  19318. 15:10:06Those go together in that tupil, right?
  19319. 15:10:09Apply these steps to this uh these
  19320. 15:10:12columns. And then apply this step. Now
  19321. 15:10:15what is the step? This is a one hot
  19322. 15:10:17encoder
  19323. 15:10:19which is going to uh uh encode um those
  19324. 15:10:24features and it's going to uh ignore um
  19325. 15:10:28basically nulls for now. That's a choice
  19326. 15:10:31but it's going to ignore um uh basically
  19327. 15:10:36ignore nles and and uh skip over them
  19328. 15:10:39for now. We now we know there's no NLES
  19329. 15:10:42because we already did an is NA from
  19330. 15:10:45before and we know there's not any NLES
  19331. 15:10:48in that ocean proximity. So this isn't
  19332. 15:10:49going to be an issue. But that's what
  19333. 15:10:52that would do.
  19334. 15:10:54But we have a one hot encoder here which
  19335. 15:10:57we're going to apply to our categorical
  19336. 15:11:00features. Now of course that's just the
  19337. 15:11:03ocean proximity feature but that but
  19338. 15:11:05again you see the pattern in the tupole
  19339. 15:11:07is apply this transform which is the one
  19340. 15:11:10hot encoding to this column apply these
  19341. 15:11:13numerical transforms which is a whole
  19342. 15:11:15pipeline. So it's two steps in a
  19343. 15:11:18pipeline of um
  19344. 15:11:22uh an imputer and a scaler are going to
  19345. 15:11:25be applied to this
  19346. 15:11:28really nice. So those are going to be
  19347. 15:11:29all together in this column transformer
  19348. 15:11:32and that is our way to signal that for
  19349. 15:11:33these numerical features use these
  19350. 15:11:36steps. For our categorical features use
  19351. 15:11:38this step and and you know if we had
  19352. 15:11:41more than one step we were applying to
  19353. 15:11:42categorical we could build a pipeline
  19354. 15:11:45for the categorical and it would and do
  19355. 15:11:47the same thing. We have more than one
  19356. 15:11:49step here and so it's good practice when
  19357. 15:11:52you have more than one step to just put
  19358. 15:11:53that in a pipeline because we have more
  19359. 15:11:55than one step. We'll just put that in
  19360. 15:11:57this list inside of the pipeline and we
  19361. 15:12:00can map that pipeline to those features.
  19362. 15:12:03Here we only have one step. So it's okay
  19363. 15:12:05to just put that there um and apply that
  19364. 15:12:09to the categorical features. But if we
  19365. 15:12:11had more than one step um it would be
  19366. 15:12:14good practice to put that in a pipeline
  19367. 15:12:17which is what we do here. Right? This
  19368. 15:12:18pipeline is being mapped to these
  19369. 15:12:20features. This step is being applied to
  19370. 15:12:23this feature.
  19371. 15:12:28Okay,
  19372. 15:12:30how about that? Are you guys able to run
  19373. 15:12:33that one? Does that make sense what we
  19374. 15:12:35have set up so far? So, we're almost
  19375. 15:12:37there. We almost have our final
  19376. 15:12:38pipeline. We have our pre-processing
  19377. 15:12:40basically done to say our numerical
  19378. 15:12:43features should be processed with that
  19379. 15:12:44other pipeline and our categorical
  19380. 15:12:47features should be one hot encoded.
  19381. 15:12:49We're getting close. The only thing
  19382. 15:12:50we're really missing here is a model.
  19383. 15:12:54The only thing we're really missing is
  19384. 15:12:56to have our final model training
  19385. 15:12:59pipeline is to actually include a model
  19386. 15:13:01which should come at the end.
  19387. 15:13:04Right? So it should we should be doing
  19388. 15:13:06these steps first
  19389. 15:13:09then doing modeling which we know right
  19390. 15:13:12we we've done that uh many times. We've
  19391. 15:13:14done our pre-processing and then we do
  19392. 15:13:15our modeling.
  19393. 15:13:19Any questions on that?
  19394. 15:13:37Okay.
  19395. 15:13:38Fantastic.
  19396. 15:13:42All right.
  19397. 15:13:45So, if we wanted to uh see if we wanted
  19398. 15:13:48to test this so far, um we could. So we
  19399. 15:13:51could run the pre-processing and
  19400. 15:13:53actually run a fit transform on our data
  19401. 15:13:56and this will um basically apply that
  19402. 15:13:59pipeline to the data. Now this would be
  19403. 15:14:01a sanity check. This is a good this is a
  19404. 15:14:04good kind of um this is a good sanity
  19405. 15:14:07check that our pre-processing
  19406. 15:14:12works. So it's doing what we expected to
  19407. 15:14:15do. It's not our final pipeline because
  19408. 15:14:17we don't have our model in there yet.
  19409. 15:14:19But this is just to ensure that all of
  19410. 15:14:21the features are kind of behaving as we
  19411. 15:14:23expect. So we can uh we can do that and
  19412. 15:14:27we can take a look at the um results.
  19413. 15:14:30This looks pretty good. This all of our
  19414. 15:14:32numerical features ended up scaled
  19415. 15:14:35which is pretty good. And we have one
  19416. 15:14:37hot encoded features for that ocean
  19417. 15:14:39proximity over here.
  19418. 15:14:42Okay. So this looks pretty this looks
  19419. 15:14:44reasonable of those steps being applied
  19420. 15:14:46to the right columns. But this is a good
  19421. 15:14:49kind of sanity check to just run our fit
  19422. 15:14:51transform on our data to ensure those
  19423. 15:14:55steps are actually happening and they
  19424. 15:14:57are. You can see here the result of the
  19425. 15:15:00scaling and the uh the one hot encoding.
  19426. 15:15:04So that that all looks pretty
  19427. 15:15:05reasonable,
  19428. 15:15:09right?
  19429. 15:15:11And uh what we should also do is make
  19430. 15:15:15sure there are no nulls in this which
  19431. 15:15:16there shouldn't be because we did the
  19432. 15:15:18imper. So we should be doing uh is na
  19433. 15:15:22dot
  19434. 15:15:25sum
  19435. 15:15:29and there is no nulls anymore. So that
  19436. 15:15:31looks pretty good right? Those got
  19437. 15:15:33filled in uh by doing our steps. our
  19438. 15:15:37pipeline steps executed really nicely on
  19439. 15:15:39our training data um and and we were off
  19440. 15:15:44and running. And there's nothing unique
  19441. 15:15:45about the training data. We could do
  19442. 15:15:47this to our test data as well
  19443. 15:15:50and verify that those steps are running
  19444. 15:15:52and they would, right? There's nothing
  19445. 15:15:54really that special about running it on
  19446. 15:15:55the training data. Um it should also
  19447. 15:15:59work on the test features as well and it
  19448. 15:16:00does. You can check that for yourself.
  19449. 15:16:05Okay.
  19450. 15:16:10All right. So, that's pretty cool. We
  19451. 15:16:12can uh verify all that's working.
  19452. 15:16:21Any questions on that?
  19453. 15:16:24We're almost there with our full
  19454. 15:16:25pipeline. This this is this is not the
  19455. 15:16:28full pipeline, but this is something
  19456. 15:16:30that will run during our full pipeline.
  19457. 15:16:32Of course, our features are going to be
  19458. 15:16:34transformed according to those steps and
  19459. 15:16:36then it will be uh put into our model to
  19460. 15:16:38either predict or train with. Um
  19461. 15:16:42so let's do that. Let's actually build
  19462. 15:16:44out our final uh model here. So it's
  19463. 15:16:48actually going to be really easy to do.
  19464. 15:16:50All we need to do is um put in our
  19465. 15:16:53model. So here we're going to import the
  19466. 15:16:56ridge model here. Now, we could use any
  19467. 15:16:59we could use linear regression, we could
  19468. 15:17:00use lasso, we could use elastic net. Um,
  19469. 15:17:03we're just going to use ridge um uh um
  19470. 15:17:07just to test it out. And um we are going
  19471. 15:17:11to uh now put in a final pipeline. So,
  19472. 15:17:15we're going to use our pipeline. And so,
  19473. 15:17:18we're going to create a new one here and
  19474. 15:17:21map our pre-processing to our
  19475. 15:17:23pre-processing that we've already built.
  19476. 15:17:25So this is a column transformer that
  19477. 15:17:27already has all of our steps. And then
  19478. 15:17:30notice what comes after it is just the
  19479. 15:17:32model. Now that's pretty pretty basic,
  19480. 15:17:34but it makes sense that it should come
  19481. 15:17:36after that model. Um and of course this
  19482. 15:17:39is a generic name. We could we can name
  19483. 15:17:41it whatever we want to. Um model ridge
  19484. 15:17:44is pretty reasonable um to because it is
  19485. 15:17:48a ridge uh regression. But uh of course
  19486. 15:17:51we could we could change that.
  19487. 15:17:55Okay. So that builds out our uh final um
  19488. 15:17:58pipeline. So now we have a pipeline. And
  19489. 15:18:02what's great about that is this signals
  19490. 15:18:05that all of these steps should be
  19491. 15:18:07completed prior to doing anything with
  19492. 15:18:09this model. So all of those processing
  19493. 15:18:12steps are going to run and then we're
  19494. 15:18:14going to do ffit or predict and that. So
  19495. 15:18:17that's really great. It ensures that all
  19496. 15:18:19those steps are running together every
  19497. 15:18:21single time we call.predict with this
  19498. 15:18:24with this model. So we're just going to
  19499. 15:18:26use the pipeline in place of the model
  19500. 15:18:30to ensure that all those steps are
  19501. 15:18:32running together. And this is our this
  19502. 15:18:35is kind of our final pipeline that we
  19503. 15:18:37would use uh with like something like
  19504. 15:18:39ffit or predict.
  19505. 15:18:42So let me make that a note of that. Now
  19506. 15:18:45we can use this final pipeline just like
  19507. 15:18:50a regular model i.e. pipeline.fit
  19508. 15:18:56or pipeline.predict.
  19509. 15:19:00So we could use it in ei in either
  19510. 15:19:01fashion uh to to train the pipeline
  19511. 15:19:06would be this guy and then use the
  19512. 15:19:08pipeline to predict would be this. And
  19513. 15:19:10what we should realize is under the
  19514. 15:19:11hood, these steps are running first and
  19515. 15:19:13then we train it or these steps run
  19516. 15:19:16first then we use it for prediction.
  19517. 15:19:24Okay,
  19518. 15:19:26questions on that. Does that make sense
  19519. 15:19:29on this final pipeline here? It's just
  19520. 15:19:32now it it's really cool because we have
  19521. 15:19:34a pipeline
  19522. 15:19:36made up of a of a pipeline really,
  19523. 15:19:39right? a pipeline made up of a pipeline.
  19524. 15:19:41But that's scikitlearn allows you to do
  19525. 15:19:42that to compose pipelines in this way.
  19526. 15:19:46That's is pretty uh pretty uh normal
  19527. 15:19:49there.
  19528. 15:19:59Okay.
  19529. 15:20:01What I want to show you is we can
  19530. 15:20:03actually use this pipeline in a grid
  19531. 15:20:06search. So that's pretty amazing. We can
  19532. 15:20:08use this pipeline in any way we can use
  19533. 15:20:10a mo like a regular model. It's just
  19534. 15:20:13that now our pre-processing steps have
  19535. 15:20:16kind of been packaged together with our
  19536. 15:20:18model to ensure that they always run
  19537. 15:20:21anytime we do any processing with this
  19538. 15:20:22model. Um so for instance we can do a
  19539. 15:20:27grid search just like we did with a
  19540. 15:20:29regular with with just a model right
  19541. 15:20:31with just this. Um we can do the same
  19542. 15:20:34thing with the whole pipeline. Um, so
  19543. 15:20:37the only catch is that you want to make
  19544. 15:20:40sure in your grid you name things in the
  19545. 15:20:44appropriate way inside of your your uh
  19546. 15:20:46keys in your dictionary. So uh for
  19547. 15:20:49instance um inside of the grid uh we're
  19548. 15:20:54going to set up the alpha that would be
  19549. 15:20:56used with this ridge regression by
  19550. 15:20:59referencing its name. So this is model
  19551. 15:21:01ridge is this is the name of the model
  19552. 15:21:04inside of the pipeline. So you want to
  19553. 15:21:06make sure that goes first.
  19554. 15:21:08And then what scikitlearn does is it
  19555. 15:21:11recognizes parameters that belong with
  19556. 15:21:13this model by using a double underscore.
  19557. 15:21:17So the so you have underscore alpha um
  19558. 15:21:22here. So the double
  19559. 15:21:25uh underscore
  19560. 15:21:27signals a parameter
  19561. 15:21:31belonging to model ridge in the in the
  19562. 15:21:37pipeline.
  19563. 15:21:39Okay, so we have a model ridge is just a
  19564. 15:21:42reference to the model in our pipeline.
  19565. 15:21:44That's the one we're going to test out
  19566. 15:21:46these parameters with. and
  19567. 15:21:48underscore_pha is just a way to say this
  19568. 15:21:51alpha belongs to this model. Okay, it
  19569. 15:21:56belongs so it's going to be used with
  19570. 15:21:58that model in our pipeline. Um otherwise
  19571. 15:22:02it's going to work exactly the same way.
  19572. 15:22:04It's just we need to line up this naming
  19573. 15:22:06convention of of scikitlearn.
  19574. 15:22:08You just have to reference this to
  19575. 15:22:11whatever name you provided here and then
  19576. 15:22:14underscore parameter. So L1 ratio alpha
  19577. 15:22:18whatever right would go there.
  19578. 15:22:21Okay. So there is a range from 0.1 to2
  19579. 15:22:26uh step size of 0.1
  19580. 15:22:29um and then we do our grid search CV. So
  19581. 15:22:32this is exactly the same setup as we had
  19582. 15:22:34before. It's just that our model is now
  19583. 15:22:38the pipeline. So our pipeline is going
  19584. 15:22:40in there. Um we have our grid going in
  19585. 15:22:43there. we have our scoring is the same,
  19586. 15:22:46you know, negative absolute error. Um,
  19587. 15:22:48we're using five-fold cross validation
  19588. 15:22:51and we're parallelizing that search. Um,
  19589. 15:22:54so we're going to search through these
  19590. 15:22:55alphas and uh basically fit this to our
  19591. 15:23:00um data and find the best um find the
  19592. 15:23:05best alpha.
  19593. 15:23:07So, it's going to try out all those
  19594. 15:23:08combinations and try to come up with the
  19595. 15:23:11best alpha.
  19596. 15:23:14So looks like the best alpha was 0.1 for
  19597. 15:23:17the ridge.
  19598. 15:23:19Okay. Is the best. So then um if we
  19599. 15:23:24wanted to we could uh then predict using
  19600. 15:23:28the model um which would be doing
  19601. 15:23:31something like this. Um, and we could
  19602. 15:23:34also go back and do something like so we
  19603. 15:23:38could
  19604. 15:23:42now use um this param. So we could do
  19605. 15:23:47model
  19606. 15:23:50um equals ridge
  19607. 15:23:55and then we could put in our alpha.
  19608. 15:23:58Um, alpha is our results, our best
  19609. 15:24:01parameters, and then we get that model
  19610. 15:24:03ridge alpha. And then we just rebuild
  19611. 15:24:05our our pipeline
  19612. 15:24:10equals um pipeline and then we uh put in
  19613. 15:24:14this new model here. So we could do
  19614. 15:24:16this. This would be going back and just
  19615. 15:24:20um putting in our best alpha here for
  19616. 15:24:24this model and then uh ensuring that's
  19617. 15:24:27part of our our pipeline. So we're just
  19618. 15:24:28overwriting that pipeline with the best
  19619. 15:24:30model there
  19620. 15:24:36to get the best model in our pipeline.
  19621. 15:24:42Okay,
  19622. 15:24:44so that's all this is doing is just
  19623. 15:24:46initializing a new um let me actually I
  19624. 15:24:49can put this code in here.
  19625. 15:24:56This is actually just getting this is
  19626. 15:24:58just getting a model with the best alpha
  19627. 15:25:00and then reinserting that into our our
  19628. 15:25:03uh we're just overwriting our final
  19629. 15:25:05pipeline there with the best model that
  19630. 15:25:07we have.
  19631. 15:25:10So pretty cool that pipeline can be used
  19632. 15:25:13basically exactly like a model, right?
  19633. 15:25:15It's it's going right here in the grid
  19634. 15:25:17search and being used uh entirely like a
  19635. 15:25:20basic model. So we do ffit
  19636. 15:25:23um and that allows us to use it. We
  19637. 15:25:26could do predict we could even do
  19638. 15:25:28pipeline.predict once we we could go
  19639. 15:25:30back and do final pipeline.fit
  19640. 15:25:33um with this and then final
  19641. 15:25:34pipeline.predict with this and evaluate
  19642. 15:25:41Okay,
  19643. 15:25:43so pretty cool that pipeline can be used
  19644. 15:25:45uh basically exactly like how a model
  19645. 15:25:47would be any way we'd use a model.f
  19646. 15:25:50model.pred predict we can use a
  19647. 15:25:51pipeline.
  19648. 15:25:54So grid search is for instance something
  19649. 15:25:56that can use a model in there. Um but
  19650. 15:25:59instead of just a model we're ensuring
  19651. 15:26:00we have our pre-processing steps kind of
  19652. 15:26:03bundled with that model in this
  19653. 15:26:04pipeline.
  19654. 15:26:06Any
  19655. 15:26:10questions on
  19656. 15:26:12uh this example so far?
  19657. 15:26:18Were you guys able to run it up to here?
  19658. 15:26:20Were you able to run the grid search?
  19659. 15:26:36Okay, great.
  19660. 15:26:45Okay.
  19661. 15:26:47Okay. So, this is this is uh just
  19662. 15:26:49showing you what's actually happening
  19663. 15:26:51underneath the hood is uh you know,
  19664. 15:26:54we're doing some scaling. We're doing
  19665. 15:26:56some one hot encoding. Um
  19666. 15:27:00and we're doing some uh we're doing a
  19667. 15:27:03model here. And that's all part of our
  19668. 15:27:06pipeline. Um, and then we can use the
  19669. 15:27:10pipeline however we want. So for
  19670. 15:27:11example, I know it's not here, but for
  19671. 15:27:14an example, we could use um once we do
  19672. 15:27:17once we have this final pipeline um we
  19673. 15:27:20can can use the final um
  19674. 15:27:25pipeline to predict. So we can do um
  19675. 15:27:29predictions
  19676. 15:27:31equals final
  19677. 15:27:34pipeline.predict
  19678. 15:27:36and then we can pass in our test data.
  19679. 15:27:38Now what happens on this is once we have
  19680. 15:27:41ran our our pipeline.fit we have a
  19681. 15:27:44trained pipeline and then when we run
  19682. 15:27:47this final pipeline.predict uh this data
  19683. 15:27:50is going to be transformed.
  19684. 15:27:52It's going to go through those
  19685. 15:27:53transformation steps and then we would
  19686. 15:27:55apply our model to it at the end uh to
  19687. 15:27:59to make those predictions and then we
  19688. 15:28:00can evaluate those predictions which is
  19689. 15:28:02what we're doing kind of here
  19690. 15:28:05right.
  19691. 15:28:12Okay.
  19692. 15:28:16All right. So in conclusion uh we have
  19693. 15:28:19gone through a lot of stuff here. Um,
  19694. 15:28:22we've gone through regression, we've
  19695. 15:28:24done the regularization on regression.
  19696. 15:28:27So hopefully we have a good foundation
  19697. 15:28:29on regression. Um, what we're going to
  19698. 15:28:31do in a little bit is actually do some
  19699. 15:28:33additional practice with regression on a
  19700. 15:28:35new problem. We're going to do a
  19701. 15:28:36capstone problem and do some additional
  19702. 15:28:39regression work with that. Um, so we'll
  19703. 15:28:43do that next. Um but the other thing we
  19704. 15:28:47learned is how to evaluate the
  19705. 15:28:48regression using things like mean
  19706. 15:28:49squared error, RMSSE which is square
  19707. 15:28:52root of that. Um which is which is
  19708. 15:28:55really cool. So we have a sense of that
  19709. 15:28:58error which is our distance from our
  19710. 15:29:00prediction to the actual value. That's
  19711. 15:29:02always what these uh that's always what
  19712. 15:29:05these things are doing like this, right?
  19713. 15:29:08This mean absolute error metric from
  19714. 15:29:10scikitlearn is computing the average
  19715. 15:29:13distance from these predictions to these
  19716. 15:29:15test labels that we have right those
  19717. 15:29:18actual values. Um and that gives us a
  19718. 15:29:20sense on average how far away are our
  19719. 15:29:23predictions
  19720. 15:29:24um to see how good of a model that we
  19721. 15:29:27have, right? And we should be evaluating
  19722. 15:29:30that error generally
  19723. 15:29:32um against the scale of our targets to
  19724. 15:29:36see, you know,
  19725. 15:29:38uh how far off we typically are.
  19726. 15:29:43Okay. Any questions at all on this
  19727. 15:29:45lesson on regression? Uh anything we
  19728. 15:29:47covered up to this point?
  19729. 15:29:49We're going to do some more practice
  19730. 15:29:50with the next. We'll do the capstone. So
  19731. 15:29:54we get So we just do some more
  19732. 15:29:56regression problems.
  19733. 15:30:07Yeah, it's a that's another bad score.
  19734. 15:30:09It's a little bit hard to interpret this
  19735. 15:30:10though because it's m ae. Um so one
  19736. 15:30:14thing we could do is is compute mean
  19737. 15:30:17squared error and then take the square
  19738. 15:30:19root of it to get the RMSSE which is a
  19739. 15:30:22much better uh evaluation metric in
  19740. 15:30:25terms of our target. Um so we could
  19741. 15:30:28actually run that. Uh if we go back here
  19742. 15:30:31and um we could generate for instance we
  19743. 15:30:34could generate the MSE which is the mean
  19744. 15:30:37squared
  19745. 15:30:39error
  19746. 15:30:41and it's it's the same exact function uh
  19747. 15:30:44of using our predictions
  19748. 15:30:49um
  19749. 15:30:50and then we could just print that out
  19750. 15:30:52mean squared error.
  19751. 15:30:56So we have mean squared error and then
  19752. 15:30:58what we can do is let's take the um MP.
  19753. 15:31:03Square root of that.
  19754. 15:31:06So that way we can generate the RMSSE.
  19755. 15:31:08So yeah that I mean that's pretty bad.
  19756. 15:31:10That's uh pretty bad. Uh now let's let's
  19757. 15:31:15go back and look at our
  19758. 15:31:18uh data though. So let's take a look at
  19759. 15:31:20the average for our Y. Um remember one
  19760. 15:31:23thing we should be doing is taking a
  19761. 15:31:25look at um what our uh let's take a look
  19762. 15:31:28at y test mean
  19763. 15:31:31to get an average value. So the average
  19764. 15:31:34value is in the 200,000s. So
  19765. 15:31:38this isn't this isn't awful. This is
  19766. 15:31:4170,000. It's still a decent amount of
  19767. 15:31:43error. It's not as bad as the models we
  19768. 15:31:45have before though, right? This is an
  19769. 15:31:47average
  19770. 15:31:49median price of the house is in the
  19771. 15:31:5226,000 range and our error is off by
  19772. 15:31:56like 70,000,
  19773. 15:31:59right?
  19774. 15:32:01So, it's not good. Um, but it's not
  19775. 15:32:07hor like as bad as the it's not as
  19776. 15:32:10horrible as we've seen so far. Right.
  19777. 15:32:12This is a little bit better of a model.
  19778. 15:32:14A little bit better. closer to zero
  19779. 15:32:16would be better, right? Um but the
  19780. 15:32:19smaller the better. But uh remember this
  19781. 15:32:22is the um these even the mean absolute
  19782. 15:32:26error is is technically in similar units
  19783. 15:32:29as the as the uh um
  19784. 15:32:33as the target. So 50,000 60,000 here
  19785. 15:32:3770,000 it's still a decent amount of
  19786. 15:32:39error in terms of 200,000.
  19787. 15:32:44Uh so far we only come up with models
  19788. 15:32:45and test their accuracy with available
  19789. 15:32:47data. We haven't used a model to make
  19790. 15:32:48completely new predictions on No, we
  19791. 15:32:50haven't done that. Uh except we know how
  19792. 15:32:53to do that. Um it would so to make
  19793. 15:32:56predictions on new data would be exactly
  19794. 15:32:58how we're making them on our available
  19795. 15:33:00data because we actually do that all the
  19796. 15:33:03time. If we go back down to our model
  19797. 15:33:05building,
  19798. 15:33:07um it's it looks just like this, right?
  19799. 15:33:09where we take so for instance we do
  19800. 15:33:13predictions all the time on test data
  19801. 15:33:15that was never involved in the training.
  19802. 15:33:18So it's it's as if this data mimics new
  19803. 15:33:22data that we've never seen before. So if
  19804. 15:33:25we had new raw data it would just it
  19805. 15:33:28would be the same exact process. the new
  19806. 15:33:31now with our pipeline it makes it a
  19807. 15:33:33little bit easier because with the
  19808. 15:33:35pipeline
  19809. 15:33:36um the raw data will go through those
  19810. 15:33:39transformations which it should right
  19811. 15:33:41the raw data should because if it's
  19812. 15:33:42missing data it needs to be filled in if
  19813. 15:33:44it has categoricals it needs to be one
  19814. 15:33:46hot encoded so that's the purpose of the
  19815. 15:33:49pipeline actually is to make sure that
  19816. 15:33:53if we're dealing with raw data um those
  19817. 15:33:56steps can happen on the data before it
  19818. 15:34:00goes into the model. Right? So we so
  19819. 15:34:05that's kind of the purpose of the
  19820. 15:34:07pipeline
  19821. 15:34:08is to ensure that we run those steps
  19822. 15:34:11ahead of using it using a model with it.
  19823. 15:34:15But but ultimately that's how it uh any
  19824. 15:34:18scikitlearn model is going to be doing
  19825. 15:34:20the predict even if it's a pipeline
  19826. 15:34:22right it's going to be uh we just go
  19827. 15:34:25back down here it's going to be um
  19828. 15:34:27predict it's always going to be that on
  19829. 15:34:30new data
  19830. 15:34:32yeah
  19831. 15:34:36okay
  19832. 15:34:39really good question uh what I wanted to
  19833. 15:34:41do next was do some practice um I wanted
  19834. 15:34:45to go over to the capstone session five.
  19835. 15:34:50So, in this course, we have some more
  19836. 15:34:52capstone sessions. So, if you have a
  19837. 15:34:54moment, you want to pull up those
  19838. 15:34:56capstone session materials, the
  19839. 15:34:58incremental capstone session materials.
  19840. 15:35:00Um, we're going to be doing session five
  19841. 15:35:03today. So, this is just um remember it's
  19842. 15:35:06just extra practice that we do after
  19843. 15:35:08we've covered some concepts. So we are
  19844. 15:35:11going to do some regression practice now
  19845. 15:35:14that we've uh covered regression um
  19846. 15:35:16pretty fully and then um this will this
  19847. 15:35:19will be good practice before we head
  19848. 15:35:21into lesson four on classification. So I
  19849. 15:35:24just want to do this practice now while
  19850. 15:35:26it's fresh while the material is kind of
  19851. 15:35:28fresh in our in our minds. Um do you
  19852. 15:35:31guys have the capstone materials? Do you
  19853. 15:35:34know where to get it in your LMS? It's
  19854. 15:35:36in your LMS and the resources the uh
  19855. 15:35:39capstone materials you want to download
  19856. 15:35:41that so you can get the the data um and
  19857. 15:35:45the the slides for the instructions
  19858. 15:35:48right the or PDF I think for you guys
  19859. 15:35:52but uh let me ask you do you have those
  19860. 15:35:55we want to pull up session five if you
  19861. 15:35:57have it
  19862. 15:36:01thank you I was just going to share that
  19863. 15:36:03appreciate that yeah so this is going to
  19864. 15:36:05be session Question five. Um, you're
  19865. 15:36:07also going to want the data that we're
  19866. 15:36:11going to use with this, which is going
  19867. 15:36:12to be the, uh, bike rental data set.
  19868. 15:36:15I'll share that with you guys now.
  19869. 15:36:18So, we're going to be using this bike
  19870. 15:36:19rentals data set for this uh, for this
  19871. 15:36:22capstone. Um, it's the one we're going
  19872. 15:36:24to build a regression model off of.
  19873. 15:36:29Okay. So, you should have that one from
  19874. 15:36:31the the capstones data sets as well.
  19875. 15:36:35All right. So, let's go through this.
  19876. 15:36:37Um, we are going to be doing uh machine
  19877. 15:36:41learning here. So, we're going to be
  19878. 15:36:42doing uh so we're talking about that
  19879. 15:36:45kind of example product that Aura
  19880. 15:36:48product that was in our original
  19881. 15:36:49capstone. Um, and in order to do uh to
  19882. 15:36:54to to um make decisions, it's going to
  19883. 15:36:57have to build some models. Um, in this
  19884. 15:37:00case, it's going to be doing some bike
  19885. 15:37:02rental modeling. um which is the data
  19886. 15:37:04set we have. So we're moving away from
  19887. 15:37:06that healthcare data set going into this
  19888. 15:37:08bike rental data set as an example of
  19889. 15:37:10the capabilities here. Um so we're going
  19890. 15:37:14to do this first capstone. Um after we
  19891. 15:37:18do classification, we'll do this
  19892. 15:37:20practice. Um after we do unsupervised
  19893. 15:37:23learning, we'll do this practice on
  19894. 15:37:24clustering. And then after we do
  19895. 15:37:26recommendation uh which is the last
  19896. 15:37:29lesson um we'll come back and do
  19897. 15:37:31practice with building a recommendation
  19898. 15:37:32engine. Okay, but we're going to do this
  19899. 15:37:34one today. Um and then we will uh do
  19900. 15:37:39these other capstones as we go along. So
  19901. 15:37:41this will be session six, session seven,
  19902. 15:37:44and session 8
  19903. 15:37:46uh later on in the course. Okay.
  19904. 15:37:51Um,
  19905. 15:37:56okay. A little bit about the data. So,
  19906. 15:37:58uh, in this capstone, we're going to be
  19907. 15:38:00working with a, um, a shop, like a
  19908. 15:38:03retail shop that rents out bikes. And
  19909. 15:38:06they have data, um, on a per day basis
  19910. 15:38:10with, um, actually on a per hour basis
  19911. 15:38:13on the number of bikes that they rented
  19912. 15:38:15in every hour. Um, so maybe one hour
  19913. 15:38:18they rented out 20 bikes, another hour
  19914. 15:38:20they rented out 30 bikes. Um, another
  19915. 15:38:23hour they rented out 15. So they have
  19916. 15:38:26that data here in in the CSV. Um, they
  19917. 15:38:29have other kinds of data like the
  19918. 15:38:31environmental data like what the
  19919. 15:38:33temperature was at that hour, the
  19920. 15:38:35humidity, if there was snowfall if it's
  19921. 15:38:37a holiday, the wind, the visibility, the
  19922. 15:38:41due point, um, solar radiation,
  19923. 15:38:43rainfall, uh, what season it was. um
  19924. 15:38:47what uh what day it was like
  19925. 15:38:51um the functional is like if it's uh I
  19926. 15:38:54believe it's like if it's a weekend or
  19927. 15:38:55weekday um which would be uh
  19928. 15:38:58nonfunctional
  19929. 15:39:00um so
  19930. 15:39:03based on those features we have a bunch
  19931. 15:39:05of tasks okay so we have um based on the
  19932. 15:39:11uh rented by count hour of the day um
  19933. 15:39:16temperature, humidity, wind speed,
  19934. 15:39:17rainfall, and whatever other features
  19935. 15:39:19we're that are in the data set. We're
  19936. 15:39:21actually going to build a model to
  19937. 15:39:22predict the bike count required for
  19938. 15:39:24every hour to have a stable supply of
  19939. 15:39:27rented bikes. So, our goal is actually
  19940. 15:39:28going to be to predict the bike rental
  19941. 15:39:31count per hour. Um and uh we are going
  19942. 15:39:36to do that using all the features we
  19943. 15:39:38have at our disposal. um like mostly
  19944. 15:39:41those environmental features and what
  19945. 15:39:43day it is, those kind of things. Um
  19946. 15:39:47so we're going to load our data. We're
  19947. 15:39:49going to do our usual check. So check
  19948. 15:39:51for any NLES, handle those missing nles.
  19949. 15:39:54Um we're actually going to practice
  19950. 15:39:56converting our date because we actually
  19951. 15:39:58do have things based on a date here. So
  19952. 15:40:00we can convert it over to a datetime
  19953. 15:40:02object, extract different uh features
  19954. 15:40:05from that um like the month or the day
  19955. 15:40:09of the week. Um we're going to check uh
  19956. 15:40:12correlation using the heat map. We're
  19957. 15:40:14going to do some plots, some very basic
  19958. 15:40:16plots. The focus is going to be on the
  19959. 15:40:18modeling. So, I probably won't spend too
  19960. 15:40:20much time on the plots today, but um
  19961. 15:40:23there's some plot tasks in here like the
  19962. 15:40:25uh the the histogram of the bike count,
  19963. 15:40:27the histogram of the numerical features.
  19964. 15:40:30Um box plot of the bikes against the
  19965. 15:40:33categoricals.
  19966. 15:40:35Um so we can do some plots like that
  19967. 15:40:38from Seabor for instance. Uh Seabor
  19968. 15:40:40category plot of rented bike count
  19969. 15:40:42against features like hour, holiday,
  19970. 15:40:44rainfall, snowfall. Um, so we can see
  19971. 15:40:47how that stacks up against like
  19972. 15:40:48different hours of the day, different
  19973. 15:40:50holidays, rainfall, different weather
  19974. 15:40:52events. Um, then we're going to do then
  19975. 15:40:55we're going to start building our model,
  19976. 15:40:56right? So encode our categorical
  19977. 15:40:58features. Um, identify target variable
  19978. 15:41:01and do the split and then do scaling and
  19979. 15:41:04do three different models. So we're
  19980. 15:41:06actually going to build a linear
  19981. 15:41:07regression. We're going to build a lasso
  19982. 15:41:09regression and build a ridge regression
  19983. 15:41:11for the hourly bike count. and we're
  19984. 15:41:15going to see which model performs the
  19985. 15:41:16best. Now,
  19986. 15:41:19we could and should um build this into a
  19987. 15:41:24pipeline. So, that could be something we
  19988. 15:41:27practice. Um but this initial
  19989. 15:41:31instructions actually doesn't require
  19990. 15:41:33doing that, but I think it's really good
  19991. 15:41:34practice to um build our model. So, we
  19992. 15:41:38could use git dummies. Like it says
  19993. 15:41:40here, hint to use git dummies. We could
  19994. 15:41:42do that and build our model that way,
  19995. 15:41:45but more practical would be doing the
  19996. 15:41:47steps we did towards the end of lesson
  19997. 15:41:49three, which is um actually putting
  19998. 15:41:51everything together into a pipeline. All
  19999. 15:41:53right, that'd be more practical and then
  20000. 15:41:55fitting the pipeline and predicting with
  20001. 15:41:57it um for evaluation.
  20002. 15:42:01So, uh we'll do that. I think we'll do
  20003. 15:42:04that instead because I think that'll be
  20004. 15:42:07more practical. The the pipelines are
  20005. 15:42:09really uh useful. So the things we'll
  20006. 15:42:12have to do when we build our pipeline
  20007. 15:42:13will be um making sure we handle the
  20008. 15:42:16missing values. So we'll want that imper
  20009. 15:42:18in there for numerical features. We'll
  20010. 15:42:20want our one hot encoding very similar
  20011. 15:42:22pipeline to the one we built earlier. Um
  20012. 15:42:26and then we'll want to basically map
  20013. 15:42:27those to the right columns using the
  20014. 15:42:28column transformer. And then um we'll
  20015. 15:42:31have our pipeline ready to go for
  20016. 15:42:33training and prediction.
  20017. 15:42:36Right.
  20018. 15:42:38Okay. So, those are going to be our uh
  20019. 15:42:41steps. Any questions on this before we
  20020. 15:42:45kind of get started on it.
  20021. 15:42:54Okay. So, let me go over to the
  20022. 15:42:58notebooks and let me actually do a new
  20023. 15:43:02notebook.
  20024. 15:43:08And I'm going to name this uh
  20025. 15:43:12capstone
  20026. 15:43:14session five.
  20027. 15:43:18Okay.
  20028. 15:43:20So, I'm going to come back here and
  20029. 15:43:22reference the steps here. Okay. So the
  20030. 15:43:24steps are to load our data set and
  20031. 15:43:27basically check for nles.
  20032. 15:43:30So we should be pretty adept at doing
  20033. 15:43:32that. We're going to import pandas as
  20034. 15:43:35pd.
  20035. 15:43:38And I need to make sure the data sets
  20036. 15:43:40available. So I need to
  20037. 15:43:45uh load that here. So we're going to use
  20038. 15:43:49our bike rental.
  20039. 15:44:04Florida bike rentals.
  20040. 15:44:22And then look at the first five rows.
  20041. 15:44:34Uh, I got a decoder error.
  20042. 15:44:39UTF8 codec can decode by in
  20043. 15:44:58one moment.
  20044. 15:45:27I think we have an error in the data
  20045. 15:45:29set. Was anybody able to get this to
  20046. 15:45:30run?
  20047. 15:45:32Hopefully they
  20048. 15:45:35or did you get the same error as me?
  20049. 15:45:50giving me the same error.
  20050. 15:45:58I think we need to set a
  20051. 15:46:14encoding
  20052. 15:46:37having other issues. What other issues?
  20053. 15:46:41Sorry, I think this data is kind of
  20054. 15:46:43corrupted.
  20055. 15:46:50I may need to open it externally.
  20056. 15:46:55Okay, let me open it. There may be just
  20057. 15:46:58a bad character that needs to be
  20058. 15:47:01removed.
  20059. 15:47:14Okay,
  20060. 15:47:22let me try a different Let me try
  20061. 15:47:24something real quick.
  20062. 15:47:31I think the data is needs to be updated.
  20063. 15:47:37Oh, saved it as the wrong file.
  20064. 15:47:48One moment.
  20065. 15:48:04Okay, let me try uploading this.
  20066. 15:48:16Okay, there. That worked better. So, let
  20067. 15:48:18me give you the data. I think it was uh
  20068. 15:48:25yeah, I think it was that the degree
  20069. 15:48:28code was giving it some issues. So, I
  20070. 15:48:32actually just removed it.
  20071. 15:48:39You can do that or you could just work
  20072. 15:48:40with this one.
  20073. 15:48:43I removed the temperature had a strange
  20074. 15:48:46like degree symbol that wasn't being
  20075. 15:48:47parsed.
  20076. 15:48:51So, I just removed that in this data
  20077. 15:48:53set. This one should work. The one I
  20078. 15:48:55just sent you guys should work. Or yeah,
  20079. 15:48:57I guess you could try the CP1252
  20080. 15:49:01with the original data. See if that
  20081. 15:49:02works. Did that work for you?
  20082. 15:49:14Okay, it works with that encoding. So
  20083. 15:49:16yeah, you could use that encoding or
  20084. 15:49:19uh
  20085. 15:49:21remove that temperature degree which is
  20086. 15:49:23what I did from that. So I use that
  20087. 15:49:26other data set.
  20088. 15:49:34Yeah, that opens the data but
  20089. 15:49:38it's not it's not in a data frame.
  20090. 15:49:43We just we want it in a dataf frame to
  20091. 15:49:45work with our models and doing all of
  20092. 15:49:46our Yeah. Like that opens the file. If
  20093. 15:49:49it just if it was a text file that's
  20094. 15:49:51fine, but it's not in a data frame. Want
  20095. 15:49:54in a structured data frame so we could
  20096. 15:49:56use it.
  20097. 15:50:04Okay. So, one of those methodologies
  20098. 15:50:07hopefully works. So you can either work
  20099. 15:50:08with the file I sent and read it like
  20100. 15:50:10this or you can use the encoding.
  20101. 15:50:13Did you guys were other people able to
  20102. 15:50:15open it once they change the encoding?
  20103. 15:50:29Okay.
  20104. 15:50:42Okay, very good.
  20105. 15:50:46Let's see what we have. Let's do
  20106. 15:50:47df.info.
  20107. 15:50:48Let's see what we have, which is pretty
  20108. 15:50:50standard step to do once we first load
  20109. 15:50:52in some data. Um, so we have about 14
  20110. 15:50:56columns here. We have a date and then we
  20111. 15:51:00have um which is an object right now but
  20112. 15:51:03we're actually going to convert that
  20113. 15:51:04over to a datetime object in a minute.
  20114. 15:51:07Um we have our bike count which is what
  20115. 15:51:09we want to ultimately predict. This is
  20116. 15:51:11the target variable in our data.
  20117. 15:51:13Remember that's going to be the problem
  20118. 15:51:14is to predict that hourly bike count. Um
  20119. 15:51:18the hour of the day that we are
  20120. 15:51:20producing that bike count um is a
  20121. 15:51:22feature. the temperature which is in
  20122. 15:51:25degrees Celsius. Uh I I removed that
  20123. 15:51:28from there but that's it was in Celsius.
  20124. 15:51:31Um
  20125. 15:51:33and then we have a bunch of different
  20126. 15:51:35features which are uh a kind of boolean
  20127. 15:51:38like yes no holiday no holiday the
  20128. 15:51:42season. So these guys are going to be
  20129. 15:51:46good um candidates for
  20130. 15:51:50uh these are going to be good candidates
  20131. 15:51:52for um doing our uh one hot encoding.
  20132. 15:51:56Right? These are probably the three that
  20133. 15:51:58we should pick to do uh some
  20134. 15:52:01transformations to for our one encoding.
  20135. 15:52:08Okay.
  20136. 15:52:15Everything else is pretty numerical
  20137. 15:52:16though, so those should be fine to keep
  20138. 15:52:18those. But they would they're just going
  20139. 15:52:20to be good candidates for scaling,
  20140. 15:52:22right? Really good candidates for
  20141. 15:52:23scaling.
  20142. 15:52:26All right, let's go back to here.
  20143. 15:52:37And let's go back to the task.
  20144. 15:52:40So we loaded the data set. Um we're we
  20145. 15:52:43are going to check for nulls. It looks
  20146. 15:52:45like there's actually not any nles. So
  20147. 15:52:47there may not be any uh imputing that we
  20148. 15:52:49really need to do for this. Um but we
  20149. 15:52:53can check. So we can use is na or is
  20150. 15:52:57null and then do the sum.
  20151. 15:53:00Um looks like we don't have any. So that
  20152. 15:53:02that's good. There's no nulls.
  20153. 15:53:07That's pretty good. So we don't have to
  20154. 15:53:08worry about filling in any blanks
  20155. 15:53:10really.
  20156. 15:53:12Um but you know that would normally um
  20157. 15:53:16that would be an important part of our
  20158. 15:53:18pipeline right is filling in any nles.
  20159. 15:53:21Looks like we don't have to worry about
  20160. 15:53:22that here.
  20161. 15:53:24So that's good.
  20162. 15:53:29So we can
  20163. 15:53:32uh
  20164. 15:53:34we can extract we can do the date uh
  20165. 15:53:38extraction. And now one thing that we're
  20166. 15:53:39going to do is uh create purposely for
  20167. 15:53:43this we're going to create a weekend or
  20168. 15:53:45weekday feature. Okay, weekday or
  20169. 15:53:50weekend which should be really easy to
  20170. 15:53:52do from the day of the week. um which we
  20171. 15:53:55should be able to extract from this uh
  20172. 15:53:58from the date.
  20173. 15:54:00Let's go back to Were you by the way,
  20174. 15:54:02were you guys able to run this?
  20175. 15:54:04Just want to make sure everyone's with
  20176. 15:54:06me. Checking if there's any nulls. We
  20177. 15:54:09did that.
  20178. 15:54:26Okay.
  20179. 15:54:35All right. So, we're following along
  20180. 15:54:37there. We checked if there's any nles.
  20181. 15:54:39Let's do um our conversion. So, let's do
  20182. 15:54:43um df
  20183. 15:54:45uh date.
  20184. 15:54:48And this is going to be um
  20185. 15:54:51PD.2
  20186. 15:54:53date time and then df
  20187. 15:54:56date.
  20188. 15:54:59All right. And then let's sanity check
  20189. 15:55:01that that it got converted by doing
  20190. 15:55:03info.
  20191. 15:55:25Oh, we might need this format.
  20192. 15:55:43Um, maybe we should do
  20193. 15:55:51Okay, so that actually worked. Let's
  20194. 15:55:54see. Let's double check the
  20195. 15:55:57date.
  20196. 15:55:59Okay, so that worked. It extracted it
  20197. 15:56:01into
  20198. 15:56:03uh it extracted it into the right dates.
  20199. 15:56:09Some weird encoding with this
  20200. 15:56:17So it goes all the way up to 2018
  20201. 15:56:20from
  20202. 15:56:22the very beginning data is in 2017 the
  20203. 15:56:24beginning of the year right. So it go
  20204. 15:56:27and then the tail is all the way in
  20205. 15:56:292018.
  20206. 15:56:32This this extracts the date.
  20207. 15:56:41If you guys want to run that
  20208. 15:56:53you guys able to run this
  20209. 15:56:57to extract the date. Okay, perfect. So,
  20210. 15:57:00this extracts the date. By the way, the
  20211. 15:57:02reason we need this is because our date
  20212. 15:57:04formats are not uniform. Uh they were
  20213. 15:57:07actually in different encodings. So some
  20214. 15:57:08of them had the year uh the string was
  20215. 15:57:11in a slightly different format where the
  20216. 15:57:13year was last or the year was first. So
  20217. 15:57:16if we do format mixed, it kind of
  20218. 15:57:17rearrang it kind of puts it in a uniform
  20219. 15:57:20arrangement with the year. It's year,
  20220. 15:57:22month, date, but it parses that out um
  20221. 15:57:26correctly. Uh so we have year, month,
  20222. 15:57:29date in there.
  20223. 15:57:31Um but the this strings were in a mixed
  20224. 15:57:34format. So we uh put that argument in
  20225. 15:57:38there to handle that case.
  20226. 15:57:43Okay. So the reason we're going to do
  20227. 15:57:45that is that we should be able to create
  20228. 15:57:48a day of the week uh feature. Um so
  20229. 15:57:52let's actually do that and add it to our
  20230. 15:57:54data frame. So we should be able to
  20231. 15:57:55create um day of week
  20232. 15:58:00And we should be able to extract um
  20233. 15:58:05the DF
  20234. 15:58:06date. And then we do our usual DT dot um
  20235. 15:58:12day
  20236. 15:58:14of week.
  20237. 15:58:19Okay. So, and then let's see what that
  20238. 15:58:21does. So, if we add that, let's actually
  20239. 15:58:24see um let's see what our new features
  20240. 15:58:27are.
  20241. 15:58:29So we add that it should go onto the end
  20242. 15:58:32of the data frame as the day of the week
  20243. 15:58:35is a numerical day of the week. So this
  20244. 15:58:37this is the uh um
  20245. 15:58:41this looks like a th or no a Wednesday.
  20246. 15:58:44I think that's the third day.
  20247. 15:58:48Sunday Monday being uh Monday being
  20248. 15:58:51actually this would be a Thursday. I
  20249. 15:58:52think Monday would be zero.
  20250. 15:58:56This would be Thursday.
  20251. 15:59:00Okay. So, we extract that day of the
  20252. 15:59:02week.
  20253. 15:59:07So, it's just we're adding a new column
  20254. 15:59:09called day of week, which extracts the
  20255. 15:59:12day of the week from the date feature,
  20256. 15:59:14the datetime feature, uh the day of the
  20257. 15:59:18week.
  20258. 15:59:23And we're verifying that here. It's now
  20259. 15:59:25a new feature called day of the week.
  20260. 15:59:27which is a number.
  20261. 15:59:40Does that make sense what this is doing?
  20262. 15:59:48Yeah, it's a Thursday. So, I think yeah,
  20263. 15:59:50Monday is zero
  20264. 15:59:53and Sunday is six. So,
  20265. 15:59:57um, Monday is
  20266. 15:59:59zero, Sunday is
  20267. 16:00:03six.
  20268. 16:00:05Yeah. So, this should be a Thursday.
  20269. 16:00:14This first five rows is a Thursday. And,
  20270. 16:00:16and by the way, the data um, this is all
  20271. 16:00:18on the Thursday. This is a different
  20272. 16:00:20hours of the day.
  20273. 16:00:23Different hours of the day. Um, and we
  20274. 16:00:27the the thing that we don't know about
  20275. 16:00:28this is the necessarily the time zone.
  20276. 16:00:30So, it may seem strange that like this
  20277. 16:00:32is zero, which is kind of like midnight.
  20278. 16:00:35Um, it could be it could be in a
  20279. 16:00:37different time zone. So, we don't know
  20280. 16:00:38that uh necessarily, but the there is um
  20281. 16:00:43you know 250 bikes, 200 bikes, 173. It
  20282. 16:00:47starts to decrease over these hours.
  20283. 16:00:51Just kind of notice that.
  20284. 16:01:11Okay. So, by the way, uh we should be
  20285. 16:01:14able to create a new feature. So let's
  20286. 16:01:16create
  20287. 16:01:18the weekend feature
  20288. 16:01:21which is um we can create using
  20289. 16:01:25uh weekend
  20290. 16:01:28and we can take our DF
  20291. 16:01:32um day of week
  20292. 16:01:37and then we can just um
  20293. 16:01:40say is this um greater than or equal to
  20294. 16:01:45uh greater than or equal to five
  20295. 16:01:50cuz that would be five or six. And we're
  20296. 16:01:52going to
  20297. 16:01:54um put this as
  20298. 16:02:11actually. Let's just leave it like that.
  20299. 16:02:12That should be fine.
  20300. 16:02:17No, let's change it as type
  20301. 16:02:20uh int.
  20302. 16:02:25So, let's see this feature.
  20303. 16:02:35Let's see if this works.
  20304. 16:02:44So this is not a weekend uh because it's
  20305. 16:02:47a Thursday, right? So anything bigger
  20306. 16:02:49than five would be bigger than or equal
  20307. 16:02:51to five would be five or six which would
  20308. 16:02:53be Saturday, Sunday. Um so we know it's
  20309. 16:02:56not a weekend day. This is uh zero.
  20310. 16:03:02Okay.
  20311. 16:03:09So just creating those features there
  20312. 16:03:14and I can paste these in. Does that make
  20313. 16:03:16sense what what I just did?
  20314. 16:03:19Any questions on that code? And this is
  20315. 16:03:22important. If you don't have this um it
  20316. 16:03:25this will be a true or a false, but we
  20317. 16:03:28want it to be a zero or a one. So when
  20318. 16:03:31it's actually true, this should be a
  20319. 16:03:33one. When it's false, it'll be a zero.
  20320. 16:03:36So, we want it to be an integer rather
  20321. 16:03:38than a true false. So, that's why I have
  20322. 16:03:40this part. That's why I did this here
  20323. 16:03:44to make sure it's um make sure it's an
  20324. 16:03:46integer.
  20325. 16:04:01Good.
  20326. 16:04:05Okay,
  20327. 16:04:07so we're pretty much uh doing that and
  20328. 16:04:09we convert it to day and extract day. Um
  20329. 16:04:14let's
  20330. 16:04:16check our heat map.
  20331. 16:04:20Let's do that.
  20332. 16:04:26So let's import Seabour
  20333. 16:04:30as SNS.
  20334. 16:04:33So, we're going to do our heat map next.
  20335. 16:04:35Let me uh make a note of that. So, we're
  20336. 16:04:38going to do
  20337. 16:04:40heat mapap
  20338. 16:04:42heat map. Um,
  20339. 16:04:45so we're going to do uh SNS
  20340. 16:04:50heat map
  20341. 16:04:52and then let's do uh df.correlation.
  20342. 16:04:56And let's make sure we do numeric
  20343. 16:05:00um only
  20344. 16:05:02equals to true
  20345. 16:05:10and then let's do annotate
  20346. 16:05:13equals to true
  20347. 16:05:15so that we get those uh correlation
  20348. 16:05:17values that are displayed on the heat
  20349. 16:05:19map. So what this is going to do is
  20350. 16:05:21create our heat map with our correlation
  20351. 16:05:23matrix um where we're making sure we
  20352. 16:05:26only do the numerical features of course
  20353. 16:05:28when we do uh the heat map.
  20354. 16:05:34Let's generate that. Um okay so we have
  20355. 16:05:39some
  20356. 16:05:41uh pretty mild um correlations. Now the
  20357. 16:05:45ones that are so the ones that are
  20358. 16:05:48correlated are the temperature and
  20359. 16:05:50dupoint temperature. Those are pretty
  20360. 16:05:52correlated.
  20361. 16:05:54Um so I think we could argue that we
  20362. 16:05:56should drop one of those. Probably just
  20363. 16:05:59the dupoint temperature we could drop.
  20364. 16:06:01Um that's a really high correlation
  20365. 16:06:04right right here.
  20366. 16:06:06Let me draw it in red. That's a this
  20367. 16:06:08this one here
  20368. 16:06:10which is the Dupoint temperature against
  20369. 16:06:12the regular temperature. That's super
  20370. 16:06:14high. 0.91
  20371. 16:06:16is nearly a onetoone correlation.
  20372. 16:06:20So, that's a good candidate to be
  20373. 16:06:22dropped. Uh, one of those guys, I would
  20374. 16:06:24argue probably the Dupoint temperature
  20375. 16:06:26we could get rid of dropping. Um, and
  20376. 16:06:30just keep the regular temperature
  20377. 16:06:31because they're nearly identical. Um, if
  20378. 16:06:36you go down here,
  20379. 16:06:38day of the week and weekend are
  20380. 16:06:40correlated. Um, and that makes sense.
  20381. 16:06:42That's a pretty strong correlation
  20382. 16:06:44because of course if it's depending on
  20383. 16:06:47what day of the week it is, it is the
  20384. 16:06:48weekend or not. So we could go and we
  20385. 16:06:52could go ahead and drop the day of the
  20386. 16:06:54week column if we wanted to because
  20387. 16:06:55we've already derived the weekend uh
  20388. 16:06:58feature which is a simpler feature. Is
  20389. 16:07:00it the weekend or is it not the weekend?
  20390. 16:07:02Um so we could probably drop day of week
  20391. 16:07:07and be okay. That's a pretty strong
  20392. 16:07:09correlation. Otherwise, it's all pretty
  20393. 16:07:12weak. I don't see any other strong
  20394. 16:07:14correlations
  20395. 16:07:16uh necessarily.
  20396. 16:07:18Um so
  20397. 16:07:21that seems pretty reasonable is that we
  20398. 16:07:23could get rid of we could get rid of
  20399. 16:07:24this one and we could get rid of day of
  20400. 16:07:27week and probably be okay,
  20401. 16:07:30right? Those are pretty strong
  20402. 16:07:32correlations. is 08 and 0.91.
  20403. 16:07:34Pretty strong.
  20404. 16:07:42Okay. Were you guys able to run that one
  20405. 16:07:43and see the the uh heat map? And does
  20406. 16:07:46that make sense based on what I'm
  20407. 16:07:48saying?
  20408. 16:07:52So, we want to run the heat map
  20409. 16:07:55and pass in that correlation. And we
  20410. 16:07:56want to make sure we turn this to true
  20411. 16:07:58to only do the numerical features. And
  20412. 16:08:00then this to true to show the value
  20413. 16:08:04on the uh on the heat map.
  20414. 16:08:25able to run that one.
  20415. 16:08:32Great.
  20416. 16:08:47Okay.
  20417. 16:08:54Um, that's a good question. Any
  20418. 16:08:56recommendation on how to choose between
  20419. 16:08:57the two? Uh,
  20420. 16:09:00not really. I think
  20421. 16:09:04I don't think it really matters. If
  20422. 16:09:05they're correlated to each other, then
  20423. 16:09:08uh including one of them,
  20424. 16:09:11um, including one of them should give
  20425. 16:09:13you the same information as the other,
  20426. 16:09:15especially if they're really correlated.
  20427. 16:09:17So, it doesn't really matter too much.
  20428. 16:09:19Um
  20429. 16:09:21the way that I would choose is to think
  20430. 16:09:24about like I'll give you an example in
  20431. 16:09:26this day of week versus weekend. Um we
  20432. 16:09:30could I would argue drop day of week
  20433. 16:09:34because the weekend is a simpler
  20434. 16:09:36feature. It's only is zero or one and
  20435. 16:09:39that's directly derived from day of
  20436. 16:09:40week.
  20437. 16:09:42So it has less um complexity to it. it's
  20438. 16:09:47a little bit simpler of a feature and I
  20439. 16:09:48think that makes it easier to work with.
  20440. 16:09:51Um,
  20441. 16:09:53but the truth is that uh we could do
  20442. 16:09:57both options and try them out and
  20443. 16:09:59evaluate the results, right? So, well to
  20444. 16:10:02be to be truly thorough, what we could
  20445. 16:10:05do is build a model where we've dropped
  20446. 16:10:07this one and do the evaluation and then
  20447. 16:10:09go back and build a model where we've
  20448. 16:10:11dropped this one and do the evaluation.
  20449. 16:10:14Right? So that's the proper way to do it
  20450. 16:10:17is to actually just build both models
  20451. 16:10:19with each one dropped and see which
  20452. 16:10:21performs better.
  20453. 16:10:28Um otherwise I I tend to prefer to go
  20454. 16:10:31for the simplicity whatever one has kind
  20455. 16:10:33of a lower range.
  20456. 16:10:39But the truth is if they're if they're
  20457. 16:10:41really strongly correlated, it's not
  20458. 16:10:42going to matter too much. Uh because
  20459. 16:10:44they're going to give you the same
  20460. 16:10:45information, right? Uh because they're
  20461. 16:10:48so strongly correlated. Like if I in
  20462. 16:10:51this data, like if I know the
  20463. 16:10:53temperature, I pretty much know what the
  20464. 16:10:54Dupoint temperature is going to be.
  20465. 16:10:57They're so correlated.
  20466. 16:10:59So it doesn't it doesn't really matter
  20467. 16:11:01which one I drop
  20468. 16:11:09but I yeah prefer to go for the
  20469. 16:11:10simplicity.
  20470. 16:11:16Okay.
  20471. 16:11:18Um going back to let's let's do this
  20472. 16:11:23plot now of the distribution of the
  20473. 16:11:25rented bike count. So that should be
  20474. 16:11:28taking a look at the
  20475. 16:11:30um
  20476. 16:11:32the rented
  20477. 16:11:35bike count and then just doing a
  20478. 16:11:38histogram
  20479. 16:11:43and we can see what that distribution
  20480. 16:11:45is. Most of it is it less than 250.
  20481. 16:11:50Um but there are some values that are
  20482. 16:11:53really high, right? There are some days
  20483. 16:11:56that turn out to be over 3,000 into the
  20484. 16:11:593500 range. Um, that's quite a bit, but
  20485. 16:12:03most of the days are stacked over here
  20486. 16:12:05in this like 250 bucket. So, by far
  20487. 16:12:08that's the most. And then it kind of
  20488. 16:12:10decreases from there. Most values are in
  20489. 16:12:14that range and then uh it kind of
  20490. 16:12:17declines.
  20491. 16:12:22So that should just be this simple
  20492. 16:12:24histogram here.
  20493. 16:12:27So this is the
  20494. 16:12:34histogram.
  20495. 16:12:38That should be a simple one to build.
  20496. 16:12:44Were
  20497. 16:12:50you guys able to run that one?
  20498. 16:13:00Sweet.
  20499. 16:13:02Okay, pretty basic. Just showing us how
  20500. 16:13:05it's distributed. And we kind of noticed
  20501. 16:13:06that most of it is in the 250 bucket or
  20502. 16:13:09below. But there's a good amount that's
  20503. 16:13:11out there beyond like in the 500,
  20504. 16:13:147500,000
  20505. 16:13:16um
  20506. 16:13:18all the way up to there's some days that
  20507. 16:13:20register with a 3,000 and above, right?
  20508. 16:13:233500.
  20509. 16:13:24So,
  20510. 16:13:31in fact, we could look, we didn't do
  20511. 16:13:33this, but it might be worth doing is we
  20512. 16:13:35could look at the describe because we
  20513. 16:13:38never looked at the maximum.
  20514. 16:13:42Um,
  20515. 16:13:44so for the bike count, there are some
  20516. 16:13:47days that are zero
  20517. 16:13:50and the maximum is 3500. 3556
  20518. 16:13:56and the median is around 500 bikes.
  20519. 16:14:00Um
  20520. 16:14:02the average around 700.
  20521. 16:14:18So that's just our usual describe
  20522. 16:14:30All right, let's see what else. Uh,
  20523. 16:14:34plot the histogram of all numerical
  20524. 16:14:36features. So, uh, let's do that.
  20525. 16:14:48So uh luckily we have a shortcut to do
  20526. 16:14:51this. If you guys remember we have our
  20527. 16:14:54SNS um pair plot and we can pass in our
  20528. 16:15:01uh our data
  20529. 16:15:03is our df right so we can pass that in.
  20530. 16:15:07Um so this is kind of a nice this is a
  20531. 16:15:10nice thing to run that will um generate
  20532. 16:15:14the uh the scatter plots of all the
  20533. 16:15:17features against each other kind of like
  20534. 16:15:19the correlation but at the same time
  20535. 16:15:21produce the histograms on the diagonal.
  20536. 16:15:24Right? So this should be a nice plot to
  20537. 16:15:26to see all the histograms
  20538. 16:15:29uh in one kind of uh grid.
  20539. 16:16:00That's just this one.
  20540. 16:16:15Okay, so pretty this is obviously quite
  20541. 16:16:18a bit of data, but this is all the
  20542. 16:16:20features against each other. Um, look at
  20543. 16:16:23this. I mean, a couple things you see
  20544. 16:16:25right away is look at the ones that are
  20545. 16:16:26really highly correlated like the
  20546. 16:16:27temperature against the Dupoint
  20547. 16:16:29temperature. Do you guys see this strong
  20548. 16:16:32correlation here?
  20549. 16:16:34So that's that's a very indicative of a
  20550. 16:16:37very strong correlation, right? That's
  20551. 16:16:39the temperature against the Dupoint
  20552. 16:16:40temperature. That kind of makes sense
  20553. 16:16:42that it's uh that was the 0.91
  20554. 16:16:45correlation. So of course it's like
  20555. 16:16:47that.
  20556. 16:16:51Of course it looks like that, right? Uh
  20557. 16:16:53just a very strong correlation on the on
  20558. 16:16:56the x-axis or sorry on the diagonal is
  20559. 16:16:58all the histograms. They're a little bit
  20560. 16:17:00zoomed out so difficult to see. Um but
  20561. 16:17:04um we can get a sense of how some of
  20562. 16:17:07these are distributed like the first one
  20563. 16:17:10uh
  20564. 16:17:12sorry this first one
  20565. 16:17:15which is the bite count we already did
  20566. 16:17:17um temperature we can see how that's
  20567. 16:17:20distributed this third one it's kind of
  20568. 16:17:23evenly distributed
  20569. 16:17:26left of due versus temperature this one
  20570. 16:17:29that one is the hours
  20571. 16:17:33versus the due point. So, it it's kind
  20572. 16:17:36of evenly spaced out. Um, which would
  20573. 16:17:40probably be which would make sense
  20574. 16:17:42because this data is across two years.
  20575. 16:17:44So, you're going to get a lot of
  20576. 16:17:45seasonal data in there, right? So, it's
  20577. 16:17:48going to it's going to vary across like
  20578. 16:17:50seasons. Yeah. So, it kind of looks it's
  20579. 16:17:52very evenly spread out.
  20580. 16:17:55Okay. Were you able to get Parpot to
  20581. 16:17:57run?
  20582. 16:17:58Takes a moment to run.
  20583. 16:18:01takes a moment, but it produces all of
  20584. 16:18:02these plots, including all the
  20585. 16:18:03histograms.
  20586. 16:18:05And again, if we wanted to zoom in on
  20587. 16:18:07any one particular histogram, we could
  20588. 16:18:09do that. We just have to basically copy
  20589. 16:18:11and paste this code and swap out this
  20590. 16:18:13feature. Just swap out this feature and
  20591. 16:18:15we can get a zoomed in uh plot of any
  20592. 16:18:18one of those uh features like the
  20593. 16:18:21temperature
  20594. 16:18:23um or visibility or whatever it is.
  20595. 16:18:27So, what I want to do, uh, I think we'll
  20596. 16:18:28take one more break here. Um, and then
  20597. 16:18:31what we'll do is we'll come back and
  20598. 16:18:33just finish up our practice. Um, I'm
  20599. 16:18:35going to start building the model. I
  20600. 16:18:37know it wants us to do some additional
  20601. 16:18:38plotting. Um, but I want to get to
  20602. 16:18:42mainly the plotting elements, including
  20603. 16:18:44doing the pipeline one more time. Um, so
  20604. 16:18:49we're going to do that. Um, we're going
  20605. 16:18:52to practice building out the pipeline in
  20606. 16:18:53a similar fashion to exactly how we
  20607. 16:18:56built the pipeline earlier and then do
  20608. 16:18:59we're going to build our models and
  20609. 16:19:00we're going to practice that coming up.
  20610. 16:19:02I'll probably leave the plotting to you
  20611. 16:19:04guys to do as kind of homework if you
  20612. 16:19:06want to do that. Um, just because I want
  20613. 16:19:08to get to the to the modeling. Hello and
  20614. 16:19:11welcome to machine learning tutorial
  20615. 16:19:13part one. This is part one of a machine
  20616. 16:19:15learning series put on by SimplyLearn.
  20617. 16:19:18My name is Richard Kersner. I'm with the
  20618. 16:19:20SimplyLearn team. That's
  20619. 16:19:21www.simplearn.com.
  20620. 16:19:23Get certified, get ahead. What's in it
  20621. 16:19:26for you today? Well, we'll start off
  20622. 16:19:28with a brief explanation of why machine
  20623. 16:19:30learning and what is machine learning.
  20624. 16:19:32And then we'll get into a few of the
  20625. 16:19:34types of machine learning. machine
  20626. 16:19:35learning algorithms, linear regression,
  20627. 16:19:38decision trees, support vector machine,
  20628. 16:19:41and finally, we'll do a use case where
  20629. 16:19:43we're going to classify whether a recipe
  20630. 16:19:44is of a cupcake or a muffin using the
  20631. 16:19:47SPM or the support vector machine.
  20632. 16:19:49Sounds like a delicious way to explore
  20633. 16:19:51machine learning. So, why machine
  20634. 16:19:53learning? Why do we even care about
  20635. 16:19:55having these computers come up and be
  20636. 16:19:57able to do all these new things for us?
  20637. 16:19:59Well, because machines can now drive
  20638. 16:20:01your car for you. still very in the
  20639. 16:20:03infant stage but it's just exploding as
  20640. 16:20:05we see with uh Google's Whimo and then
  20641. 16:20:08Uber had their program which
  20642. 16:20:10unfortunately crashed. They know that
  20643. 16:20:12this is huge. This is going to be the
  20644. 16:20:13huge industry to change our whole
  20645. 16:20:15transportation infrastructure. Machine
  20646. 16:20:18learning is now used to detect over 50
  20647. 16:20:20eye diseases. Do you know how amazing
  20648. 16:20:22that is to have a computer that doublech
  20649. 16:20:24checkcks for the doctor for things they
  20650. 16:20:25might miss? That's just huge in the
  20651. 16:20:27health industry. pretty soon they
  20652. 16:20:29actually do already have that with in
  20653. 16:20:30some areas where maybe not for eyes but
  20654. 16:20:33for other diseases where they're using
  20655. 16:20:34the camera on your phone to help
  20656. 16:20:36pre-diagnose before you go in and see
  20657. 16:20:38the doctor. And because the machine can
  20658. 16:20:40now unlock your phone with your face, I
  20659. 16:20:43mean, that's just cool having it being
  20660. 16:20:44able to identify your face or your voice
  20661. 16:20:47and be able to turn stuff on and off for
  20662. 16:20:49you depending on where you're at and
  20663. 16:20:50what you need. Talk about an ultimate
  20664. 16:20:52automation our world we live in. And as
  20665. 16:20:55we dig in deeper, we have a nice example
  20666. 16:20:57of Facebook. As you can see here, they
  20667. 16:20:59have the Facebook post with Halloween.
  20668. 16:21:01Comment yes if you want it order here.
  20669. 16:21:04Nobody likes spam posts on Facebook that
  20670. 16:21:07annoy them into interacting with likes,
  20671. 16:21:09shares, comments, and other actions. I
  20672. 16:21:12remember the original ones were all if
  20673. 16:21:13you don't click on here, you will have
  20674. 16:21:16bad luck or some kind of fear factor.
  20675. 16:21:18Well, this is a huge thing in a social
  20676. 16:21:20media when people are getting spammed.
  20677. 16:21:22And so this tactic known as engagement
  20678. 16:21:25bait takes advantage of Facebook's
  20679. 16:21:27newsfeed algorithm by choosing
  20680. 16:21:29engagement in order to get the greater
  20681. 16:21:31reach. To eliminate engagement bait, the
  20682. 16:21:34company reviewed and categorized
  20683. 16:21:36hundreds of thousands of posts to train
  20684. 16:21:38a machine learning model that detects
  20685. 16:21:39different types of engagement bait. So
  20686. 16:21:41in this case, we have we're using
  20687. 16:21:42Facebook, but this is of course across
  20688. 16:21:44all the different social media. they
  20689. 16:21:46have different tools are building and
  20690. 16:21:47the Facebook scroll gif will be replaced
  20691. 16:21:50kind of like a virus coming in there and
  20692. 16:21:52notices that there's a certain setup
  20693. 16:21:54with Facebook and it's able to replace
  20694. 16:21:56it and they have like vote baiting react
  20695. 16:21:59baiting share baiting they have all
  20696. 16:22:01these different these are kind of
  20697. 16:22:03general titles but there certainly are a
  20698. 16:22:05lot of way of baiting you to go in there
  20699. 16:22:06and click on something so they fed all
  20700. 16:22:08this this data was fed into the machine
  20701. 16:22:10and then they have the new post the new
  20702. 16:22:12post comes up that takes over part of
  20703. 16:22:14the Facebook setup up and that's what
  20704. 16:22:16you're looking at. You're looking at
  20705. 16:22:17this new post that's replaced like a
  20706. 16:22:19virus has replaced that. So what
  20707. 16:22:20Facebook did to eliminate this is they
  20708. 16:22:22start scanning for keywords and phrases
  20709. 16:22:24like this and checks the click-through
  20710. 16:22:26rate. So it starts looking for people
  20711. 16:22:28who are clicking through it without even
  20712. 16:22:30looking at it or clicking through it and
  20713. 16:22:32it's not something that normally would
  20714. 16:22:33be clicked through. Once Facebook has
  20715. 16:22:35scanned for these keywords and phrases,
  20716. 16:22:37it is now able to identify the spam
  20717. 16:22:40coming in and this makes your life
  20718. 16:22:42easier. So you're not getting spammed.
  20719. 16:22:44It's not like walking through an airport
  20720. 16:22:45and in a lot of countries you have like
  20721. 16:22:47hundreds of people trying to sell you
  20722. 16:22:48time share. Come join us. Sign up for
  20723. 16:22:51this. Eliminates that annoyingness. So
  20724. 16:22:52now you can just enjoy your Facebook and
  20725. 16:22:54your cat pictures. Or maybe it's your
  20726. 16:22:56family pictures. Mine is family.
  20727. 16:22:58Certainly people like their cat pictures
  20728. 16:23:00too. Another good example is Google's
  20729. 16:23:02Deep Mind project Alph Go. A computer
  20730. 16:23:05program that plays a board game Go has
  20731. 16:23:07defeated the world's number one go
  20732. 16:23:09player and I hope I say his name right.
  20733. 16:23:12Kijiji the ultimate go challenge game a
  20734. 16:23:14three of three was on May 27th 2017 so
  20735. 16:23:18that was just last year that this
  20736. 16:23:20happened and what makes this so
  20737. 16:23:22important is that you know go is just is
  20738. 16:23:25a game so it's not like you're driving a
  20739. 16:23:26car or something in our real world but
  20740. 16:23:29they are using games to learn how to get
  20741. 16:23:32the machine learning program to learn
  20742. 16:23:35they want it to learn how to learn and
  20743. 16:23:36that is a huge step a lot of this is
  20744. 16:23:39still in its infant stage as far as
  20745. 16:23:40development
  20746. 16:23:41as we saw what happened with the as I
  20747. 16:23:44referred to earlier the Uber cars. They
  20748. 16:23:46lost their whole division because they
  20749. 16:23:48jumped ahead too fast. So still an
  20750. 16:23:50infant stage, but boy is this like the
  20751. 16:23:52beginning of just an amazing world that
  20752. 16:23:55is automated in ways we can't even
  20753. 16:23:57imagine what tomorrow's going to look
  20754. 16:23:59like. We've looked at a lot of examples
  20755. 16:24:01of machine learning. So let's see if we
  20756. 16:24:03can give a little bit more of a concrete
  20757. 16:24:05definition. What is machine learning?
  20758. 16:24:08Machine learning is the science of
  20759. 16:24:09making computers learn and act like
  20760. 16:24:11humans by feeding data and information
  20761. 16:24:13without being explicitly programmed. And
  20762. 16:24:16we see here we have a nice little
  20763. 16:24:17diagram where we have our ordinary
  20764. 16:24:19system, your computer nowadays, you can
  20765. 16:24:22even run a lot of the stuff on a cell
  20766. 16:24:24phone because cell phones have advanced
  20767. 16:24:25so much. And then with artificial
  20768. 16:24:27intelligence and machine learning, it
  20769. 16:24:29now takes the data and it learns from
  20770. 16:24:32what happened before and then it
  20771. 16:24:33predicts what's going to come next. And
  20772. 16:24:36then really the biggest part right now
  20773. 16:24:38in machine learning that's going on is
  20774. 16:24:39it improves on that. How do we find a
  20775. 16:24:42new solution? So we go from descriptive
  20776. 16:24:45where it's learning about stuff and
  20777. 16:24:46understanding how it fits together to
  20778. 16:24:48predicting what it's going to do to
  20779. 16:24:50postcripting coming up with a new
  20780. 16:24:52solution. And when we're working on
  20781. 16:24:55machine learning, there's a number of
  20782. 16:24:56different diagrams that people have
  20783. 16:24:58posted for what steps to go through. A
  20784. 16:25:00lot of it might be very domain specific.
  20785. 16:25:03So if you're working on photo
  20786. 16:25:05identification versus language versus
  20787. 16:25:08medical or physics, some of these are
  20788. 16:25:11switched around a little bit or new
  20789. 16:25:12things are put in. They're very specific
  20790. 16:25:14to the domain. This is kind of a very
  20791. 16:25:15general diagram. First, you want to
  20792. 16:25:17define your objective. Very important to
  20793. 16:25:20know what it is you're wanting to
  20794. 16:25:21predict. Then you're going to be
  20795. 16:25:22collecting the data. So once you've
  20796. 16:25:24defined an objective, you need to
  20797. 16:25:25collect the data that matches. You spend
  20798. 16:25:28a lot of time in data science collecting
  20799. 16:25:30data and the next step preparing the
  20800. 16:25:32data. You got to make sure that your
  20801. 16:25:34data is clean going in. There's the old
  20802. 16:25:36saying, bad data in, bad answer out or
  20803. 16:25:40bad data out. And then once you've gone
  20804. 16:25:43through and we've cleaned all this stuff
  20805. 16:25:44coming in, then you're going to select
  20806. 16:25:47the algorithm. Which algorithm are you
  20807. 16:25:49going to use? You're going to train that
  20808. 16:25:50algorithm. In this case, I think we're
  20809. 16:25:52going to be working with SVM, the
  20810. 16:25:54support vector machine. Then you have to
  20811. 16:25:56test the model. Does this model work? Is
  20812. 16:25:58this a valid model for what we're doing?
  20813. 16:26:00And then once you've tested it, you want
  20814. 16:26:02to run your prediction. You want to run
  20815. 16:26:04your prediction or your choice or
  20816. 16:26:06whatever output it's going to come up
  20817. 16:26:07with. And then once everything is set
  20818. 16:26:09and you've done lots of testing, then
  20819. 16:26:12you want to go ahead and deploy the
  20820. 16:26:13model. And remember I said domain
  20821. 16:26:15specific. This is very general as far as
  20822. 16:26:17the scope of doing something. A lot of
  20823. 16:26:19models you get halfway through and you
  20824. 16:26:21realize that your data is missing
  20825. 16:26:23something and you have to go collect new
  20826. 16:26:24data because you've run a test in here
  20827. 16:26:26someplace along the line. You're saying,
  20828. 16:26:28"Hey, I'm not really getting the answers
  20829. 16:26:29I need." So, there's a lot of things
  20830. 16:26:30that are domain specific that become
  20831. 16:26:32part of this model. This is a very
  20832. 16:26:34general model, but it's a very good
  20833. 16:26:35model to start with. And we do have some
  20834. 16:26:38basic divisions of what machine learning
  20835. 16:26:40does that's important to know. For
  20836. 16:26:42instance, do you want to predict a
  20837. 16:26:44category? Well, if you're categorizing
  20838. 16:26:46thing, that's classification. For
  20839. 16:26:48instance, whether the stock price will
  20840. 16:26:50increase or decrease. So in other words,
  20841. 16:26:52I'm looking for a yes no answer. Is it
  20842. 16:26:54going up or is it going down? And in
  20843. 16:26:56that case, we'd actually say, is it
  20844. 16:26:57going up? True. If it's not going up,
  20845. 16:26:59it's false, meaning it's going down.
  20846. 16:27:01This way, it's a yes, no. 01. Do you
  20847. 16:27:04want to predict a quantity? That's
  20848. 16:27:06regression. So remember, we just did
  20849. 16:27:08classification. Now we're looking at
  20850. 16:27:10regression. These are the two major
  20851. 16:27:12divisions in what data is doing. For
  20852. 16:27:14instance, predicting the age of a person
  20853. 16:27:16based on the height, weight, health, and
  20854. 16:27:18other factors. So based on these
  20855. 16:27:20different factors, you might guess how
  20856. 16:27:21old a person is. And then there are a
  20857. 16:27:23lot of domain specific things like do
  20858. 16:27:26you want to detect an anomaly? That's
  20859. 16:27:28anomaly detection. This is actually very
  20860. 16:27:31popular right now. For instance, you
  20861. 16:27:32want to detect money withdrawal
  20862. 16:27:33anomalies. You want to know when
  20863. 16:27:34someone's making a withdrawal that might
  20864. 16:27:36not be their own account. We've actually
  20865. 16:27:38brought this up because this is really
  20866. 16:27:40big right now. If you're predicting the
  20867. 16:27:42stock whether to buy stock or not, you
  20868. 16:27:44want to be able to know if what's going
  20869. 16:27:45on in the stock market is an anomaly,
  20870. 16:27:48use a different prediction model because
  20871. 16:27:49something else is going on. You got to
  20872. 16:27:51pull out new information in there or is
  20873. 16:27:53this just the norm? I'm going to get my
  20874. 16:27:55normal return on my money invested. So
  20875. 16:27:58being able to detect anomalies is very
  20876. 16:27:59big in data science these days. Another
  20877. 16:28:02question that comes up which is on what
  20878. 16:28:04we call untrained data is do you want to
  20879. 16:28:07discover structure in unexplored data
  20880. 16:28:10and that's called clustering. For
  20881. 16:28:12instance, finding groups of customers
  20882. 16:28:14with similar behavior given a large
  20883. 16:28:16database of customer data containing
  20884. 16:28:18their demographics and past buying
  20885. 16:28:20records. And in this case, we might
  20886. 16:28:23notice that anybody who's wearing
  20887. 16:28:25certain set of shoes goes shopping at
  20888. 16:28:27certain stores or whatever it is. are
  20889. 16:28:29going to make certain purchases. By
  20890. 16:28:31having that information, it helps us to
  20891. 16:28:33market or group people together. So then
  20892. 16:28:35we can now explore that group and find
  20893. 16:28:37out what it is we want to market to them
  20894. 16:28:39if you're in the marketing world. And
  20895. 16:28:40that might also work in just about any
  20896. 16:28:42arena. You might want to group people
  20897. 16:28:44together whether they're uh based on
  20898. 16:28:47their different areas and investments
  20899. 16:28:49and financial background, whether you're
  20900. 16:28:52going to give them a loan or not. before
  20901. 16:28:54you even start looking at whether
  20902. 16:28:55they're a valid customer for the bank,
  20903. 16:28:57you might want to look at all these
  20904. 16:28:58different areas and group them together
  20905. 16:29:00based on unknown data. So, you're not
  20906. 16:29:02you don't know what the data is going to
  20907. 16:29:03tell you, but you want to cluster people
  20908. 16:29:04together that come together. Let's take
  20909. 16:29:07a quick detour for quiz time. Oh, my
  20910. 16:29:10favorite. So, we're going to have a
  20911. 16:29:12couple questions here under quiz time
  20912. 16:29:15and um we'll be posting the answers in
  20913. 16:29:17these part two of this tutorial. So,
  20914. 16:29:20let's go ahead and take a look at these
  20915. 16:29:22quiz times questions and hopefully
  20916. 16:29:23you'll get them all right and it'll get
  20917. 16:29:25you thinking about how to process data
  20918. 16:29:27and what's going on. Can you tell what's
  20919. 16:29:29happening in the following cases? Of
  20920. 16:29:31course, you're sitting there with your
  20921. 16:29:33cup of coffee and you have your checkbox
  20922. 16:29:34and your pen trying to figure out what's
  20923. 16:29:36your next step in your data science
  20924. 16:29:38analysis. So, the first one is grouping
  20925. 16:29:41documents into different categories
  20926. 16:29:44based on the topic and content of each
  20927. 16:29:46document. Very big these days. you know,
  20928. 16:29:49you have legal documents, you have uh
  20929. 16:29:52maybe it's a sports group documents,
  20930. 16:29:53maybe you're analyzing newspaper
  20931. 16:29:55postings, but certainly having that
  20932. 16:29:58automated is a huge thing in today's
  20933. 16:30:00world. B, identifying handwritten digits
  20934. 16:30:03in images correctly. So, we want to know
  20935. 16:30:06whether uh they're writing an A or
  20936. 16:30:07capital A, B, C, what are they writing
  20937. 16:30:10out in their hand digit, their
  20938. 16:30:11handwriting. C behavior of a website
  20939. 16:30:14indicating that the site is not working
  20940. 16:30:17as designed. D, predicting salary of an
  20941. 16:30:21individual based on his or her years of
  20942. 16:30:24experience with HR hiring uh setup
  20943. 16:30:27there. So stay tuned for part two. We'll
  20944. 16:30:29go ahead and answer these questions when
  20945. 16:30:31we get to the part two of this tutorial
  20946. 16:30:33or you can just simply write at the
  20947. 16:30:35bottom and send a note to SimplyLearn
  20948. 16:30:36and they'll follow up with you on it.
  20949. 16:30:39Back to our regular content. Now these
  20950. 16:30:41last few bring us into the next topic
  20951. 16:30:44which is another way of dividing our
  20952. 16:30:45types of machine learning and that is
  20953. 16:30:47with supervised unsupervised
  20954. 16:30:51and reinforcement learning. Supervised
  20955. 16:30:54learning is a method used to enable
  20956. 16:30:56machines to classify predict objects,
  20957. 16:30:58problems or situations based on labeled
  20958. 16:31:01data fed to the machine. And in here you
  20959. 16:31:03see we have a jumble of data with
  20960. 16:31:05circles, triangles and squares. And we
  20961. 16:31:08label them. We have what's a circle,
  20962. 16:31:09what's a triangle, what's a square and
  20963. 16:31:11we have our model training and it trains
  20964. 16:31:13it. So we know the answer. Very
  20965. 16:31:15important when you're doing supervised
  20966. 16:31:16learning, you already know the answer to
  20967. 16:31:18a lot of your information coming in. So
  20968. 16:31:21you have a huge group of data coming in
  20969. 16:31:23and then you have new data coming in. So
  20970. 16:31:25we've trained our model. The model now
  20971. 16:31:27knows the difference between a circle, a
  20972. 16:31:29square, a triangle. And now that we've
  20973. 16:31:31trained it, we can send in in this case
  20974. 16:31:33a square and a circle goes in and it
  20975. 16:31:35predicts that the top one's a square and
  20976. 16:31:37the next one's a circle. And you can see
  20977. 16:31:39that this is uh being able to predict
  20978. 16:31:41whether someone's going to default on a
  20979. 16:31:42loan because I was talking about banks
  20980. 16:31:44earlier. Supervised learning on stock
  20981. 16:31:46market whether you're going to make
  20982. 16:31:48money or not. That's always important.
  20983. 16:31:50And if you are looking to make a fortune
  20984. 16:31:52in the stock market, keep in mind it is
  20985. 16:31:54very difficult to get all the data
  20986. 16:31:56correct on the stock market. It is very
  20987. 16:31:58uh it fluctuates in ways you really hard
  20988. 16:32:00to predict. So it's quite a roller
  20989. 16:32:03coaster ride. If you're running machine
  20990. 16:32:04learning on the stock market, you start
  20991. 16:32:06realizing you really have to dig for new
  20992. 16:32:08data. So we have supervised learning.
  20993. 16:32:10And if you have supervised, we need
  20994. 16:32:12unsupervised learning. In unsupervised
  20995. 16:32:15learning, machine learning model finds
  20996. 16:32:17the hidden pattern in an unlabeled data.
  20997. 16:32:20So in this case, instead of telling it
  20998. 16:32:22what the circle is and what a triangle
  20999. 16:32:24is and what a square is, it goes in
  21000. 16:32:26there, looks at them, and says for
  21001. 16:32:27whatever reason, it groups them
  21002. 16:32:29together. Maybe it'll group it by the
  21003. 16:32:30number of corners. And it notices that a
  21004. 16:32:33number of them all have three corners, a
  21005. 16:32:35number of them all have four corners,
  21006. 16:32:36and a number of them all have no
  21007. 16:32:38corners. And it's able to filter those
  21008. 16:32:40through and group them together. We
  21009. 16:32:41talked about that earlier with looking
  21010. 16:32:43at a group of people who are out
  21011. 16:32:44shopping. We want to group them together
  21012. 16:32:46to find out what they have in common.
  21013. 16:32:48And of course, once you understand what
  21014. 16:32:50people have in common, maybe you have
  21015. 16:32:52one of them who's a customer at your
  21016. 16:32:54store, or you have five of them are
  21017. 16:32:55customer at your store, and they have a
  21018. 16:32:57lot in common with five others who are
  21019. 16:32:59not customers at your store. How do you
  21020. 16:33:01market to those five who aren't
  21021. 16:33:02customers at your store yet? They fit
  21022. 16:33:04the demographs of who's going to shop
  21023. 16:33:05there, and you'd like them to shop at
  21024. 16:33:07your store, not the one next door. Of
  21025. 16:33:09course, this is a simplified version.
  21026. 16:33:10You can see very easily the difference
  21027. 16:33:12between a triangle and a circle, which
  21028. 16:33:13is might not be so easy in marketing.
  21029. 16:33:15Reinforcement learning. Reinforcement
  21030. 16:33:17learning is an important type of machine
  21031. 16:33:19learning where an agent learns how to
  21032. 16:33:21behave in an environment by performing
  21033. 16:33:23actions and seeing the result. And we
  21034. 16:33:25have here where the in this case a baby.
  21035. 16:33:28It's actually great that they used an
  21036. 16:33:29infant for this slide because the
  21037. 16:33:31reinforcement learning is very much in
  21038. 16:33:33its infant stages. But it's also
  21039. 16:33:35probably the biggest machine learning
  21040. 16:33:38demand out there right now or in the
  21041. 16:33:40future. It's going to be coming up over
  21042. 16:33:41the next few years is reinforcement
  21043. 16:33:43learning and how to make that work for
  21044. 16:33:45us. And you can see here where we have
  21045. 16:33:47our action. In the action in this one,
  21046. 16:33:49it goes into the fire. Hopefully, the
  21047. 16:33:51baby didn't it's just a little candle,
  21048. 16:33:53not a giant fire pit like it looks like
  21049. 16:33:54here. When the baby comes out and the
  21050. 16:33:56new state is the baby is sad and crying
  21051. 16:33:59because they got burned on the fire. And
  21052. 16:34:00then maybe they take another action. The
  21053. 16:34:02baby's called the agent because it's the
  21054. 16:34:04one taking the actions. And in this
  21055. 16:34:06case, they didn't go into the fire. They
  21056. 16:34:07went a different direction. And now the
  21057. 16:34:09baby's happy and laughing and playing.
  21058. 16:34:11Reinforcement learning is very easy to
  21059. 16:34:13understand because that's how as humans
  21060. 16:34:15that's one of the ways we learn. We
  21061. 16:34:17learn whether it is you burn yourself on
  21062. 16:34:19the stove, don't do that anymore. Don't
  21063. 16:34:21touch the stove. In the big picture,
  21064. 16:34:23being able to have machine learning
  21065. 16:34:25programming or an AI be able to do this
  21066. 16:34:27is huge because now we're starting to
  21067. 16:34:29learn how to learn. That's a big jump in
  21068. 16:34:33the world of computer and machine
  21069. 16:34:34learning. And we're going to go back and
  21070. 16:34:36just kind of go back over supervised
  21071. 16:34:38versus unsupervised learning.
  21072. 16:34:40Understanding this is huge because this
  21073. 16:34:42is going to come up in any project
  21074. 16:34:44you're working on. We have in supervised
  21075. 16:34:47learning, we have labeled data. We have
  21076. 16:34:49direct feedback. So someone's already
  21077. 16:34:51gone in there and said, "Yes, that's a
  21078. 16:34:53triangle. No, that's not a triangle."
  21079. 16:34:55And then you predict an outcome. So you
  21080. 16:34:56have a nice prediction. This is this
  21081. 16:34:58this new set of data is coming in and we
  21082. 16:35:00know what it's going to be. And then
  21083. 16:35:01with unsupervised trading, it's not
  21084. 16:35:03labeled. So we really don't know what it
  21085. 16:35:06is. There's no feedback. So, we're not
  21086. 16:35:08telling it whether it's right or wrong.
  21087. 16:35:10We're not telling it whether it's a
  21088. 16:35:12triangle or a square. We're not telling
  21089. 16:35:14it to go left or right. All we do is
  21090. 16:35:16we're finding hidden structure in the
  21091. 16:35:18data, grouping the data together to find
  21092. 16:35:20out what connects to each other. And
  21093. 16:35:23then you can use these together. So,
  21094. 16:35:25imagine you have an image and you're not
  21095. 16:35:27sure what you're looking for. So, you go
  21096. 16:35:29in and you have the unstructured data.
  21097. 16:35:32Find all these things that are connected
  21098. 16:35:34together and then somebody looks at
  21099. 16:35:35those and labels them. Now you can take
  21100. 16:35:38that label data and program something to
  21101. 16:35:40predict what's in the picture. So you
  21102. 16:35:42can see how they go back and forth and
  21103. 16:35:44you can start connecting all these
  21104. 16:35:46different tools together to make a
  21105. 16:35:47bigger picture. There are many
  21106. 16:35:49interesting machine learning algorithms.
  21107. 16:35:51Let's have a look at a few of them.
  21108. 16:35:53Hopefully this gave you a little flavor
  21109. 16:35:54of what's out there and these are some
  21110. 16:35:56of the most important ones that are
  21111. 16:35:57currently being used. We'll take a look
  21112. 16:35:59at linear regression, decision tree and
  21113. 16:36:02the support vector machine. Let's start
  21114. 16:36:04with a closer look at linear regression.
  21115. 16:36:07Linear regression is perhaps one of the
  21116. 16:36:09most well-known and well understood
  21117. 16:36:10algorithms in statistics and machine
  21118. 16:36:12learning. Linear regression is a linear
  21119. 16:36:15model. For example, a model that assumes
  21120. 16:36:17a linear relationship between the input
  21121. 16:36:19variables x and the single output
  21122. 16:36:22variable y. And you'll see this if you
  21123. 16:36:24remember from your algebra classes, y =
  21124. 16:36:27mx + c. Imagine we are predicting
  21125. 16:36:30distance traveled y from speed x. Our
  21126. 16:36:33linear regression model representation
  21127. 16:36:35for this problem would be y = m * x + c
  21128. 16:36:38or distance = m * speed + c where m is
  21129. 16:36:43the coefficient and c is the y
  21130. 16:36:45intercept. And we're going to look at
  21131. 16:36:47two different variations of this. First,
  21132. 16:36:49we're going to start with time is
  21133. 16:36:50constant. And you can see we have a
  21134. 16:36:52bicyclist. He's got a safety gear on,
  21135. 16:36:54thank goodness. Speed equals 10
  21136. 16:36:56meters/s. And so over a certain amount
  21137. 16:36:59of time, his distance equals 36 km. We
  21138. 16:37:02have a second bicyclist who's going
  21139. 16:37:04twice the speed or 20 m/s. And you can
  21140. 16:37:08guess if he's going twice the speed and
  21141. 16:37:09time is a constant, then he's going to
  21142. 16:37:11go twice the distance. And that's easy
  21143. 16:37:14to compute. 36 * 2, you get 72 km. And
  21144. 16:37:18so if you had the question of how fast
  21145. 16:37:21would somebody going three times that
  21146. 16:37:22speed or 30 m/s is, you can easily
  21147. 16:37:25compute the distance in our head. We can
  21148. 16:37:27do that without needing a computer, but
  21149. 16:37:29we want to do this for more complicated
  21150. 16:37:31data. So, it's kind of nice to compare
  21151. 16:37:32the two. But, let's just take a look at
  21152. 16:37:34that and what that looks like in a
  21153. 16:37:35graph. So, in a linear regression model,
  21154. 16:37:38we have our distance to the speed and we
  21155. 16:37:40have our m equals the ve slope of the
  21156. 16:37:44line. And we'll notice that the line has
  21157. 16:37:45a plus slope. And as the speed
  21158. 16:37:47increases, distance also increases.
  21159. 16:37:49Hence, the variables have a positive
  21160. 16:37:52relationship. And so your speed of the
  21161. 16:37:54person which equals y = mx plus c
  21162. 16:37:56distance traveled in a fixed interval of
  21163. 16:37:58time. And we could very easily compute
  21164. 16:38:00either following the line or just
  21165. 16:38:02knowing it's 3 * 10 m/s that this is
  21166. 16:38:05roughly 102 km distance that this third
  21167. 16:38:07bicycle has traveled. One of the key
  21168. 16:38:10definitions on here is positive
  21169. 16:38:13relationship. So the slope of the line
  21170. 16:38:16is positive. As distance increase so
  21171. 16:38:18does speed increase. Let's take a look
  21172. 16:38:20at our second example where we put
  21173. 16:38:21distance is a constant. So we have speed
  21174. 16:38:24equals 10 m/s. They have a certain
  21175. 16:38:26distance to go and it takes him 100
  21176. 16:38:29seconds to travel that distance. And we
  21177. 16:38:31have our second bicyclist who's still
  21178. 16:38:32doing 20 m/s. Since he's going twice the
  21179. 16:38:35speed, we can guess he'll cover the
  21180. 16:38:37distance in about half the time, 50
  21181. 16:38:39seconds. And of course, you could
  21182. 16:38:40probably guess on the third one, 100
  21183. 16:38:42divided by 30 since he's going three
  21184. 16:38:44times the speed. You can easily guess
  21185. 16:38:46that this is 33.3333
  21186. 16:38:49seconds time. We put that into a linear
  21187. 16:38:51regression model or a graph. If the
  21188. 16:38:53distance is assumed to be constant,
  21189. 16:38:55let's see the relationship between speed
  21190. 16:38:57and time. And as time goes up, the
  21191. 16:38:59amount of speed to go that same distance
  21192. 16:39:01goes down. So now your m equals a minus
  21193. 16:39:04v slope of the line. As the speed
  21194. 16:39:06increases, time decreases. Hence, the
  21195. 16:39:08variable has a negative relationship.
  21196. 16:39:11Again, there's our definition. positive
  21197. 16:39:13relationship and negative relationship
  21198. 16:39:15dependent on the slope of the line and
  21199. 16:39:17with a simple formula like this um and
  21200. 16:39:20even a significant amount of data. Let's
  21201. 16:39:23uh see what the mathematical
  21202. 16:39:24implementation of linear regression and
  21203. 16:39:26we'll take this data. So suppose we have
  21204. 16:39:28this data set where we have xyx= 1 2 3 4
  21205. 16:39:325 standard series and the y value is 3
  21206. 16:39:3622 43. When we take that and we go ahead
  21207. 16:39:39and plot these points on a graph, you
  21208. 16:39:42can see there's kind of a nice
  21209. 16:39:43scattering and you could probably
  21210. 16:39:44eyeball a line through the middle of it.
  21211. 16:39:47But we're going to calculate that exact
  21212. 16:39:48line for linear regression. And the
  21213. 16:39:50first thing we do is we come up here and
  21214. 16:39:52we have the mean of Xi. And remember
  21215. 16:39:55mean is basically the average. So we
  21216. 16:39:57added five plus 4 plus 3 plus 2 plus 1
  21217. 16:40:00and divide by five. And that simply
  21218. 16:40:02comes out as three. And then we'll do
  21219. 16:40:04the same for y. We'll go ahead and add
  21220. 16:40:06up all those numbers and divide by five.
  21221. 16:40:09And we end up with a mean value of y of
  21222. 16:40:11i equals 2.8 where the x i references
  21223. 16:40:15it's an average or means value. And the
  21224. 16:40:16yi also equals a means value of y. And
  21225. 16:40:19when we plot that, you'll see that we
  21226. 16:40:21can put in the y= 2.8 and the x= 3 in
  21227. 16:40:25there on our graph. We kind of gave it a
  21228. 16:40:27little different color so you could sort
  21229. 16:40:28it out with the dashed lines on it. And
  21230. 16:40:30it's important to note that when we do
  21231. 16:40:32the linear regression, the linear
  21232. 16:40:34regression model should go through that
  21233. 16:40:36dot. Now, let's find our regression
  21234. 16:40:38equation to find the best fit line.
  21235. 16:40:40Remember, we go ahead and take our y= mx
  21236. 16:40:42plus c. So, we're looking for m and c.
  21237. 16:40:44So, to find this equation for our data,
  21238. 16:40:47we need to find our slope of m and our
  21239. 16:40:50coefficient of c. And we have y = mx + c
  21240. 16:40:55where m equals the sum of x - x average
  21241. 16:40:59* y - y average or y means and x means
  21242. 16:41:02over the sum of x - x means squared.
  21243. 16:41:06That's how we get the slope of the value
  21244. 16:41:07of the line. And we can easily do that
  21245. 16:41:09by creating some columns here. We have
  21246. 16:41:11xy. Computers are really good about
  21247. 16:41:14iterating through data. And so we can
  21248. 16:41:16easily compute this and fill in a graph
  21249. 16:41:18of data. And in our graph you can easily
  21250. 16:41:21see that if we have our x value of 1 and
  21251. 16:41:24if you remember the x i or the means
  21252. 16:41:26value is 3. 1 - 3 equals a -2 and 2 - 3
  21253. 16:41:32= a -1 so on and so forth. And we can
  21254. 16:41:35easily fill in the column of x - x i y -
  21255. 16:41:38yi. And then from those we can compute x
  21256. 16:41:42- x i^ 2 and x - x i * y - yi. And you
  21257. 16:41:47can guess it that the next step is to go
  21258. 16:41:49ahead and sum the different columns for
  21259. 16:41:50the answers we need. So we get a total
  21260. 16:41:52of 10 for our x - x i^2 and a total of 2
  21261. 16:41:56for x - x i * y - yi. And we plug those
  21262. 16:42:01in, we get 2/10, which equals2. So now
  21263. 16:42:04we know the slope of our line equals2.
  21264. 16:42:06So we can calculate the value of c.
  21265. 16:42:09That'd be the next step is we need to
  21266. 16:42:10know where it crosses the y ais. And if
  21267. 16:42:13you remember, I mentioned earlier that
  21268. 16:42:15the linear regression line has to pass
  21269. 16:42:18through the means value, the one that we
  21270. 16:42:20showed earlier. We can just flip back up
  21271. 16:42:22there to that graph. And you can see
  21272. 16:42:24right here, there's our means value,
  21273. 16:42:26which is 3 x= 3 and y= 2.8. And since we
  21274. 16:42:30know that value, we can simply plug that
  21275. 16:42:33into our formula. Y =2x + c. So we plug
  21276. 16:42:38that in, we get 2.8 8 =2 * 3 + C. And
  21277. 16:42:42you can just solve for C. So now we know
  21278. 16:42:44that our coefficient equals 2.2. And
  21279. 16:42:47once we have all that, we can go ahead
  21280. 16:42:50and plot our regression line. Y =2 * X +
  21281. 16:42:542.2. And then from this equation, we can
  21282. 16:42:57compute new values. So let's predict the
  21283. 16:43:00values of Y using X= 1 2 3 4 5 and plot
  21284. 16:43:04the points. Remember the 1 2 3 4 5 was
  21285. 16:43:07our original x values. So now we're
  21286. 16:43:09going to see what y thinks they are, not
  21287. 16:43:11what they actually are. And we plug
  21288. 16:43:13those in, we get y of designated with y
  21289. 16:43:16of p. You can see that x= 1 = 2.4, x= 2=
  21290. 16:43:202.6, and so on and so on. So we have our
  21291. 16:43:23y predicted values of what we think it's
  21292. 16:43:26going to be when we plug those numbers
  21293. 16:43:27in. And when we plot the predicted
  21294. 16:43:29values along with the actual values, we
  21295. 16:43:31can see the difference. And this is one
  21296. 16:43:33of the things that's very important with
  21297. 16:43:34linear regression in any of these models
  21298. 16:43:36is to understand the error. And so we
  21299. 16:43:38can calculate the error on all of our
  21300. 16:43:40different values. And you can see over
  21301. 16:43:41here we plotted um x and y and y
  21302. 16:43:45predict. And we draw a little line so
  21303. 16:43:46you can sort of see what the error looks
  21304. 16:43:48like there between the different points.
  21305. 16:43:50So our goal is to reduce this error. We
  21306. 16:43:52want to minimize that error value on our
  21307. 16:43:54linear regression model. Minimizing the
  21308. 16:43:57distance. There are lots of ways to
  21309. 16:43:59minimize the distance between the line
  21310. 16:44:01and the data points like sum of squared
  21311. 16:44:03errors, sum of absolute errors, root
  21312. 16:44:06mean square error, etc. We keep moving
  21313. 16:44:08this line through the data points to
  21314. 16:44:10make sure the best fit line has the
  21315. 16:44:12least squared distance between the data
  21316. 16:44:14points and the regression line. So to
  21317. 16:44:16recap with a very simple linear
  21318. 16:44:18regression model, we first figure out
  21319. 16:44:20the formula of our line through the
  21320. 16:44:22middle and then we slowly adjust the
  21321. 16:44:24line to minimize the error. Keep in mind
  21322. 16:44:27this is a very simple formula. The math
  21323. 16:44:29gets even though the math is very much
  21324. 16:44:31the same, it gets much more complex as
  21325. 16:44:33we add in different dimensions. So this
  21326. 16:44:35is only two dimensions. Y equals MX + C.
  21327. 16:44:38But you can take that out to X ZQ all
  21328. 16:44:42the different features in there and they
  21329. 16:44:44can plot a linear regression model on
  21330. 16:44:46all of those using the different
  21331. 16:44:47formulas to minimize the error. Let's go
  21332. 16:44:50ahead and take a look at decision trees.
  21333. 16:44:52A very different way to solve problems
  21334. 16:44:54in the linear regression model. Decision
  21335. 16:44:56tree is a treeshaped algorithm used to
  21336. 16:44:58determine a course of action. Each
  21337. 16:45:00branch of a tree represents a possible
  21338. 16:45:02decision, occurrence, or reaction. We
  21339. 16:45:05have data which tells us if it is a good
  21340. 16:45:07day to play golf. And if we were to open
  21341. 16:45:10this data up in a general spreadsheet,
  21342. 16:45:12you can see we have the outlook, whether
  21343. 16:45:14it's rainy, overcast, sunny,
  21344. 16:45:17temperature, hot, mild, cool, humidity,
  21345. 16:45:20windy, and did I like to play golf that
  21346. 16:45:23day? Yes or no. So, we're taking a
  21347. 16:45:25census. And certainly, I wouldn't want a
  21348. 16:45:27computer telling me when I should go
  21349. 16:45:29play golf or not. But you could imagine
  21350. 16:45:30if you got up in the night before,
  21351. 16:45:32you're trying to plan your day and it
  21352. 16:45:34comes up and says, "Tomorrow would be a
  21353. 16:45:36good day for golf for you in the morning
  21354. 16:45:38and not a good day in the afternoon or
  21355. 16:45:40something like that." This becomes very
  21356. 16:45:41beneficial and we see this in a lot of
  21357. 16:45:43applications coming out now where it
  21358. 16:45:44gives you suggestions and lets you know
  21359. 16:45:47what what would uh fit the match for you
  21360. 16:45:49for the next day or the next purchase or
  21361. 16:45:51the next uh whatever you know next mail
  21362. 16:45:53out in this case is tomorrow a good day
  21363. 16:45:56for playing golf based on the weather
  21364. 16:45:57coming in. And so we come up and let's
  21365. 16:46:00uh determine if you should play golf
  21366. 16:46:02when the day is sunny and windy. So we
  21367. 16:46:04found out the forecast tomorrow is going
  21368. 16:46:05to be sunny and windy. And suppose we
  21369. 16:46:08draw our tree like this. We're going to
  21370. 16:46:10have our humidity. And then we have our
  21371. 16:46:12normal, which is uh if it's if you have
  21372. 16:46:14a normal humidity, you're going to go
  21373. 16:46:16play golf. And if the humidity is really
  21374. 16:46:18high, then we look at the outlook. And
  21375. 16:46:20if the outlook is sunny, overcast, or
  21376. 16:46:22rainy, it's going to change what you
  21377. 16:46:24choose to do. So if you know that it's a
  21378. 16:46:26very high humidity and it's sunny,
  21379. 16:46:29you're probably not going to play golf
  21380. 16:46:30cuz you're going to be out there
  21381. 16:46:31miserable, fighting off the mosquitoes
  21382. 16:46:33that are out joining you to play golf
  21383. 16:46:35with you. Maybe if it's rainy, you
  21384. 16:46:36probably don't want to play in the rain.
  21385. 16:46:38But if it's slightly overcast and you
  21386. 16:46:39get just the right shadow, that's a good
  21387. 16:46:42day to play golf and be outside out on
  21388. 16:46:44the green. Now, in this example, you can
  21389. 16:46:47probably make your own tree pretty
  21390. 16:46:49easily cuz it's a very simple set of
  21391. 16:46:51data going in. But the question is, how
  21392. 16:46:53do you know what to split? Where do you
  21393. 16:46:54split your data? What if this is much
  21394. 16:46:56more complicated data where it's not
  21395. 16:46:58something that you would particularly
  21396. 16:47:00understand? like studying cancer, they
  21397. 16:47:03take about 36 measurements of the
  21398. 16:47:05cancerous cells and then each one of
  21399. 16:47:07those measurements represents how
  21400. 16:47:10bulbous it is, how extended it is, how
  21401. 16:47:12sharp the edges are, something that as a
  21402. 16:47:14human we would have no understanding of.
  21403. 16:47:16So how do we decide how to split that
  21404. 16:47:18data up and is that the right decision
  21405. 16:47:20tree? But so that's a question that's
  21406. 16:47:21going to come up. Is this the right
  21407. 16:47:23decision tree? For that we should
  21408. 16:47:25calculate entropy and information gain.
  21409. 16:47:28Two important vocabulary words there are
  21410. 16:47:31the entropy and the information gain.
  21411. 16:47:33Entropy. Entropy is a measure of
  21412. 16:47:35randomness or impurity in the data set.
  21413. 16:47:38Entropy should be low. So we want the
  21414. 16:47:40chaos to be as low as possible. We don't
  21415. 16:47:43want to look at it and be confused by
  21416. 16:47:45the images or what's going on there with
  21417. 16:47:46mixed data. And the information gain, it
  21418. 16:47:49is a measure of decrease in entropy
  21419. 16:47:51after the data set is split. Also known
  21420. 16:47:54as entropy reduction. information gain
  21421. 16:47:57should be high. So we want our
  21422. 16:47:59information that we get out of the split
  21423. 16:48:01to be as high as possible. Let's take a
  21424. 16:48:03look at entropy from the mathematical
  21425. 16:48:06side. In this case, we're going to
  21426. 16:48:08denote entropy as I of P of and N where
  21427. 16:48:12P is the probability that you're going
  21428. 16:48:14to play a game of golf and N is the
  21429. 16:48:17probability where you're not going to
  21430. 16:48:19play the game of golf. Now, you don't
  21431. 16:48:20really have to memorize these formulas.
  21432. 16:48:22There's a few of them out there
  21433. 16:48:23depending on what you're working with.
  21434. 16:48:25But it's important to note that this is
  21435. 16:48:26where this formula is coming from. So
  21436. 16:48:28when you see it, you're not lost when
  21437. 16:48:29you're running your programming, unless
  21438. 16:48:31you're building your own decision tree
  21439. 16:48:32code in the back. And we simply have a
  21440. 16:48:35log 2 of p + n minus n / p + n * the log
  21441. 16:48:40squar of n of p plus n. But let's break
  21442. 16:48:43that down and see what actually looks
  21443. 16:48:44like when we're computing that from the
  21444. 16:48:46computer script side. Entropy of a
  21445. 16:48:49target class of the data set is the
  21446. 16:48:51whole entropy. So we have entropy play
  21447. 16:48:53golf. And we look at this. If we go back
  21448. 16:48:56to the data, you can simply count how
  21449. 16:48:58many yeses and no in our complete data
  21450. 16:49:00set for playing golf days. In our
  21451. 16:49:03complete set, we find we have five days
  21452. 16:49:06we did play golf and nine days we did
  21453. 16:49:08not play golf. And so our I equals, if
  21454. 16:49:11you add those together, 9 + 5 is 14. And
  21455. 16:49:13so our I equals 5 over 14 and 9 over 14.
  21456. 16:49:17That's our PNN values that we plug into
  21457. 16:49:19that formula. And you can go 5 over
  21458. 16:49:2214=.36.
  21459. 16:49:249 over4=64.
  21460. 16:49:26And when you do the whole equation, you
  21461. 16:49:28get the -.36
  21462. 16:49:30log<unk>^ 2 of.36 minus.64 log<unk> of
  21463. 16:49:3664. And we get a set value. We get 94.
  21464. 16:49:40So we now have a full entropy value for
  21465. 16:49:42the whole set of data that we're working
  21466. 16:49:44with. And we want to make that entropy
  21467. 16:49:47go down. And just like we calculated the
  21468. 16:49:49entropy out for the whole set, we can
  21469. 16:49:51also calculate entropy for playing golf
  21470. 16:49:54and the outlook. Is it going to be
  21471. 16:49:55overcast or rainy or sunny? And so we
  21472. 16:49:58look at the entropy. We have P of sunny
  21473. 16:50:00times E of three of two. And that just
  21474. 16:50:03comes out how many sunny days yes and
  21475. 16:50:06how many sunny days no over the total,
  21476. 16:50:08which is five. Don't forget to put the
  21477. 16:50:10we'll divide that five out later on.
  21478. 16:50:11equals P overcast = 4 comma 0 plus rainy
  21479. 16:50:16= 2a 3 and then when you do the whole
  21480. 16:50:18setup we have 5 over4 remember I said
  21481. 16:50:22there was a total of five 5 over 14 *
  21482. 16:50:25the i of 3 of 2 + 4 over 14 * the 4 0
  21483. 16:50:30and 514 over i of 23 and so we can now
  21484. 16:50:34compute the entropy of just the part
  21485. 16:50:37that has to do with the forecast and we
  21486. 16:50:39get 693 similar We can calculate the
  21487. 16:50:42entropy of other predictors like
  21488. 16:50:44temperature, humidity and wind. And so
  21489. 16:50:46we look at the gain outlook. How much
  21490. 16:50:48are we going to gain from this entropy
  21491. 16:50:50play golf minus entropy play golf
  21492. 16:50:52outlook? And we can take the original
  21493. 16:50:550.94 for the whole set minus the entropy
  21494. 16:50:58of just the rainy day and temperature
  21495. 16:51:01and we end up with a gain of.247.
  21496. 16:51:04So this is our information gain.
  21497. 16:51:06Remember we define entropy and we define
  21498. 16:51:09information gain. The higher the
  21499. 16:51:10information gain, the lower the entropy,
  21500. 16:51:13the better. The information gain of the
  21501. 16:51:15other three attributes can be calculated
  21502. 16:51:16in the same way. So we have our gain for
  21503. 16:51:19temperature equals 0.029.
  21504. 16:51:22We have our gain for humidity
  21505. 16:51:23equals.152.
  21506. 16:51:25And our gain for a windy day equals
  21507. 16:51:270048. And if you do a quick comparison,
  21508. 16:51:30you'll see the 247 is the greatest gain
  21509. 16:51:34of information. So that's the split we
  21510. 16:51:36want. Now let's build the decision tree.
  21511. 16:51:38So, we have the outlook. Is it going to
  21512. 16:51:40be sunny, overcast, or rainy? That's our
  21513. 16:51:42first split because that gives us the
  21514. 16:51:44most information gain. And we can
  21515. 16:51:45continue to go down the tree using the
  21516. 16:51:47different information gains with the
  21517. 16:51:49largest information. We can continue
  21518. 16:51:51down the nodes of the tree where we
  21519. 16:51:53choose the attribute with the largest
  21520. 16:51:54information gain as the root node and
  21521. 16:51:56then continue to split each subnode with
  21522. 16:51:59the largest information gain that we can
  21523. 16:52:01compute. And although it's a little bit
  21524. 16:52:02of a tongue twister to say all that, you
  21525. 16:52:05can see that it's a very easy to view
  21526. 16:52:07visual model. We have our outlook. We
  21527. 16:52:09split it three different directions. If
  21528. 16:52:11the outlook is overcast, we're going to
  21529. 16:52:13play. And then we can split those
  21530. 16:52:15further down if we want. So if the over
  21531. 16:52:17outlook is sunny, but then it's also
  21532. 16:52:19windy. If it's uh windy, we're not going
  21533. 16:52:22to play. If it's uh not windy, we'll
  21534. 16:52:24play. So, we can easily build a nice
  21535. 16:52:26decision tree to guess what we would
  21536. 16:52:28like to do tomorrow and give us a nice
  21537. 16:52:30recommendation for the day. So, we want
  21538. 16:52:32to know if it's a good day to play golf
  21539. 16:52:34when it's sunny and windy. Remember the
  21540. 16:52:35original question that came out,
  21541. 16:52:37tomorrow's weather report is sunny and
  21542. 16:52:38windy. You can see by going down the
  21543. 16:52:40tree, we go outlook sunny, outlook
  21544. 16:52:42windy. We're not going to play golf
  21545. 16:52:44tomorrow. So, our little smartwatch pops
  21546. 16:52:45up and says, I'm sorry, tomorrow's not a
  21547. 16:52:48good day for golf. It's going to be
  21548. 16:52:50sunny and windy. And if you're a huge
  21549. 16:52:52golf fan, you might go, "Uh oh, it's not
  21550. 16:52:55a good day to play golf." We can go in
  21551. 16:52:57and watch a golf game at home. So, we'll
  21552. 16:52:59sit in front of the TV instead of being
  21553. 16:53:00out playing golf in the wind. Now that
  21554. 16:53:02we looked at our decision tree, let's
  21555. 16:53:04look at the third one of our algorithms
  21556. 16:53:06we're investigating. Support vector
  21557. 16:53:08machine. Support vector machine is a
  21558. 16:53:10widely used classification algorithm.
  21559. 16:53:12The idea of support vector machine is
  21560. 16:53:14simple. The algorithm creates a
  21561. 16:53:16separation line which divides the
  21562. 16:53:18classes in the best possible manner. For
  21563. 16:53:20example, dog or cat, disease or no
  21564. 16:53:22disease. Suppose we have a labeled
  21565. 16:53:24sample data which tells height and
  21566. 16:53:27weight of males and females. A new data
  21567. 16:53:30point arrives and we want to know
  21568. 16:53:31whether it's going to be a male or a
  21569. 16:53:33female. So we start by drawing a line.
  21570. 16:53:36We draw decision lines. But if we
  21571. 16:53:37consider decision line one, then we will
  21572. 16:53:39classify the individual as a male. And
  21573. 16:53:42if we consider decision line two, then
  21574. 16:53:44it'll be a female. So you can see this
  21575. 16:53:46person kind of lies in the middle of the
  21576. 16:53:48two groups. So it's a little confusing
  21577. 16:53:49trying to figure out which line they
  21578. 16:53:50should be under. We need to know which
  21579. 16:53:52line divides the classes correctly. But
  21580. 16:53:54how the goal is to choose a hyper plane
  21581. 16:53:57and that is one of the key words they
  21582. 16:53:59use when we talk about support vector
  21583. 16:54:01machines. Choose a hyper plane with the
  21584. 16:54:04greatest possible margin between the
  21585. 16:54:06decision line and the nearest point
  21586. 16:54:07within the training set. So you can see
  21587. 16:54:10here we have our support vector. We have
  21588. 16:54:12the two nearest points to it and we draw
  21589. 16:54:14a line between those two points. And the
  21590. 16:54:17distance margin is the distance between
  21591. 16:54:19the hyper plane and the nearest data
  21592. 16:54:21point from either set. So we actually
  21593. 16:54:23have a value and it should be equal
  21594. 16:54:26distant between the two points that
  21595. 16:54:28we're comparing it to. When we draw the
  21596. 16:54:30hyperplanes, we observe that line one
  21597. 16:54:32has a maximum distance. So we observe
  21598. 16:54:35that line one has a maximum distance
  21599. 16:54:37margin. So we'll classify the new data
  21600. 16:54:39point correctly. And our result on this
  21601. 16:54:41one is going to be that the new data
  21602. 16:54:43point is MEL. One of the reasons we call
  21603. 16:54:45it a hyper plane versus a line is that a
  21604. 16:54:49lot of times we're not looking at just
  21605. 16:54:51weight and height. We might be looking
  21606. 16:54:53at 36 different features or dimensions.
  21607. 16:54:56And so when we cut it with a hyper
  21608. 16:54:58plane, it's more of a three-dimensional
  21609. 16:55:00cut in the data, multi-dimensional that
  21610. 16:55:03cuts the data a certain way. And each
  21611. 16:55:05plane continues to cut it down until we
  21612. 16:55:07get the best fit or match. Let's
  21613. 16:55:09understand this with the help of an
  21614. 16:55:11example. Problem statement. You always
  21615. 16:55:12start with a problem statement when
  21616. 16:55:14you're going to put some code together.
  21617. 16:55:15We're going to do some coding now.
  21618. 16:55:16Classifying muffin and cupcake recipes
  21619. 16:55:18using support vector machines. So the
  21620. 16:55:21cupcake versus the muffin. Let's have a
  21621. 16:55:24look at our data set. And we have the
  21622. 16:55:26different recipes here. We have a muffin
  21623. 16:55:28recipe that has so much flour. I'm not
  21624. 16:55:30sure what measurement 55 is in, but it
  21625. 16:55:33has 55, maybe it's ounces, but it has a
  21626. 16:55:36certain amount of flour, certain amount
  21627. 16:55:38of milk, sugar, butter, egg, baking
  21628. 16:55:41powder, vanilla, and salt. And so based
  21629. 16:55:43on these measurements, we want to guess
  21630. 16:55:45whether we're making a muffin or a
  21631. 16:55:47cupcake. And you can see in this one, we
  21632. 16:55:49don't have just two features. We don't
  21633. 16:55:51just have height and weight as we did
  21634. 16:55:53before between the male and female. In
  21635. 16:55:55here, we have a number of features. In
  21636. 16:55:57fact, in this, we're looking at eight
  21637. 16:55:59different features to guess whether it's
  21638. 16:56:01a muffin or a cupcake. What's the
  21639. 16:56:04difference between a muffin and a
  21640. 16:56:05cupcake? Turns out muffins have more
  21641. 16:56:08flour, while cupcakes have more butter
  21642. 16:56:10and sugar. So, basically, the cupcakes a
  21643. 16:56:12little bit more of a dessert, where the
  21644. 16:56:14muffin's a little bit more of a fancy
  21645. 16:56:15bread. But how do we do that in Python?
  21646. 16:56:18How do we code that to go through
  21647. 16:56:20recipes and figure out what the recipe
  21648. 16:56:21is? And I really just want to say
  21649. 16:56:24cupcakes versus muffins like some big
  21650. 16:56:28professional wrestling thing. Before we
  21651. 16:56:30start in our cupcakes versus muffins, we
  21652. 16:56:32are going to be working in Python.
  21653. 16:56:34There's many versions of Python, many
  21654. 16:56:36different editors. That is one of the
  21655. 16:56:38strengths and weaknesses of Python is it
  21656. 16:56:41just has so much stuff attached to it.
  21657. 16:56:43It's one of the more popular data
  21658. 16:56:45science programming packages you can
  21659. 16:56:47use. In this case, we're going to go
  21660. 16:56:49ahead and use Anaconda in Jupyter
  21661. 16:56:52Notebook. The Anaconda Navigator has all
  21662. 16:56:55kinds of fun tools. Once you're into the
  21663. 16:56:58Anaconda Navigator, you can change
  21664. 16:57:00environments. I actually have a number
  21665. 16:57:02of environments on here. We'll be using
  21666. 16:57:04Python 36 environment. So, this is in
  21667. 16:57:07Python version 36. Although, it doesn't
  21668. 16:57:09matter too much which version you use. I
  21669. 16:57:12usually try to stay with the 3x because
  21670. 16:57:14they're current unless you have a
  21671. 16:57:15project that's very specifically in
  21672. 16:57:17version 2x 27 I think is usually what
  21673. 16:57:19most people use in the version two. And
  21674. 16:57:22then once we're in our um Jupiter
  21675. 16:57:24notebook editor, I can go up and create
  21676. 16:57:26a new file and we'll just jump in here.
  21677. 16:57:30In this case, we're doing SPM muffin
  21678. 16:57:33versus cupcake. And then let's start
  21679. 16:57:35with our packages for data analysis.
  21680. 16:57:40And we almost always use a couple
  21681. 16:57:41there's a few very standard packages we
  21682. 16:57:43use. We use import oops import
  21683. 16:57:50numpy
  21684. 16:57:52that's for number python. They usually
  21685. 16:57:54denote it as np that's very comma that's
  21686. 16:57:57very common. And then we're going to
  21687. 16:57:59import pandas as pd. And numpy deals
  21688. 16:58:03with number arrays. There's a lot of
  21689. 16:58:05cool things you can do with the numpy uh
  21690. 16:58:07setup as far as multiplying all the
  21691. 16:58:10values in an array in a numpy array data
  21692. 16:58:12array. Pandas I can't remember if we're
  21693. 16:58:15using it actually in this data set. I
  21694. 16:58:17think we do as an import it makes a nice
  21695. 16:58:19data frame. And the difference between a
  21696. 16:58:21data frame and a numpy array is that a
  21697. 16:58:24data frame is more like your Excel
  21698. 16:58:25spreadsheet. You have columns, you have
  21699. 16:58:28indexes. So you have different ways of
  21700. 16:58:30referencing it easily viewing it. And
  21701. 16:58:32there's additional features you can run
  21702. 16:58:34on a data frame. And pandas kind of sits
  21703. 16:58:36on numpy. So they you need them both in
  21704. 16:58:38there. And then finally, we're working
  21705. 16:58:41with the support vector machine. So from
  21706. 16:58:45sklearn, we're going to use the sklearn
  21707. 16:58:47model. Import SVM support vector
  21708. 16:58:51machine.
  21709. 16:58:53And then as a data scientist, you should
  21710. 16:58:56always try to visualize your data. Some
  21711. 16:58:59data obviously is too complicated or
  21712. 16:59:01doesn't make any sense to the human. But
  21713. 16:59:03if it's possible, it's good to take a
  21714. 16:59:05second look at it so that you can
  21715. 16:59:06actually see what you're doing. Now, for
  21716. 16:59:08that, we're going to use two packages.
  21717. 16:59:10We're going to import mapplot
  21718. 16:59:12library.pipplot as plt. Again, very
  21719. 16:59:15common. And we're going to import seabor
  21720. 16:59:18as sns. And we'll go ahead and set the
  21721. 16:59:21font scale in the SNS right in our
  21722. 16:59:23import line. That's what this U
  21723. 16:59:25semicolon followed by a line of data.
  21724. 16:59:28We're going to set the SNS. And these
  21725. 16:59:30are great because the the seabour sits
  21726. 16:59:32on top of map plot library just like
  21727. 16:59:35pandas sits on numpy. So it adds a lot
  21728. 16:59:37more features and uses and control.
  21729. 16:59:40We're obviously not going to get into
  21730. 16:59:41mattplot library and seabour. It' be its
  21731. 16:59:43own tutorial. We're really just focusing
  21732. 16:59:45on the SVM, the support vector machine
  21733. 16:59:48from sklearn. And since we're in Jupyter
  21734. 16:59:52notebook, uh we have to add a special
  21735. 16:59:54line in here for our mattplot library.
  21736. 16:59:57And that's your percentage sign or amber
  21737. 17:00:00sign mattplot library in line. Now, if
  21738. 17:00:04you're doing this in just a straight
  21739. 17:00:06code project, a lot of times I use like
  21740. 17:00:08Notepad++
  21741. 17:00:09and I'll run it from there. You don't
  21742. 17:00:11have to have that line in there because
  21743. 17:00:13it'll just pop up as its own window on
  21744. 17:00:14your computer depending on how your
  21745. 17:00:16computer's set up because we're running
  21746. 17:00:18this in the Jupyter notebook as a
  21747. 17:00:20browser setup. This tells it to display
  21748. 17:00:23all of our graphics right below on the
  21749. 17:00:26page. So that's what that line is for.
  21750. 17:00:29Remember the first time I ran this, I
  21751. 17:00:30didn't know that and I had to go look
  21752. 17:00:31that up years ago. It's quite a
  21753. 17:00:33headache. So mattplot library inline is
  21754. 17:00:36just because we're running this on the
  21755. 17:00:38web setup and we can go ahead and run
  21756. 17:00:40this. make sure all our modules are in.
  21757. 17:00:42They're all imported, which is great. If
  21758. 17:00:44you don't have them import, you'll need
  21759. 17:00:46to go ahead and pip. Use the pip or
  21760. 17:00:48however you do it. There's a lot of
  21761. 17:00:49other install packages out there,
  21762. 17:00:51although pip is the most common. And you
  21763. 17:00:53have to make sure these are all
  21764. 17:00:54installed on your Python setup. The next
  21765. 17:00:57step, of course, is we got to look at
  21766. 17:00:58the data. You can't run a model for
  21767. 17:01:01predicting data if you don't have actual
  21768. 17:01:02data. So, to do that, let me go ahead
  21769. 17:01:04and open this up and take a look. And we
  21770. 17:01:07have our uh cupcakes versus muffins. and
  21771. 17:01:10it's a CSV file or CSV meaning that it's
  21772. 17:01:13commaepparated variable
  21773. 17:01:15and it's going to open it up in a nice
  21774. 17:01:17uh spreadsheet for me. And you can see
  21775. 17:01:19up here we have the type we have muffin
  21776. 17:01:21muffin muffin cupcake cupcake cupcake
  21777. 17:01:24and then it's broken up into flour,
  21778. 17:01:25milk, sugar, butter, egg, baking powder,
  21779. 17:01:28vanilla and salt. So we can do is we can
  21780. 17:01:31go ahead and look at this data also in
  21781. 17:01:33our Python.
  21782. 17:01:36Let us create a variable recipes equals
  21783. 17:01:40we're going to use our pandas module
  21784. 17:01:43read CSV. Remember is a commaepparated
  21785. 17:01:46variable
  21786. 17:01:48and the file name happened to be
  21787. 17:01:50cupcakes versus muffins. Oops, I got
  21788. 17:01:52double brackets there.
  21789. 17:01:57Do it this way.
  21790. 17:02:01There we go. cupcakes versus muffins.
  21791. 17:02:05Because the program I loaded or the the
  21792. 17:02:08place I saved this particular Python
  21793. 17:02:10program is in the same folder, we can
  21794. 17:02:12get by with just the file name. But
  21795. 17:02:14remember, if you're storing it in a
  21796. 17:02:15different location, you have to also put
  21797. 17:02:16down the full path on there.
  21798. 17:02:20And then because we're in pandas, we're
  21799. 17:02:22going to go ahead and you can actually
  21800. 17:02:25in line you can do this, but let me do
  21801. 17:02:27the full print. You can just type in
  21802. 17:02:29recipes.head head in the Jupyter
  21803. 17:02:32notebook. But if you're running in code
  21804. 17:02:34in a different script, you'd need to go
  21805. 17:02:36ahead and type out the whole print
  21806. 17:02:37recipes.
  21807. 17:02:39And Pandanda's knows that's going to do
  21808. 17:02:41the first five lines of data. And if we
  21809. 17:02:44flip back on over to the spreadsheet
  21810. 17:02:47where we opened up our CSV file,
  21811. 17:02:50uh you can see where it starts on line
  21812. 17:02:52two. This one calls it zero. And then 2
  21813. 17:02:553 4 5 6 is going to match. Go and close
  21814. 17:02:58that out because we don't need that
  21815. 17:02:59anymore. And it always starts at zero.
  21816. 17:03:02And these are it automatically indexes
  21817. 17:03:04it since we didn't tell it to use an
  21818. 17:03:06index in here. So that's the index
  21819. 17:03:08number for the left hand side. And it
  21820. 17:03:10automatically took the top row as
  21821. 17:03:13labels. So pandas using it to read a CSV
  21822. 17:03:17is just really slick and fast. One of
  21823. 17:03:20the reasons we love our pandas, not just
  21824. 17:03:22because they're cute and cuddly teddy
  21825. 17:03:23bears.
  21826. 17:03:26And let's go ahead and plot our data.
  21827. 17:03:29And I'm not going to plot all of it. I'm
  21828. 17:03:31just going to plot the uh sugar and
  21829. 17:03:34flour. Now, obviously, you can see where
  21830. 17:03:37they get really complicated if we have
  21831. 17:03:39tons of different features. And so,
  21832. 17:03:42you'll break them up and maybe look at
  21833. 17:03:43just two of them at a time to see how
  21834. 17:03:45they connect.
  21835. 17:03:48And to plot them, we're going to go
  21836. 17:03:49ahead and use Seabor. So, that's our
  21837. 17:03:51SNS. And the command for that is SNS.LM
  21838. 17:03:56plot. And then the two different
  21839. 17:03:58variables I'm going to plot is flour and
  21840. 17:04:00sugar.
  21841. 17:04:02Data equals recipes. The hue equals
  21842. 17:04:05type. And this is a lot of fun because
  21843. 17:04:07it knows that this is pandas coming in.
  21844. 17:04:10So this is one of the powerful things
  21845. 17:04:12about pandas mixed with seabor and doing
  21846. 17:04:17graphing. And then we're going to use a
  21847. 17:04:19pallet set one. There's a lot of
  21848. 17:04:21different sets in there. You can go look
  21849. 17:04:23them up for seabor. We do a regular fit
  21850. 17:04:26regular equals false. So, we're not
  21851. 17:04:27really trying to fit anything. And it's
  21852. 17:04:30a scatter KWS.
  21853. 17:04:32A lot of these settings you can look up
  21854. 17:04:33in Seabor. Half of these you could
  21855. 17:04:35probably leave off when you run them.
  21856. 17:04:37Somebody played with this and found out
  21857. 17:04:38that these were the best settings for
  21858. 17:04:40doing a Seabor plot. And let's go ahead
  21859. 17:04:43and run that. And because it does it in
  21860. 17:04:45line, it just puts it right on the page.
  21861. 17:04:49And you can see right here that just
  21862. 17:04:52based on sugar and flour alone, there's
  21863. 17:04:55a definite split. And we use these
  21864. 17:04:58models because you can actually look at
  21865. 17:04:59it and say, "Hey, if I drew a line right
  21866. 17:05:01between the middle of the blue dots and
  21867. 17:05:03the red dots, we'd be able to do an SVM
  21868. 17:05:07and and a hyper plane right there in the
  21869. 17:05:09middle.
  21870. 17:05:12Then the next step is to format or
  21871. 17:05:16pre-process
  21872. 17:05:20our data.
  21873. 17:05:22And we're going to break that up into
  21874. 17:05:23two parts.
  21875. 17:05:26We need a type label. And remember,
  21876. 17:05:29we're going to decide whether it's a
  21877. 17:05:31muffin or a cupcake. Well, a computer
  21878. 17:05:32doesn't know muffin or cupcake. It knows
  21879. 17:05:34zero and one. So, what we're going to do
  21880. 17:05:37is we're going to create a type label.
  21881. 17:05:39And from this we'll create a numpy array
  21882. 17:05:42nump where and this is where we can do
  21883. 17:05:46some logic. We take our recipes from our
  21884. 17:05:49panda and wherever type equals muffin
  21885. 17:05:52it's going to be zero. And then if it
  21886. 17:05:55doesn't equal muffin which is cupcakes
  21887. 17:05:57it's going to be one. So we create our
  21888. 17:05:59type label. This is the answer. So when
  21889. 17:06:02we're doing our training model remember
  21890. 17:06:04we have to have a a training data. This
  21891. 17:06:06is what we're going to train it with. Is
  21892. 17:06:07that it's zero or one? it's a muffin or
  21893. 17:06:09it's not.
  21894. 17:06:12And then we're going to create our
  21895. 17:06:13recipe features.
  21896. 17:06:16And if you remember correctly from right
  21897. 17:06:17up here, the first column is type.
  21898. 17:06:21So we really don't need the type column
  21899. 17:06:22because that's our muffin or cupcake.
  21900. 17:06:24And in pandas, we can easily sort that
  21901. 17:06:27out.
  21902. 17:06:29We take our value recipes
  21903. 17:06:33columns. That's a pandas function built
  21904. 17:06:36into pandas.
  21905. 17:06:39values converting them to values. So
  21906. 17:06:41it's just the column titles going across
  21907. 17:06:43the top and we don't want the first one.
  21908. 17:06:46So what we do is since it always starts
  21909. 17:06:47at zero, we want one
  21910. 17:06:51colon till the end.
  21911. 17:06:54And then we want to go ahead and make
  21912. 17:06:56this a list. And this converts it to a
  21913. 17:06:59list of strings.
  21914. 17:07:02And then we can go ahead and just take a
  21915. 17:07:04look and see what we're looking at for
  21916. 17:07:05the features. Make sure it looks right.
  21917. 17:07:08Me go ahead and run that.
  21918. 17:07:12And I forgot the S on recipes. So, we'll
  21919. 17:07:14go ahead and add the S in there and then
  21920. 17:07:16run that. And we can see we have flour,
  21921. 17:07:18milk, sugar, butter, egg, baking powder,
  21922. 17:07:22vanilla, and salt. And that matches what
  21923. 17:07:24we have up here, right? Where we printed
  21924. 17:07:26out everything but the type. So, we have
  21925. 17:07:28our features and we have our label.
  21926. 17:07:32Now, the recipe features is just the
  21927. 17:07:34titles of the columns. We actually need
  21928. 17:07:37the ingredients.
  21929. 17:07:41And at this point, we have a couple
  21930. 17:07:43options. One, we could run it over all
  21931. 17:07:46the ingredients.
  21932. 17:07:48And when you're doing this, usually you
  21933. 17:07:50do. But for our example, we want to
  21934. 17:07:51limit it so you can easily see what's
  21935. 17:07:53going on because if we did all the
  21936. 17:07:55ingredients, we have, you know, that's
  21937. 17:07:57what, um, seven, eight different
  21938. 17:08:00hyperplanes that would be built into it.
  21939. 17:08:02We only want to look at one. So you can
  21940. 17:08:04see what the SVM is doing.
  21941. 17:08:06And so we'll take our recipes and we'll
  21942. 17:08:08do just flour and sugar. Again, you can
  21943. 17:08:12replace that with your recipe features
  21944. 17:08:14and do all of them, but we're going to
  21945. 17:08:15do just flour and sugar. And we're going
  21946. 17:08:17to convert that to values. We don't need
  21947. 17:08:19to make a list out of it because it's
  21948. 17:08:21not string values. These are actual
  21949. 17:08:23values on there. And we can go ahead and
  21950. 17:08:26just print
  21951. 17:08:28ingredients. And you can see what that
  21952. 17:08:30looks like.
  21953. 17:08:32Uh, and so we have just the nanoflower
  21954. 17:08:34and sugar, just the two sets of plots.
  21955. 17:08:38And just for fun, let's go ahead and
  21956. 17:08:40take this over here and take our recipe
  21957. 17:08:43features.
  21958. 17:08:46And so if we decided to use all the
  21959. 17:08:48recipe features, you'll see that it
  21960. 17:08:50makes a nice column of different data.
  21961. 17:08:52So it just strips out all the labels and
  21962. 17:08:54everything. We just have just the
  21963. 17:08:55values. But because we want to be able
  21964. 17:08:57to view this easily in a plot later on,
  21965. 17:09:01we'll go ahead and take that and just do
  21966. 17:09:02flour and sugar.
  21967. 17:09:05And we'll run that. And you'll see it's
  21968. 17:09:07just the two columns.
  21969. 17:09:10So the next step is to go ahead and fit
  21970. 17:09:12our model.
  21971. 17:09:14We'll go ahead and just call it model.
  21972. 17:09:16And it's a SVM. We're using a package
  21973. 17:09:19called SVC.
  21974. 17:09:23In this case, we're going to go ahead
  21975. 17:09:24and set the kernel equals linear. So,
  21976. 17:09:27it's using a specific setup on there.
  21977. 17:09:29And if we go to the reference on their
  21978. 17:09:31website for the SVM,
  21979. 17:09:34you'll see that there's about there's
  21980. 17:09:36eight of them here. Three of them are
  21981. 17:09:38for regression.
  21982. 17:09:39Three are for classification. The SVC,
  21983. 17:09:43support vector classification, is
  21984. 17:09:45probably one of the most commonly used.
  21985. 17:09:47And then there's also one for detecting
  21986. 17:09:49outliers and another one that has to do
  21987. 17:09:51with something a little bit more
  21988. 17:09:52specific on the model. But SVC and SVR
  21989. 17:09:55are the two most commonly used standing
  21990. 17:09:57for support vector classifier and
  21991. 17:10:00support vector regression. Remember
  21992. 17:10:02regression is an actual value, a float
  21993. 17:10:05value or whatever you're trying to work
  21994. 17:10:07on. And SBC is a classifier. So it's a
  21995. 17:10:10yes, no, true, false.
  21996. 17:10:13But for this we want to know 01 muffin
  21997. 17:10:15cupcake. If we go ahead and create our
  21998. 17:10:17model and once we have our model
  21999. 17:10:19created, we're going to do model.fit.
  22000. 17:10:22And this is very common, especially in
  22001. 17:10:23the sklearn. All their models are
  22002. 17:10:25followed with the fit command.
  22003. 17:10:28And what we put into the fit, what we're
  22004. 17:10:30training with it is we're putting in the
  22005. 17:10:32ingredients, which in this case we
  22006. 17:10:34limited to just flour and sugar, and the
  22007. 17:10:36type label. Is it a muffin or cupcake?
  22008. 17:10:40Now, in more complicated data science
  22009. 17:10:43series, you'd want to split into, we
  22010. 17:10:46won't get into that today, where you
  22011. 17:10:47split it into training data and test
  22012. 17:10:50data. And they even do something where
  22013. 17:10:52they split it into thirds, where a third
  22014. 17:10:54is used for where you switch between
  22015. 17:10:55which one's training and test. There's
  22016. 17:10:57all kinds of things go into that. It
  22017. 17:10:59gets very complicated when you get to
  22018. 17:11:00the higher end. Not overly complicated,
  22019. 17:11:02just an extra step, which we're not
  22020. 17:11:04going to do today because this is a very
  22021. 17:11:06simple set of data.
  22022. 17:11:08And let's go ahead and run this. And now
  22023. 17:11:09we have our model fit. And uh I got an
  22024. 17:11:12error here. So let me fix that real
  22025. 17:11:14quick. It's capital SBC. It turns out
  22026. 17:11:18I did it lowercase.
  22027. 17:11:20Support vector
  22028. 17:11:23classifier. There we go. Let's go ahead
  22029. 17:11:25and run that. And you'll see it comes up
  22030. 17:11:27with all this information that it prints
  22031. 17:11:29out automatically. These are the
  22032. 17:11:32defaults of the model. You notice that
  22033. 17:11:34we changed the kernel to linear. And
  22034. 17:11:36there's our kernel linear on the
  22035. 17:11:37printout. And there's other different
  22036. 17:11:39settings you can mess with.
  22037. 17:11:42We're going to just to leave that alone
  22038. 17:11:44for right now. For this, we don't really
  22039. 17:11:45need to mess with any of those.
  22040. 17:11:49So, next we're going to dig a little bit
  22041. 17:11:51into our newly trained model. And we're
  22042. 17:11:55going to do this so we can show you on a
  22043. 17:11:56graph.
  22044. 17:11:58And let's go ahead and get the
  22045. 17:12:00separating.
  22046. 17:12:04and we're going to say uh we're going to
  22047. 17:12:05use a W for our variable on here and
  22048. 17:12:08we're going to do model.coreeficient_0.
  22049. 17:12:14So what the heck is that? Again, we're
  22050. 17:12:15digging into the model. So we've already
  22051. 17:12:18got a prediction and a train. This is a
  22052. 17:12:21math behind it that we're looking at
  22053. 17:12:23right now. And so the w is going to
  22054. 17:12:27represent two different coefficients.
  22055. 17:12:30And if you remember, we had y = mx + c.
  22056. 17:12:33So these coefficients are connected to
  22057. 17:12:35that but in two-dimensional it's a
  22058. 17:12:38plane.
  22059. 17:12:40We don't want to spend too much time on
  22060. 17:12:42this because you can get lost in the
  22061. 17:12:44confusion of the math. So if you're a
  22062. 17:12:46math wiz this is great. You can go
  22063. 17:12:48through here and you'll see that we have
  22064. 17:12:50a= minus w of 0 over w of 1. Remember
  22065. 17:12:54there's two different values there. And
  22066. 17:12:56that's basically the slope that we're
  22067. 17:12:59generating.
  22068. 17:13:01And then we're going to build an xx.
  22069. 17:13:03What is xx? We're going to set it up to
  22070. 17:13:06a numpy array. There's our np line
  22071. 17:13:09space. So we're creating a line
  22072. 17:13:12of values between 30 and 60. So it just
  22073. 17:13:15creates a set of numbers for x. And then
  22074. 17:13:18if you remember correctly, we have our
  22075. 17:13:21formula y equals the slope * x
  22076. 17:13:26plus the intercept. Well, to make this
  22077. 17:13:29work, we can do this as y
  22078. 17:13:32equals the slope times each value in
  22079. 17:13:36that array. That's the neat thing about
  22080. 17:13:38numpy. So, when I do a * xx, which is a
  22081. 17:13:41whole numpy array of values, it
  22082. 17:13:43multiplies a across all of them. And
  22083. 17:13:45then it takes those same values and we
  22084. 17:13:47subtract the model intercept. That's
  22085. 17:13:50your uh we had mx plus c. So, that'd be
  22086. 17:13:53the c from the formula y mx plus c.
  22087. 17:13:57And that's where all these numbers come
  22088. 17:13:59from. A little bit confusing because
  22089. 17:14:00it's digging out of these different
  22090. 17:14:02arrays. And then what we want to do is
  22091. 17:14:04we're going to take this and we're going
  22092. 17:14:06to go ahead and plot it. So plot the
  22093. 17:14:09parallels to separating hyper plane that
  22094. 17:14:11pass through the support vectors. And so
  22095. 17:14:14we're going to create B equals a model
  22096. 17:14:17support vectors. Pulling our support
  22097. 17:14:19vectors out there. Here's our y, which
  22098. 17:14:22we now know is a set of data. And we
  22099. 17:14:24have uh we're going to create y down = a
  22100. 17:14:27* xx + b1 - a * b 0. And then model
  22101. 17:14:33support vector b is going to be set that
  22102. 17:14:35to a new value the minus1 setup. And y y
  22103. 17:14:38up = a * xx + b1 - a * b 0. And we can
  22104. 17:14:45go ahead and just run this to load these
  22105. 17:14:46variables up. If you wanted to know
  22106. 17:14:49understand a little bit more of what's
  22107. 17:14:50going on, you can see if we print
  22108. 17:14:55y, let me just run that. You can see
  22109. 17:14:58it's an array. This is a line. It's
  22110. 17:15:00going to have in this case between 30
  22111. 17:15:02and 60. So there's going to be 30
  22112. 17:15:04variables in here. And the same thing
  22113. 17:15:06with y y up y y y y y y y y y y y y y y
  22114. 17:15:08y y y y y y y y y y y y y y y y y y y y
  22115. 17:15:08y y y y y y y y y y y y y y y y y y y y
  22116. 17:15:08y y y y y y y y y y y y y y y y y y y y
  22117. 17:15:08y y y y y y down and we'll we'll plot
  22118. 17:15:10those in just a minute on a graph so you
  22119. 17:15:12can see what those look like.
  22120. 17:15:14Just go ahead and delete that out of
  22121. 17:15:15here and run that. So, it loads up the
  22122. 17:15:18variables. Nice clean slate. I'm just
  22123. 17:15:21going to copy this from before. Remember
  22124. 17:15:23this? Our SNS, our Seabor plot, LM plot,
  22125. 17:15:27flower, sugar. And I'll just go and run
  22126. 17:15:30that real quick so you can see what
  22127. 17:15:31remember what that looks like. It's just
  22128. 17:15:32a straight graph on there. And then one
  22129. 17:15:35of the neat things is because Seabour
  22130. 17:15:37sits on top of piplot,
  22131. 17:15:40we can do the piplot for the line going
  22132. 17:15:42through. And that is simply plt.plot
  22133. 17:15:47And that's our xx and y are two
  22134. 17:15:51corresponding values xy. And then
  22135. 17:15:53somebody played with this to figure out
  22136. 17:15:55that the line width equals 2 and the
  22137. 17:15:57color black would look nice. So let's go
  22138. 17:16:00ahead and run this whole thing with the
  22139. 17:16:01pi plot on there. And you can see when
  22140. 17:16:04we do this, it's just doing flour and
  22141. 17:16:06sugar on here.
  22142. 17:16:08Corresponding line between the sugar and
  22143. 17:16:10the flour and the muffin versus cupcake.
  22144. 17:16:16Um, and then we generated the support
  22145. 17:16:18vectors, the y down and y up. So let's
  22146. 17:16:21take a look and see what that looks
  22147. 17:16:23like.
  22148. 17:16:24So we'll do our plot.
  22149. 17:16:27And again, this is all against xx,
  22150. 17:16:30our x value, but this time we have y
  22151. 17:16:34down.
  22152. 17:16:36And let's do something a little fun with
  22153. 17:16:37this. We can put in a k dash dash. That
  22154. 17:16:42just tells it to make it a dotted line.
  22155. 17:16:46And if we're going to do the down one,
  22156. 17:16:49we also want to do the up one. So here's
  22157. 17:16:52our y
  22158. 17:16:55up. And when we run that, it adds both
  22159. 17:16:58sets of line. And so here's our support.
  22160. 17:17:01And this is what you expect. You expect
  22161. 17:17:03these two lines to go through the
  22162. 17:17:04nearest data point. So the dash lines go
  22163. 17:17:07through the nearest muffin and the
  22164. 17:17:08nearest cupcake when it's plotting it.
  22165. 17:17:11And then your SVM goes right down the
  22166. 17:17:13middle. So it gives it a nice split in
  22167. 17:17:14our data. And you can see how easy it is
  22168. 17:17:16to see based just on sugar and flour
  22169. 17:17:19which one's a muffin or a cupcake.
  22170. 17:17:23Let's go ahead and create a function
  22171. 17:17:28to predict
  22172. 17:17:31muffin or cupcake.
  22173. 17:17:34I've got my uh recipes. I pulled off the
  22174. 17:17:37um internet and I want to see the
  22175. 17:17:40difference between a
  22176. 17:17:42muffin or a cupcake. And so we need a
  22177. 17:17:44function to push that through. And uh we
  22178. 17:17:46create a function with deaf. And let's
  22179. 17:17:49call it muffin or cupcake. And remember,
  22180. 17:17:51we're just doing flour and sugar today.
  22181. 17:17:53We're not doing all the ingredients. And
  22182. 17:17:55that actually is a pretty good split.
  22183. 17:17:56You really don't need all the
  22184. 17:17:57ingredients to know it's flour and
  22185. 17:17:59sugar. And let's go ahead and do an if
  22186. 17:18:02else statement. So if model predict
  22187. 17:18:07is of flower and sugar equals zero. So
  22188. 17:18:11we take our model and we do run a
  22189. 17:18:13predict. It's very common in sklearn
  22190. 17:18:14where you have a predict. You put the
  22191. 17:18:17data in and it's going to return a
  22192. 17:18:19value. In this case if it equals zero
  22193. 17:18:21then print you're looking at a muffin
  22194. 17:18:23recipe. Else if it's not zero that means
  22195. 17:18:26it's one and you're looking at a cupcake
  22196. 17:18:29recipe. That's pretty straightforward
  22197. 17:18:31for
  22198. 17:18:32function or def for definition. Deaf is
  22199. 17:18:35how you do that in Python. And of
  22200. 17:18:38course, if you're going to create a
  22201. 17:18:38function, you should run something in
  22202. 17:18:40it. And so, let's run a cupcake. And
  22203. 17:18:42we're going to send it values 50 and 20.
  22204. 17:18:44A muffin or a cupcake. I don't know what
  22205. 17:18:45it is. And let's run this and just see
  22206. 17:18:48what it gives us. It says, "Oh, it's a
  22207. 17:18:50muffin. You're looking at a muffin
  22208. 17:18:52recipe." So, it very easily predicts
  22209. 17:18:54whether we're looking at a muffin or a
  22210. 17:18:55cupcake recipe. Let's plot this. There
  22211. 17:19:00we go. Plot this on the graph so we can
  22212. 17:19:02see what that actually looks like. And
  22213. 17:19:04I'm just going to copy and paste it from
  22214. 17:19:06below where we plotting all the points
  22215. 17:19:08in there.
  22216. 17:19:09So, this is nothing different than we
  22217. 17:19:11did before. If I run it, you'll see it
  22218. 17:19:13has all the points and the lines on
  22219. 17:19:15there. And what we want to do is we want
  22220. 17:19:17to add another point. And we'll do
  22221. 17:19:20pltot.
  22222. 17:19:23And if you remember correctly, we did
  22223. 17:19:24for our test we did 50
  22224. 17:19:27and 20. And then somebody went in here
  22225. 17:19:30and decided we'll do yo for yellow or
  22226. 17:19:33it's kind of a orangeish yellow color is
  22227. 17:19:34going to come out. Marker size nine.
  22228. 17:19:36Those are settings you can play with.
  22229. 17:19:38Somebody else played with them to come
  22230. 17:19:40up with the right setup so it looks
  22231. 17:19:41good. And you can see there it is
  22232. 17:19:43graphed clearly a muffin.
  22233. 17:19:46In this case in cupcakes versus muffins,
  22234. 17:19:50the muffin has won. And if you'd like to
  22235. 17:19:53do your own muffin cupcake contender
  22236. 17:19:56series, you certainly can send a note
  22237. 17:19:59down below and the team at SimplyLearn
  22238. 17:20:01will send you over the data they use for
  22239. 17:20:04the muffin and cupcake. And that's true
  22240. 17:20:05of any of the data. We didn't actually
  22241. 17:20:07run a plot on it earlier. We had men
  22242. 17:20:09versus women. You can also request that
  22243. 17:20:12information to run it on your data
  22244. 17:20:14setup. So you can test that out.
  22245. 17:20:17So to go back over our setup, we went
  22246. 17:20:19ahead for our support vector machine
  22247. 17:20:21code. We did a predict 40 parts flour,
  22248. 17:20:2420 parts sugar. I think it was different
  22249. 17:20:26than the one we did whether it's a
  22250. 17:20:28muffin or a cupcake. Hence, we have
  22251. 17:20:30built a classifier using SVM which is
  22252. 17:20:33able to classify if a recipe is of a
  22253. 17:20:36cupcake or a muffin. Which wraps up our
  22254. 17:20:39cupcake versus muffin. So the key
  22255. 17:20:41takeaways, what is machine learning? We
  22256. 17:20:44discussed that with some of the
  22257. 17:20:46different aspects of machine learning on
  22258. 17:20:48there. We went into types of machine
  22259. 17:20:50learning. If you memorize we have
  22260. 17:20:52supervised, unsupervised and
  22261. 17:20:54reinforcement learning. We discussed
  22262. 17:20:56regression line or best fit and we did
  22263. 17:20:59the building a decision tree and what
  22264. 17:21:02the logic is behind that. And finally we
  22265. 17:21:04did classification using SVM support
  22266. 17:21:07vector machine and we did the code in
  22267. 17:21:10there. Today we are diving into machine
  22268. 17:21:12learning, the technology behind things
  22269. 17:21:14like Netflix recommendations, CD, and
  22270. 17:21:17even the face unlock of your phone.
  22271. 17:21:19Machine learning helps devices get
  22272. 17:21:21smarter by learning from data and
  22273. 17:21:23predicting what we might like or need.
  22274. 17:21:25And here's why machine learning is huge
  22275. 17:21:27for your career. Right now, machine
  22276. 17:21:29learning jobs are among the fastest
  22277. 17:21:31growing roles worldwide. Companies in
  22278. 17:21:33every industry, tech, healthcare,
  22279. 17:21:36finance, and more, are looking for
  22280. 17:21:38people with machine learning skills to
  22281. 17:21:40improve their products, automate tasks,
  22282. 17:21:42and make smarter decisions. Machine
  22283. 17:21:45learning engineers in the US earn around
  22284. 17:21:47$112,000 on average with plenty of room
  22285. 17:21:50for growth as you gain experience. So,
  22286. 17:21:52if you want to jump into this exciting
  22287. 17:21:54field, learning machine learning can
  22288. 17:21:56open doors to highpaying in- demand
  22289. 17:21:58jobs. So in this video I'll guide you
  22290. 17:22:00through the ultimate road map to master
  22291. 17:22:02machine learning in 2025 one step at a
  22292. 17:22:05time. So let's get started. So in the
  22293. 17:22:08first month start with the foundations
  22294. 17:22:10of programming. So programming is a
  22295. 17:22:12language you'll use to communicate with
  22296. 17:22:14your computer and bring machine learning
  22297. 17:22:16algorithms to life. So this month is all
  22298. 17:22:19about Python, the language of choice for
  22299. 17:22:21most machine learning practitioners. So
  22300. 17:22:23here's what to focus on. First, learn
  22301. 17:22:26Python basics. Begin with Python's
  22302. 17:22:28fundamentals like variables, data types,
  22303. 17:22:31loops and functions. So spend time
  22304. 17:22:34writing small programs daily to get
  22305. 17:22:36comfortable. After that explore the key
  22306. 17:22:38libraries like numpy, pandas and
  22307. 17:22:41scikitlearn. So numpy is for numerical
  22308. 17:22:43operations. It makes handling large data
  22309. 17:22:46sets faster and easier. And pandas is to
  22310. 17:22:49manipulate and analyze data. So pandas
  22311. 17:22:52allow you to filter, sort and reshape
  22312. 17:22:54data in a breeze. And then scikitlearn
  22313. 17:22:56is for implementing algorithms in just a
  22314. 17:22:58few lines of code. So now you might have
  22315. 17:23:00heard about R, another language used in
  22316. 17:23:02machine learning. But don't stress about
  22317. 17:23:04it now. Python will serve you well,
  22318. 17:23:06especially as a beginner, because it's
  22319. 17:23:08simpler and more flexible. So aim to
  22320. 17:23:10spend an hour or two each day coding. By
  22321. 17:23:12the end of this month, you'll have a
  22322. 17:23:14solid base to build on. Now, in the
  22323. 17:23:16second month, get organized with version
  22324. 17:23:18control and data structures. So this
  22325. 17:23:20month is about learning how to organize
  22326. 17:23:22and manage your code effectively and
  22327. 17:23:24sharpening your problem solving skills
  22328. 17:23:26with data structures and algorithms. So
  22329. 17:23:28first is version control with git. So
  22330. 17:23:31think of git as your project history
  22331. 17:23:32tracker. So imagine working on a big
  22332. 17:23:35project and making changes then
  22333. 17:23:37realizing something went wrong. You want
  22334. 17:23:39to go back to an earlier version, right?
  22335. 17:23:41So that's where git comes in. And here's
  22336. 17:23:44what you should practice. Number one is
  22337. 17:23:46committing changes. So save different
  22338. 17:23:49versions of your work as you progress.
  22339. 17:23:51And then branching which means work on
  22340. 17:23:53separate features without affecting your
  22341. 17:23:55main code. And then comes merging which
  22342. 17:23:58means combining changes from different
  22343. 17:24:00versions once they are ready. So you
  22344. 17:24:02have to set up an account on GitHub or
  22345. 17:24:04GitLab to store your projects online. So
  22346. 17:24:06not only will this be super useful, but
  22347. 17:24:09it'll also start building your
  22348. 17:24:10portfolio. Now next is data structures
  22349. 17:24:13and algorithm. So think of data
  22350. 17:24:15structures like tools in a toolkit. So
  22351. 17:24:17each one like arrays, stacks, cues, etc.
  22352. 17:24:20serves a specific purpose. So here's how
  22353. 17:24:22to approach them. Number one, arrays and
  22354. 17:24:24lists. Now arrays and lists are for
  22355. 17:24:26storing data in sequence. After that,
  22356. 17:24:29you can get familiar with stacks and
  22357. 17:24:30cues. So stacks and cues are for tasks
  22358. 17:24:33that need ordered data access. And then
  22359. 17:24:35you have sorting and searching
  22360. 17:24:36algorithms. So these make your programs
  22361. 17:24:39more efficient. And that's super
  22362. 17:24:41important in machine learning where data
  22363. 17:24:42can get massive. So the goal here is to
  22364. 17:24:44build up your problem solving skills
  22365. 17:24:46which are key to machine learning
  22366. 17:24:48success. So take it slow, practice daily
  22367. 17:24:50and you'll see progress. Now in the
  22368. 17:24:53third month, learn to access data with
  22369. 17:24:55SQL. So in machine learning, a lot of
  22370. 17:24:57work involves accessing and organizing
  22371. 17:24:59data from databases. So SQL, a
  22372. 17:25:02structured query language, is your
  22373. 17:25:04ticket to getting the data you need for
  22374. 17:25:05training ML models. So here's what you
  22375. 17:25:08should focus on. Select and where. So
  22376. 17:25:10these commands help you pull specific
  22377. 17:25:12pieces of data and then you can move on
  22378. 17:25:14to joins. Joins usually combine data
  22379. 17:25:17from different tables. So this is so
  22380. 17:25:19powerful that you'll use it all the
  22381. 17:25:20time. And then comes group by and
  22382. 17:25:22aggregate functions. They are great for
  22383. 17:25:24summarizing data to find patterns. So
  22384. 17:25:27spend time working with sample databases
  22385. 17:25:29you can find online and practice writing
  22386. 17:25:32queries. Being comfortable with SQL will
  22387. 17:25:34save you time when preparing data for
  22388. 17:25:36your models. Now after completing the
  22389. 17:25:38third month you can move on to
  22390. 17:25:40mathematics which is building your
  22391. 17:25:42analytical mind. So this month we are
  22392. 17:25:44tackling the math behind machine
  22393. 17:25:45learning. So don't worry you don't need
  22394. 17:25:47to be a math genius but understanding
  22395. 17:25:49certain concepts will make everything
  22396. 17:25:51feel less mysterious. So in this month
  22397. 17:25:53you have to focus on linear algebra. So
  22398. 17:25:56this is the math behind how models see
  22399. 17:25:58data. So you can study vectors, matrices
  22400. 17:26:00and operations like multiplication. Next
  22401. 17:26:03comes calculus. So you'll use calculus
  22402. 17:26:05to help your models learn. So you have
  22403. 17:26:07to focus on derivatives and gradients
  22404. 17:26:09which help minimize errors in your
  22405. 17:26:11model. And then you can move on to
  22406. 17:26:13probability and statistics. So
  22407. 17:26:15understanding probability helps you make
  22408. 17:26:17sense of data. So learn about
  22409. 17:26:19distributions like normal distribution,
  22410. 17:26:20bormal distribution and then variance
  22411. 17:26:23and standard deviation. So once you have
  22412. 17:26:25learned maths, next you'll be moving on
  22413. 17:26:27to data handling and visualization which
  22414. 17:26:30is the heart of machine learning as you
  22415. 17:26:31all know. So with Matt under your belt,
  22416. 17:26:34it's time to dig into data handling and
  22417. 17:26:36visualization. So data preparation is
  22418. 17:26:38vital because your model is only as good
  22419. 17:26:40as the data you feed it. So number one
  22420. 17:26:43comes data manipulation. So using pandas
  22421. 17:26:45and numpy, you'll clean and organize
  22422. 17:26:47your data. You might be removing missing
  22423. 17:26:49values like clean up messy data so it
  22424. 17:26:51doesn't confuse your model. And then
  22425. 17:26:53you'll learn transforming variables like
  22426. 17:26:55converting data into formats that work
  22427. 17:26:57for models. And then you will move on to
  22428. 17:26:59encoding categorical data like changing
  22429. 17:27:02text data like female or male into
  22430. 17:27:04numbers. Now once you're done with data
  22431. 17:27:06manipulation, next comes data
  22432. 17:27:07visualization. So visualization is how
  22433. 17:27:10you get to see your data before training
  22434. 17:27:12a model. So here you have to learn
  22435. 17:27:13mattplot lip and seabboard. So you can
  22436. 17:27:16create line charts, histograms, scatter
  22437. 17:27:18plots and heat maps. So this lets you
  22438. 17:27:21explore patterns and spot outliers. So
  22439. 17:27:23understanding these patterns in your
  22440. 17:27:25data is crucial for building effective
  22441. 17:27:27models. Now in the sixth month you'll be
  22442. 17:27:29moving on to the machine learning
  22443. 17:27:31fundamentals. So now it's time to start
  22444. 17:27:33building your own models. So you will
  22445. 17:27:35focus on two main types of machine
  22446. 17:27:37learning this month. Number one comes
  22447. 17:27:39the supervised learning. So this is when
  22448. 17:27:41you train a model on label data where
  22449. 17:27:43the outcome is already known. So you'll
  22450. 17:27:46work with algorithms like linear
  22451. 17:27:47regression which predicts a continuous
  22452. 17:27:49outcome. Then you'll work with decision
  22453. 17:27:51trees which breaks down decisions into a
  22454. 17:27:53tree structure. And then you have
  22455. 17:27:55support vector machines under supervised
  22456. 17:27:56learning which updates data into
  22457. 17:27:58classes. Now after supervised learning
  22458. 17:28:00comes unsupervised learning. So here
  22459. 17:28:03your model identifies patterns in data
  22460. 17:28:05without labeled outcomes. So two popular
  22461. 17:28:08techniques in unsupervised learning is
  22462. 17:28:09number one clustering like K means
  22463. 17:28:11clustering which means group similar
  22464. 17:28:13data points and then you have
  22465. 17:28:15dimensionality reduction. This reduces
  22466. 17:28:17data complexity by focusing on key
  22467. 17:28:19features. So you can use scikitle learn
  22468. 17:28:21to try out these algorithms on sample
  22469. 17:28:23data sets. So this will give you
  22470. 17:28:25hands-on experience with model training
  22471. 17:28:27and you will learn to fine-tune them to
  22472. 17:28:29get better results. Now before moving
  22473. 17:28:31on, if you are interested in advancing
  22474. 17:28:33your career in the field of AI and
  22475. 17:28:35machine learning, simple learns
  22476. 17:28:36post-graduate program delivered in
  22477. 17:28:38collaboration with Purdue University and
  22478. 17:28:40IBM is a perfect opportunity. This
  22479. 17:28:43highly ranked program offers a
  22480. 17:28:45comprehensive curriculum covering
  22481. 17:28:47essential topics like machine learning,
  22482. 17:28:49deep learning, NLP, computer vision,
  22483. 17:28:52reinforcement learning, generative AI,
  22484. 17:28:54prompt engineering, and many more. With
  22485. 17:28:56hands-on experience to 25 plus projects
  22486. 17:28:58and access to 20 plus cutting edge
  22487. 17:29:00tools, you will gain the skills needed
  22488. 17:29:02to excel in today's competitive job
  22489. 17:29:04market. So join now and elevate your
  22490. 17:29:07expertise with the backing of Produce
  22491. 17:29:08academic excellence and IBM's
  22492. 17:29:10industry-leading insight. You can find
  22493. 17:29:12the course link in the description box
  22494. 17:29:14and pin comments. Now moving on to the
  22495. 17:29:16seventh month, you'll be building and
  22496. 17:29:18training models with advanced libraries.
  22497. 17:29:20So by now you have experimented with
  22498. 17:29:22some basic models. So let's step it up
  22499. 17:29:24with advanced tools like TensorFlow and
  22500. 17:29:26PyTorch. So these libraries offer more
  22501. 17:29:29flexibility and power. So TensorFlow and
  22502. 17:29:31PyTorch. So here you can start with
  22503. 17:29:33simple models and work your way up. So
  22504. 17:29:35these libraries allow for building
  22505. 17:29:37neural networks which you'll be studying
  22506. 17:29:39more on the next month. Now once you
  22507. 17:29:41have become familiar with TensorFlow and
  22508. 17:29:43PyTorch, you can move on to model
  22509. 17:29:44training and evaluation. So you have to
  22510. 17:29:46learn to split data into training and
  22511. 17:29:48testing sets and evaluate models using
  22512. 17:29:51metrics like accuracy and precision. So
  22513. 17:29:53your goal this month should be to get
  22514. 17:29:55comfortable with these libraries and
  22515. 17:29:57understand how they handle data and
  22516. 17:29:59model training behind the scenes. So
  22517. 17:30:01once you are done with this, you'll be
  22518. 17:30:02moving on to the eighth month where
  22519. 17:30:04you'll be dealing with advanced machine
  22520. 17:30:06learning. So this month's concept will
  22521. 17:30:08be number one on n symbol learning which
  22522. 17:30:10means combining multiple models to get
  22523. 17:30:12better predictions. So here you'll be
  22524. 17:30:14learning about bagging for example
  22525. 17:30:17random forests here multiple decision
  22526. 17:30:19trees make predictions and then you have
  22527. 17:30:21boosting like ada boost xg boost so
  22528. 17:30:24models learn from each other's mistakes
  22529. 17:30:26over here and after ensemble learning
  22530. 17:30:28comes deep learning. So here you explore
  22531. 17:30:31neural networks which mimic the human
  22532. 17:30:33brain. So you'll learn about neural
  22533. 17:30:34network basics. So you can start with
  22534. 17:30:36simple fully connected networks and then
  22535. 17:30:38you can move on to back propagation and
  22536. 17:30:40gradient descent. So these helps your
  22537. 17:30:42model learn and improve. So you can use
  22538. 17:30:45TensorFlow or PyTorch to practice
  22539. 17:30:47building neural networks. So you can
  22540. 17:30:48work on projects to reinforce these
  22541. 17:30:50concepts. Now moving on, you have two
  22542. 17:30:52specialize on topics like NLP and
  22543. 17:30:55computer vision. So machine learning
  22544. 17:30:57applications are so powerful and here
  22545. 17:30:59you'll get a taste of two major fields
  22546. 17:31:01which is NLP or natural language
  22547. 17:31:02processing. So here they work with text
  22548. 17:31:05data with tasks like sentiment analysis
  22549. 17:31:07and text classification. So you can
  22550. 17:31:09start with basic pre-processing like
  22551. 17:31:12tokenization, stop word removal and move
  22552. 17:31:14to building simple NLP models. After
  22553. 17:31:16that you can try computer vision. So for
  22554. 17:31:18image data you have to learn CNN
  22555. 17:31:21convolutional neural networks. So these
  22556. 17:31:23network analyze visual patterns making
  22557. 17:31:26them ideal for image classification. So
  22558. 17:31:28you practice with open data sets like
  22559. 17:31:30text, documents or images and apply the
  22560. 17:31:32concepts you will learn to see results
  22561. 17:31:34in real world applications. Now in the
  22562. 17:31:3610th month you'll be dealing with model
  22563. 17:31:38deployment which is bringing your models
  22564. 17:31:40to life. So here you'll be using Flask
  22565. 17:31:43or Django. So you can use these
  22566. 17:31:44frameworks to create a web API so users
  22567. 17:31:47can interact with your model. For
  22568. 17:31:49example, build a web app that lets
  22569. 17:31:50people upload images for classification.
  22570. 17:31:53And then you can also try out Docker. So
  22571. 17:31:55package your model and its dependencies
  22572. 17:31:57so it can run on any machine. So this is
  22573. 17:31:59super helpful for deploying models
  22574. 17:32:01without compatibility issues. So by the
  22575. 17:32:03end of this month, you'll be able to
  22576. 17:32:04share your models with the world. So
  22577. 17:32:06moving on to the 11th month, you'll be
  22578. 17:32:08starting with cloud and production. So
  22579. 17:32:11this month, you'll learn how to deploy
  22580. 17:32:12models on the cloud and ensure they
  22581. 17:32:14perform well in real world environments.
  22582. 17:32:16So you'll be dealing with cloud
  22583. 17:32:17platforms like AWS, Google Cloud or
  22584. 17:32:20Azure. So you have to learn to deploy
  22585. 17:32:22models of the cloud provider
  22586. 17:32:23accessibility and scalability. And then
  22587. 17:32:26comes monitoring and maintenance. So
  22588. 17:32:28understand how to track your models
  22589. 17:32:29performance over time and update it as
  22590. 17:32:32needed. So these skills are essential
  22591. 17:32:34for maintaining models in production and
  22592. 17:32:36ensuring they stay reliable. And finally
  22593. 17:32:38you will be creating real world projects
  22594. 17:32:40and portfolio building. So here you have
  22595. 17:32:42to choose topics that interest you and
  22596. 17:32:45showcase your skills. So first you can
  22597. 17:32:47start with full projects. So complete
  22598. 17:32:49projects that go from data cleaning and
  22599. 17:32:51model building to deployment. So ideas
  22600. 17:32:53could be a sentiment analysis tool or an
  22601. 17:32:55image recognition app. And then you have
  22602. 17:32:57to build your portfolio. So organize and
  22603. 17:32:59document your projects, host them on
  22604. 17:33:01GitHub and create an online portfolio to
  22605. 17:33:03share with potential employers or
  22606. 17:33:05collaborators. So by following this road
  22607. 17:33:07map, you'll be well prepared to handle
  22608. 17:33:09real world machine learning challenges
  22609. 17:33:11and have an impressive portfolio to show
  22610. 17:33:13for it.
  22611. 17:33:13>> Welcome to machine learning tutorial
  22612. 17:33:16part two. My name is Richard Kersner
  22613. 17:33:18with the SimplyLearn team. That is
  22614. 17:33:20www.simplearn.com.
  22615. 17:33:23Get certified, get ahead. Today in our
  22616. 17:33:26second tutorial, we're going to cover K
  22617. 17:33:28means linear regression along with going
  22618. 17:33:31over the quiz questions we had during
  22619. 17:33:33our first tutorial. What's in it for
  22620. 17:33:36you? We're going to cover clustering.
  22621. 17:33:38What is clustering? K means clustering
  22622. 17:33:41which is one of the most common used
  22623. 17:33:43clustering tools out there including a
  22624. 17:33:45flowchart to understand K means
  22625. 17:33:47clustering and how it functions and then
  22626. 17:33:48we'll do an actual Python live demo on
  22627. 17:33:51clustering of cars based on brands. Then
  22628. 17:33:54we're going to cover logistic
  22629. 17:33:55regression. What is logistic regression?
  22630. 17:33:58Logistic regression curve and sigmoid
  22631. 17:34:00function. And then we'll do another
  22632. 17:34:02Python code demo to classify a tumor as
  22633. 17:34:05malignant or benign based on features.
  22634. 17:34:08And let's start with clustering. Suppose
  22635. 17:34:10we have a pile of books of different
  22636. 17:34:12genres. Now we divide them into
  22637. 17:34:14different groups like fiction, horror,
  22638. 17:34:17education, and as we can see from this
  22639. 17:34:20young lady, she definitely is into heavy
  22640. 17:34:22horror. You can just tell by those eyes
  22641. 17:34:23and the maple Canadian leaf on her
  22642. 17:34:25shirt. But we have fiction, horror, and
  22643. 17:34:27education. And we want to go ahead and
  22644. 17:34:29divide our books up. Well, organizing
  22645. 17:34:31objects into groups based on similarity
  22646. 17:34:33is clustering. And in this case, as
  22647. 17:34:36we're looking at the books, we're
  22648. 17:34:37talking about clustering things with
  22649. 17:34:39known categories. But you can also use
  22650. 17:34:41it to explore data. So you might not
  22651. 17:34:43know the categories. You just know that
  22652. 17:34:45you need to divide it up in some way to
  22653. 17:34:47conquer the data and to organize it
  22654. 17:34:49better. But in this case, we're going to
  22655. 17:34:51be looking at clustering in specific
  22656. 17:34:52categories. And let's just take a deeper
  22657. 17:34:54look at that. We're going to use K means
  22658. 17:34:57clustering. K means clustering is
  22659. 17:34:59probably the most commonly used
  22660. 17:35:00clustering tool in the machine learning
  22661. 17:35:02library. K means clustering is an
  22662. 17:35:05example of unsupervised learning. If you
  22663. 17:35:07remember from our previous thing, it is
  22664. 17:35:10used when you have unlabeled data. So we
  22665. 17:35:12don't know the answer yet. We have a
  22666. 17:35:14bunch of data that we want to cluster to
  22667. 17:35:16different groups. Define clusters in the
  22668. 17:35:18data based on feature similarity. So
  22669. 17:35:21we've introduced a couple terms here.
  22670. 17:35:23We've already talked about unsupervised
  22671. 17:35:25learning and unlabeled data. So we don't
  22672. 17:35:28know the answer yet. We're just going to
  22673. 17:35:30group stuff together and see if we can
  22674. 17:35:31find an unanswer
  22675. 17:35:33connect. We've also introduced feature
  22676. 17:35:36similarity. Features being different
  22677. 17:35:38features of the data. Now, with books,
  22678. 17:35:40we can easily see fiction and horror and
  22679. 17:35:44history books. But a lot of times with
  22680. 17:35:46data, some of that information isn't so
  22681. 17:35:48easy to see right when we first look at
  22682. 17:35:50it. And so, K means is one of those
  22683. 17:35:51tools where we can start finding things
  22684. 17:35:53that connect that match with each other.
  22685. 17:35:55Suppose we have these data points and
  22686. 17:35:57want to assign them into a cluster. Now
  22687. 17:36:00when I look at these data points, I
  22688. 17:36:01would probably group them into two
  22689. 17:36:03clusters just by looking at them. I'd
  22690. 17:36:04say two of these group of data kind of
  22691. 17:36:06come together. But in K means we pick K
  22692. 17:36:09clusters and assign random centrids to
  22693. 17:36:12clusters where the K clusters represents
  22694. 17:36:15two different clusters. We pick K
  22695. 17:36:17clusters and say random centroidids to
  22696. 17:36:19the clusters. Then we compute distance
  22697. 17:36:21from objects to the centrids. Now we
  22698. 17:36:24form new clusters based on minimum
  22699. 17:36:26distances and calculate the centrids. So
  22700. 17:36:29we figure out what the best distance is
  22701. 17:36:32for the centrid. Then we move the
  22702. 17:36:33centrid and recalculate those distances.
  22703. 17:36:35Repeat previous two steps iteratively
  22704. 17:36:38till the cluster centroid stop changing
  22705. 17:36:40their positions and become static.
  22706. 17:36:42Repeat previous two steps iteratively
  22707. 17:36:44till the cluster centroid stop changing
  22708. 17:36:46and the positions become static. Once
  22709. 17:36:48the clusters become static, then K means
  22710. 17:36:50clustering algorithm is said to be
  22711. 17:36:52converged. And there's another term we
  22712. 17:36:54see throughout machine learning is
  22713. 17:36:56converged. That means whatever math
  22714. 17:36:58we're using to figure out the answer has
  22715. 17:37:00come to a solution or it's converged on
  22716. 17:37:02an answer. Shall we see the flowchart to
  22717. 17:37:04understand make a little bit more sense
  22718. 17:37:06by putting it into a nice easy step by
  22719. 17:37:08step? So we start, we choose K. We'll
  22720. 17:37:11look at the elbow method in just a
  22721. 17:37:13moment. We assign random centrids to
  22722. 17:37:16clusters and sometimes you pick the
  22723. 17:37:18centrids because you might look at the
  22724. 17:37:20data in a in a graph and say ah these
  22725. 17:37:21are probably the central points. Then we
  22726. 17:37:24compute the distance from the objects to
  22727. 17:37:26the centrids. We take that and we form
  22728. 17:37:29new clusters based on minimum distance
  22729. 17:37:31and calculate their centrids. Then we
  22730. 17:37:33compute the distance from objects to the
  22731. 17:37:35new centrids. And then we go back and
  22732. 17:37:37repeat those last two steps. We
  22733. 17:37:39calculate the distances. So as we're
  22734. 17:37:42doing it, it brings into the new centrid
  22735. 17:37:44and then we move the centrid around and
  22736. 17:37:46we figure out what the best which
  22737. 17:37:48objects are closest to each centrid. So
  22738. 17:37:50the objects can switch from one centroid
  22739. 17:37:52to the other as the centroidids are
  22740. 17:37:53moved around and we continue that until
  22741. 17:37:55it is converged. Let's see an example of
  22742. 17:37:58this. Suppose we have this data set of
  22743. 17:38:01seven individuals and their score on two
  22744. 17:38:03topics A and B. Uh so here's our subject
  22745. 17:38:07in this case referring to the person
  22746. 17:38:09taking the uh test and then we have
  22747. 17:38:12subject A where we see what they've
  22748. 17:38:14scored on their first subject and we
  22749. 17:38:15have subject B and we can see what they
  22750. 17:38:17score on the second subject. Now let's
  22751. 17:38:19take two farthest apart points as
  22752. 17:38:21initial cluster centroidids. Now
  22753. 17:38:23remember we talked about selecting them
  22754. 17:38:25randomly or we can also just put them in
  22755. 17:38:27different points and pick the furthest
  22756. 17:38:28one apart so they move together. Either
  22757. 17:38:30one works okay depending on what kind of
  22758. 17:38:32data you're working on and what you know
  22759. 17:38:34about it. So we took the two furthest
  22760. 17:38:36points one and one and five and seven.
  22761. 17:38:40And now let's take the two farthest
  22762. 17:38:41apart points as initial cluster
  22763. 17:38:43centrids. Each point is then assigned to
  22764. 17:38:46the closest cluster with respect to the
  22765. 17:38:48distance from the centrids. So we take
  22766. 17:38:51each one of these points in there. We
  22767. 17:38:52measure that distance. And you can see
  22768. 17:38:53that if we measured each of those
  22769. 17:38:55distances and you use the the
  22770. 17:38:57Pythagorean theorem for a triangle in
  22771. 17:38:59this case because you know the x and the
  22772. 17:39:01y and you can figure out the diagonal
  22773. 17:39:04line from that or you can just take a
  22774. 17:39:05ruler and put it on your monitor. That'd
  22775. 17:39:07be kind of silly but it would work if
  22776. 17:39:08you're just eyeballing it. You can see
  22777. 17:39:10how they naturally come together in
  22778. 17:39:12certain areas. Now we again calculate
  22779. 17:39:15the centroidids of each cluster. So
  22780. 17:39:17cluster one and then cluster two and we
  22781. 17:39:19look at each individual dot. There's
  22782. 17:39:22one, two, three. We're in one cluster.
  22783. 17:39:24Uh the centrid then moves over. It
  22784. 17:39:26becomes 1.8 comma 2.3. So remember it
  22785. 17:39:30was at 1 and one. Well, the very center
  22786. 17:39:32of the data we're looking at would put
  22787. 17:39:33it at the one point roughly 22, but 1.8
  22788. 17:39:36and 2.3. And the second one, if we
  22789. 17:39:38wanted to make the overall mean vector,
  22790. 17:39:41the average vector of all the different
  22791. 17:39:42distances to that centrid, we come up
  22792. 17:39:44with 4, 1, and 54. So we've now moved
  22793. 17:39:48the centrids. We compare each
  22794. 17:39:50individual's distance to its own cluster
  22795. 17:39:52mean and to that of the opposite cluster
  22796. 17:39:54and we find build a nice chart on here
  22797. 17:39:57that the as we move that centrid around
  22798. 17:39:59we now have a new different kind of
  22799. 17:40:01clustering of groups and using uklidian
  22800. 17:40:03distance between the points and the mean
  22801. 17:40:05we get the same formula you see new
  22802. 17:40:07formulas coming up. So we have our
  22803. 17:40:09individual dots distance to the mean
  22804. 17:40:11centrid of the cluster and distance to
  22805. 17:40:12the mean centrid of the cluster. Only
  22806. 17:40:14individual three is nearer to the mean
  22807. 17:40:16of the opposite cluster cluster two than
  22808. 17:40:19its own cluster one. And you can see
  22809. 17:40:21here in the diagram where we've kind of
  22810. 17:40:23circled that one in the middle. So when
  22811. 17:40:25we've moved the clust the centroidids of
  22812. 17:40:27the clusters over one of the points
  22813. 17:40:29shifted to the other cluster because
  22814. 17:40:30it's closer to that group of
  22815. 17:40:32individuals. Thus, individual 3 is
  22816. 17:40:34relocated to cluster two, resulting in a
  22817. 17:40:37new partition. And we regenerate all
  22818. 17:40:39those numbers of how close they are to
  22819. 17:40:41the different clusters. For the new
  22820. 17:40:43clusters, we will find the actual
  22821. 17:40:44cluster centroidids. So now we move the
  22822. 17:40:47centrids over. And you can see that
  22823. 17:40:49we've now formed two very distinct
  22824. 17:40:50clusters on here. On comparing the
  22825. 17:40:52distance of each individual's distance
  22826. 17:40:54to its own cluster mean and to that of
  22827. 17:40:56the opposite cluster, we find that the
  22828. 17:40:58data points are stable. Hence, we have
  22829. 17:41:00our final clusters. Now if you remember
  22830. 17:41:03I brought up a concept earlier K mean on
  22831. 17:41:05the K means algorithm choosing the right
  22832. 17:41:08value of K will help in less number of
  22833. 17:41:10iterations and to find the appropriate
  22834. 17:41:12number of clusters in a data set we use
  22835. 17:41:14the elbow method and within sum of
  22836. 17:41:18squares WSS is defined as the sum of the
  22837. 17:41:20squared distance between each member of
  22838. 17:41:23the cluster and its centrid and so you
  22839. 17:41:25see we've done here is we have the
  22840. 17:41:27number of clusters and as you do the
  22841. 17:41:30same K means algorithm over the
  22842. 17:41:33different clusters and you calculate
  22843. 17:41:35what that centrid looks like and you
  22844. 17:41:36find the optimal you can actually find
  22845. 17:41:38the optimal number of clusters using the
  22846. 17:41:40elbow the graph is called as the elbow
  22847. 17:41:42method and on this we guessed at two
  22848. 17:41:44just by looking at the data but as you
  22849. 17:41:46can see the slope you actually just look
  22850. 17:41:48for right there where the elbow is in
  22851. 17:41:50the slope and you have a clear answer
  22852. 17:41:51that we want two different to start with
  22853. 17:41:54k means equals two a lot of times people
  22854. 17:41:56end up computing k means equals 2 3 four
  22855. 17:41:59five until they find the value which
  22856. 17:42:02fits on the elbow joint. Sometimes you
  22857. 17:42:04can just look at the data and if you're
  22858. 17:42:06really good with that specific domain
  22859. 17:42:08remember domain I mentioned that last
  22860. 17:42:10time you'll know that that where to pick
  22861. 17:42:12those numbers and where to start
  22862. 17:42:13guessing at what that k value is. So
  22863. 17:42:15let's take this and we're going to use a
  22864. 17:42:17use case using k means clustering to
  22865. 17:42:20cluster cars into brands using
  22866. 17:42:22parameters such as horsepower, cubic
  22867. 17:42:24inches, make, year, etc. So, we're going
  22868. 17:42:27to use the data set cars data having
  22869. 17:42:30information about three brands of cars,
  22870. 17:42:32Toyota, Honda, and Nissan. We'll go back
  22871. 17:42:35to my favorite tool, the Anaconda
  22872. 17:42:37Navigator with the Jupiter notebook. And
  22873. 17:42:41let's go ahead and flip over to our
  22874. 17:42:42Jupyter notebook. And in our Jupyter
  22875. 17:42:45Notebook, I'm going to go ahead and just
  22876. 17:42:46paste the uh basic code that we usually
  22877. 17:42:49start a lot of these off with. We're not
  22878. 17:42:51going to go too much into this code
  22879. 17:42:52because we've already discussed numpy.
  22880. 17:42:54We've already discussed mapplot library
  22881. 17:42:56and pandas. Numpy being the number
  22882. 17:42:58array, pandas being the pandas data
  22883. 17:43:00frame and mattplot for the graphing. And
  22884. 17:43:03don't forget uh since if you're using
  22885. 17:43:04the Jupyter notebook, you do need the
  22886. 17:43:06mattplot library in line so that it
  22887. 17:43:08plots everything on the screen. If
  22888. 17:43:10you're using a different Python editor,
  22889. 17:43:13then you probably don't need that
  22890. 17:43:14because it'll have a popup window on
  22891. 17:43:16your computer. And we'll go ahead and
  22892. 17:43:18run this just to load our libraries and
  22893. 17:43:20our setup into here. The next step is of
  22894. 17:43:23course to look at our data which I've
  22895. 17:43:25already opened up in a spreadsheet. And
  22896. 17:43:28you can see here we have the miles per
  22897. 17:43:29gallon, cylinders, cubic inches,
  22898. 17:43:32horsepower, weight pounds, how you know
  22899. 17:43:34how heavy it is, time it takes to get to
  22900. 17:43:3660. My card is probably on this one at
  22901. 17:43:39about 80 or 90. What year it is? So this
  22902. 17:43:42is you can actually see this is kind of
  22903. 17:43:43older cars and then the brand Toyota,
  22904. 17:43:46Honda, Nissan. So the different cars are
  22905. 17:43:48coming from all the way from 1971 if we
  22906. 17:43:51scroll down to uh the 80s. We have
  22907. 17:43:54between the 70s and 80s a number of cars
  22908. 17:43:55that they've put out. And let's uh we
  22909. 17:43:58come back here. We're going to do
  22910. 17:43:59importing the data. So we'll go ahead
  22911. 17:44:01and do data set equals and we'll use
  22912. 17:44:04pandas to read this in. And it's uh from
  22913. 17:44:06a CSV file. Remember, you can always
  22914. 17:44:08post this in the comments and request
  22915. 17:44:11the data files for these either in the
  22916. 17:44:13comments here on the YouTube video or go
  22917. 17:44:15to simplylearn.com and request that. The
  22918. 17:44:18car CSV, I put it in the same folder as
  22919. 17:44:21the code that I've stored. So, my Python
  22920. 17:44:23code is stored in the same folder, so I
  22921. 17:44:25don't have to put the full path. If you
  22922. 17:44:27store them in different folders, you do
  22923. 17:44:28have to change this and double check
  22924. 17:44:30your name variables. And we'll go ahead
  22925. 17:44:31and run this. And uh we've chosen data
  22926. 17:44:34set arbitrarily because, you know, it's
  22927. 17:44:35a data set we're importing. And we've
  22928. 17:44:37now imported our car CSV into the data
  22929. 17:44:39set. As you know, you have to prep the
  22930. 17:44:42data. So, we're going to create the X
  22931. 17:44:43data. This is the one that we're going
  22932. 17:44:45to try to figure out what's going on
  22933. 17:44:47with. And then there is a number of ways
  22934. 17:44:49to do this, but we'll do it in a simple
  22935. 17:44:51loop so you can actually see what's
  22936. 17:44:53going on. So, we'll do for i and x.c
  22937. 17:44:57columns. So, we're going to go through
  22938. 17:44:58each of the columns. And a lot of times
  22939. 17:45:01it's important I I'll make lists of the
  22940. 17:45:04columns and do this because I might
  22941. 17:45:06remove certain columns or there might be
  22942. 17:45:08columns that I want to be processed
  22943. 17:45:10differently. But for this we can go
  22944. 17:45:12ahead and take x of i and we want to go
  22945. 17:45:16fill na and that's a pandas command. But
  22946. 17:45:19the question is what are we going to
  22947. 17:45:20fill the missing data with? We
  22948. 17:45:22definitely don't want to just put in a
  22949. 17:45:24number that doesn't actually mean
  22950. 17:45:25something. And so one of the tricks you
  22951. 17:45:27can do with this is we can take x of i.
  22952. 17:45:31And in addition to that, we want to go
  22953. 17:45:33ahead and turn this into an integer
  22954. 17:45:35because a lot of these are integers. So
  22955. 17:45:37we'll go ahead and keep it integers. And
  22956. 17:45:39me add the bracket here. And a lot of
  22957. 17:45:41editors will do this. They'll think that
  22958. 17:45:42you're closing one bracket. Make sure
  22959. 17:45:44you get that second bracket in there if
  22960. 17:45:45it's a double bracket. That's always
  22961. 17:45:47something that happens regularly. So
  22962. 17:45:49once we have our integer of x of yi,
  22963. 17:45:51this is going to fill in any missing
  22964. 17:45:53data with the average. And I was so busy
  22965. 17:45:55closing one set of brackets, I forgot
  22966. 17:45:57that the mean is also has brackets in
  22967. 17:45:59there for the pandas. So we can see
  22968. 17:46:01here, we're going to fill in all the
  22969. 17:46:02data with the average value for that
  22970. 17:46:04column. So if there's missing data is in
  22971. 17:46:06the average of the data it does have.
  22972. 17:46:08Then once we've done that, we'll go
  22973. 17:46:10ahead and loop through it again
  22974. 17:46:12and just check and see to make sure
  22975. 17:46:15everything is filled in correctly. And
  22976. 17:46:17we'll print and then we take x is null.
  22977. 17:46:21And this returns a set of the null value
  22978. 17:46:23or the how many lines are null. And
  22979. 17:46:25we'll just sum that up to see what that
  22980. 17:46:27looks like. And so when I run this and
  22981. 17:46:29so with the X, what we want to do is we
  22982. 17:46:31want to remove the last column because
  22983. 17:46:33that had the models. That's what we're
  22984. 17:46:35trying to see if we can cluster these
  22985. 17:46:36things and figure out the models. There
  22986. 17:46:38is so many different ways to sort the X
  22987. 17:46:41out. For one, we could take the X and we
  22988. 17:46:44could go data set, our variable we're
  22989. 17:46:47using, and use the eyelocation, one of
  22990. 17:46:50the features that's in pandas, and we
  22991. 17:46:53could take that and then take all the
  22992. 17:46:55rows and all but the last column of the
  22993. 17:46:58data set. And at this time, we could do
  22994. 17:47:01values. We just convert it to values.
  22995. 17:47:03So, that's one way to do this. And if I
  22996. 17:47:05let me just put this down here and print
  22997. 17:47:08X, it's a capital X we chose. and I run
  22998. 17:47:11this, you can see it's just the values.
  22999. 17:47:13We could also take out the values and
  23000. 17:47:16it's not going to return anything
  23001. 17:47:17because there's no values connected to
  23002. 17:47:19it. What I like to do with this is
  23003. 17:47:21instead of doing the location which does
  23004. 17:47:24integers more common is to come in here
  23005. 17:47:26and we have our data set and we're going
  23006. 17:47:29to do data set dot or data set columns.
  23007. 17:47:34And remember that lists all the columns.
  23008. 17:47:36So if I come in here, let me just mark
  23009. 17:47:40that as red and I print data set.c
  23010. 17:47:45columns.
  23011. 17:47:48You can see that I have my index here. I
  23012. 17:47:49have my MPG cylinders everything
  23013. 17:47:52including the brand which we don't want.
  23014. 17:47:54So the way to get rid of the brand would
  23015. 17:47:56be to do data columns of everything but
  23016. 17:47:59the last one minus one. So now if I
  23017. 17:48:01print this, you'll see the brand
  23018. 17:48:03disappears. And so I can actually just
  23019. 17:48:05take data set columns minus one and I'll
  23020. 17:48:10put it right in here for the columns
  23021. 17:48:12we're going to look at.
  23022. 17:48:14And let's unmark this.
  23023. 17:48:17And unmark this.
  23024. 17:48:20And now if I do an x.ad
  23025. 17:48:23I now have a new data frame. And you can
  23026. 17:48:26see right here we have all the different
  23027. 17:48:28columns except for the brand at the end
  23028. 17:48:29of the year. And it turns out when you
  23029. 17:48:33start playing with the data set, you're
  23030. 17:48:35going to get an error later on and it'll
  23031. 17:48:36say cannot convert string to float
  23032. 17:48:40value. And that's because it for some
  23033. 17:48:42reason these things the way they
  23034. 17:48:43recorded them must have been recorded as
  23035. 17:48:44strings. So we have a neat feature in
  23036. 17:48:47here on pandas to convert. And it is
  23037. 17:48:50simply convert objects.
  23038. 17:48:55And for this we're going to do convert
  23039. 17:48:57oops convert underscore
  23040. 17:49:01numeric numeric equals true. And yes, I
  23041. 17:49:05did have to go look that up. I don't
  23042. 17:49:07have it memorized the convert numeric in
  23043. 17:49:09there. If I'm working with a lot of
  23044. 17:49:10these things, I remember them, but um
  23045. 17:49:13depending on where I'm at, what I'm
  23046. 17:49:14doing, I usually have to look it up. And
  23047. 17:49:16we run that. Oops, I must have missed
  23048. 17:49:18something in here. Let me double check
  23049. 17:49:19my spelling. And when I double check my
  23050. 17:49:21spilling, you'll see I missed the first
  23051. 17:49:23underscore in the convert objects. And
  23052. 17:49:25when I run this, it now has everything
  23053. 17:49:27converted into a numeric value because
  23054. 17:49:31that's what we're going to be working
  23055. 17:49:32with is numeric values down here.
  23056. 17:49:35And the next part is that we need to go
  23057. 17:49:38through the data and eliminate null
  23058. 17:49:40values. Most people when they're doing
  23059. 17:49:42small amounts, you working with small
  23060. 17:49:44data pools discover afterwards that they
  23061. 17:49:46have a null value and they have to go
  23062. 17:49:47back and do this. So, you know, be aware
  23063. 17:49:50whenever we're formatting this data,
  23064. 17:49:52things are going to pop up and sometimes
  23065. 17:49:54you go backwards to fix it. And that's
  23066. 17:49:56fine. That's just part of exploring the
  23067. 17:49:58data and understanding what you have.
  23068. 17:50:01And I should have done this earlier, but
  23069. 17:50:03let me go ahead and increase the size of
  23070. 17:50:05my window one notch.
  23071. 17:50:09There we go. Easier to see.
  23072. 17:50:12So, we'll do 4 I in working with X dot
  23073. 17:50:16columns. will page through all the
  23074. 17:50:17columns. And we want to take X of I and
  23075. 17:50:21we're going to change that. We're going
  23076. 17:50:22to alter it. And so with this, we want
  23077. 17:50:25to go ahead and fill in X of I. Pandas
  23078. 17:50:29has the fill in a. And that just fills
  23079. 17:50:32in any non-existent missing data. And
  23080. 17:50:36we'll put my brackets up. And there's a
  23081. 17:50:38lot of different ways to fill this data.
  23082. 17:50:41If you have a really large data set,
  23083. 17:50:43some people just void out that data
  23084. 17:50:45because if and then look at it later in
  23085. 17:50:46a separate exploration of data. One of
  23086. 17:50:49the tricks we can do is we can take our
  23087. 17:50:53column and we can find the means
  23088. 17:50:56and the means is in there or quotation
  23089. 17:50:59marks. So we take the columns, we're
  23090. 17:51:01going to fill in the non-existing one
  23091. 17:51:03with the means. The problem is that
  23092. 17:51:05returns a decimal float. So some of
  23093. 17:51:08these aren't decimals. Certainly, you
  23094. 17:51:11may need to be a little careful of doing
  23095. 17:51:12this, but for this example, we're just
  23096. 17:51:14going to fill it in with the integer
  23097. 17:51:16version of this. Keeps it on par with
  23098. 17:51:18the other data that isn't a decimal
  23099. 17:51:20point.
  23100. 17:51:23And then what we also want to do is we
  23101. 17:51:24want to double check. A lot of times you
  23102. 17:51:27do this first part first to double
  23103. 17:51:29check, then you do the fill, and then
  23104. 17:51:30you do it again just to make sure you
  23105. 17:51:31did it right. So, we're going to go
  23106. 17:51:33through and test for missing data. And
  23107. 17:51:37one of the re ways you can do that is
  23108. 17:51:40simply go in here and take our X of I
  23109. 17:51:44column. So it's going to go through the
  23110. 17:51:46X of I column. It says is null. So it's
  23111. 17:51:48going to return any any place there's a
  23112. 17:51:50null value. It actually goes through all
  23113. 17:51:52the rows of each column is null. And
  23114. 17:51:55then we want to go ahead and sum that.
  23115. 17:51:57So we take that, we add the sum value.
  23116. 17:51:59And these are all pandas. So is null is
  23117. 17:52:01a panda command and so is sum. And if we
  23118. 17:52:04go through that and we go ahead and run
  23119. 17:52:05it
  23120. 17:52:09and we go ahead and take and run that,
  23121. 17:52:10you'll see that all the columns have
  23122. 17:52:12zero null values. So we've now tested
  23123. 17:52:15and double checked and our data is nice
  23124. 17:52:16and clean. We have no null values.
  23125. 17:52:18Everything is now a number value. We
  23126. 17:52:20turned it into numeric and we've removed
  23127. 17:52:23the last column in our data. And at this
  23128. 17:52:26point, we're actually going to start
  23129. 17:52:28using the elbow method to find the
  23130. 17:52:30optimal number of clusters. So, we're
  23131. 17:52:32now actually getting into the sklearn
  23132. 17:52:34part. Uh, the K means clustering on
  23133. 17:52:37here. I guess we'll go ahead and zoom it
  23134. 17:52:40up one more notch so you can see what
  23135. 17:52:41I'm typing in here.
  23136. 17:52:44And then from sklearn going to or
  23137. 17:52:48sklearn
  23138. 17:52:50cluster, we're going to import K means.
  23139. 17:52:57I always forget to capitalize the K and
  23140. 17:52:59the M when I do this. So it's capital K,
  23141. 17:53:01capital M K means.
  23142. 17:53:05And we'll go and create a um array WCSS
  23143. 17:53:09equals we'll make it an empty array. If
  23144. 17:53:11you remember from the elbow method from
  23145. 17:53:14our slide
  23146. 17:53:16within the sums of squares, WSS is
  23147. 17:53:18defined as the sum of squared distance
  23148. 17:53:21between each member of the cluster and
  23149. 17:53:23it centrid. So we're looking at that
  23150. 17:53:25change in differences as far as a
  23151. 17:53:27squared distance. And we're going to run
  23152. 17:53:29this over a number of K mean values.
  23153. 17:53:33In fact, let's go for I in range. We'll
  23154. 17:53:36do 11 of them.
  23155. 17:53:39Range zero of 11.
  23156. 17:53:41And the first thing we're going to do is
  23157. 17:53:42we're going to create the actual we'll
  23158. 17:53:45do it all lowercase.
  23159. 17:53:51And so we're going to create this object
  23160. 17:53:54from the K means that we just imported.
  23161. 17:53:58And the variable that we want to put
  23162. 17:54:00into this is in clusters. We're going to
  23163. 17:54:05set that equals to I. That's the most
  23164. 17:54:07important one because we're looking at
  23165. 17:54:08how increasing the number of clusters
  23166. 17:54:11changes our answer. There are a lot of
  23167. 17:54:14settings to the K means. Our guys in the
  23168. 17:54:17back did a great job just kind of
  23169. 17:54:19playing with some of them. The most
  23170. 17:54:21common ones that you see in a lot of
  23171. 17:54:23stuff is how you enit your K means. So
  23172. 17:54:26we have K means plus plus. This is just
  23173. 17:54:30a tool to let the model itself be smart
  23174. 17:54:33how it picks it centrids to start with
  23175. 17:54:35its initial centroidids. We only want to
  23176. 17:54:37iterate no more than 300 times. We have
  23177. 17:54:39a max iteration we put in there. We have
  23178. 17:54:42the infinite the random state equals
  23179. 17:54:44zero. You really don't need to worry too
  23180. 17:54:46much about these when you're first
  23181. 17:54:48learning this. As you start digging in
  23182. 17:54:50deeper, you start finding that these are
  23183. 17:54:51shortcuts that will speed up the process
  23184. 17:54:55as far as a setup. But the big one that
  23185. 17:54:57we're working with is the inclusters
  23186. 17:54:59equals I. So, we're going to literally
  23187. 17:55:02train our K means 11 times. We're going
  23188. 17:55:04to do this process 11 times. And if
  23189. 17:55:08you're working with big data, you know,
  23190. 17:55:10the first thing you do is you run a
  23191. 17:55:12small sample of the data so you can test
  23192. 17:55:13all your stuff on it. And you can
  23193. 17:55:15already see the problem that if I'm
  23194. 17:55:17going to iterate through a terabyte of
  23195. 17:55:19data 11 times and then the K means
  23196. 17:55:22itself is iterating through the data
  23197. 17:55:23multiple times. That's a heck of a
  23198. 17:55:25process. So you got to be a little
  23199. 17:55:27careful with this. A lot of times though
  23200. 17:55:29you can find your elbow using the elbow
  23201. 17:55:32method. Find your optimal number on a
  23202. 17:55:34sample of data especially if you're
  23203. 17:55:35working with larger data sources. So we
  23204. 17:55:38want to go ahead and take our K means
  23205. 17:55:39and we're just going to fit it. If
  23206. 17:55:41you're looking at any of the sklearn,
  23207. 17:55:43very common that you fit your model. And
  23208. 17:55:45if you remember correctly, our variable
  23209. 17:55:46we're using is the capital X. And once
  23210. 17:55:49we fit this value, we go back to the um
  23211. 17:55:53array we made. And we want to go and
  23212. 17:55:54just append that value on the end.
  23213. 17:55:57And it's not the actual fit we're
  23214. 17:55:59pinning in there. It's when it generates
  23215. 17:56:01it, it generates the value you're
  23216. 17:56:03looking for is inertia. So k
  23217. 17:56:05means.inertia will pull that specific
  23218. 17:56:07value out that we need.
  23219. 17:56:10And let's get a visual on this. We'll do
  23220. 17:56:13our PLT plot. And what we're plotting
  23221. 17:56:16here
  23222. 17:56:17is first the x axis, which is range 0
  23223. 17:56:2111. So that will generate a nice little
  23224. 17:56:23plot there. And the wcss for our y axis.
  23225. 17:56:30It's always nice to give our uh plot a
  23226. 17:56:32title.
  23227. 17:56:34And let's see, we'll just give it the
  23228. 17:56:36elbow method for the title. And let's
  23229. 17:56:38get some labels. So let's go ahead and
  23230. 17:56:40do PLT X label.
  23231. 17:56:43And what we'll do, we'll do number of
  23232. 17:56:45clusters for that. And PLT Y label. And
  23233. 17:56:50for that, we can do oops, there we go.
  23234. 17:56:52WCSS since that's what we're doing on
  23235. 17:56:54the plot on there. And finally, we want
  23236. 17:56:56to go ahead and display our graph, which
  23237. 17:56:58is simply plt. Oops.
  23238. 17:57:02Show. There we go. And because we have
  23239. 17:57:04it set to inline, it'll appear inline.
  23240. 17:57:07Hopefully I didn't make a type error on
  23241. 17:57:09there.
  23242. 17:57:13And you can see we get a very nice
  23243. 17:57:14graph. You can see a very nice elbow
  23244. 17:57:16joint there at uh two and again right
  23245. 17:57:19around three and four. And then after
  23246. 17:57:21that there's not very much. Now as a
  23247. 17:57:24data scientist, if I was looking at
  23248. 17:57:26this, I would do either three or four.
  23249. 17:57:29And I'd actually try both of them to see
  23250. 17:57:31what the u output look like. And they've
  23251. 17:57:33already tried this in the back. So,
  23252. 17:57:35we're just going to use three as a setup
  23253. 17:57:36on here. And let's go ahead and see what
  23254. 17:57:38that looks like when we actually use
  23255. 17:57:39this to show the different kinds of
  23256. 17:57:42cars.
  23257. 17:57:45And so, let's go ahead and apply the K
  23258. 17:57:47means to the cars data set. And
  23259. 17:57:50basically, we're going to copy the code
  23260. 17:57:52that we loop through up above where K
  23261. 17:57:54means equals K means number of clusters.
  23262. 17:57:56And we're just going to set the number
  23263. 17:57:57of clusters to three since that's what
  23264. 17:58:00we're going to look for. And you could
  23265. 17:58:01do three and four on this and graph them
  23266. 17:58:04just to see how they come up
  23267. 17:58:05differently. It'd be kind of curious to
  23268. 17:58:06look at that. But for this, we're just
  23269. 17:58:08going to set it to three. Go ahead and
  23270. 17:58:10create our own variable Y k means for
  23271. 17:58:13our answers. And we're going to set that
  23272. 17:58:16equal to Whoops, my double equal there
  23273. 17:58:19to K means. But we're not going to do a
  23274. 17:58:22fit. We're going to do a fit predict is
  23275. 17:58:25the setup you want to use. And when
  23276. 17:58:27you're using untrained models, you'll
  23277. 17:58:29see um a slightly different because
  23278. 17:58:31usually you see fit and then you see
  23279. 17:58:32just the predict. But we want to both
  23280. 17:58:34fit and predict the k means on this. And
  23281. 17:58:38that's fit underscore predict. And then
  23282. 17:58:40our capital x is the data we're working
  23283. 17:58:42with.
  23284. 17:58:43And before we plot this data, we're
  23285. 17:58:45going to do a little pandas trick. We're
  23286. 17:58:47going to take our x value and we're
  23287. 17:58:49going to set x as matrix. So we're
  23288. 17:58:51converting this into a nice rows and
  23289. 17:58:54columns kind of setup. But we want the
  23290. 17:58:56we're going to have columns equals none.
  23291. 17:58:57So it's just going to be a matrix of
  23292. 17:58:59data in here. And let's go ahead and run
  23293. 17:59:02that.
  23294. 17:59:04A little warning. You'll see this
  23295. 17:59:05warnings pop up because things are
  23296. 17:59:06always being updated. So there's like
  23297. 17:59:08minor changes in the versions and future
  23298. 17:59:11versions. Let's set a matrix. Now that
  23299. 17:59:13it's more common to set it values
  23300. 17:59:16instead of doing as matrix, but mass
  23301. 17:59:18matrix works just fine for right now and
  23302. 17:59:20you'll want to update that later on. But
  23303. 17:59:22let's go ahead and dive in and plot this
  23304. 17:59:23and see what that looks like. And before
  23305. 17:59:27we dive into plotting this data, I
  23306. 17:59:29always like to take a look and see what
  23307. 17:59:30I am plotting. So let's take a look at
  23308. 17:59:33why K means. I'm just going to print
  23309. 17:59:35that out down here. And we see we have
  23310. 17:59:38an array of answers. We have 2 1 0 2 1
  23311. 17:59:412. So it's clustering these different
  23312. 17:59:45rows of data based on the three
  23313. 17:59:47different spaces it thinks it's going to
  23314. 17:59:49be.
  23315. 17:59:51And then let's go ahead and print X and
  23316. 17:59:53see what we have for X. And we'll see
  23317. 17:59:55that X is an array. It's a matrix. So we
  23318. 17:59:59have our different values in the array.
  23319. 18:00:01And what we're going to do, it's very
  23320. 18:00:03hard to plot all the different values in
  23321. 18:00:05the array. So we're only going to be
  23322. 18:00:07looking at the first two or positions
  23323. 18:00:10zero and one. And if you were doing a
  23324. 18:00:13full presentation in front of the board
  23325. 18:00:16meeting, you might actually do a little
  23326. 18:00:18different and and dig a little deeper
  23327. 18:00:20into the different aspects because this
  23328. 18:00:22is all the different columns we looked
  23329. 18:00:23at. But we'll only look at columns one
  23330. 18:00:25and two for this to make it easy. So
  23331. 18:00:28let's go ahead and clear this data out
  23332. 18:00:29of here and let's bring up our plot. And
  23333. 18:00:32we're going to do a scatter plot here.
  23334. 18:00:34So pl scatter.
  23335. 18:00:37And
  23336. 18:00:38this looks a little complicated. So
  23337. 18:00:40let's explain what's going on with this.
  23338. 18:00:42We're going to take the x values
  23339. 18:00:46and we're only interested in y of k
  23340. 18:00:48means equals 0, the first cluster. Okay?
  23341. 18:00:52And then we're going to take value zero
  23342. 18:00:54for the x-axis. And then we're going to
  23343. 18:00:56do the same thing here. We're only
  23344. 18:00:58interested in k means equals 0, but
  23345. 18:01:01we're going to take the second column.
  23346. 18:01:02So we're only looking at the first two
  23347. 18:01:04columns in our answer or in the data.
  23348. 18:01:07And then the guys in the back played
  23349. 18:01:09with this a little bit to make it
  23350. 18:01:10pretty.
  23351. 18:01:12And they discovered that it looks good
  23352. 18:01:14with a size equals 100. That's the size
  23353. 18:01:16of the dots. We're going to use red for
  23354. 18:01:19this one. And when they were looking at
  23355. 18:01:22the data and what came out, it was
  23356. 18:01:24definitely the Toyota on this. We're
  23357. 18:01:26just going to go ahead and label it
  23358. 18:01:27Toyota. Again, that's something you
  23359. 18:01:29really have to explore in here as far as
  23360. 18:01:32playing with those numbers and see what
  23361. 18:01:34looks good. We'll go ahead and hit enter
  23362. 18:01:35in there. And I'm just going to paste in
  23363. 18:01:37the next two lines, which is the next
  23364. 18:01:40two cars. And this is our Nissa and
  23365. 18:01:43Honda. And you'll see with our scatter
  23366. 18:01:46plot, we're now looking at where Y_K
  23367. 18:01:48means equals 1. And we want the zero
  23368. 18:01:51column and YK means equals 2. Again,
  23369. 18:01:53we're looking at just the first two
  23370. 18:01:55columns, zero and one. And each of these
  23371. 18:01:57rows then corresponds to Nissan and
  23372. 18:02:00Honda.
  23373. 18:02:02And I'll go ahead and hit enter on
  23374. 18:02:03there. And uh finally, let's take a look
  23375. 18:02:05and put the centrids on there. Again,
  23376. 18:02:08we're going to do a scatter plot.
  23377. 18:02:11And on the centrids, you can just pull
  23378. 18:02:13that from our K means, the uh model we
  23379. 18:02:16created cluster centers. And we're going
  23380. 18:02:19to just do um
  23381. 18:02:22all of them in the first number and all
  23382. 18:02:25of them in the second number, which is
  23383. 18:02:2601 because you always start with zero
  23384. 18:02:28and one.
  23385. 18:02:30And then they were playing with the size
  23386. 18:02:32and everything to make it look good.
  23387. 18:02:34We'll do a size of 300. We're going to
  23388. 18:02:36make the color yellow. And we'll label
  23389. 18:02:38them. It's always good to have some good
  23390. 18:02:39labels. Centroidids.
  23391. 18:02:42And then we do want to do a title. PLT
  23392. 18:02:45title.
  23393. 18:02:47And pop up there. PLT title. So you
  23394. 18:02:50always make want to make your graphs
  23395. 18:02:51look pretty. And we'll call it clusters
  23396. 18:02:52of car make. And one of the features of
  23397. 18:02:56the plot library is you can add a
  23398. 18:03:01legend. It'll automatically bring in it
  23399. 18:03:03since we've already labeled the
  23400. 18:03:05different aspects of the legend with
  23401. 18:03:06Toyota, Nissan, and Honda.
  23402. 18:03:09And finally, we want to go ahead and
  23403. 18:03:10show so we can actually see it. And
  23404. 18:03:13remember, it's in line. Uh so if you're
  23405. 18:03:15using a different editor that's not the
  23406. 18:03:16Jupyter notebook, you'll get a popup of
  23407. 18:03:19this. And you should have a nice set of
  23408. 18:03:21clusters here. So we can look at this
  23409. 18:03:22and we have a clusters of Honda in
  23410. 18:03:25green, Toyota in red, Nissan in purple.
  23411. 18:03:29And you can see where they put the
  23412. 18:03:30centroidids to separate them.
  23413. 18:03:32Now when we're looking at this, we can
  23414. 18:03:34also plot a lot of other different data
  23415. 18:03:37on here as far because we only looked at
  23416. 18:03:39the first two columns. This is just
  23417. 18:03:40column one and two or 01 as as you label
  23418. 18:03:44them in computer scripting. But you can
  23419. 18:03:46see here we have a nice clusters of car
  23420. 18:03:47making. and we were able to pull out the
  23421. 18:03:49data and you can see how just these two
  23422. 18:03:51columns form very distinct clusters of
  23423. 18:03:54data. So if you were exploring new data
  23424. 18:03:57you might take a look and say well what
  23425. 18:03:58makes these different almost going in
  23426. 18:04:01reverse you start looking at the data
  23427. 18:04:03and pulling apart the columns to find
  23428. 18:04:04out why is the first group set up the
  23429. 18:04:07way it is. Maybe you're doing loans and
  23430. 18:04:09you want to go, well, why is this group
  23431. 18:04:11not defaulting on their loans and why is
  23432. 18:04:13the last group defaulting on their
  23433. 18:04:14loans? And why is the middle group 50%
  23434. 18:04:16defaulting on their bank loans? And you
  23435. 18:04:19start finding ways to manipulate the
  23436. 18:04:21data and pull out the answers you want.
  23437. 18:04:26So now that you've seen how to use K
  23438. 18:04:28mean for clustering, let's move on to
  23439. 18:04:31the next topic. Now let's look into
  23440. 18:04:34logistic regression. The logistic
  23441. 18:04:36regression algorithm is the simplest
  23442. 18:04:38classification algorithm used for binary
  23443. 18:04:41or multiclassification problems. And we
  23444. 18:04:43can see we have our little girl from
  23445. 18:04:45Canada who's into horror books is back.
  23446. 18:04:47That's actually really scary when you
  23447. 18:04:49think about that with those big eyes. In
  23448. 18:04:51the previous tutorial, we learned about
  23449. 18:04:52linear regression, dependent and
  23450. 18:04:55independent variables. So to brush up,
  23451. 18:04:58y= mx + c. Very basic algebraic function
  23452. 18:05:03of uh y and x. The dependent variable is
  23453. 18:05:06the target class variable we are going
  23454. 18:05:08to predict. The independent variables X1
  23455. 18:05:12all the way up to XN are the features or
  23456. 18:05:15attributes we're going to use to predict
  23457. 18:05:17the target class. We know what a linear
  23458. 18:05:19regression looks like. But using the
  23459. 18:05:21graph, we cannot divide the outcome into
  23460. 18:05:23categories. It's really hard to
  23461. 18:05:25categorize 1.5, 3.6, 9.8. Uh for
  23462. 18:05:30example, a linear regression graph can
  23463. 18:05:32tell us that with increase in number of
  23464. 18:05:35hours studied, the marks of a student
  23465. 18:05:37will increase, but it will not tell us
  23466. 18:05:39whether the student will pass or not. In
  23467. 18:05:41such cases where we need the output as
  23468. 18:05:44categorical value, we will use logistic
  23469. 18:05:46regression. And for that, we're going to
  23470. 18:05:48use the sigmoid function. So you can see
  23471. 18:05:50here we have our marks 0 to 100, number
  23472. 18:05:53of hours studied. That's going to be
  23473. 18:05:54what they're comparing it to in this
  23474. 18:05:56example. And we usually form a line that
  23475. 18:05:58says y = mx + c. And when we use the
  23476. 18:06:01sigmoid function, we have p = 1 / 1 + e
  23477. 18:06:06the minus y, it generates a sigmoid
  23478. 18:06:09curve. And so you can see right here
  23479. 18:06:11when you take the ln, which is the
  23480. 18:06:13natural logarithm. I always thought it
  23481. 18:06:16should be nl, not ln. That's just the
  23482. 18:06:18inverse of uh e your e to the minus y.
  23483. 18:06:22And so we do this, we get ln of p 1 - p
  23484. 18:06:26= m * x + c. That's the sigmoid curve
  23485. 18:06:29function we're looking for. And we can
  23486. 18:06:31zoom in on the function and you'll see
  23487. 18:06:33that the function as it deres goes to
  23488. 18:06:36one or to zero depending on what your x
  23489. 18:06:39value is. And the probability if it's
  23490. 18:06:41greater than 0.5, the value is
  23491. 18:06:44automatically rounded off to one
  23492. 18:06:46indicating that the student will pass.
  23493. 18:06:47So if they're doing a certain amount of
  23494. 18:06:49studying, they will probably pass. Then
  23495. 18:06:51you have a threshold value at the 0.5.
  23496. 18:06:54It automatically puts that right in the
  23497. 18:06:55middle usually. And your probability if
  23498. 18:06:57it's less than 0.5, the value run it off
  23499. 18:06:59to zero indicating the student will
  23500. 18:07:01fail. So if they're not studying very
  23501. 18:07:02hard, they're probably going to fail.
  23502. 18:07:04This, of course, is ignoring the
  23503. 18:07:06outliers of that one student who's just
  23504. 18:07:07a natural genius and doesn't need any
  23505. 18:07:09studying to memorize everything. That's
  23506. 18:07:11not me, unfortunately. Have to study
  23507. 18:07:14hard to learn new stuff. problem
  23508. 18:07:16statement to classify whether a tumor is
  23509. 18:07:19malignant or B9. And this is actually
  23510. 18:07:22one of my favorite data sets to play
  23511. 18:07:24with because it has so many features and
  23512. 18:07:27when you look at them, you really are
  23513. 18:07:29hard to understand. You can't just look
  23514. 18:07:31at them and know the answer. So it gives
  23515. 18:07:33you a chance to kind of dive into what
  23516. 18:07:34data looks like when you aren't able to
  23517. 18:07:36understand the specific domain of the
  23518. 18:07:38data. But I also want you to remind you
  23519. 18:07:40that in the domain of medicine, if I
  23520. 18:07:42told you that my probability was really
  23521. 18:07:45good at classified things that say 90%
  23522. 18:07:48or 95% and I'm classifying whether
  23523. 18:07:51you're going to have a malignant or a B9
  23524. 18:07:54tumor, I'm guessing that you're going to
  23525. 18:07:56go get it tested anyways. So you got to
  23526. 18:07:57remember the domain we're working with.
  23527. 18:07:59So why would you want to do that if you
  23528. 18:08:01know you're just going to go get a
  23529. 18:08:02biopsy? Because you know it's that
  23530. 18:08:04serious. This is like an all or nothing.
  23531. 18:08:07just referencing the domain. It's
  23532. 18:08:08important. It might help the doctor know
  23533. 18:08:11where to look just by understanding what
  23534. 18:08:14kind of tumor it is. So it might help
  23535. 18:08:17them or aid them on something they
  23536. 18:08:18missed from before. So let's go ahead
  23537. 18:08:20and dive into the code and I'll come
  23538. 18:08:22back to the domain part of it in just a
  23539. 18:08:24minute. So use case and we're going to
  23540. 18:08:26do our normal imports here where we're
  23541. 18:08:28importing numpy, pandas, seabour, the
  23542. 18:08:30mattplot library and we're going to do
  23543. 18:08:33mattplot library in line since I'm going
  23544. 18:08:34to switch over to Anaconda. So, let's go
  23545. 18:08:36ahead and flip over there and get this
  23546. 18:08:38started. So, I've opened up a new window
  23547. 18:08:40in my Anaconda Jupyter Notebook. And by
  23548. 18:08:44the way, Jupyter Notebook, uh, you don't
  23549. 18:08:46have to use Anaconda for the Jupyter
  23550. 18:08:48Notebook. I just love the interface and
  23551. 18:08:49all the tools that Anaconda brings. So,
  23552. 18:08:52we got our import numpy aspy
  23553. 18:08:55number array. We have our pandas pd.
  23554. 18:08:58We're going to bring in Seabor to help
  23555. 18:09:00us with our graphs as SNS. So many
  23556. 18:09:03really nice tools in both Seabour and
  23557. 18:09:05Mattplot library. And we'll do our
  23558. 18:09:06mapplot library.pipplot as plt. And then
  23559. 18:09:10of course we want to let it know to do
  23560. 18:09:11it in line. And let's go and just run
  23561. 18:09:13that. So it's all set up. And we're just
  23562. 18:09:16going to call our data data. Not
  23563. 18:09:18creative today. Uh equals pd. And this
  23564. 18:09:20happens to be in a CSV file. So we'll
  23565. 18:09:25use a pdread_csv.
  23566. 18:09:28And I happen to name the file. renamed
  23567. 18:09:31it data forp2.csv.
  23568. 18:09:33You can of course um write in the
  23569. 18:09:35comments below the YouTube and request
  23570. 18:09:37for the data set itself or go to the
  23571. 18:09:38SimplyLearn website and we'll be happy
  23572. 18:09:40to supply that for you. And let's just
  23573. 18:09:43um open up the data before we go any
  23574. 18:09:45further and let's just see what it looks
  23575. 18:09:46like in a spreadsheet.
  23576. 18:09:48So when I pop it open in a local
  23577. 18:09:50spreadsheet, this is just a CSV file,
  23578. 18:09:53comma separated variables. We have an
  23579. 18:09:55ID. So I guess the U categorizes for
  23580. 18:09:58reference or what ID which test was
  23581. 18:10:00done. The diagnosis M for malignant, B
  23582. 18:10:04for B9. So there's two different options
  23583. 18:10:06on there. And that's what we're going to
  23584. 18:10:07try to predict is the M and B and test
  23585. 18:10:09it. And then we have like the radius
  23586. 18:10:12mean or average the texture average,
  23587. 18:10:14perimeter mean, area mean, smoothness. I
  23588. 18:10:17don't know about you, but unless you're
  23589. 18:10:20a doctor in the field, most of the
  23590. 18:10:22stuff, I mean, you can guess what
  23591. 18:10:23concave means just by the term concave,
  23592. 18:10:26but I really wouldn't know what that
  23593. 18:10:28means in the measurements they're
  23594. 18:10:29taking. So, they have all kinds of stuff
  23595. 18:10:30like how smooth it is, uh, the symmetry,
  23596. 18:10:33and these are all float values. You just
  23597. 18:10:35page through them real quick, and you'll
  23598. 18:10:37see there's, I believe, 36, if I
  23599. 18:10:39remember correctly, in this one.
  23600. 18:10:42So there's a lot of different values
  23601. 18:10:43they take and all these measurements
  23602. 18:10:45they take when they go in there and they
  23603. 18:10:46take a look at the different growth, the
  23604. 18:10:48tumorous growth. So back in our data and
  23605. 18:10:52I put this in the same folder as a code.
  23606. 18:10:54So I saved this code in that folder.
  23607. 18:10:57Obviously if you have it in a different
  23608. 18:10:58location, you want to put the full path
  23609. 18:11:00in there and we'll just do uh pandas
  23610. 18:11:05first five lines of data with the data
  23611. 18:11:07head. And we run that. We can see that
  23612. 18:11:10we have pretty much what we just looked
  23613. 18:11:12at. We have an ID. We have a diagnosis.
  23614. 18:11:15If we go all the way across, you'll see
  23615. 18:11:17all the different columns coming across
  23616. 18:11:19displayed nicely for our data.
  23617. 18:11:23And while we're exploring the data, our
  23618. 18:11:26uh Seabor, which we referenced as SNS,
  23619. 18:11:29makes it very easy to go in here and do
  23620. 18:11:31a joint plot. You'll notice the very
  23621. 18:11:34similar to because it is sitting on top
  23622. 18:11:36of the U plot library. So, the joint
  23623. 18:11:39plot does a lot of work for us. And
  23624. 18:11:41we're just going to look at the first
  23625. 18:11:42two columns that we're interested in,
  23626. 18:11:44the radius mean and the texture mean.
  23627. 18:11:46We'll just look at those two columns and
  23628. 18:11:49data equals data. So that tells it which
  23629. 18:11:52two columns we're plotting and that
  23630. 18:11:53we're going to use the data that we
  23631. 18:11:55pulled in. Let's just run that. And it
  23632. 18:11:58generates a really nice graph on here.
  23633. 18:12:00And there's all kinds of cool things on
  23634. 18:12:02this graph to look at. I mean, we have
  23635. 18:12:03the texture mean and the radius mean
  23636. 18:12:05obviously the axes. You can also see
  23637. 18:12:10and uh one of the cool things on here is
  23638. 18:12:12you can also see the histogram. They
  23639. 18:12:13show that for the radius mean where is
  23640. 18:12:15the most common radius mean come up and
  23641. 18:12:17where the most common texture is. So
  23642. 18:12:20we're looking at the tech the on each
  23643. 18:12:22growth it's average texture and on each
  23644. 18:12:25radius it's average uh radius on there
  23645. 18:12:28gets a little confusing because we're
  23646. 18:12:29talking about the individual objects
  23647. 18:12:31average. And then we can also look over
  23648. 18:12:33here and see the the histogram showing
  23649. 18:12:36us the median or how common each
  23650. 18:12:39measurement is. And that's only two
  23651. 18:12:42columns. So let's dig a little deeper
  23652. 18:12:44into Seabor. They also have a heat map.
  23653. 18:12:47And if you're not familiar with heat
  23654. 18:12:49maps, a heat map just means it's in
  23655. 18:12:51color. That's all that means. Heat map.
  23656. 18:12:53I guess the original ones were plotting
  23657. 18:12:55heat density on something. And so ever
  23658. 18:12:57since then it's just called a heat map.
  23659. 18:12:58And we're going to take our data and get
  23660. 18:13:00our corresponding numbers to put that
  23661. 18:13:02into the heat map. And that's simply
  23662. 18:13:04data.coR
  23663. 18:13:06for that. That's a pandas expression.
  23664. 18:13:09Let's remember we're working in a pandas
  23665. 18:13:11data frame. So that's one of the cool
  23666. 18:13:12tools in pandas for our data. And let's
  23667. 18:13:15just pull that information into a heat
  23668. 18:13:17map and see what that looks like. And
  23669. 18:13:19you'll see that we're now looking at all
  23670. 18:13:21the different features. We have our ID.
  23671. 18:13:24We have our texture. We have our area,
  23672. 18:13:26our compactness, concave points. And if
  23673. 18:13:29you look down the middle of this chart
  23674. 18:13:31diagonal going from the upper left to
  23675. 18:13:32bottom right, it's all white. That's
  23676. 18:13:35because when you compare texture to
  23677. 18:13:38texture, they're identical. So they're
  23678. 18:13:40100% or in this case perfect one in
  23679. 18:13:43their correspondence.
  23680. 18:13:45And you'll see that when you look at say
  23681. 18:13:48area or right below it, it has almost a
  23682. 18:13:51black on there. when you compare it to
  23683. 18:13:53texture. So these have almost no
  23684. 18:13:54corresponding data. They don't really
  23685. 18:13:56form a linear graph or something that
  23686. 18:13:58you can look at and say how connected
  23687. 18:13:59they are. They're very scattered data.
  23688. 18:14:02This is really just a really nice graph
  23689. 18:14:04to get a quick look at your data.
  23690. 18:14:06Doesn't so much change what you do, but
  23691. 18:14:08it changes verifying. So when you get an
  23692. 18:14:11answer or something like that or you
  23693. 18:14:12start looking at some of these
  23694. 18:14:13individual pieces, you might go, "Hey,
  23695. 18:14:15that doesn't match. according to showing
  23696. 18:14:18our heat map, this should not correlate
  23697. 18:14:21with each other. And if it is, you're
  23698. 18:14:22going to have to start asking, well,
  23699. 18:14:23why? What's going on? What else is
  23700. 18:14:25coming in there? But it does show some
  23701. 18:14:27really cool information on here. I mean,
  23702. 18:14:30we can see from the ID, there's no real
  23703. 18:14:33one feature that just says if you go
  23704. 18:14:36across the top line that lights up.
  23705. 18:14:39There's no one feature that says, hey,
  23706. 18:14:40if the area is a certain size, then it's
  23707. 18:14:42going to be B9 or malignant. It says
  23708. 18:14:44there's some that sort of add up and
  23709. 18:14:46that's a big hint in the data that we're
  23710. 18:14:49trying to ID this whether it's malignant
  23711. 18:14:51or B9. That's a big hint to us as data
  23712. 18:14:54scientists to go okay we can't solve
  23713. 18:14:56this with any one feature. It's going to
  23714. 18:14:59be something that includes all the
  23715. 18:15:00features or many of the different
  23716. 18:15:02features to come up with a solution for
  23717. 18:15:03it. And while we're exploring the data
  23718. 18:15:07let's explore one more area and let's
  23719. 18:15:09look at data isnull. We want to check
  23720. 18:15:12for null values in our data. If you
  23721. 18:15:15remember from earlier in this tutorial,
  23722. 18:15:18we did it a little differently where we
  23723. 18:15:19added stuff up and sum them up. You can
  23724. 18:15:22actually with pandas do it really
  23725. 18:15:23quickly. Data.isnull and summit. And
  23726. 18:15:25it's going to go across all the columns.
  23727. 18:15:27So when I run this,
  23728. 18:15:30you're going to see all the columns come
  23729. 18:15:32up with no null data.
  23730. 18:15:36So we've just just to rehash these last
  23731. 18:15:39few steps. We've done a lot of
  23732. 18:15:42exploration. We have looked at the first
  23733. 18:15:44two columns and seen how they plot with
  23734. 18:15:47the seabour with a joint plot which
  23735. 18:15:49shows both the histogram and the data
  23736. 18:15:52plotted on the XY coordinates. And
  23737. 18:15:54obviously you can do that more in detail
  23738. 18:15:58with different columns and see how they
  23739. 18:15:59plot together. And then we took and did
  23740. 18:16:02the Seabor heat map the SNS
  23741. 18:16:05heat mapap of the data. And you can see
  23742. 18:16:07right here where it did a nice job
  23743. 18:16:09showing us some bright spots where stuff
  23744. 18:16:11correlates with each other and forms a
  23745. 18:16:13very nice combination or points of
  23746. 18:16:16scattering points. And you can also see
  23747. 18:16:18areas that don't.
  23748. 18:16:20And then finally, we went ahead and
  23749. 18:16:22checked the data. Is the data null
  23750. 18:16:24value? Do we have any missing data in
  23751. 18:16:26there? Very important step because it'll
  23752. 18:16:28crash later on. If you forget to do this
  23753. 18:16:31step, it will remind you when you get
  23754. 18:16:33that nice error code that says null
  23755. 18:16:35values. Okay. So, not a big deal if you
  23756. 18:16:38miss it, but it it's no fun having to go
  23757. 18:16:40back when you're when you're in a huge
  23758. 18:16:42process and you've missed this step and
  23759. 18:16:44now you're 10 steps later and you got to
  23760. 18:16:45go remember where you were pulling the
  23761. 18:16:47data in.
  23762. 18:16:49So, we need to go ahead and pull out our
  23763. 18:16:51X and our Y. So, we just put that down
  23764. 18:16:54here and we'll set the X equal to. And
  23765. 18:16:57there's a lot of different options here.
  23766. 18:16:59Certainly we could do X equals all the
  23767. 18:17:01columns except for the first two because
  23768. 18:17:04if you remember the first two is the ID
  23769. 18:17:05and the diagnosis. So that certainly
  23770. 18:17:08would be an option. But what we're going
  23771. 18:17:10to do is we're actually going to focus
  23772. 18:17:12on the worst. The worst radius, the
  23773. 18:17:14worst texture, parameter area,
  23774. 18:17:17smoothness, compactness, and so on. One
  23775. 18:17:20of the reasons to start dividing your
  23776. 18:17:22data up when you're looking at this
  23777. 18:17:24information is sometimes the data will
  23778. 18:17:28be the same data coming in. So if I have
  23779. 18:17:30two measurements coming into my model,
  23780. 18:17:33it might overweigh them. It might
  23781. 18:17:35overpower the other measurements because
  23782. 18:17:37it's measuring it's basically taking
  23783. 18:17:38that information in twice. That's a
  23784. 18:17:40little bit past the scope of this
  23785. 18:17:41tutorial. I want you to take away from
  23786. 18:17:43this though is that we are dividing the
  23787. 18:17:45data up into pieces and our team in the
  23788. 18:17:47back went ahead and said hey let's just
  23789. 18:17:49look at the worst. So I'm going to
  23790. 18:17:51create a an array and you'll see this
  23791. 18:17:54array radius worst texture worst
  23792. 18:17:56perimeter worst. We've just taken the
  23793. 18:17:58worst of the worst and I'm just going to
  23794. 18:18:00put that in my X. So this X is still a
  23795. 18:18:02pandas data frame but it's just those
  23796. 18:18:05columns. And our Y, if you remember
  23797. 18:18:08correctly, is going to be Oops, hold on
  23798. 18:18:10one second. It's not X. is data. There
  23799. 18:18:13we go. So, x equals data and then it's a
  23800. 18:18:16list of the different columns, the worst
  23801. 18:18:17of the worst. And if we're going to take
  23802. 18:18:19that, then we have to have our answer
  23803. 18:18:21for our y for the stuff we know. And if
  23804. 18:18:24you remember correctly, we're just going
  23805. 18:18:25to be looking at
  23806. 18:18:28the diagnosis. That's all we care about
  23807. 18:18:30is what is it diagnosed? Is it B9 or
  23808. 18:18:32malignant? And since it's a single
  23809. 18:18:35column, we can just do diagnosis. Oh, I
  23810. 18:18:37forgot to put the brackets. There we go.
  23811. 18:18:39Okay. So, it's just diagnosis on there.
  23812. 18:18:42And we can also real quickly do like an
  23813. 18:18:44X do. If you want to see what that looks
  23814. 18:18:46like and Y head
  23815. 18:18:50and run this and you'll see um it only
  23816. 18:18:53does the last one. I forgot about that.
  23817. 18:18:55If you don't do print, you can see that
  23818. 18:18:57the the Y.D is just mm because the first
  23819. 18:19:00ones are all malignant. And if I run
  23820. 18:19:02this, the X do head is just the first
  23821. 18:19:04five values of radius worst, texture
  23822. 18:19:07worst, parameter worst, area worst, and
  23823. 18:19:09so on. I'll go ahead and take that out.
  23824. 18:19:13So, moving down to the next step, we've
  23825. 18:19:17built our two data sets, our answer and
  23826. 18:19:20then the features we want to look at.
  23827. 18:19:23In data science, it's very important to
  23828. 18:19:26test your model. So we do that by
  23829. 18:19:29splitting the data
  23830. 18:19:32and from sklearn model selection we're
  23831. 18:19:34going to import train test split. So
  23832. 18:19:37we're going to split it into two groups.
  23833. 18:19:39There are so many ways to do this. I
  23834. 18:19:41noticed in one of the more modern ways
  23835. 18:19:43they actually split it into three groups
  23836. 18:19:45and then you model each group and test
  23837. 18:19:48it against the other groups. So you have
  23838. 18:19:50all kinds and there's reasons for that
  23839. 18:19:51which is past the scope of this and for
  23840. 18:19:53this particular example isn't necessary
  23841. 18:19:56for this. We're just going to split it
  23842. 18:19:57into two groups. one to train our data
  23843. 18:19:59and one to test our data. And the
  23844. 18:20:02sklearn uh.mmodel selection we have
  23845. 18:20:05train tests split. You could write your
  23846. 18:20:07own quick code to do this where you just
  23847. 18:20:09randomly divide the data up into two
  23848. 18:20:11groups but they do it for us nicely
  23849. 18:20:14and we actually can almost we can
  23850. 18:20:16actually do it in one statement with
  23851. 18:20:17this where we're going to generate four
  23852. 18:20:19variables capital X train capital X
  23853. 18:20:23test. So we have our training data we're
  23854. 18:20:25going to use to fit the model and then
  23855. 18:20:27we need something to test it and then we
  23856. 18:20:29have our y train. So we're going to
  23857. 18:20:30train the answer and then we have our
  23858. 18:20:32test. So this is the stuff we want to
  23859. 18:20:33see how good it did on our model. And
  23860. 18:20:36we'll go ahead and take our train test
  23861. 18:20:38split that we just imported.
  23862. 18:20:41And we're going to do X and our Y, our
  23863. 18:20:43two different data that's going in for
  23864. 18:20:45our split. And then the guys in the back
  23865. 18:20:48came up and wanted us to go ahead and
  23866. 18:20:49use a test size equals.3.
  23867. 18:20:52That's test size. Random state. It's
  23868. 18:20:55always nice to kind of switch a random
  23869. 18:20:57state around, but not that important.
  23870. 18:20:59What this means is that the test size is
  23871. 18:21:01we're going to take 30% of the data and
  23872. 18:21:04we're going to put that into our test
  23873. 18:21:06variables, our Y test and our X test.
  23874. 18:21:09And we're going to do 70% into the X
  23875. 18:21:11train and the Y train. So, we're going
  23876. 18:21:13to use 70% of the data to train our
  23877. 18:21:15model and 30% to test it. Let's go ahead
  23878. 18:21:18and run that and load those up. So now
  23879. 18:21:21we have all our stuff split up and all
  23880. 18:21:23our data ready to go. And now we get to
  23881. 18:21:25the actual logistics part. We're
  23882. 18:21:27actually going to do our create our
  23883. 18:21:28model. So let's go ahead and bring that
  23884. 18:21:30in from sklearn. We're going to bring in
  23885. 18:21:33our linear model and we're going to
  23886. 18:21:34import logistic regression. That's the
  23887. 18:21:37actual model we're using. And let's
  23888. 18:21:39we'll call it log model.
  23889. 18:21:42Oops, there we go. Model. And let's just
  23890. 18:21:44set this equal to our logistic
  23891. 18:21:46regression that we just imported. So now
  23892. 18:21:49we have a variable log model set to that
  23893. 18:21:51class for us to use. And with most the
  23894. 18:21:55uh models in the sklearn, we just need
  23895. 18:21:58to go ahead and fix it. Fit do a fit on
  23896. 18:22:01there. And we use our x train that we
  23897. 18:22:04separated out with our y train. And
  23898. 18:22:07let's go ahead and run this. So once
  23899. 18:22:08we've run this, we'll have a model that
  23900. 18:22:10fits this data that 70% of our training
  23901. 18:22:13data.
  23902. 18:22:15Uh, and of course it prints this out
  23903. 18:22:17that tells us all the different
  23904. 18:22:18variables that you can set on there.
  23905. 18:22:20There's a lot of different choices you
  23906. 18:22:21can make, but for Word do, we're just
  23907. 18:22:23going to let all the defaults set. We
  23908. 18:22:25don't really need to mess with those on
  23909. 18:22:26this particular example. And there's
  23910. 18:22:28nothing in here that really stands out
  23911. 18:22:29as super important until you start
  23912. 18:22:32fine-tuning it. But for what we're
  23913. 18:22:34doing, the basics will work just fine.
  23914. 18:22:36And then let's we need to go ahead and
  23915. 18:22:38test out our model. Is it working? So
  23916. 18:22:41let's create a variable Y predict. And
  23917. 18:22:43this is going to be equal to our log
  23918. 18:22:46model. And we want to do a predict.
  23919. 18:22:49Again, very standard format for the
  23920. 18:22:52sklearn library is taking your model and
  23921. 18:22:54doing a predict on it. And we're going
  23922. 18:22:56to test y predict against the y test. So
  23923. 18:22:59we want to know what the model thinks
  23924. 18:23:01it's going to be. That's what our y
  23925. 18:23:02predict is. And with that, we want the
  23926. 18:23:05capital xx test. So we have our train
  23927. 18:23:08set and our test set. And now we're
  23928. 18:23:10going to do our y predict. And let's go
  23929. 18:23:12ahead and run that.
  23930. 18:23:15And if we uh print
  23931. 18:23:18y predict, let me go ahead and run that.
  23932. 18:23:22You'll see it comes up and it predents a
  23933. 18:23:25prints a nice array of uh B and M for B9
  23934. 18:23:28and malignant
  23935. 18:23:30for all the different test data we put
  23936. 18:23:32in there. So, it does pretty good. We're
  23937. 18:23:34not sure exactly how good it does, but
  23938. 18:23:36we can see that it actually works and is
  23939. 18:23:38functional. Was very easy to create.
  23940. 18:23:40You'll always discover with our data
  23941. 18:23:42science that as you explore this, you
  23942. 18:23:45spend a significant amount of time
  23943. 18:23:47prepping your data and making sure your
  23944. 18:23:50data coming in is good. Uh there's a
  23945. 18:23:52saying, good data in, good answers out.
  23946. 18:23:56Bad data in, bad answers out. That's
  23947. 18:23:59only half the thing. That's only half of
  23948. 18:24:02it. Selecting your models becomes the
  23949. 18:24:04next part as far as how good your models
  23950. 18:24:06are. and then of course fine-tuning it
  23951. 18:24:08depending on what model you're using. So
  23952. 18:24:11we come in here, we want to know how
  23953. 18:24:12good this came out. So we have our Y
  23954. 18:24:14predict here, log model.predict X test.
  23955. 18:24:20So for deciding how good our model is,
  23956. 18:24:23we're going to go from the
  23957. 18:24:24sklearn.metrics,
  23958. 18:24:26we're going to import classification
  23959. 18:24:28report. And that just reports how good
  23960. 18:24:30our model is doing. And then we're going
  23961. 18:24:31to feed it the model data. And let's
  23962. 18:24:33just print this out. and we'll take our
  23963. 18:24:36uh classification report
  23964. 18:24:39and we're going to put into there
  23965. 18:24:43our test our actual data. So this is
  23966. 18:24:46what we actually know is true and our
  23967. 18:24:49prediction what our model predicted for
  23968. 18:24:51that data on the test side. And let's
  23969. 18:24:54run that and see what that does.
  23970. 18:24:57So we pull that up. You'll see that we
  23971. 18:24:59have um a precision for B9 and malignant
  23972. 18:25:03B and M. And we have a precision of 93
  23973. 18:25:06and 91, a total of 92. So it's kind of
  23974. 18:25:10the average between these two of 92.
  23975. 18:25:12There's all kinds of different
  23976. 18:25:13information on here. Your F1 score,
  23977. 18:25:16your recall, your support coming through
  23978. 18:25:19on this. And for this, I'll go ahead and
  23979. 18:25:22just flip back to our slides that they
  23980. 18:25:23put together for describing it. And so
  23981. 18:25:26here we're going to look at the
  23982. 18:25:26precision using the classification
  23983. 18:25:28report. And you see this is the same
  23984. 18:25:30print out I had up above. Some of the
  23985. 18:25:32numbers might be different because it
  23986. 18:25:34does randomly pick out which data we're
  23987. 18:25:36using. So this model is able to predict
  23988. 18:25:39the type of tumor with 91% accuracy. So
  23989. 18:25:43we look back here that's you will see
  23990. 18:25:45where we have uh B9 and malignant. It
  23991. 18:25:47actually has 92 coming up here. We're
  23992. 18:25:49looking about a 92 91% precision. And
  23993. 18:25:52remember I reminded you about domain.
  23994. 18:25:54So, when we're talking about the domain
  23995. 18:25:55of a medical domain with a very
  23996. 18:25:58catastrophic outcome, you know, at 91 or
  23997. 18:26:0092% precision, you're still going to go
  23998. 18:26:03in there and have somebody do a biopsy
  23999. 18:26:06on it. Very different than if you're
  24000. 18:26:08investing money and there's a 92% chance
  24001. 18:26:10you're going to earn 10% and 8% chance
  24002. 18:26:14you're going to lose 8%, you're probably
  24003. 18:26:16going to bet the money because at that
  24004. 18:26:17odds, it's pretty good that you'll make
  24005. 18:26:19some money. And in the long run, you do
  24006. 18:26:20that enough, you definitely will make
  24007. 18:26:22money. And also with this domain, I've
  24008. 18:26:24actually seen them use this to identify
  24009. 18:26:26different forms of cancer. That's one of
  24010. 18:26:29the things that they're starting to use
  24011. 18:26:30these models for because then it helps a
  24012. 18:26:32doctor know what to investigate. So that
  24013. 18:26:34wraps up this section. We're finally
  24014. 18:26:37we're going to go in there and let's
  24015. 18:26:38discuss the answers to the quiz asked in
  24016. 18:26:40machine learning tutorial part one. Can
  24017. 18:26:43you tell what's happening in the
  24018. 18:26:44following cases? Grouping documents into
  24019. 18:26:47different categories based on the topic
  24020. 18:26:50and content of each document. This is an
  24021. 18:26:52example of clustering where K means
  24022. 18:26:54clustering can be used to group the
  24023. 18:26:55documents by topics using bag of words
  24024. 18:26:58approach. So if you gotten in there that
  24025. 18:27:00you're looking for clustering and
  24026. 18:27:02hopefully you had at least one or two
  24027. 18:27:04examples like K means that are used for
  24028. 18:27:06clustering different things then give
  24029. 18:27:08yourself a two thumbs up. B identifying
  24030. 18:27:11handwritten digits in images correctly.
  24031. 18:27:14This is an example of classification.
  24032. 18:27:17The traditional approach to solving this
  24033. 18:27:18would be to extract digit dependent
  24034. 18:27:20features like curvature of different
  24035. 18:27:22digits etc. and then use a classifier
  24036. 18:27:24like SVM to distinguish between images.
  24037. 18:27:27Again, if you got the fact that it's a
  24038. 18:27:29classification example, give yourself a
  24039. 18:27:31thumb up. And if you're able to go, hey,
  24040. 18:27:33let's use SVM or another model for this,
  24041. 18:27:36give yourself those two thumbs up on it.
  24042. 18:27:38C. Behavior of a website indicating that
  24043. 18:27:41the site is not working as designed.
  24044. 18:27:44This is an example of anomaly detection.
  24045. 18:27:47In this case, the algorithm learns what
  24046. 18:27:49is normal and what is not normal,
  24047. 18:27:51usually by observing the logs of the
  24048. 18:27:53website. Give yourself a thumbs up if
  24049. 18:27:55you got that one. And just for a bonus,
  24050. 18:27:57can you think of another example of
  24051. 18:27:59anomaly detection? One of the ones I use
  24052. 18:28:01it for in my own business is detecting
  24053. 18:28:03anomalies in stock markets. Stock
  24054. 18:28:06markets are very fickled and they behave
  24055. 18:28:08very erratic. So finding those erratic
  24056. 18:28:10areas and then finding ways to track
  24057. 18:28:12down why they're erratic. Was something
  24058. 18:28:14released in social media? Was something
  24059. 18:28:16released you can see where knowing where
  24060. 18:28:18that anomaly is can help you to figure
  24061. 18:28:21out what the answer is to it in another
  24062. 18:28:23area. D predicting salary of an
  24063. 18:28:25individual based on his or her years of
  24064. 18:28:28experience. This is an example of
  24065. 18:28:30regression. This problem can be
  24066. 18:28:32mathematically defined as a function
  24067. 18:28:33between independent years of experience
  24068. 18:28:35and dependent variables salary of an
  24069. 18:28:38individual. And if you guess that this
  24070. 18:28:40was a regression model, give yourself a
  24071. 18:28:42thumbs up. And if you were able to
  24072. 18:28:43remember that it was between independent
  24073. 18:28:46and dependent variables and that terms,
  24074. 18:28:49give yourself two thumbs up. Summary. So
  24075. 18:28:52to wrap it up, we went over what is K
  24076. 18:28:55means and we went through also the chart
  24077. 18:28:58of choosing your elbow method and
  24078. 18:29:00assigning a random centrid to the
  24079. 18:29:02clusters, computing the distance and
  24080. 18:29:04then going in there and figuring out
  24081. 18:29:06what the minimum centroidids is and
  24082. 18:29:08computing the distance and going through
  24083. 18:29:09that loop until it gets the perfect
  24084. 18:29:11centrid. And we looked into the elbow
  24085. 18:29:13method to choose K based on running our
  24086. 18:29:16clusters across a number of variables
  24087. 18:29:17and finding the best location for that.
  24088. 18:29:19We did a nice example of clustering cars
  24089. 18:29:22with K means even though we only looked
  24090. 18:29:23at the first two columns to make it
  24091. 18:29:25simple and easy to graph. You can easily
  24092. 18:29:27extrapolate that and look at all the
  24093. 18:29:29different columns and see how they all
  24094. 18:29:31fit together. And we looked at what is
  24095. 18:29:33logistic regression. We discussed the
  24096. 18:29:35sigmoid function. What is logistic
  24097. 18:29:38regression? And then we went into an
  24098. 18:29:40example of classifying tumors with
  24099. 18:29:42logistics. I hope you enjoyed part two
  24100. 18:29:45of machine learning. So in today's
  24101. 18:29:47session we will discuss what RNN model
  24102. 18:29:49is. Moving ahead we will see why should
  24103. 18:29:52we use RNN. After that we will see how
  24104. 18:29:55does RNN work recurrent neural network.
  24105. 18:29:58After covering these topics we will move
  24106. 18:30:00forward and see types of RNN recurrent
  24107. 18:30:03neural network and applications of RNN.
  24108. 18:30:06At the end we will do a hands-off lab
  24109. 18:30:09demo of sentiment analysis using RNN. So
  24110. 18:30:11before starting let us have a simple
  24111. 18:30:13question to brush our knowledge. So
  24112. 18:30:15question is what are the application of
  24113. 18:30:17RNN? Okay, NLP,
  24114. 18:30:20time series, image captioning and all of
  24115. 18:30:24the above. Please answer in the comment
  24116. 18:30:25section below and we will update the
  24117. 18:30:27correct answer in the pin comments or
  24118. 18:30:29you can pause this video, give it a
  24119. 18:30:31thought and answer in the comment
  24120. 18:30:33section. Before we move on to the
  24121. 18:30:34programming part, let's discuss what RNN
  24122. 18:30:37is and proceed further for the same. So
  24123. 18:30:39what is RNN? Recurrent neural network.
  24124. 18:30:41So RNN work on the principle of saving
  24125. 18:30:44output on a particular layer and feeding
  24126. 18:30:46this back to the input in order to
  24127. 18:30:47predict the output of the layer. This is
  24128. 18:30:50how can convert a feed neural network
  24129. 18:30:52into a recurrent neural network RN. The
  24130. 18:30:54node in different layers of neural
  24131. 18:30:56network are compressed to form a single
  24132. 18:30:58layer of recurrent neural network. A B
  24133. 18:31:00and C are the parameters of neural
  24134. 18:31:03network. Now that you understand what
  24135. 18:31:05RNN is, let's look at the way why RNN.
  24136. 18:31:08Okay. So why RNN? RNN were created
  24137. 18:31:11because there are few issues in the feed
  24138. 18:31:13forward neural network cannot handle the
  24139. 18:31:14sequential data considers only the
  24140. 18:31:16current input cannot memorize previous
  24141. 18:31:19input. Okay. So the solution of these
  24142. 18:31:21issues is RNN and RNN can handle
  24143. 18:31:24sequential data accepting the current
  24144. 18:31:26input data and previously received input
  24145. 18:31:28data. So RNN can memorize previous input
  24146. 18:31:31due to their internal memory. So moving
  24147. 18:31:33forward let's see how does RNN networks
  24148. 18:31:37work. Okay. So the input layer X takes
  24149. 18:31:40an input to the neural network and
  24150. 18:31:42process it and the passes it into the
  24151. 18:31:44middle layer. The middle layer edge can
  24152. 18:31:46consist of multiple hidden layers each
  24153. 18:31:49with its own activation function and
  24154. 18:31:51weight and biases. If you have a neural
  24155. 18:31:53network where the various parameters of
  24156. 18:31:55different hidden layers are not affected
  24157. 18:31:57by the previous layer that is the neural
  24158. 18:32:00network does not have the memory then
  24159. 18:32:02you can use RNN. So the RNN will
  24160. 18:32:06standardize the different activation
  24161. 18:32:07function and weights and biases so that
  24162. 18:32:09each hidden layer has the same
  24163. 18:32:11parameter. Then instead of creating
  24164. 18:32:13multiple hidden layers, it will create
  24165. 18:32:15one end loop over it as many time it has
  24166. 18:32:18required. So moving forward let's see
  24167. 18:32:20types of RNN. So there are four types of
  24168. 18:32:23RNN
  24169. 18:32:25one to one,
  24170. 18:32:27one to many, many to many and many to
  24171. 18:32:30one.
  24172. 18:32:32So let's see one to one RNN. So this
  24173. 18:32:35type of neural network is known as the
  24174. 18:32:37vanilla neural network. It is used for
  24175. 18:32:39general machine learning problem which
  24176. 18:32:41has a single input and a single output.
  24177. 18:32:43Now see
  24178. 18:32:45one to many RNN. This type of neural
  24179. 18:32:48network has a single input and multiple
  24180. 18:32:50outputs. An example of this is a image
  24181. 18:32:52captioning. Now let's see many to one
  24182. 18:32:55RNN. This RNN take a sequence of input
  24183. 18:32:58and generates a single output. Sentiment
  24184. 18:33:00analysis is a good example of this kind
  24185. 18:33:02of neural network where a given sentence
  24186. 18:33:04can be classified as expressing positive
  24187. 18:33:06or negative sentiment. And the last one
  24188. 18:33:08is many to many RNN. This RNN takes a
  24189. 18:33:12sequence of inputs and generates a
  24190. 18:33:13sequence of output. Machine translation
  24191. 18:33:15is the one of the example. So moving
  24192. 18:33:17forward, let's see application of
  24193. 18:33:19recurrent neural network. First one is
  24194. 18:33:22image captioning. RNNs are used to
  24195. 18:33:25caption an image by analyzing the
  24196. 18:33:28activities present. The second one is
  24197. 18:33:30time series prediction. Any time series
  24198. 18:33:32problem like predicting the prices of
  24199. 18:33:34stocks in a particular month can be
  24200. 18:33:36solved using RNN. And the third one is
  24201. 18:33:39natural language processing. Text mining
  24202. 18:33:41and sentiment analysis can be carried
  24203. 18:33:43out using RNN or NLP. Natural language
  24204. 18:33:47processing. The fourth one is machine
  24205. 18:33:49translation. Given an input in one
  24206. 18:33:52language, RNNs can be used to translate
  24207. 18:33:54the input into different language as
  24208. 18:33:56output. So now let's move to the
  24209. 18:33:58programming part. First we will import
  24210. 18:34:00some libraries major libraries for the
  24211. 18:34:02first we will import for the data frame.
  24212. 18:34:05So I will write import
  24213. 18:34:12pd.
  24214. 18:34:13The second one is import numpy
  24215. 18:34:17as np.
  24216. 18:34:20So pandas is a software library written
  24217. 18:34:22for the python programming language for
  24218. 18:34:24data manipulation and analysis. In
  24219. 18:34:26particular, it offers a data structure
  24220. 18:34:27and operations for manipulating
  24221. 18:34:29numerical tables and the time series.
  24222. 18:34:32And this numpy numpy is a library for
  24223. 18:34:34the Python programming language adding
  24224. 18:34:36support to four large multi-dimensional
  24225. 18:34:39array and matrices along with a large
  24226. 18:34:41collection of highle mathematical
  24227. 18:34:44function to operate on these arrays.
  24228. 18:34:46Okay. So for plotting we will import
  24229. 18:34:49some libraries like seabon
  24230. 18:34:54as
  24231. 18:34:56SNS. This is nothing just a short form
  24232. 18:34:59of we don't have to write again and
  24233. 18:35:01again CON c we can write SNS. So then
  24234. 18:35:04another one is from
  24235. 18:35:08wordcloud
  24236. 18:35:12port
  24237. 18:35:24mattplot lib
  24238. 18:35:27dot
  24239. 18:35:28pip plot
  24240. 18:35:30s plt. library.
  24241. 18:35:34Okay, so Seabone is a library that uses
  24242. 18:35:36Matt plot lib underneath to plot graphs.
  24243. 18:35:39It will be used to visualize zandom
  24244. 18:35:41distribution and the word cloud is a
  24245. 18:35:44visual representations
  24246. 18:35:46of words. Cloud creators are used to
  24247. 18:35:48highlight popular words and phrases
  24248. 18:35:50based on frequency and relevance. They
  24249. 18:35:52provide you with quick and simple visual
  24250. 18:35:55insights that can lead to more in-depth
  24251. 18:35:57analysis. And this mattplot lil mattplot
  24252. 18:36:00lib is a plotting library for the python
  24253. 18:36:01programming language and its numerical
  24254. 18:36:03mathematic
  24255. 18:36:04ext extension numpy. It provides an
  24256. 18:36:07object- oriented API for embedding plots
  24257. 18:36:09into application using general purpose
  24258. 18:36:12UI.
  24259. 18:36:13Okay. Like tinker wxython QT or gtk.
  24260. 18:36:19So let's import some
  24261. 18:36:22NLTK
  24262. 18:36:25natural language toolkit.
  24263. 18:36:28So
  24264. 18:36:29I will write import
  24265. 18:36:32NLTK.
  24266. 18:36:36Okay. from
  24267. 18:36:38NLTK
  24268. 18:36:40dot stem
  24269. 18:36:43importizer
  24270. 18:36:54then from analytic dot corpus
  24271. 18:36:58imports
  24272. 18:37:06and
  24273. 18:37:08from
  24274. 18:37:10NL ticket dot tokenize
  24275. 18:37:14port
  24276. 18:37:21tokenize
  24277. 18:37:25NLTK the natural language toolkit or
  24278. 18:37:28more commonly NLTK is a suit of
  24279. 18:37:30libraries and programs for symbolic and
  24280. 18:37:33statical natural language processing for
  24281. 18:37:35English written in Python programming
  24282. 18:37:37language and this is stop words. Stop
  24283. 18:37:39words are words that are so common they
  24284. 18:37:41are basically ignored by typical
  24285. 18:37:44tokenizers and this word tokenize is a
  24286. 18:37:46function in Python that splits a given
  24287. 18:37:48sentence into words using the analytical
  24288. 18:37:51library. Okay. So let's import some
  24289. 18:37:54scikitlearn
  24290. 18:37:56library. So for that I will write from
  24291. 18:37:59skarn
  24292. 18:38:02dot model
  24293. 18:38:05collection
  24294. 18:38:08import
  24295. 18:38:10train
  24296. 18:38:12test.
  24297. 18:38:17Okay. Then from skarn
  24298. 18:38:25dot feature
  24299. 18:38:31extraction
  24300. 18:38:34dot text import
  24301. 18:38:39vectorizer.
  24302. 18:38:44And then from
  24303. 18:38:47skarn dot matrices
  24304. 18:38:51matrix
  24305. 18:38:52import
  24306. 18:38:54confusion metric
  24307. 18:39:01classification.
  24308. 18:39:08Okay.
  24309. 18:39:10So, scikitlarn is a free source software
  24310. 18:39:13machine learning library for Python
  24311. 18:39:15programming language. It features
  24312. 18:39:17various classification, regression and
  24313. 18:39:18clustering algorithms including support
  24314. 18:39:21vector machine learning, logistic
  24315. 18:39:22regression and many others like random
  24316. 18:39:25forest classifier. And this train test
  24317. 18:39:28split method is used to split our data
  24318. 18:39:30into train and test set. First, we need
  24319. 18:39:32to divide our data into features like X
  24320. 18:39:34and Y labels. And this TF ID vectorzer
  24321. 18:39:38converts a collection of raw documents
  24322. 18:39:40into a matrix of TF features. The fast
  24323. 18:39:44text or what to vectorizer what
  24324. 18:39:47embedding Python implementation and this
  24325. 18:39:49confusion matrix. A confusion matrix is
  24326. 18:39:51a table that is used to define the
  24327. 18:39:53performance of a classification
  24328. 18:39:54algorithm. Okay.
  24329. 18:39:57Then we'll import some libraries like
  24330. 18:40:00prom skarn
  24331. 18:40:08do linear model
  24332. 18:40:15port
  24333. 18:40:17logistic
  24334. 18:40:21regression.
  24335. 18:40:23So then from
  24336. 18:40:26colonm
  24337. 18:40:32port
  24338. 18:40:40and from
  24339. 18:40:51import
  24340. 18:40:54random
  24341. 18:40:56forests classifier.
  24342. 18:41:01Okay. Then from
  24343. 18:41:04skarn dot name base
  24344. 18:41:10portoli
  24345. 18:41:15base.
  24346. 18:41:19Okay.
  24347. 18:41:22So everything is correct. You will see
  24348. 18:41:26while running. So logistic regression
  24349. 18:41:28estimate the probability of an event
  24350. 18:41:31occurring such as voted or didn't vote
  24351. 18:41:34based on a given data set of the
  24352. 18:41:36independent variable.
  24353. 18:41:40L SVC logistic regression estimate
  24354. 18:41:44sorry linear support vector machine SVC
  24355. 18:41:47is an algorithm that attempts to find a
  24356. 18:41:49hyper plane to maximize the distance
  24357. 18:41:52between classified samples and this
  24358. 18:41:54random forest classifier creates a set
  24359. 18:41:56of decision trees from a randomly
  24360. 18:41:59selected subset of the training set
  24361. 18:42:03and this Bernoli NBoli
  24362. 18:42:05name base is a part of the name base
  24363. 18:42:07family it is based on Bernoli
  24364. 18:42:09distribution ution and accept only
  24365. 18:42:11binary values that is zero or one.
  24366. 18:42:15So let's import some tensorflow. So
  24367. 18:42:17import
  24368. 18:42:20tensorflow
  24369. 18:42:22dot
  24370. 18:42:24compad dot v2 and
  24371. 18:42:29then import
  24372. 18:42:32tensorflow
  24373. 18:42:35data sets
  24374. 18:42:37as tfds.
  24375. 18:42:40So, TensorFlow is a free and open-source
  24376. 18:42:42library for machine learning and
  24377. 18:42:43artificial intelligence across a range
  24378. 18:42:46of task but has a particular focus on
  24379. 18:42:48training and inference of deep neural
  24380. 18:42:50networks. Okay, let's import warnings.
  24381. 18:42:57Nothing. Warning.
  24382. 18:43:02The warnings
  24383. 18:43:11import
  24384. 18:43:18string
  24385. 18:43:21import.
  24386. 18:43:24So everything is basic. Just let's see
  24387. 18:43:25the pickle. Typically is a Python is
  24388. 18:43:28primarily used in serializing and
  24389. 18:43:31deserializing a Python object structure.
  24390. 18:43:33Okay, let's run it. Let's see how many
  24391. 18:43:37error
  24392. 18:43:42after that we will load the data set and
  24393. 18:43:45uh we will go through data
  24394. 18:43:46visualization. Okay. Word cloud cannot
  24395. 18:43:49import name word cloud. Okay. C C will
  24396. 18:43:52be capital here.
  24397. 18:43:57random forest.
  24398. 18:44:06Okay, it's still loading here. Let's
  24399. 18:44:08see. Okay, so loading is done. So now
  24400. 18:44:12let's load the data set. So we'll write
  24401. 18:44:14data equals to PD dot
  24402. 18:44:19read
  24403. 18:44:21CSV
  24404. 18:44:26name
  24405. 18:44:37test.
  24406. 18:44:48So you can find this data set on the
  24407. 18:44:50description box below.
  24408. 18:44:53According to question
  24409. 18:45:04polarity
  24410. 18:45:16ID,
  24411. 18:45:18comma date,
  24412. 18:45:20comma
  24413. 18:45:22query
  24414. 18:45:46per you forget comma
  24415. 18:46:01polarity. Okay.
  24416. 18:46:14Seems fine.
  24417. 18:46:19Let me change this first
  24418. 18:46:28is using RNN.
  24419. 18:46:33Okay.
  24420. 18:46:36So here I will write data plus data dot
  24421. 18:46:42sample.
  24422. 18:46:45Let's
  24423. 18:46:49do one.
  24424. 18:47:05Okay. So, let me like brief uh tell you
  24425. 18:47:09that what we are going. Okay.
  24426. 18:47:14Let me brief you like what we will do in
  24427. 18:47:16this sentiment analysis using RNA. So in
  24428. 18:47:20this demo like you will see uh text
  24429. 18:47:22processing on Twitter data set and after
  24430. 18:47:26that we will perform different machine
  24431. 18:47:27learning algorithms on the data such as
  24432. 18:47:29logistic regression random forest
  24433. 18:47:32classifier SVC nas to classify positive
  24434. 18:47:35and negative dudes. After that I will
  24435. 18:47:38also build RNN recurrent neural network
  24436. 18:47:40which is the best fit for such textual
  24437. 18:47:43sentiment analysis. Okay. Since it's a
  24438. 18:47:46sequential data set which is requirement
  24439. 18:47:48for the RNN network. So let's dive into.
  24440. 18:47:52So now
  24441. 18:47:54we will see the data data visualization
  24442. 18:47:56data set details target like the
  24443. 18:47:58polarity of the tweets zero negative.
  24444. 18:48:01Okay. then the date like date of the
  24445. 18:48:04tweet and the polarity and the user that
  24446. 18:48:07what tweeted then the text okay so I
  24447. 18:48:10will write print
  24448. 18:48:15data set
  24449. 18:48:21data
  24450. 18:48:23shape
  24451. 18:48:25okay
  24452. 18:48:30let me first do like this. Yeah.
  24453. 18:48:35So there are 20
  24454. 18:48:40or you can say two like rows and six
  24455. 18:48:43number of columns. Okay. So it is a huge
  24456. 18:48:46data. I will you can find this data set
  24457. 18:48:49from the description box below. So here
  24458. 18:48:52let's see the data
  24459. 18:48:57and why I use head. Head is used for
  24460. 18:49:01like
  24461. 18:49:03for showing
  24462. 18:49:07top 10 rows of the data set. If you will
  24463. 18:49:10use tail instead of head, it will show
  24464. 18:49:12the last 10 rows of the data set. Okay.
  24465. 18:49:15Here polarity zero. Zero means negative
  24466. 18:49:18and four means positive. Okay. Like you
  24467. 18:49:21can consider 01.
  24468. 18:49:25This is ID, date, then query. then user
  24469. 18:49:29then the text.
  24470. 18:49:35Okay. So
  24471. 18:49:38here I will do data
  24472. 18:49:41clarity.
  24473. 18:49:51Okay. These are the 04. Okay.
  24474. 18:49:54Uniqueness. Zero means negative and the
  24475. 18:49:56four means positive. replacing the value
  24476. 18:49:58four as one for the ease of
  24477. 18:50:00understanding what I said to you you can
  24478. 18:50:02consider as 01. So data
  24479. 18:50:05polarity
  24480. 18:50:09to data
  24481. 18:50:12polarity
  24482. 18:50:19to one
  24483. 18:50:22and then data.
  24484. 18:50:30So now you can see 0 1 0 1 0 1 1 0.
  24485. 18:50:34Okay.
  24486. 18:50:36So if you will write only head it will
  24487. 18:50:38show the top five rows only. Okay.
  24488. 18:50:43So now let's use one Python function
  24489. 18:50:46describe
  24490. 18:50:48data dotribe.
  24491. 18:50:59So as you can see here count is two lakh
  24492. 18:51:02and the mean of the particular row is
  24493. 18:51:05this and the ID is this standard
  24494. 18:51:08deviation minimum value the 25% the 50%
  24495. 18:51:12and the 75% and the maximum
  24496. 18:51:15okay let's see the number of positive
  24497. 18:51:18versus negative tagged sentence okay
  24498. 18:51:27so here I I will write positives
  24499. 18:51:32to data
  24500. 18:51:34polarity
  24501. 18:51:41data dot polarity
  24502. 18:51:44= 1.
  24503. 18:51:47Then it is
  24504. 18:51:50data
  24505. 18:51:54polarity
  24506. 18:51:58data dot polarity
  24507. 18:52:03is equals to zero.
  24508. 18:52:08Print
  24509. 18:52:11total
  24510. 18:52:14length of the data is
  24511. 18:52:27dot format
  24512. 18:52:30data
  24513. 18:52:31dot shape. Yep.
  24514. 18:52:50Now I will print
  24515. 18:52:54the total length, the negative and the
  24516. 18:52:56positive. Okay. So number of positive
  24517. 18:53:09Okay.
  24518. 18:53:14Format
  24519. 18:53:18positives.
  24520. 18:53:24So I will copy
  24521. 18:53:29this and paste it here.
  24522. 18:53:31And here I will do the changes for the
  24523. 18:53:34negatives.
  24524. 18:53:43Okay. Now let's see.
  24525. 18:53:50So here polarity is not defined.
  24526. 18:54:07So as you can see the total length of
  24527. 18:54:09the data is two lakh and the number of
  24528. 18:54:11positive sentences is like one lakh 46
  24529. 18:54:18and number of negatives okay spelling
  24530. 18:54:21this
  24531. 18:54:25the number of negative text sentences
  24532. 18:54:2899,954.
  24533. 18:54:30Okay. So now we have a brief data.
  24534. 18:54:53So now let's get a word count p of text.
  24535. 18:54:56So for this I will write
  24536. 18:55:05count
  24537. 18:55:06words
  24538. 18:55:13done
  24539. 18:55:16length of
  24540. 18:55:18start split.
  24541. 18:55:23Okay.
  24542. 18:55:29And now let's plot a word count
  24543. 18:55:31distribution for both positive and
  24544. 18:55:33negative. So I will create a bar plot.
  24545. 18:55:36So for that I will write it
  24546. 18:55:40word
  24547. 18:55:42count.
  24548. 18:55:45data
  24549. 18:55:48text
  24550. 18:55:50dot apply
  24551. 18:55:54but count.
  24552. 18:56:00Okay, then I will write P positive= data
  24553. 18:56:05then
  24554. 18:56:08count
  24555. 18:56:15data dot polarity
  24556. 18:56:19is equals to 1
  24557. 18:56:23and
  24558. 18:56:25let me copy this Here
  24559. 18:56:35I will write zero
  24560. 18:56:41and okay then
  24561. 18:56:44plt dot figure
  24562. 18:56:48and figure size
  24563. 18:56:52equ= to
  24564. 18:56:5312 Thanks.
  24565. 18:57:07Okay. Then plt
  24566. 18:57:12LT dota
  24567. 18:57:1845
  24568. 18:57:23then plt dot x label
  24569. 18:57:29word count
  24570. 18:57:33plt dot y label
  24571. 18:57:38and frequency
  24572. 18:57:44we'll write uh g
  24573. 18:57:49dot
  24574. 18:57:55comma n
  24575. 18:57:58Uh
  24576. 18:58:19alpha also 0.5 Five
  24577. 18:58:30positive.
  24578. 18:58:39Okay. Then let's make a legend also.
  24579. 18:58:46Location should be
  24580. 18:58:53Right.
  24581. 18:59:00False.
  24582. 18:59:03Data word count equals to
  24583. 18:59:15Okay, my bad.
  24584. 18:59:27So as you can see the positive and the
  24585. 18:59:29negatives.
  24586. 18:59:31Okay.
  24587. 18:59:33So these are the like word count
  24588. 18:59:35distribution for both positive and
  24589. 18:59:37negative. Okay.
  24590. 18:59:41Now let's uh what we can do we can do
  24591. 18:59:45the get like get the common words in
  24592. 18:59:47training data set for the training data
  24593. 18:59:49set. So for that I will do
  24594. 18:59:53from
  24595. 18:59:55collections
  24596. 18:59:57import
  24597. 19:00:01counter
  24598. 19:00:04or
  24599. 19:00:07words
  24600. 19:00:08to
  24601. 19:00:12for
  24602. 19:00:18test
  24603. 19:00:22data
  24604. 19:00:24text
  24605. 19:00:30line
  24606. 19:00:32dot split
  24607. 19:00:37forward. Word
  24608. 19:00:40and words.
  24609. 19:00:48If length of
  24610. 19:00:51word
  24611. 19:00:53than two
  24612. 19:00:58all
  24613. 19:01:01dot
  24614. 19:01:05one dot lower
  24615. 19:01:13here I can write counter
  24616. 19:01:16all words
  24617. 19:01:19dot most
  24618. 19:01:22common then I need 20.
  24619. 19:01:31So as you can see these are the most
  24620. 19:01:33common word used like in every sentence
  24621. 19:01:36the and you for have that I am but just
  24622. 19:01:40like this out over all.
  24623. 19:01:43So these are the most common words like
  24624. 19:01:46it used the is used like 64,000 times
  24625. 19:01:49and like this UR is used for 8,000 times
  24626. 19:01:56something like that. So now we will do
  24627. 19:01:58some data pro data processing. Okay. Now
  24628. 19:02:02let's do the data processing.
  24629. 19:02:05So
  24630. 19:02:08div
  24631. 19:02:17and SNS dot current plot
  24632. 19:02:24data
  24633. 19:02:26polarity.
  24634. 19:02:29Okay,
  24635. 19:02:34these are the uh negatives and this
  24636. 19:02:37positives.
  24637. 19:02:41There is a slight change I guess that is
  24638. 19:02:45why it's not looking
  24639. 19:02:49so much of different like there's a
  24640. 19:02:51slight
  24641. 19:02:5346 different so that is why it's looking
  24642. 19:02:56almost same. Okay.
  24643. 19:03:01So now removing the unnecessary columns
  24644. 19:03:04like query, user, word count, data dot
  24645. 19:03:08drop,
  24646. 19:03:13date
  24647. 19:03:17query
  24648. 19:03:25and word count.
  24649. 19:03:35X = 1
  24650. 19:03:38comma
  24651. 19:03:40place= to true.
  24652. 19:03:45Okay.
  24653. 19:03:47Uh A will be true.
  24654. 19:03:57So here I will write data
  24655. 19:04:01what is this? No. Okay my bad.
  24656. 19:04:15So here I will write data dot drop
  24657. 19:04:19id
  24658. 19:04:21comma
  24659. 19:04:23one
  24660. 19:04:36then data dot head
  24661. 19:04:42the data see we have only the to the
  24662. 19:04:44polarity and the text. Okay.
  24663. 19:04:52So
  24664. 19:04:55now uh let's see the null values.
  24665. 19:05:00So data
  24666. 19:05:02dot
  24667. 19:05:09um
  24668. 19:05:15data
  24669. 19:05:20print. Okay.
  24670. 19:05:43So there is no null values. So now
  24671. 19:05:45converting pandas's object to a string
  24672. 19:05:48type.
  24673. 19:05:50For that we have to write
  24674. 19:05:52text
  24675. 19:05:55to data
  24676. 19:05:59text.
  24677. 19:06:07Yeah.
  24678. 19:06:15Get as type.
  24679. 19:06:22Yeah. So now download the stop words
  24680. 19:06:26NLTK.
  24681. 19:06:29Download
  24682. 19:06:37words.
  24683. 19:06:40words
  24684. 19:06:43as you said
  24685. 19:06:45stop words
  24686. 19:06:51it's in English
  24687. 19:07:00stop
  24688. 19:07:16This
  24689. 19:07:37These are some, you know, stop words.
  24690. 19:07:42So moving forward, let's download
  24691. 19:07:45NLTK dot download.net.
  24692. 19:08:10So the pre-processing steps taken are
  24693. 19:08:12like lower casting each text is
  24694. 19:08:14converted to lower case then remover of
  24695. 19:08:17URLs will do this we will do okay links
  24696. 19:08:20starting with http or https or ww are
  24697. 19:08:24replaced by like commas and removing
  24698. 19:08:28usernames removing short words removing
  24699. 19:08:30stop words like limitization is the
  24700. 19:08:32process will do of for the converting a
  24701. 19:08:35word to its base Okay. So for that
  24702. 19:08:41what I will do
  24703. 19:08:44we'll just copy the whole code for you.
  24704. 19:08:49We'll explain you one by one what I've
  24705. 19:08:51done.
  24706. 19:08:53Okay.
  24707. 19:08:58So this is a course for the URL pattern
  24708. 19:09:00for removing all the WW, HTTPS and HTTP
  24709. 19:09:05type of thing and removing
  24710. 19:09:08them. Then I have used pattern for the
  24711. 19:09:12lower casting removing all the URLs.
  24712. 19:09:15Okay.
  24713. 19:09:16Then removing all the usernames like at
  24714. 19:09:19the red and removing punctuations
  24715. 19:09:22and stop words.
  24716. 19:09:25Okay. Like this.
  24717. 19:09:30So now what we have to do data
  24718. 19:09:40processed
  24719. 19:09:43weights
  24720. 19:09:48then data
  24721. 19:09:51Next
  24722. 19:10:04dot apply
  24723. 19:10:08lambda
  24724. 19:10:10x
  24725. 19:10:12process
  24726. 19:10:17then tweets.
  24727. 19:10:24Okay.
  24728. 19:10:35Then print
  24729. 19:10:39next.
  24730. 19:10:41reprocessing.
  24731. 19:10:50It is taking time
  24732. 19:11:17It will be completed. It will return
  24733. 19:11:19here the text prep-processing is done.
  24734. 19:11:22Okay.
  24735. 19:11:24As you can see the text prep-processing
  24736. 19:11:26is done. So now let's check
  24737. 19:11:31data dot add
  24738. 19:11:3410.
  24739. 19:11:37As you can see see the at the rate and
  24740. 19:11:40this slices are gone.
  24741. 19:11:43Okay.
  24742. 19:11:44So now the text is pre-processed.
  24743. 19:11:48So now what we will do? We will analyze
  24744. 19:11:50the data. So now we are going to analyze
  24745. 19:11:52the pre-processed data to get an
  24746. 19:11:54understanding of it. We will plot word
  24747. 19:11:56clouds for positive and negative dudes
  24748. 19:11:58from our data set and see which words
  24749. 19:12:01occurs the most. Okay. First we will uh
  24750. 19:12:05create for the negative words or
  24751. 19:12:08negative tweets you can say. So I will
  24752. 19:12:10write pl dot figure
  24753. 19:12:13then figure size
  24754. 19:12:1815.
  24755. 19:12:23Okay. Then word cloud also
  24756. 19:12:27word cloud
  24757. 19:12:32x words
  24758. 19:12:362,00 comma
  24759. 19:12:39width = to 1,600
  24760. 19:12:45comma
  24761. 19:12:46height = to 800
  24762. 19:12:52rate
  24763. 19:12:58dot join the data dot polarity
  24764. 19:13:05and I will write here polarity
  24765. 19:13:10okay equals equals to zero
  24766. 19:13:17again
  24767. 19:13:24then
  24768. 19:13:27processed tweets.
  24769. 19:13:33Okay.
  24770. 19:13:36Then here
  24771. 19:13:38I have to write plt dot show
  24772. 19:13:44me show
  24773. 19:13:49wcolation
  24774. 19:14:00linear.
  24775. 19:14:04Perhaps you forget the comma here.
  24776. 19:14:13So 2000
  24777. 19:14:16then comma width
  24778. 19:14:19dot generate
  24779. 19:14:32here. what I can do.
  24780. 19:14:44Let me run now. Let's see. Hope this
  24781. 19:14:48time it will work.
  24782. 19:14:55Guess is still loading.
  24783. 19:15:16As you can see this is
  24784. 19:15:22okay like today I am and work don't wish
  24785. 19:15:27they need much. These are the most
  24786. 19:15:30negative tweets. Okay, words from
  24787. 19:15:33negative tweets you can say,
  24788. 19:15:36right? So, let's see the positive
  24789. 19:15:40tweets. Okay,
  24790. 19:15:46so the thing will be same.
  24791. 19:15:50Let me copy
  24792. 19:15:53paste it here. So for this I will do one
  24793. 19:16:01it will take a little bit of time
  24794. 19:16:05to come loading like as you can see hit
  24795. 19:16:09can't. Okay sorry
  24796. 19:16:14these are the negative words.
  24797. 19:16:18Okay still loading. So let's wait for
  24798. 19:16:21like few seconds.
  24799. 19:16:26Now you can see the positive words like
  24800. 19:16:28love, okay, good, lol
  24801. 19:16:33and awesome something like that. Okay,
  24802. 19:16:37so these are some
  24803. 19:16:39positive words. So now let's do the
  24804. 19:16:42vectorzation and splitting the data like
  24805. 19:16:44storing into input variable process to X
  24806. 19:16:47and output variable polarity to Y. Okay,
  24807. 19:16:50we'll do that.
  24808. 19:16:57So x = to data
  24809. 19:17:02possessed
  24810. 19:17:07with
  24811. 19:17:10values
  24812. 19:17:13and pi= to data
  24813. 19:17:18entity
  24814. 19:17:21dot
  24815. 19:17:22values.
  24816. 19:17:24Okay.
  24817. 19:17:26Now I will write here print
  24818. 19:17:31dot shape
  24819. 19:17:35print y dot.shape.
  24820. 19:17:41Okay cool. So now what we will do we
  24821. 19:17:43will convert text to word frequency
  24822. 19:17:45vectors. Okay. TF to IDF. So this is an
  24823. 19:17:48acronym that stand for term frequency to
  24824. 19:17:52inverse document frequency which are the
  24825. 19:17:54components of the resulting scores
  24826. 19:17:56assigned to each word. Okay. So term
  24827. 19:17:58frequency this summarize how often a
  24828. 19:18:00given word appears within a document and
  24829. 19:18:03inverse document frequency this
  24830. 19:18:05downscales word that appear a lot across
  24831. 19:18:07documents. Okay. So now here we will
  24832. 19:18:11convert a collection of raw documents to
  24833. 19:18:12a matrix of TF to IDF features. Okay.
  24834. 19:18:16And then I will write
  24835. 19:18:19enter
  24836. 19:18:21kazut
  24837. 19:18:23riser
  24838. 19:18:30and sublinear
  24839. 19:18:40x = to
  24840. 19:18:45dot with
  24841. 19:18:48transform
  24842. 19:18:55printed.
  24843. 19:19:06Okay.
  24844. 19:19:11Print
  24845. 19:19:15number of feature
  24846. 19:19:24comma length
  24847. 19:19:28vector
  24848. 19:19:31do get
  24849. 19:19:34their names.
  24850. 19:19:41So number of feature words are like 1703
  24851. 19:19:44to 1.
  24852. 19:19:46Okay. Now we will do like
  24853. 19:19:51now let's print the shape.
  24854. 19:20:10So now we will do the split uh spread to
  24855. 19:20:13train and test. So the pre-provised data
  24856. 19:20:15is divided into two sets of data
  24857. 19:20:17training data and the testing data. So
  24858. 19:20:19data set upon which the model would be
  24859. 19:20:21trained on contains 80% data and the
  24860. 19:20:25test data is the data set upon which
  24861. 19:20:27model would be tested again contains 20%
  24862. 19:20:30of data. So for that I will write extra
  24863. 19:20:37test
  24864. 19:20:39comma
  24865. 19:20:50Test
  24866. 19:21:00test size
  24867. 19:21:02= to 0.20 2
  24868. 19:21:07random
  24869. 19:21:09state
  24870. 19:21:11101.
  24871. 19:21:16Okay. Random state.
  24872. 19:21:23So what I will do? I will do the you
  24873. 19:21:25know print the shape of X train, Y
  24874. 19:21:27train, X test, Y test like how many
  24875. 19:21:29columns are there? Rows not column
  24876. 19:21:32exactly the rows are there. Okay.
  24877. 19:21:35So we'll paste there. So see
  24878. 19:21:39extra train like this is a total was
  24879. 19:21:41like two lakh.
  24880. 19:21:44Okay. So 1 lakh 60,000 in training as we
  24881. 19:21:49discussed earlier like 80% in training
  24882. 19:21:51and 20% in testing. Okay.
  24883. 19:21:56So now let's do the model building.
  24884. 19:21:59Okay. Model evaluating functions. So now
  24885. 19:22:03let's make a model.
  24886. 19:22:07Okay.
  24887. 19:22:08And first I will do I will write and
  24888. 19:22:12then I will explain you the whole. Okay.
  24889. 19:22:16So here what I did uh this will tell you
  24890. 19:22:19the accuracy of the model of training
  24891. 19:22:20data and the testing data. Okay. Then we
  24892. 19:22:24will predict the values for test data
  24893. 19:22:26set and the evaluation for the data set.
  24894. 19:22:28Then we will compute and plot the
  24895. 19:22:30confusion matrix.
  24896. 19:22:32Okay, the both the categories negative
  24897. 19:22:34positives. Okay, group name will be true
  24898. 19:22:36negative and the false positive. Okay,
  24899. 19:22:39so there's nothing that's let's run it.
  24900. 19:22:46So now what we will do? We will do first
  24901. 19:22:48for the logistic regression. So here I
  24902. 19:22:51will write LG equals to
  24903. 19:22:54logistic
  24904. 19:22:58regression.
  24905. 19:23:01Okay. Then history
  24906. 19:23:04equals to LG do fit
  24907. 19:23:09X train,
  24908. 19:23:13Y train
  24909. 19:23:16with model
  24910. 19:23:19evaluate
  24911. 19:23:23LG. Now let's see
  24912. 19:23:30this is for the logistic regression.
  24913. 19:23:34Okay,
  24914. 19:23:38as you can see the accuracy of the
  24915. 19:23:40training data is 83% the testing data is
  24916. 19:23:4277%.
  24917. 19:23:43Okay.
  24918. 19:23:48So this is the confidence matrix the
  24919. 19:23:51predictive value like these are the
  24920. 19:23:53categories.
  24921. 19:23:55Now let's see for the linear SPM. For
  24922. 19:23:57that I will write SPM
  24923. 19:24:02equals to
  24924. 19:24:06SVC
  24925. 19:24:11then SVM
  24926. 19:24:14dot fit
  24927. 19:24:18train.
  24928. 19:24:23Then model
  24929. 19:24:27evaluate
  24930. 19:24:30of SVM.
  24931. 19:24:37Okay.
  24932. 19:24:39And after that we will do for random
  24933. 19:24:40forest and the N base. Okay. Then we
  24934. 19:24:43will start with the RNN.
  24935. 19:24:45So as you can see the accuracy of
  24936. 19:24:47training data is very pretty good 93%
  24937. 19:24:50and logation is 83% and the testing is
  24938. 19:24:55less than
  24939. 19:24:57regression model. Let's see for the
  24940. 19:25:00random forest. So I will write here RF
  24941. 19:25:03equals to
  24942. 19:25:06random forest
  24943. 19:25:10fire
  24944. 19:25:15m=
  24945. 19:25:18to 20
  24946. 19:25:23criterion = to
  24947. 19:25:29tropy Okay.
  24948. 19:25:32Then max
  24949. 19:25:34depth equals to 50.
  24950. 19:25:38Then RF dot fit
  24951. 19:25:44X train,
  24952. 19:25:48Y train
  24953. 19:25:51and model
  24954. 19:25:58evaluate.
  24955. 19:26:05Okay,
  24956. 19:26:14loading. Let's see the accuracy how it
  24957. 19:26:16will come.
  24958. 19:26:26After this we will do for the name base
  24959. 19:26:28and after that we will move on to the
  24960. 19:26:31our main model RNN recurrent neural
  24961. 19:26:34networks.
  24962. 19:26:35It's still loading.
  24963. 19:26:41Guess it will take little bit of time.
  24964. 19:26:51So as you can see the confusion matrix.
  24965. 19:26:56Okay. So training data accuracy is 75%
  24966. 19:27:01very less. So now let's see the last
  24967. 19:27:04model name base. Okay. So, NB equals to
  24968. 19:27:12NB
  24969. 19:27:16NB dot fit
  24970. 19:27:21SP,
  24971. 19:27:26wide train.
  24972. 19:27:33Okay. Then model
  24973. 19:27:37evaluate.
  24974. 19:27:45Oh, NAB base training 867.
  24975. 19:27:50So, as for
  24976. 19:27:52linear SEC has the best
  24977. 19:27:55test training uh accuracy you can say
  24978. 19:27:58and the best testing accuracy is 7670
  24979. 19:28:0376.45 4 five
  24980. 19:28:06see logistic regression. So now let's
  24981. 19:28:08move to the our main model RNN. So what
  24982. 19:28:12is RNN recurrent neural network at the
  24983. 19:28:14start
  24984. 19:28:16are the state-ofthe-art algorithm for
  24985. 19:28:18sequential data and are used by Apple CD
  24986. 19:28:21and Google search voice. It is the first
  24987. 19:28:24algorithm that remembers its input due
  24988. 19:28:27to an internal memory which make it
  24989. 19:28:28perfectly suited for machine learning
  24990. 19:28:30problem that involve sequential data.
  24991. 19:28:33And there is one more thing embedding
  24992. 19:28:35layer. Embedding layer is one of the
  24993. 19:28:36available layers in KAS. This is mainly
  24994. 19:28:39used in natural language processing
  24995. 19:28:41related applications such as language
  24996. 19:28:43modeling but it can also be used with
  24997. 19:28:45other tasks that involve neural networks
  24998. 19:28:47while dealing with NLP problems. We can
  24999. 19:28:49use pre-trained word embedding such as
  25000. 19:28:51glow.
  25001. 19:28:53Alternately we can also train our own
  25002. 19:28:56embeddings using kas emitting layer.
  25003. 19:29:00LSTM layer long short-term memory
  25004. 19:29:02networks usually called LSTMs I have
  25005. 19:29:05made already many videos you can check
  25006. 19:29:08it out were introduced by Skyer these
  25007. 19:29:11have widely been used for speech
  25008. 19:29:13recognition language processing
  25009. 19:29:15sentiment analysis and text prediction
  25010. 19:29:17before going deep into LSTM we should
  25011. 19:29:20first understand the need of LSTM which
  25012. 19:29:22can be explained by the drawback of
  25013. 19:29:24practical use of RNN so let's start with
  25014. 19:29:26RNA
  25015. 19:29:28Okay.
  25016. 19:29:31So here I will importing some libraries.
  25017. 19:29:36Okay. So after that I will write import
  25018. 19:29:41kas
  25019. 19:29:48version
  25020. 19:29:512.110. Okay fine.
  25021. 19:29:57So now let's
  25022. 19:30:03paint
  25023. 19:30:19X test, comma,
  25024. 19:30:23white train.
  25025. 19:30:29Let's do train
  25026. 19:30:33test
  25027. 19:30:41weights, comma, data dot polarity
  25028. 19:30:46dot values.
  25029. 19:30:49Then test
  25030. 19:30:52size equals to 0.2. Test size 0.2 means
  25031. 19:30:59like 80 and 20%
  25032. 19:31:02thing 80 to training and then 20% to
  25033. 19:31:08testing.
  25034. 19:31:13Okay.
  25035. 19:31:15and let's
  25036. 19:31:20the model evaluation. Okay.
  25037. 19:31:24So I will these are relu sigmoid all the
  25038. 19:31:28you know the layers.
  25039. 19:31:32So now this epoch it will run till 5,000
  25040. 19:31:35like count will go till 5,000. Okay see
  25041. 19:31:40the 5,000 and it will go to 1 to 10. So
  25042. 19:31:42it will take time. So I will get back to
  25043. 19:31:44you after this completing this. Okay.
  25044. 19:31:48Now as you can see uh the box
  25045. 19:31:52ran successfully. Okay. So what should I
  25046. 19:31:56do? But I will give some space here. So
  25047. 19:31:59now we will see the positive and
  25048. 19:32:01negative outcome. Okay. This is
  25049. 19:32:04something like testing. Okay. We will
  25050. 19:32:06test. We will predict. we will give one
  25051. 19:32:10uh a sentence and then we will predict
  25052. 19:32:13it is coming right or wrong. The
  25053. 19:32:15accuracy is giving a right or wrong.
  25054. 19:32:17Okay. So here I will write
  25055. 19:32:20sequence
  25056. 19:32:22equals to tokenizer
  25057. 19:32:26dot text
  25058. 19:32:31to
  25059. 19:32:33sequences.
  25060. 19:32:36Okay, then I'll write this
  25061. 19:32:42data science
  25062. 19:32:45article.
  25063. 19:32:48This was
  25064. 19:32:51okay.
  25065. 19:32:53So here I will write test equals to P
  25066. 19:32:58sequences
  25067. 19:33:02and here I will write sequence
  25068. 19:33:07Comma max length
  25069. 19:33:13to
  25070. 19:33:15max length.
  25071. 19:33:24Then I will write here prediction equals
  25072. 19:33:28to model.
  25073. 19:33:37We write model
  25074. 19:33:43then we'll write model 12
  25075. 19:33:49dot predict
  25076. 19:33:53then test.
  25077. 19:33:55Okay.
  25078. 19:33:57If diction
  25079. 19:34:02is greater than 0.5 means 50%.
  25080. 19:34:06Then
  25081. 19:34:08it should print
  25082. 19:34:11positive.
  25083. 19:34:14Okay.
  25084. 19:34:16Else
  25085. 19:34:22negative.
  25086. 19:34:26Okay. Let me run this.
  25087. 19:34:30Okay. Sequential
  25088. 19:34:32object has no okay spellic
  25089. 19:34:40see the negative because here is the
  25090. 19:34:42word worst it is showing correct. Now
  25091. 19:34:45check from the RNN model. So model
  25092. 19:34:48equals to kas dot models dot load
  25093. 19:34:54models. Here we will load RNN model. RNN
  25094. 19:34:58model
  25095. 19:35:00SG file. It is pre-trained model. Okay.
  25096. 19:35:04Pretend RNN model. So sequence
  25097. 19:35:10tokenizer
  25098. 19:35:14dot text
  25099. 19:35:17to sequences.
  25100. 19:35:22Then
  25101. 19:35:24I will write here this this
  25102. 19:35:31ML
  25103. 19:35:34course
  25104. 19:35:37is best.
  25105. 19:35:40Okay. S equals to P sequences
  25106. 19:35:47sequence
  25107. 19:36:00X
  25108. 19:36:04okay then prediction equals to model dot
  25109. 19:36:09predict
  25110. 19:36:11Then test
  25111. 19:36:15if prediction is greater than 0.5
  25112. 19:36:23in
  25113. 19:36:24positive 0.5 means 50% more than 50%.
  25114. 19:36:29Else
  25115. 19:36:31print negative
  25116. 19:36:42attribute load models
  25117. 19:36:50positive because this ML course is best.
  25118. 19:36:53So there is no negative word.
  25119. 19:36:57Okay.
  25120. 19:37:05So what we will do now we will do model
  25121. 19:37:08saving loading and prediction. Okay. So
  25122. 19:37:11for that uh I will write import pickle
  25123. 19:37:17file = to open
  25124. 19:37:23vectorzer
  25125. 19:37:32then
  25126. 19:37:40here I will Pickle
  25127. 19:37:42dot dump
  25128. 19:37:45dump
  25129. 19:37:46vector file vector.
  25130. 19:38:03Okay.
  25131. 19:38:04So like this I have to write for name
  25132. 19:38:07base logitation SVM and random forest.
  25133. 19:38:12So
  25134. 19:38:13what I will do
  25135. 19:38:16right here.
  25136. 19:38:17Okay. Let's run this.
  25137. 19:38:20Okay.
  25138. 19:38:35Now what we have to do? We have to
  25139. 19:38:37predict using saved model. Okay.
  25140. 19:38:47What we will do here? We will load model
  25141. 19:38:49first and we will predict. Okay. So
  25142. 19:38:55first I will write the function name
  25143. 19:38:57load
  25144. 19:38:59models
  25145. 19:39:02and we will load the vectorzer. So file
  25146. 19:39:05equals to
  25147. 19:39:07open
  25148. 19:39:09vectorzer
  25149. 19:39:11dot pickle
  25150. 19:39:16IBizer
  25151. 19:39:31file.
  25152. 19:39:33file dot close.
  25153. 19:39:39Now I'm loading the logistic regression
  25154. 19:39:41model. So for that we have to write open
  25155. 19:39:57B
  25156. 19:39:59LG to pick
  25157. 19:40:04code
  25158. 19:40:07file
  25159. 19:40:09then file dot close
  25160. 19:40:13then
  25161. 19:40:18riser
  25162. 19:40:20LG.
  25163. 19:40:22Okay.
  25164. 19:40:27Yeah. So now we will predict the
  25165. 19:40:30sentiment. So for that I will write here
  25166. 19:40:33predict
  25167. 19:40:37riser
  25168. 19:40:47text.
  25169. 19:40:49Okay. Uh so here we will predict the
  25170. 19:40:52sentiment. So for that
  25171. 19:40:59text equals to
  25172. 19:41:02process
  25173. 19:41:11then demands for
  25174. 19:41:15sentiment in
  25175. 19:41:17text.
  25176. 19:41:21Then text
  25177. 19:41:24data
  25178. 19:41:30dot transform.
  25179. 19:41:37This is
  25180. 19:41:42okay. Then sentiment
  25181. 19:41:49model dot predict.
  25182. 19:41:57So here I will make a list of text with
  25183. 19:41:59sentiment. So for that I will write data
  25184. 19:42:01equals to empty array. Then for text
  25185. 19:42:07prediction
  25186. 19:42:09and zip
  25187. 19:42:12text x
  25188. 19:42:14sentiment
  25189. 19:42:18dt
  25190. 19:42:20append
  25191. 19:42:23text prediction.
  25192. 19:42:26Okay.
  25193. 19:42:29Then we will convert the list into pas
  25194. 19:42:31data frames. So for that I will write df
  25195. 19:42:33= to ad dot data plane
  25196. 19:42:41comma columns
  25197. 19:42:44person
  25198. 19:42:46next
  25199. 19:42:48comma
  25200. 19:42:50sentiment
  25201. 19:42:58then df equals to df dot
  25202. 19:43:02Replace
  25203. 19:43:07comma 1
  25204. 19:43:18positive
  25205. 19:43:24and here.
  25206. 19:43:30Okay.
  25207. 19:43:32So at last I will write here if
  25208. 19:43:41to
  25209. 19:43:47then here we will loading the model
  25210. 19:43:50vectorzer
  25211. 19:43:52comma ng plus load.
  25212. 19:44:00Here we text to classify like what
  25213. 19:44:03should be in the list. So like text
  25214. 19:44:08here I will like I love machine
  25215. 19:44:13name.
  25216. 19:44:21So
  25217. 19:44:33John
  25218. 19:44:37be so
  25219. 19:44:52So here df equals to date
  25220. 19:45:02the command text
  25221. 19:45:06then print
  25222. 19:45:08df.
  25223. 19:45:19I love machine learning. Positive. B is
  25224. 19:45:21so active. Positive. J I feel so good.
  25225. 19:45:24Negative. Okay. There is
  25226. 19:45:39and
  25227. 19:45:41yeah. See now it's coming. Okay.
  25228. 19:45:47This is how you can do the sentiment
  25229. 19:45:50analysis using uh RNN model. Here we
  25230. 19:45:53have loaded RNN model. So it is showing
  25231. 19:45:55right. So let's do them as poses
  25232. 19:46:04right here.
  25233. 19:46:09Add
  25234. 19:46:11one.
  25235. 19:46:25negative. Okay,
  25236. 19:46:29RNN model is working. So right, today we
  25237. 19:46:32are going to explore K nearest neighbors
  25238. 19:46:34or KN&N which is one of the most popular
  25239. 19:46:37algorithms in data science. Python is a
  25240. 19:46:40powerful tool for data science and KNN
  25241. 19:46:42is great for classifying data by
  25242. 19:46:44predicting the category of sample base
  25243. 19:46:46on its closest neighbor. This algorithms
  25244. 19:46:49is used in many field like healthcare,
  25245. 19:46:52finance and agriculture helping us make
  25246. 19:46:55decision based on data. The best part it
  25247. 19:46:58is really easy to use. You just need to
  25248. 19:47:00pick a number for K and choose a
  25249. 19:47:03distance function to compare data
  25250. 19:47:05points. However, KN&N has it downsides.
  25251. 19:47:08It doesn't work well with the large data
  25252. 19:47:10set and it require proper scaling of the
  25253. 19:47:12data to get accurate result. In this
  25254. 19:47:15video, we will show you how KNN work
  25255. 19:47:17with real data set, the Iris data set.
  25256. 19:47:19We'll walk you through simple Python
  25257. 19:47:21code and demonstrate how to find the
  25258. 19:47:23best K value to maximize your model's
  25259. 19:47:26accuracy. So stay tuned to seekn in
  25260. 19:47:28action. So welcome to the demo part. So
  25261. 19:47:30I'm here using Google Collab. So you can
  25262. 19:47:33use any of your favorite ID like Jupyter
  25263. 19:47:36notebook, Intelligi, Visual Code Studio,
  25264. 19:47:39anything. Okay. So let me rename this
  25265. 19:47:43file as KNN classification.
  25266. 19:47:51Okay, cool. So let me tell you that KN&N
  25267. 19:47:55can be used for the classification
  25268. 19:47:56regression predictive problems. So KN
  25269. 19:47:59falls in the supervised learning family
  25270. 19:48:00of algorithms. Okay, so we will measure
  25271. 19:48:03the distance between the K neighbors and
  25272. 19:48:06the first step will be we will choose
  25273. 19:48:08the number of K of neighbors. Then uh uh
  25274. 19:48:11we'll take the k nearest neighbors of
  25275. 19:48:14the new data point according to your
  25276. 19:48:15distance metric. And the step three will
  25277. 19:48:18be our among these case neighbors count
  25278. 19:48:21the number of data points of each
  25279. 19:48:23category. Okay. Step four we will assign
  25280. 19:48:25the new data points to the category
  25281. 19:48:27where you counted the most neighbors.
  25282. 19:48:29Okay. So let's start. First let's import
  25283. 19:48:32some library. import
  25284. 19:48:35numpy
  25285. 19:48:37as np and let's import
  25286. 19:48:42pandas sp. So everyone knows what is
  25287. 19:48:45numpy and the pandas. Okay, so numpy is
  25288. 19:48:49a library for the python programming
  25289. 19:48:51language adding support for uh you know
  25290. 19:48:53large multi- dimensional arrays. Okay.
  25291. 19:48:56Along with the large collection of uh
  25292. 19:48:58what to say uh highlevel mathematical
  25293. 19:49:01functions and various pandas is a
  25294. 19:49:03software library written for the Python
  25295. 19:49:04programming language for the data
  25296. 19:49:06manipulation and analysis uh all the
  25297. 19:49:09data frames and uh data structures it
  25298. 19:49:12offers for the manipulating numerical
  25299. 19:49:14tables. Okay. And the time series you
  25300. 19:49:16can see. So moving forward uh we'll
  25301. 19:49:18import our data set. So you can download
  25302. 19:49:21the data set from the description box
  25303. 19:49:23below. Okay. data set
  25304. 19:49:26equals to pb dot read
  25305. 19:49:30csv. The data name is iris dot csv.
  25306. 19:49:34Okay. This is how you read uh your data
  25307. 19:49:38in python. Okay. Yeah. Data set is
  25308. 19:49:42loaded. So data set dot shape
  25309. 19:49:47shape. Okay. Yeah. So data uh set dot
  25310. 19:49:52shape is used for how many numbers of
  25311. 19:49:57rows and columns present in your data
  25312. 19:49:59set. Okay, 150 rows and six columns. So
  25313. 19:50:01let me tell you brief about data set
  25314. 19:50:04this data set. So this data set include
  25315. 19:50:08three Iris species with uh 50 samples
  25316. 19:50:11each as well as some properties about
  25317. 19:50:14each flower. So one flower species is
  25318. 19:50:17linearly separable from the other two
  25319. 19:50:19you can say but the other two are not
  25320. 19:50:22linearly separable from each other.
  25321. 19:50:24Okay. And this shape I told you we can
  25322. 19:50:27get a quick idea of how many instances
  25323. 19:50:28of rows and columns are present in our
  25324. 19:50:30data set. So let's see our data set.
  25325. 19:50:34Data set dot
  25326. 19:50:36head. So head is used for uh you know by
  25327. 19:50:41you can see top five rows of your data
  25328. 19:50:44set using head and if you will use tail
  25329. 19:50:47you instead of head you can see the last
  25330. 19:50:50five rows of your data set. Cool. Okay.
  25331. 19:50:52So columns are ID sample length sample
  25332. 19:50:55width petal length petal width and the
  25333. 19:50:58species. Cool. Then moving forward let's
  25334. 19:51:01describe our data sets. So these are the
  25335. 19:51:03basics uh basic function. Okay.
  25336. 19:51:08of Python you can say.
  25337. 19:51:12So data set.escribed what describes do
  25338. 19:51:16is it will give you count of all the
  25339. 19:51:18rows mean value standard deviation value
  25340. 19:51:21minimum value what is the 25% okay of
  25341. 19:51:25all the values in the particular row.
  25342. 19:51:28What is the 50%? What is the 75%? What
  25343. 19:51:32is the maximum? Maximum is 150 you can
  25344. 19:51:34say. Okay. and 25% of 150 is 38.25 25
  25345. 19:51:37this okay of all the columns if it is
  25346. 19:51:40normal uh you know character so it won't
  25347. 19:51:42give you any data okay cool yeah so
  25348. 19:51:46moving forward uh let's now take a look
  25349. 19:51:49at the number of instances row belong to
  25350. 19:51:52each classes okay so we will write data
  25351. 19:51:56set dot
  25352. 19:51:59group by
  25353. 19:52:03species
  25354. 19:52:04dot size. Okay. Uh spec S is capital
  25355. 19:52:10that's why it's showing the error. Yeah.
  25356. 19:52:14So you can see Iris Satossa are 50, Iris
  25357. 19:52:17verical are 50 and virginica is 50.
  25358. 19:52:20Okay. And the data type type is integer.
  25359. 19:52:23Cool. So as you can see data set
  25360. 19:52:25contains six columns like ID, sample
  25361. 19:52:28length, sample width and petal length,
  25362. 19:52:30petal width and spacing. The actual
  25363. 19:52:32features are described by columns 1 to
  25364. 19:52:34four. the last columns labels or
  25365. 19:52:37samples. Okay. So firstly we need to
  25366. 19:52:39split data into two arrays like X
  25367. 19:52:41features and Y labels. So how we will do
  25368. 19:52:44this? By writing code like feature
  25369. 19:52:48columns equals to
  25370. 19:52:51sample length sample width petal length
  25371. 19:52:54petal width. Just remember you are
  25372. 19:52:56writing correct name. Okay. Then x = to
  25373. 19:53:01data set
  25374. 19:53:03feature
  25375. 19:53:06columns.
  25376. 19:53:09Okay. Dot
  25377. 19:53:12values.
  25378. 19:53:13Then y = to
  25379. 19:53:18data set
  25380. 19:53:20species dot values. Cool. Then let me
  25381. 19:53:25run it. So this is uh how we can split
  25382. 19:53:29the data set okay into two arrays X and
  25383. 19:53:31the Y. Okay. In X there are feature
  25384. 19:53:34columns. These four columns are there
  25385. 19:53:36and in Y species column is there. Okay.
  25386. 19:53:39And what is the species column this
  25387. 19:53:41Satossa venica and all this. Cool.
  25388. 19:53:45Then now we will do label encoding. So
  25389. 19:53:47as you can see labels are categorical
  25390. 19:53:49Kverse classifier does not accept string
  25391. 19:53:51labels. So we need to use label encoder
  25392. 19:53:55to transform them into numbers. Okay.
  25393. 19:53:57Then iris satossa correspond to zero.
  25394. 19:54:00Iris vericy color correspond to one and
  25395. 19:54:04I is virginica correspond to two. Okay.
  25396. 19:54:07012. Cool. So how we can write from
  25397. 19:54:13skarn dot pre-processing
  25398. 19:54:18import
  25399. 19:54:21label
  25400. 19:54:23encoder okay so what we'll do label
  25401. 19:54:26encoder transform them into numbers okay
  25402. 19:54:29so I will write
  25403. 19:54:31l equals to
  25404. 19:54:34label encoder
  25405. 19:54:36okay my bad label encoder. Okay. Then y
  25406. 19:54:41= to ele alate transform y. Okay. So now
  25407. 19:54:44I will run it. Yeah. Correct. So
  25408. 19:54:47splitting data set into training set and
  25409. 19:54:49the test set now. Okay. So now we'll
  25410. 19:54:51split data set into training set and
  25411. 19:54:53test set to check later on whether or
  25412. 19:54:56not a classifier work correctly or not.
  25413. 19:54:58Okay. So here I will write from skarn
  25414. 19:55:02dot not cross. I will write it here.
  25415. 19:55:06Skarn domodel selection
  25416. 19:55:12import
  25417. 19:55:14train test split. So what is train test
  25418. 19:55:17split? Train test split is a model
  25419. 19:55:19validation procedure that reveals how
  25420. 19:55:21your model performs on your new data.
  25421. 19:55:24Okay. And what is X train access Y train
  25422. 19:55:28Y test in Python. Okay. Let me first
  25423. 19:55:30write it and then I will let you know.
  25424. 19:55:32Okay. Then I will write here
  25425. 19:55:35x train comma x test comma y train comma
  25426. 19:55:43y test. Okay
  25427. 19:55:48train test
  25428. 19:55:51split
  25429. 19:55:52then x comma y
  25430. 19:55:56comma test size
  25431. 19:55:580.2 and the random is this. Okay. Okay.
  25432. 19:56:02Some error came. Okay. Underscore model
  25433. 19:56:05selection.
  25434. 19:56:08Yeah. Cool. So what is X train X test? Y
  25435. 19:56:12train Y test. Okay. So X train and Y
  25436. 19:56:15train sets are used for training and
  25437. 19:56:17fitting the model. Okay. So the X test
  25438. 19:56:20and the Y test are the set used for
  25439. 19:56:23testing the model and it's predicting
  25440. 19:56:26the right outputs level. Okay. So here
  25441. 19:56:30you can see test size is 0.2. into means
  25442. 19:56:3380% is for testing or sorry 80% is for
  25443. 19:56:37training and 20% is for testing for the
  25444. 19:56:41new data. Cool. Yeah. So now we will see
  25445. 19:56:44some uh let's do some data
  25446. 19:56:46visualization. Okay. So here I will
  25447. 19:56:48write import
  25448. 19:56:50mattplot lib
  25449. 19:56:53dotpipplot
  25450. 19:56:56as plt
  25451. 19:56:59then import
  25452. 19:57:01cb
  25453. 19:57:03as sns
  25454. 19:57:05then here I will write person mattplot
  25455. 19:57:08lib
  25456. 19:57:10in line. So what is mattplot lib?
  25457. 19:57:12Mattplot lib is a plotting library for
  25458. 19:57:14the python programming language and it's
  25459. 19:57:17numerical mathematic extension. Okay. So
  25460. 19:57:19numpy it provides an object API for
  25461. 19:57:22embedding plots into application using
  25462. 19:57:25generating purpose GUI toolkits like
  25463. 19:57:27kintter and python or gtk. Okay. Whereas
  25464. 19:57:32seabon seabon is a library for making
  25465. 19:57:34statical graph in python. It builds on
  25466. 19:57:37the top of mattplot lip and integrates
  25467. 19:57:39closely with pandas data structure.
  25468. 19:57:42Okay, seborn helps us to explore and
  25469. 19:57:45understand the data. Okay, so I will run
  25470. 19:57:49it here. I will write from
  25471. 19:57:54pandas dot plotting
  25472. 19:57:58import
  25473. 19:58:01parallel
  25474. 19:58:04coordinates.
  25475. 19:58:07Then plt dot figure
  25476. 19:58:11size should be
  25477. 19:58:1515, 10. Okay.
  25478. 19:58:19Then parallel coordinates.
  25479. 19:58:25Okay. Then data set dot drop. I don't
  25480. 19:58:30need id,
  25481. 19:58:33x is one.
  25482. 19:58:36Okay. then comma spaces
  25483. 19:58:41then plt dot title
  25484. 19:58:44and let parallel
  25485. 19:58:48coordinates
  25486. 19:58:50plot okay and
  25487. 19:58:53you can give some font size
  25488. 19:58:56equals to 20
  25489. 19:58:59then font
  25490. 19:59:02weight
  25491. 19:59:03equals to
  25492. 19:59:05okay Let's add bold only bold then plt
  25493. 19:59:10dot x label
  25494. 19:59:14then
  25495. 19:59:15features
  25496. 19:59:17comma
  25497. 19:59:20font size
  25498. 19:59:22to 15 then plt
  25499. 19:59:25dot y label then I'll write here
  25500. 19:59:29features
  25501. 19:59:31values
  25502. 19:59:33comma font size
  25503. 19:59:37equals to 15. Okay. then plt dot legend
  25504. 19:59:45then locals to 1 comma
  25505. 19:59:53I will write have frame on
  25506. 19:59:56equals to true comma shadow
  25507. 20:00:00equals to true comma face color
  25508. 20:00:06equals to White
  25509. 20:00:11T should be capital
  25510. 20:00:15and comma edge color
  25511. 20:00:19equals to
  25512. 20:00:25okay plt dot show okay some error is
  25513. 20:00:28there plig
  25514. 20:00:30size okay spelling mistake Take.
  25515. 20:00:37Okay. One more error. PLT. Legend. Okay.
  25516. 20:00:41Face color. Okay. Some spelling stick.
  25517. 20:00:45So yeah, let me make it output in full
  25518. 20:00:48screen. Yeah. So parallel coordinates is
  25519. 20:00:50a plotting technique for plotting you
  25520. 20:00:52know multivariate data. So it allows one
  25521. 20:00:56to see clusters in the data and to
  25522. 20:00:59estimate other stat visually. So using
  25523. 20:01:02parallel coordinate points uh you know
  25524. 20:01:04are represented as the connected line
  25525. 20:01:06segments as you can see. Okay. And each
  25526. 20:01:10vertical line represent one attribute
  25527. 20:01:12and one set of connected line segments
  25528. 20:01:15represent one data point. Okay. And
  25529. 20:01:18points that tend to cluster will appear
  25530. 20:01:21closer together. Okay. So this uh this
  25531. 20:01:25color is iris satossa and this is irisy
  25532. 20:01:29color and this ba one is iris virginica
  25533. 20:01:33as you can see in the legend. Okay,
  25534. 20:01:35cool. So moving forward let's create
  25535. 20:01:38another graph and curves. Okay so here I
  25536. 20:01:42will write from
  25537. 20:01:44p and dotplotting import scatter. Oh,
  25538. 20:01:48this uh plotting import
  25539. 20:01:53Andreo curves.
  25540. 20:01:56Okay, then plt
  25541. 20:01:58dot figure. It's AI based. So it's
  25542. 20:02:02giving me suggestions. Suggestions.
  25543. 20:02:04Suggestions. Okay. So sometimes
  25544. 20:02:05suggestions are good but not always.
  25545. 20:02:08Yeah. Let's carry on. Figure size is 15,
  25546. 20:02:1210.
  25547. 20:02:15Then Andrew
  25548. 20:02:20curves
  25549. 20:02:22then I will add data set dot drop id
  25550. 20:02:25access spaces and plt and curves plot.
  25551. 20:02:29Okay fine. Okay, let me add we don't
  25552. 20:02:34need X label and all. Let me add legend
  25553. 20:02:38plot
  25554. 20:02:40legend.
  25555. 20:02:41Then same LOC equals to 1. Then
  25556. 20:02:47proposition size.
  25557. 20:02:50Then I will write here size
  25558. 20:02:54is 15
  25559. 20:02:56of frame on equals to two comma shadow
  25560. 20:03:03equals to true. So this is uh truly
  25561. 20:03:07based upon you if you want to add legend
  25562. 20:03:10or not or you can skip. If you want to
  25563. 20:03:13skip you can skip. Okay. Face color
  25564. 20:03:16equals to white. Then
  25565. 20:03:20edge color equals to black. Okay. Then
  25566. 20:03:26plt dot show.
  25567. 20:03:29Yeah. So no error is there. So let me
  25568. 20:03:32first view output in full screen. Okay.
  25569. 20:03:36So Andrew curves these are the endoc
  25570. 20:03:39curves. Okay. You can see the graph in
  25571. 20:03:41the curve. So and curve allow one to
  25572. 20:03:44plot multivariate data as a large number
  25573. 20:03:47of curves. So that are created using the
  25574. 20:03:50attributes of samples okay as
  25575. 20:03:52coefficient for 4year series. Okay. So
  25576. 20:03:54by coloring these curves differently for
  25577. 20:03:57each class. It is possible to visualize
  25578. 20:04:00data clustering. So curves belongs
  25579. 20:04:03belonging to samples of the same class
  25580. 20:04:06will usually be closer together and they
  25581. 20:04:08form large structure. Okay. As you can
  25582. 20:04:11see here and this is the legend why we
  25583. 20:04:13are setting the four color is should be
  25584. 20:04:15white and yeah and the edge color should
  25585. 20:04:19be black okay and frame on shadow should
  25586. 20:04:22be there you can see the shadow okay
  25587. 20:04:24like this if you want to skip you can
  25588. 20:04:26skip this part legend part legend part
  25589. 20:04:28okay but like it's good to have it's
  25590. 20:04:33good practice to have this cool so let's
  25591. 20:04:36create one small pair plot okay so I
  25592. 20:04:39will write here plt dot
  25593. 20:04:42figure
  25594. 20:04:46then I will write SNS dot pair plot
  25595. 20:04:51then I will write data set
  25596. 20:04:53dot drop
  25597. 20:04:56we'll drop again id we don't need comma
  25598. 20:05:00xis is one then I will write u equals to
  25599. 20:05:07species
  25600. 20:05:09size equals to three.
  25601. 20:05:13Then markers equals to
  25602. 20:05:17O SD. Yeah. Cool. Then plt dot show.
  25603. 20:05:23It's running. Yeah. So first let me make
  25604. 20:05:27it to full screen. Yeah. So here pair
  25605. 20:05:30wise is useful when you want to
  25606. 20:05:32visualize the distribution of the
  25607. 20:05:34variable or the relationship between the
  25608. 20:05:36multiple variable separately within
  25609. 20:05:39subsets of your data set. Okay. So here
  25610. 20:05:42this blue one is stoa verol and the
  25611. 20:05:46virginica. Okay. So these are some uh
  25612. 20:05:49graph you can say or here you can see
  25613. 20:05:51sample length. Okay. So green one is
  25614. 20:05:55virginica and versol are almost having
  25615. 20:05:58same saple length. Okay. And here are
  25616. 20:06:01the different types of graphs. Okay. If
  25617. 20:06:04you don't want to uh let's say if you
  25618. 20:06:07don't know how to read this graph you
  25619. 20:06:10can use this graph or you can use this
  25620. 20:06:12graph. Either you can use this graph.
  25621. 20:06:14Okay. That's the power of pair pair wise
  25622. 20:06:17plots you can say. And the sample width
  25623. 20:06:20cm. Okay. Satossa. No. Okay, virginica
  25624. 20:06:24and verticola again having almost same
  25625. 20:06:28sample width. Okay, and here as well.
  25626. 20:06:32Okay, this stoa is in different form.
  25627. 20:06:35Okay. Yeah. So, we'll uh create one
  25628. 20:06:39small uh pair of u one small graph then
  25629. 20:06:42we'll move forward. Okay. Uh I will
  25630. 20:06:45write here plt dot. So now we are
  25631. 20:06:48creating box plot figure.
  25632. 20:06:54Okay. Then data set dot drop.
  25633. 20:06:58Then again ID commais should be one
  25634. 20:07:05dot box plot. Okay. Then figure size I
  25635. 20:07:09chose this. Yeah. Let's run it. Yeah.
  25636. 20:07:12Again see this is the box plot. Petal
  25637. 20:07:15length same petal width sap length sle
  25638. 20:07:17width okay according to width it is
  25639. 20:07:20showing like from this to this width we
  25640. 20:07:23have the veric color and from this to
  25641. 20:07:26this width this length sorry sample
  25642. 20:07:28length we have uh that virginica lies
  25643. 20:07:32here Iris virginica and from here to
  25644. 20:07:35here iristosa lies okay the sample cool
  25645. 20:07:39and like you can use 3D models you can
  25646. 20:07:43use different types of charts. Okay, I
  25647. 20:07:46did four and four are enough to read the
  25648. 20:07:49data set. Okay, so now we will do uh
  25649. 20:07:52some cannon classification. Okay, we
  25650. 20:07:54will make prediction and we will see the
  25651. 20:07:56accuracy of our data set. Okay, so now
  25652. 20:07:59what I will do? I will write here from
  25653. 20:08:04skarn
  25654. 20:08:07dot
  25655. 20:08:09neighbors
  25656. 20:08:11import k neighbors. Okay.
  25657. 20:08:15Write from skarn dot
  25658. 20:08:21neighbors
  25659. 20:08:23import
  25660. 20:08:26k neighbor
  25661. 20:08:29classifier.
  25662. 20:08:31Okay.
  25663. 20:08:33Then I will write here from
  25664. 20:08:36skarn
  25665. 20:08:37dot matrix
  25666. 20:08:40import
  25667. 20:08:43confusion
  25668. 20:08:46matrix
  25669. 20:08:48comma accuracy
  25670. 20:08:50accuracy score. Okay, then I will write
  25671. 20:08:53from
  25672. 20:08:55skarn dot model selection
  25673. 20:09:01import
  25674. 20:09:04cross
  25675. 20:09:05value score. Okay, so here what we I did
  25676. 20:09:10uh we are fitting the classifier to the
  25677. 20:09:12training set and loading the libraries
  25678. 20:09:14basically. Okay, then we'll initiate a
  25679. 20:09:18learning model K equals to three. Okay.
  25680. 20:09:21So here I will write classifier.
  25681. 20:09:24Let me give one space. Classifier equals
  25682. 20:09:26to
  25683. 20:09:30neighbors.
  25684. 20:09:31Classifier
  25685. 20:09:34and
  25686. 20:09:37neighbors
  25687. 20:09:39to three. Okay. Then fitting the model.
  25688. 20:09:42I will write here classifier
  25689. 20:09:45dot fit
  25690. 20:09:50x train,
  25691. 20:09:53y train.
  25692. 20:09:56Okay. Then uh now we will predicting the
  25693. 20:09:59test uh test set results. Okay. So y
  25694. 20:10:02prediction equals to
  25695. 20:10:06classifier dot predict
  25696. 20:10:11x test. Cool. Okay. From Okay. Spelling
  25697. 20:10:16mistake.
  25698. 20:10:19Okay.
  25699. 20:10:23Okay.
  25700. 20:10:28Yeah. I guess it's fine now. So now uh
  25701. 20:10:32let's evaluate the prediction. So here I
  25702. 20:10:35will uh build the confusion matrix.
  25703. 20:10:38Okay. So here I will write cm confusion
  25704. 20:10:40matrix equals to confusion
  25705. 20:10:44matrix
  25706. 20:10:46uh y test
  25707. 20:10:48y prediction. Okay. Then here I will
  25708. 20:10:51write cm. Okay. So what is confusion
  25709. 20:10:54matrix basically? So I confusion matrix
  25710. 20:10:56uh you know is a table that shows how
  25711. 20:10:59well a model performs by comparing its
  25712. 20:11:02prediction to the actual values. Okay.
  25713. 20:11:05So what it show is a confusion uh matric
  25714. 20:11:08display the number of correct and
  25715. 20:11:10incorrect prediction of each class in
  25716. 20:11:13models whose either it can give you true
  25717. 20:11:17positive true negative false positive
  25718. 20:11:19false negative either it can give you
  25719. 20:11:21zero or one. Okay. So yeah moving
  25720. 20:11:24forward let's calculate the model
  25721. 20:11:25accuracy. This is the main part. If the
  25722. 20:11:27model accuracy is low it means uh your
  25723. 20:11:31analysis of you know classification is
  25724. 20:11:34not good. Okay. So accuracy it should be
  25725. 20:11:38more than 80 at least. accuracy
  25726. 20:11:41equals to
  25727. 20:11:43accuracy
  25728. 20:11:45score
  25729. 20:11:48y test
  25730. 20:11:50comma y prediction into 100 otherwise it
  25731. 20:11:54will give me in points so I will write
  25732. 20:11:56print
  25733. 20:11:58model
  25734. 20:12:01canon model accuracy
  25735. 20:12:05is
  25736. 20:12:07oh I will write plus STR I will round
  25737. 20:12:10it. Okay. Accuracy
  25738. 20:12:20I will add the person. So why I wrote
  25739. 20:12:23this accuracy to round? So I don't want
  25740. 20:12:26after points I need only two numbers.
  25741. 20:12:28Okay. I don't want like 6 7 8 9 10 11 12
  25742. 20:12:30like this. Okay. I need only like 80.20
  25743. 20:12:34like this. Cool. Let's run this. see
  25744. 20:12:38KN&N model accuracy is 96.67 67 that's
  25745. 20:12:41why I wrote two here and the percent
  25746. 20:12:43should be there so 96 which is very very
  25747. 20:12:46very good okay so this is how you can
  25748. 20:12:50find uh the accuracy so now let's find
  25749. 20:12:53the optimal number of neighbors in K
  25750. 20:12:56okay basically finding the best K so we
  25751. 20:12:59will use uh using cross validation
  25752. 20:13:02parameter okay tuning so first I will
  25753. 20:13:05create the list of K for KN okay so here
  25754. 20:13:08I will write K
  25755. 20:13:10list
  25756. 20:13:12equals to list
  25757. 20:13:15range
  25758. 20:13:181 comma 50A 2.
  25759. 20:13:22Here I'm I will create the list of CV
  25760. 20:13:25score. Okay. So here I will write CV
  25761. 20:13:29scores.
  25762. 20:13:32Okay. Okay. I have to give brackets.
  25763. 20:13:34Yeah.
  25764. 20:13:36So we'll here perform the 10fold cross
  25765. 20:13:39validation. Okay, I will explain you
  25766. 20:13:41what is cross validation. Don't worry.
  25767. 20:13:43So first let me write for K N K listN
  25768. 20:13:52equals to
  25769. 20:13:55K neighbor classifier
  25770. 20:13:58and here I will add N neighbor
  25771. 20:14:04neighbors equals to K. Then scores
  25772. 20:14:08equals to cross
  25773. 20:14:12value score
  25774. 20:14:15KN&N then X train
  25775. 20:14:19Y train
  25776. 20:14:21okay
  25777. 20:14:23I will write here cross validation
  25778. 20:14:25equals to 10
  25779. 20:14:28comma scoring equals to I will write
  25780. 20:14:31here accuracy
  25781. 20:14:35Okay. And then here cv
  25782. 20:14:40course dotappend
  25783. 20:14:43to course dome.
  25784. 20:14:46Okay. So now yeah let me run it. Okay.
  25785. 20:14:52Comma some error came. Okay. The error
  25786. 20:14:56is the scoring parameter.
  25787. 20:15:00Yeah. Why? Because here accuracy you can
  25788. 20:15:02see and I did the spelling mistake.
  25789. 20:15:05Okay. So now what is cross validation?
  25790. 20:15:07Of course cross validation uh you know
  25791. 20:15:09determine the accuracy of your machine
  25792. 20:15:11learning model by partitioning the data
  25793. 20:15:14into two different groups. Okay called
  25794. 20:15:17training set and testing set. You can
  25795. 20:15:18see train test and the testing set. Okay
  25796. 20:15:21X and Y. So the data is randomly
  25797. 20:15:24separated into a certain number of
  25798. 20:15:25groups or subsets called folds. Okay you
  25799. 20:15:29can see the 10 folds we have wrote. Each
  25800. 20:15:32fold contains about the same amount of
  25801. 20:15:34okay and there is one more thing
  25802. 20:15:36validation. So validation is a technique
  25803. 20:15:39for assessing the accuracy of the model
  25804. 20:15:41on data set. Okay. And this cross
  25805. 20:15:43validation we did on new data set. Cool.
  25806. 20:15:46So now let's find the best K. Okay. So
  25807. 20:15:49here I will write best K equals to K
  25808. 20:15:56list. Okay. Before that I will write one
  25809. 20:15:59thing. I will write here MSE was
  25810. 20:16:03changing to mclassification error. Okay.
  25811. 20:16:06So equals to one
  25812. 20:16:09that's X for X in for X in
  25813. 20:16:16CV scores.
  25814. 20:16:18Okay. Okay, list I will write MSE dot
  25815. 20:16:24index
  25816. 20:16:26minimum
  25817. 20:16:28MSE.
  25818. 20:16:30I will use this bracket square bracket.
  25819. 20:16:33Okay. So then I will write print
  25820. 20:16:37the best
  25821. 20:16:38optimal
  25822. 20:16:40number of
  25823. 20:16:45neighbors
  25824. 20:16:46is person
  25825. 20:16:53best K. Okay, let's run it. See the best
  25826. 20:16:58optimal number of neighbors. Okay. N
  25827. 20:17:04neighbors is 9. Okay. So in the K
  25828. 20:17:08nearest neighbor KN algorithms. Okay. So
  25829. 20:17:11K represent the number of neighbors that
  25830. 20:17:13are considered when classifying a query
  25831. 20:17:15point. Okay. See the if we'll classify
  25832. 20:17:19this particular point you will get the
  25833. 20:17:22six. Okay. 1 2 3 4 5 6. Okay. And uh let
  25834. 20:17:26me show you. If you will classify this
  25835. 20:17:28portion only, so you will get 1 2 3 4 5
  25836. 20:17:326 points like this. Okay. So the best
  25837. 20:17:35optimum number of neighbor is nine.
  25838. 20:17:37>> What is Python? Python is a high-level
  25839. 20:17:41object-oriented programming language
  25840. 20:17:43developed by Guido Van Roum in 1989 and
  25841. 20:17:46was first released in 1991.
  25842. 20:17:49Python is often called a batteries
  25843. 20:17:51included language due to its
  25844. 20:17:54comprehensive standard library. A fun
  25845. 20:17:56fact about Python is that the name
  25846. 20:17:59Python was actually taken from the
  25847. 20:18:00popular BBC comedy show of that time
  25848. 20:18:03Montipython's Flying Circus. Now let's
  25849. 20:18:07look at the top features of Python
  25850. 20:18:12first. So Python has a simple structure
  25851. 20:18:14and a clearly defined syntax. This
  25852. 20:18:16allows the learners to pick up the
  25853. 20:18:18language quickly. So it is easy to learn
  25854. 20:18:20and use.
  25855. 20:18:22Python can run on different operating
  25856. 20:18:24systems such as Windows, Linux, and Mac,
  25857. 20:18:27making it a portable language. It
  25858. 20:18:29enables programmers to develop the
  25859. 20:18:30software for several competing platforms
  25860. 20:18:32by writing a program only once.
  25861. 20:18:36Third, Python is freely available at the
  25862. 20:18:38official website since it is open
  25863. 20:18:41source. This means that source code is
  25864. 20:18:43also available to the public.
  25865. 20:18:45Now, Python uses an object-oriented
  25866. 20:18:47approach that encapsulates code within
  25867. 20:18:49objects.
  25868. 20:18:51Python provides a collection of
  25869. 20:18:53libraries for various tasks such as
  25870. 20:18:54machine learning, web development, and
  25871. 20:18:56data analysis. And finally, in Python,
  25872. 20:19:00you don't need to assign the data type
  25873. 20:19:02of the variable. When you assign some
  25874. 20:19:04value to the variable, it automatically
  25875. 20:19:06allocates the memory to the variable at
  25876. 20:19:08runtime.
  25877. 20:19:10Now, with that, let's move on to the
  25878. 20:19:12uses of Python programming.
  25879. 20:19:15So, Python programming language is used
  25880. 20:19:17to develop desktop applications and
  25881. 20:19:19build web applications too. It is
  25882. 20:19:22popularly used in the field of data
  25883. 20:19:24science, machine learning and artificial
  25884. 20:19:26intelligence to analyze data, build
  25885. 20:19:28predictive models and make business
  25886. 20:19:30decisions. Python is also widely used in
  25887. 20:19:32game development. Now, let's see some of
  25888. 20:19:35the popular Python frameworks and
  25889. 20:19:37libraries.
  25890. 20:19:38Python can be used for web development
  25891. 20:19:40using frameworks like Zango, Flask,
  25892. 20:19:43Pyramid and Churi.
  25893. 20:19:45Now you can build graphical user
  25894. 20:19:47interfaces using libraries and
  25895. 20:19:48frameworks such as Tkinter or just KER.
  25896. 20:19:52You can also use PI GTK, PIQT or PYJS or
  25897. 20:19:57Python JavaScript.
  25898. 20:20:00Now, Python is also used to perform
  25899. 20:20:01machine learning tasks using libraries
  25900. 20:20:03such as TensorFlow, PyTorch,
  25901. 20:20:06Scikitlearn, Mattplot Lib, and Scypi.
  25902. 20:20:09You can also perform mathematical
  25903. 20:20:11computations using numpy and pandas.
  25904. 20:20:15Now, let's look at the best ids that you
  25905. 20:20:17can use to write programs in Python and
  25906. 20:20:19perform specific tasks. So, we have
  25907. 20:20:22Jupyter notebook, which is part of the
  25908. 20:20:24Anaconda distribution that is widely
  25909. 20:20:26used these days. Even for our demo in
  25910. 20:20:29this video, we'll be using Jupyter
  25911. 20:20:30Notebook. I'll show you in a while. Then
  25912. 20:20:32we have the visual code editor from
  25913. 20:20:34Microsoft. This is also one of the
  25914. 20:20:36preferred IDEs by learners and
  25915. 20:20:39companies. Then we also have the popular
  25916. 20:20:42text editor called Sublime Text editor.
  25917. 20:20:45Then we also have PyCharm followed by
  25918. 20:20:48Python and Spider as our top idees. Now
  25919. 20:20:52let's look at the top companies that are
  25920. 20:20:54using Python in our day-to-day work.
  25921. 20:20:58So we have Google, Kora, Facebook, even
  25922. 20:21:02Netflix, Spotify, and Instagram. Now
  25923. 20:21:04there are other top product- based,
  25924. 20:21:06service- based and startups that also
  25925. 20:21:08use Python programming. So what really
  25926. 20:21:11is Python programming language?
  25927. 20:21:14Python is an object-oriented highle
  25928. 20:21:16programming language that supports
  25929. 20:21:18built-in data structures and dynamic
  25930. 20:21:20semantics.
  25931. 20:21:21It supports multiple programming
  25932. 20:21:23paradigms such as structured,
  25933. 20:21:25object-oriented and functional
  25934. 20:21:26programming.
  25935. 20:21:28Python is often described as batteries
  25936. 20:21:30included language because it has a
  25937. 20:21:32comprehensive collection of standard
  25938. 20:21:34libraries. Python supports different
  25939. 20:21:36modules and packages which allows
  25940. 20:21:38program modularity and code reuse.
  25941. 20:21:43Python was developed by Guido Van Rosum
  25942. 20:21:45and its implementation started in
  25943. 20:21:47December 1989.
  25944. 20:21:50Python 1.0 version was released in the
  25945. 20:21:53year 1994. Python 2.0 came out in
  25946. 20:21:56October 2000 while Python 3.0 was
  25947. 20:21:59released in December 2008.
  25948. 20:22:03Now that you have got an understanding
  25949. 20:22:05of the Python programming language,
  25950. 20:22:07let's now look at the top 10 reasons why
  25951. 20:22:09you should learn Python.
  25952. 20:22:12So at number 10, we have ease of use.
  25953. 20:22:16One of the most common reasons to like
  25954. 20:22:18Python is that it is quite easy to learn
  25955. 20:22:20and code. It provides a simple syntax
  25956. 20:22:23that improves readability and makes it
  25957. 20:22:25easier to understand. So developers can
  25958. 20:22:27create any desktop or machine based
  25959. 20:22:29application using this language. Python
  25960. 20:22:31is very versatile and is instrumental in
  25961. 20:22:33artificial intelligence and machine
  25962. 20:22:35learning. We will talk about this later
  25963. 20:22:38in the session. Compared to Java or C++,
  25964. 20:22:41it has fewer lines of codes.
  25965. 20:22:45In the example here, we are printing a
  25966. 20:22:47hello world program in Java. As you can
  25967. 20:22:50see,
  25968. 20:22:52if you have to write a program in Java,
  25969. 20:22:54you first have to declare the class name
  25970. 20:22:57along with its scope.
  25971. 20:23:02Next, using curly braces, you need to
  25972. 20:23:05pass the main method along with its
  25973. 20:23:07arguments. And then using
  25974. 20:23:09system.out.print print len method you
  25975. 20:23:12can print hello world that's quite a
  25976. 20:23:14tedious task isn't it
  25977. 20:23:19the same task of printing hello world
  25978. 20:23:20can be done using just one line of code
  25979. 20:23:22in python as shown here you can write
  25980. 20:23:25the print function and pass whatever you
  25981. 20:23:28want to display inside the brackets and
  25982. 20:23:30that will print the output it is so
  25983. 20:23:32simple
  25984. 20:23:36that is why Python is considered as a
  25985. 20:23:37highle language and it's open source
  25986. 20:23:41You can just download it from the
  25987. 20:23:42website and start using it.
  25988. 20:23:49At nine, we have active community.
  25989. 20:23:53You need a community to learn new
  25990. 20:23:55technology and friends are your best
  25991. 20:23:57asset when it comes to learning a
  25992. 20:23:59programming language. Python has large
  25993. 20:24:02community support.
  25994. 20:24:06It has an extensive and active community
  25995. 20:24:08to assist engineers, developers,
  25996. 20:24:10analysts, and data scientists with
  25997. 20:24:12expert support in case of programming
  25998. 20:24:14errors or issues with the software. You
  25999. 20:24:17can just go ahead and put your queries
  26000. 20:24:19in the community forum. The community
  26001. 20:24:21members will address your queries in
  26002. 20:24:22real quick time. Communities like Stack
  26003. 20:24:25Overflow also brings many Python experts
  26004. 20:24:27together to help learners.
  26005. 20:24:30Python enhancement proposals or PEP is
  26006. 20:24:33where the proposals and the improvements
  26007. 20:24:35are announced. Also, there are a set of
  26008. 20:24:38recommendations or core values called
  26009. 20:24:40the Zen of Python written by Tim Peters
  26010. 20:24:43that represents the guiding principles
  26011. 20:24:45for Python development.
  26012. 20:24:50Up next at 8, we have portable and
  26013. 20:24:52extensible.
  26014. 20:24:53Multiple cross- language operations can
  26015. 20:24:55be performed effectively because Python
  26016. 20:24:58is portable and extensible in nature.
  26017. 20:25:01For example, if the users have a Python
  26018. 20:25:03code written on Windows and they want to
  26019. 20:25:06execute on a Mac operating system or
  26020. 20:25:08Linux operating system or Solaris, they
  26021. 20:25:10can easily do it without any amendment.
  26022. 20:25:13They can also run this code on any
  26023. 20:25:15platform flawlessly and without any
  26024. 20:25:17interrupt.
  26025. 20:25:20Due to its extensibility feature, you
  26026. 20:25:22can integrate other programming
  26027. 20:25:24languages such as Java,Net, C and C++
  26028. 20:25:27codes with Python. The components of
  26029. 20:25:29other programming languages can be used
  26030. 20:25:31with Python and thus it can be used to
  26031. 20:25:33make a crossplatform suitable
  26032. 20:25:35application too. So it is a really good
  26033. 20:25:38feature that Python provides.
  26034. 20:25:45The next reason to learn Python is
  26035. 20:25:47testing frameworks.
  26036. 20:25:51Python supports several built-in
  26037. 20:25:53flawless testing tools and frameworks
  26038. 20:25:55that help in debugging and speeding of
  26039. 20:25:57workflows.
  26040. 20:26:00Some of the tools and frameworks
  26041. 20:26:02supported by Python are Piest, Selenium
  26042. 20:26:04and Splinter. This is the reason for
  26043. 20:26:06which every tester tries to use Python
  26044. 20:26:09based tools and frameworks to test any
  26045. 20:26:11application or code or to validate it in
  26046. 20:26:13an easier manner. Piest is the most
  26047. 20:26:16recommended testing framework for
  26048. 20:26:17functional, integrational, and unit
  26049. 20:26:19testing. You can run Selenium test
  26050. 20:26:22scripts using Python programming
  26051. 20:26:24language to automate various tasks. And
  26052. 20:26:26Splinter is an open-source tool for
  26053. 20:26:28testing web applications using Python.
  26054. 20:26:31It lets you automate browser actions
  26055. 20:26:33such as visiting URLs and interacting
  26056. 20:26:35with their items.
  26057. 20:26:39At number six, we have libraries and
  26058. 20:26:41packages.
  26059. 20:26:44Another reason why Python has become so
  26060. 20:26:46popular in the industry these days is
  26061. 20:26:48that it has a massive collection of
  26062. 20:26:49libraries and packages that make your
  26063. 20:26:51task simple and easy. It has a range of
  26064. 20:26:54libraries, packages, frameworks, and
  26065. 20:26:56modules for data manipulation,
  26066. 20:26:58statistical calculation, web
  26067. 20:27:00development, machine learning, and data
  26068. 20:27:03science.
  26069. 20:27:08Python programmers have developed tons
  26070. 20:27:10of free and open-source libraries that
  26071. 20:27:11you can use. You can find many of them
  26072. 20:27:14via Python package index, the repository
  26073. 20:27:17of Python software. Python provides the
  26074. 20:27:19default package called pip. Anaconda is
  26075. 20:27:22a third party Python ecosystem. Other
  26076. 20:27:24examples include numpy, sci and zango.
  26077. 20:27:31Then we have scripting and automation.
  26078. 20:27:37Python is not just a programming
  26079. 20:27:38language. It can also be used for
  26080. 20:27:40writing scripts for automating tasks and
  26081. 20:27:42workflows without human intervention.
  26082. 20:27:47The code can be written in the form of
  26083. 20:27:49scripts and executed later. Further, it
  26084. 20:27:52is interpreted by the machine and
  26085. 20:27:54checked for errors at runtime. The
  26086. 20:27:56machine is used to read and interpret
  26087. 20:27:58the code. Once the developer checks the
  26088. 20:28:01code, it can further run or be used
  26089. 20:28:03several times without any interruption.
  26090. 20:28:06This allows you to automate a set of
  26091. 20:28:08certain tasks within a program or the
  26092. 20:28:10same code can be used with other
  26093. 20:28:12applications as well.
  26094. 20:28:16At number four, we have web development.
  26095. 20:28:22Another reason to learn Python is that
  26096. 20:28:24it makes the web development process so
  26097. 20:28:26much easier.
  26098. 20:28:30It provides a wide collection of
  26099. 20:28:31frameworks that make it easier for
  26100. 20:28:33developers to develop web applications.
  26101. 20:28:36Some of the examples are Zango, Flask,
  26102. 20:28:39Pyramid, Turbo Gears, CherryPie, etc.
  26103. 20:28:42These frameworks are written in Python
  26104. 20:28:44which makes the code a lot faster and
  26105. 20:28:46stable.
  26106. 20:28:48The task which used to take hours in PHP
  26107. 20:28:50can be finished in minutes using Python.
  26108. 20:28:53Python is also used for web scraping.
  26109. 20:28:56Django offers many elements of intricate
  26110. 20:28:59programs such as template design,
  26111. 20:29:01management panel, signing in, signing
  26112. 20:29:04up, signing out, URL routing, etc.
  26113. 20:29:10Once the user establishes the framework,
  26114. 20:29:12all these features become ready to use.
  26115. 20:29:16Flask is a microwave framework written
  26116. 20:29:18in Python.
  26117. 20:29:21of all the components that are part of
  26118. 20:29:23this module, they are all ready to
  26119. 20:29:25execute in the server context.
  26120. 20:29:27Pinterest and LinkedIn use Flask.
  26121. 20:29:30Pyramid offers more attributes than
  26122. 20:29:32Flask. It will assist users with URL
  26123. 20:29:35routing and authentication support.
  26124. 20:29:38Turbo Gears is a highly recommended and
  26125. 20:29:41scalable framework that supports
  26126. 20:29:42features such as authentication,
  26127. 20:29:44caching, identification, management of
  26128. 20:29:47sessions, and pluggable applications.
  26129. 20:29:54Up next at number three, we have machine
  26130. 20:29:57learning.
  26131. 20:29:59The growth of machine learning has been
  26132. 20:30:00phenomenal in the last 5 years and it's
  26133. 20:30:03rapidly changing the world around us.
  26134. 20:30:05Python is one of the most preferred
  26135. 20:30:07programming languages for machine
  26136. 20:30:08learning because of its simple syntax
  26137. 20:30:10and support for several machine learning
  26138. 20:30:12libraries.
  26139. 20:30:15Using different libraries and functions
  26140. 20:30:16in Python, the system can learn and
  26141. 20:30:19train itself from past data.
  26142. 20:30:23Once the system is trained, it can then
  26143. 20:30:25learn to adjust itself to new inputs.
  26144. 20:30:29Finally, it can make predictions and
  26145. 20:30:31perform humanlike tasks automatically.
  26146. 20:30:37At number two, we have data science.
  26147. 20:30:40Machine learning and data science go
  26148. 20:30:42hand in hand. Python is robust, scalable
  26149. 20:30:45and provides extensible visualization
  26150. 20:30:47and graphics options. Hence, it is
  26151. 20:30:49widely used in data science.
  26152. 20:30:52Python has libraries such as numpy for
  26153. 20:30:55numerical computation of data, pandas
  26154. 20:30:57for operations to manipulate data on
  26155. 20:30:59numerical tables and time series. It
  26156. 20:31:01also provides simply for symbolic
  26157. 20:31:04computation and sci for technical and
  26158. 20:31:06scientific computations.
  26159. 20:31:08It has another library called pyrain
  26160. 20:31:10which is sought for python based
  26161. 20:31:12reinforcement learning, artificial
  26162. 20:31:14intelligence and neural network library.
  26163. 20:31:16Scikitlearn is the machine learning
  26164. 20:31:17library for creating classification,
  26165. 20:31:19regression and clustering algorithms.
  26166. 20:31:21And finally, it provides PyTorch and
  26167. 20:31:24TensorFlow for deep learning.
  26168. 20:31:28Finally coming to the most important and
  26169. 20:31:30the top reason to learn Python which is
  26170. 20:31:32career opportunities and salary.
  26171. 20:31:36Python language provides a variety of
  26172. 20:31:38job opportunities and promises a high
  26173. 20:31:40growth graph with huge salary prospects.
  26174. 20:31:43It is been used by most of the tech
  26175. 20:31:45giants.
  26176. 20:31:47Industry leaders using Python are
  26177. 20:31:49Amazon, Google, Facebook, IBM, NASA,
  26178. 20:31:53Netflix and YouTube.
  26179. 20:31:57Next, you can see the Google trends
  26180. 20:32:01but I have considered three programming
  26181. 20:32:02languages Python, Java and C++. I have
  26182. 20:32:06compared them for the past 12 months.
  26183. 20:32:09You can see it clearly on your screens
  26184. 20:32:10that Python has become a frontr runner
  26185. 20:32:12in terms of popularity and web search
  26186. 20:32:14volume. It means people are interested
  26187. 20:32:17in Python. They want to learn it and use
  26188. 20:32:19it in their work. You can also check for
  26189. 20:32:22the YouTube search.
  26190. 20:32:25There also you will find that Python
  26191. 20:32:26programming language is the most
  26192. 20:32:28searched language on YouTube.
  26193. 20:32:32Now on your screens you can see the
  26194. 20:32:35report of PPL which is popularity of
  26195. 20:32:38programming language index. It is
  26196. 20:32:41created by analyzing how often language
  26197. 20:32:43tutorials are searched on Google. It is
  26198. 20:32:45a leading indicator. The raw data comes
  26199. 20:32:48from Google trends. The bar graph
  26200. 20:32:50depicts that Python is the most popular
  26201. 20:32:53and widely used programming language
  26202. 20:32:54across the globe followed by Java then
  26203. 20:32:57JavaScript and C.
  26204. 20:33:00The popularity of programming language
  26205. 20:33:02index can help you decide which language
  26206. 20:33:04to study or which one to use in a new
  26207. 20:33:07software project.
  26208. 20:33:12The next graph shows the popularity of
  26209. 20:33:14Python and Java over the years starting
  26210. 20:33:17from 2004 till the current period which
  26211. 20:33:20is 2020. Worldwide, Python is the most
  26212. 20:33:23popular language. Python grew the most
  26213. 20:33:26in the last 5 years by 19.4%. 4% and
  26214. 20:33:29Java lost the most by minus 7.2%.
  26215. 20:33:37Now let's talk about the different
  26216. 20:33:38career opportunities and the job roles
  26217. 20:33:40that you can get into if you learn
  26218. 20:33:42Python language.
  26219. 20:33:44First, you can become a Python developer
  26220. 20:33:47where you will be asked to write and
  26221. 20:33:49test codes, debug programs, and
  26222. 20:33:51integrate applications with third party
  26223. 20:33:53web services.
  26224. 20:33:55Second, you can become a web developer.
  26225. 20:33:58Here you will be responsible for writing
  26226. 20:34:01serverside web application logic. Python
  26227. 20:34:03web developers usually develop back-end
  26228. 20:34:05components, connect the application with
  26229. 20:34:08third party services and support the
  26230. 20:34:10front-end developers by integrating
  26231. 20:34:12their work with the Python application.
  26232. 20:34:16You can also become a data analyst if
  26233. 20:34:17you know Python. As a data analyst, you
  26234. 20:34:20have to gather data from multiple
  26235. 20:34:22sources using scripts. analyze that
  26236. 20:34:24data, develop and implement databases
  26237. 20:34:27and data collection systems.
  26238. 20:34:31You can become a data scientist. As a
  26239. 20:34:34data scientist, you need to understand
  26240. 20:34:36the challenges in business and come up
  26241. 20:34:38with the best solutions using modern
  26242. 20:34:40tools and techniques to analyze,
  26243. 20:34:41visualize, and build prediction models
  26244. 20:34:43to make business decisions.
  26245. 20:34:46Lastly, you can be a machine learning
  26246. 20:34:48engineer where you can develop
  26247. 20:34:50intelligent machines that can learn from
  26248. 20:34:52vast volumes of data and apply knowledge
  26249. 20:34:54without human intervention.
  26250. 20:34:56So there's a lot of scopes if you learn
  26251. 20:34:58Python. But before we move on, let's
  26252. 20:35:00understand first what is Jupyter
  26253. 20:35:02Notebook. So guys, as you can see all
  26254. 20:35:03over here that Jupyter Notebook is a
  26255. 20:35:06popular open-source tool that basically
  26256. 20:35:08allows you to create and share documents
  26257. 20:35:11which contains codes, equations, you can
  26258. 20:35:14have visualizations also. Basically, it
  26259. 20:35:17is used for data analysis, machine
  26260. 20:35:19learning and scientific research which
  26261. 20:35:21makes it a very essential tools for
  26262. 20:35:23developers like data scientists and
  26263. 20:35:25researchers alike. Now before installing
  26264. 20:35:28Jupyter notebook I request you that you
  26265. 20:35:30have Python installed in your system. So
  26266. 20:35:33the requirement should be Python 3.6 or
  26267. 20:35:36greater. So now let us officially
  26268. 20:35:38navigate to the Python's website. So
  26269. 20:35:40guys as you can see all over here. So on
  26270. 20:35:42python.org if I click on download
  26271. 20:35:45Python. So we're going to see that all
  26272. 20:35:47over here download Python 3.125. So as I
  26273. 20:35:51already told you that the requirement of
  26274. 20:35:53Python should be greater than 3.6. So
  26275. 20:35:55just you can click all over here and you
  26276. 20:35:58can see the download has started.
  26277. 20:36:03So guys as you can see all over here
  26278. 20:36:05that we have installed the Python. Now
  26279. 20:36:07let us open the file. So you can see the
  26280. 20:36:10given software is going to installed on
  26281. 20:36:11this directory. Okay. So just click all
  26282. 20:36:14over here. So guys as you can see all
  26283. 20:36:17over here the Python installation of
  26284. 20:36:203.125 is in progress. Let's wait for
  26285. 20:36:23some time till it gets installed.
  26286. 20:36:27So as you can see guys all over here
  26287. 20:36:29that we have successfully installed our
  26288. 20:36:31Python. Now let us open our terminal and
  26289. 20:36:33let us check whether Python is correctly
  26290. 20:36:35installed. So we are going to type
  26291. 20:36:38python
  26292. 20:36:40/ version.
  26293. 20:36:42So as you can see all over here we have
  26294. 20:36:44successfully installed our Python. So
  26295. 20:36:46guys that was our prerequisite. Now
  26296. 20:36:49there are two ways to install Jupyter
  26297. 20:36:52notebook. The first one can be pip.
  26298. 20:36:54Okay, pip is a package manager or using
  26299. 20:36:58Anocanda distribution. So let us see
  26300. 20:37:00with pip first. So guys, pip is a
  26301. 20:37:03package manager which is used to install
  26302. 20:37:04and manage software packages libraries
  26303. 20:37:07written in Python. So you can see all
  26304. 20:37:09over here that the Python with version
  26305. 20:37:11greater than 3.6 have default pip
  26306. 20:37:14installed in them. Okay. So we can use
  26307. 20:37:16pip command to install our Jupyter
  26308. 20:37:19notebook. So guys as you can see all
  26309. 20:37:21over here we have come to the official
  26310. 20:37:23documentation of jupitter.org and it is
  26311. 20:37:26saying that installing Jupyter lab with
  26312. 20:37:28pip command. So what you can do guys you
  26313. 20:37:30can just copy all over here. You can go
  26314. 20:37:32right all over here and click on this.
  26315. 20:37:36Now as you can see all over here it has
  26316. 20:37:38started downloading the Jupyter lab.
  26317. 20:37:57So guys, we are going to install our
  26318. 20:37:59Jupyter lab with the pip command. So
  26319. 20:38:02this is the official documentation of
  26320. 20:38:04Jupyter notebook. Okay? And just all you
  26321. 20:38:07have to do is copy this and type on your
  26322. 20:38:10terminal. So as you can see all over
  26323. 20:38:12here it has started downloading the
  26324. 20:38:14packages which is required to download
  26325. 20:38:16the Jupyter notebook. Let us wait for
  26326. 20:38:18some time.
  26327. 20:38:23Okay guys, so we have successfully
  26328. 20:38:25completed this step. Now let us move on
  26329. 20:38:27to our next step. So as you can see all
  26330. 20:38:30over here. So we have installed. Okay.
  26331. 20:38:34Then what we have to do then you can
  26332. 20:38:36type this. We can launch the Jupyter lab
  26333. 20:38:39with this command on the terminal. Now
  26334. 20:38:41let us wait. So as you can see all over
  26335. 20:38:44here guys, we have successfully
  26336. 20:38:46installed our Jupyter notebook. So you
  26337. 20:38:48can go all over here and just create a
  26338. 20:38:51new notebook and you can also choose
  26339. 20:38:53your kernel and you can start working on
  26340. 20:38:56your Jupyter notebook. Suppose I'll show
  26341. 20:38:58you one snippet. So 3 + 5. Let us try to
  26342. 20:39:02run this notebook. So as you can see it
  26343. 20:39:05is giving us the eight as answer. So it
  26344. 20:39:07is following the Python syntax and in
  26345. 20:39:09this way we have successfully installed
  26346. 20:39:12our Jupyter notebook using the pip
  26347. 20:39:14command. So now as you can also see all
  26348. 20:39:18over here you can also install Jupyter
  26349. 20:39:20notebook with this command pip install
  26350. 20:39:22notebook and then you can just open it.
  26351. 20:39:24This is also an another alternative.
  26352. 20:39:27Similarly, you can install with VA also
  26353. 20:39:29same command and just open the VA. Now,
  26354. 20:39:34if you are using any other operating
  26355. 20:39:35system like Mac OS or Linux, then you
  26356. 20:39:38can install by brew install Jupyter Lab.
  26357. 20:39:41So, home will be the package manager for
  26358. 20:39:44Mac OS and Linux. So, I hope so you are
  26359. 20:39:47pretty clear with how to install Jupyter
  26360. 20:39:49notebook with the pep command. Now, I
  26361. 20:39:51have downloaded Anacondas from this
  26362. 20:39:54official website. So as you can see all
  26363. 20:39:57over here this is the official website
  26364. 20:39:59of Anaconda. Okay. Now just type your
  26365. 20:40:02email and you can just download it. So
  26366. 20:40:04similarly as you can see after
  26367. 20:40:07installing I'm going to launch my
  26368. 20:40:09installer and let us click next. Okay.
  26369. 20:40:12Let us click agree. Okay. And let us
  26370. 20:40:16install this on the given directory.
  26371. 20:40:20Let us wait for some time till the
  26372. 20:40:21installation gets complete.
  26373. 20:40:26So guys as you can see all over here we
  26374. 20:40:29have completed our installation of
  26375. 20:40:30Anoconda. So just click on finish and
  26376. 20:40:35you can say we have successfully
  26377. 20:40:36installed our Anocanda. Now let us open
  26378. 20:40:39our Anaconda navigator. So just click
  26379. 20:40:43on.
  26380. 20:40:51So as you can see all over here just
  26381. 20:40:54right click on this and our Anocanda
  26382. 20:40:56navigator will be opened. So as you can
  26383. 20:40:58see all over here this is our Anocanda
  26384. 20:41:00navigator and it is loading the packages
  26385. 20:41:03and for us to install the Jupyter
  26386. 20:41:06notebook. So as you can see all over
  26387. 20:41:08here just click on launch. So guys if
  26388. 20:41:10you click on launch it is going to open
  26389. 20:41:12our Jupyter notebook. So as you can see
  26390. 20:41:14all over here it is saying launching the
  26391. 20:41:17Jupyter notebook and it is hosted on
  26392. 20:41:19localhost 8889. So this is our hosted
  26393. 20:41:23Jupyter notebook and in similarly you
  26394. 20:41:25can create a new notebook all over here
  26395. 20:41:27and in this way you can start working
  26396. 20:41:29>> LLMs. If you ever wondered how machine
  26397. 20:41:32learning can now understand and generate
  26398. 20:41:34humanlike text, you are in the right
  26399. 20:41:36place. From chatboards like Chat GPT to
  26400. 20:41:38AI assistant that powers search engines,
  26401. 20:41:40LLMs are transforming how we interact
  26402. 20:41:43with technology. One of the most
  26403. 20:41:44exciting advancement in this space is
  26404. 20:41:46Google's Gemini or OpenAI Charging large
  26405. 20:41:50language model designed to push the
  26406. 20:41:52boundaries of what AI can achieve. In
  26407. 20:41:54this video, we will explore what LLMs
  26408. 20:41:56are, how they work, and why models like
  26409. 20:41:58Geminy are critical for the future of
  26410. 20:42:01AI. Google Gemini is part of a new wave
  26411. 20:42:04of AI models that are smarter, faster,
  26412. 20:42:06and more efficient. It is designed to
  26413. 20:42:08understand context better, offer more
  26414. 20:42:11accurate responses and integrate deeply
  26415. 20:42:13into service like Google search and
  26416. 20:42:16Google Assistant, providing more
  26417. 20:42:17humanlike interactions. So we will break
  26418. 20:42:19down the science behind LLMs including
  26419. 20:42:21their massive training data set,
  26420. 20:42:23transformer architecture and how models
  26421. 20:42:25like Gemini use deep learning innovation
  26422. 20:42:28to change industries. Plus we will
  26423. 20:42:30compare Google Gemini to other popular
  26424. 20:42:32LMS such as OpenAI Chity models showing
  26425. 20:42:35how each of these technologies is used
  26426. 20:42:36to power chat bots, virtual assistants
  26427. 20:42:39and other AIdriven application. By end
  26428. 20:42:41of this video, you will have a clear
  26429. 20:42:42understanding of how large language
  26430. 20:42:44models like Gemini work, their key
  26431. 20:42:46features, and what they mean for their
  26432. 20:42:48future AI. Don't forget to like,
  26433. 20:42:50subscribe, and hit the bell icon to
  26434. 20:42:52never miss any update from Simply Learn.
  26435. 20:42:54So, what are the large language models?
  26436. 20:42:56Large language models like Chargen
  26437. 20:42:58pre-trained transformer 4 o and Google
  26438. 20:43:02Gemini are sophisticated AI system
  26439. 20:43:04designed to comprehend and generate
  26440. 20:43:06humanlike text. These models are built
  26441. 20:43:08using deep learning techniques and are
  26442. 20:43:09trained on vast data set collected from
  26443. 20:43:12the internet. They leverage self
  26444. 20:43:13attention mechanism to analyze
  26445. 20:43:15relationship between words or tokens
  26446. 20:43:17allowing them to capture context and
  26447. 20:43:19produce coherent relevant responses.
  26448. 20:43:21LLMs have significant application
  26449. 20:43:23including powering virtual assistant
  26450. 20:43:25chatboards, content creation, language
  26451. 20:43:27translation and supporting research and
  26452. 20:43:29decision making. Their ability to
  26453. 20:43:31generate fluent and contextually
  26454. 20:43:32appropriate text has advanced natural
  26455. 20:43:34language processing and improved human
  26456. 20:43:36computer interaction. So now let's see
  26457. 20:43:38what are large language model used for.
  26458. 20:43:40Large language models are utilized in
  26459. 20:43:42scenarios with limited or no domain
  26460. 20:43:45specific data available for training.
  26461. 20:43:46These scenarios include both few short
  26462. 20:43:48and zero short training approaches which
  26463. 20:43:50rely on the model's strong inductive
  26464. 20:43:52bias and its capability to derive
  26465. 20:43:54meaningful representation from a small
  26466. 20:43:56amount of data or even no data at all.
  26467. 20:43:59So now let's see how are large language
  26468. 20:44:01models trained. Large language models
  26469. 20:44:03typically undergo pre-training on a
  26470. 20:44:06board. All encompassing data set that
  26471. 20:44:08shares statical similarities with the
  26472. 20:44:10data set specific to the target task.
  26473. 20:44:12The objective of pre-training is to
  26474. 20:44:14enable the model to require highlevel
  26475. 20:44:16feature that can later be applied during
  26476. 20:44:18the finetuning phase for specific task.
  26477. 20:44:21So there are some training processes of
  26478. 20:44:23LLM which involves several steps. The
  26479. 20:44:25first one is text prep-processing. The
  26480. 20:44:27textual data is transformed into a
  26481. 20:44:28numerical representation that the LLM
  26482. 20:44:30model can effectively process. This
  26483. 20:44:32conversion may be involve techniques
  26484. 20:44:34like tokenization encoding and creating
  26485. 20:44:36input sequences. The second one is
  26486. 20:44:38random parameter initialization. The
  26487. 20:44:39model's parameter are initialized
  26488. 20:44:41randomly before the training process
  26489. 20:44:42begins. The third one is input numerical
  26490. 20:44:44data. The numerical representation of
  26491. 20:44:46the text data is fed into the model of
  26492. 20:44:48processing. The model's architecture
  26493. 20:44:50typically based on transformers allows
  26494. 20:44:52it to capture the conceptual
  26495. 20:44:53relationship between the words or tokens
  26496. 20:44:56in the next. The fourth one is loss
  26497. 20:44:58function calculation. A loss function
  26498. 20:44:59calculation measures the discrepancy
  26499. 20:45:01between the model's prediction and the
  26500. 20:45:02actual next word or token in a syntax.
  26501. 20:45:05The LLM model aims to minimize this loss
  26502. 20:45:08during training. The fifth one is
  26503. 20:45:09parameter optimization. The model's
  26504. 20:45:11parameter are registered through
  26505. 20:45:12optimization technique. This involves
  26506. 20:45:14calculating gradient and updating the
  26507. 20:45:16parameters accordingly gradually
  26508. 20:45:18improving the model's performance. The
  26509. 20:45:20last one is iterative training. The
  26510. 20:45:21training process is repeated over
  26511. 20:45:23multiple iteration or epox until the
  26512. 20:45:26model's output achieve a satisfactory
  26513. 20:45:28level of accuracy on that given task or
  26514. 20:45:30data set. By following this training
  26515. 20:45:32process, large language model learn to
  26516. 20:45:34capture linguistic patterns, understand
  26517. 20:45:35context and generate coherent responses
  26518. 20:45:38enabling them to excel at various
  26519. 20:45:39language related tasks. The next topic
  26520. 20:45:42is how do large language models work. So
  26521. 20:45:44large language models leverage deep
  26522. 20:45:46neural network to generate output based
  26523. 20:45:48on patterns learned from the training
  26524. 20:45:50data. Typically a large language model
  26525. 20:45:52adopts a transformer architecture which
  26526. 20:45:54enables the model to identify
  26527. 20:45:55relationship between words in a sentence
  26528. 20:45:58irrespective of their position in the
  26529. 20:45:59sequence. In contrast to RNAs that rely
  26530. 20:46:02on recurrence to capture token
  26531. 20:46:04relationship transformer neural network
  26532. 20:46:07employ self attention as their primary
  26533. 20:46:09mechanism. Self attention calculates
  26534. 20:46:10attention scores that determine the
  26535. 20:46:12importance of each token with respect to
  26536. 20:46:14the other token in the text sequence
  26537. 20:46:16facilitating the modeling of intricate
  26538. 20:46:18relationship within the data. Next,
  26539. 20:46:20let's see application of large language
  26540. 20:46:22models. Large language models have a
  26541. 20:46:24wide range of application across various
  26542. 20:46:26domains. So here are some notable
  26543. 20:46:27application. The first one is natural
  26544. 20:46:29language processing NLP. Large language
  26545. 20:46:31models are used to improve natural
  26546. 20:46:33language understanding tasks such as
  26547. 20:46:35sentiment analysis, named entity
  26548. 20:46:37recognition, text classification, and
  26549. 20:46:39language modeling. The second one is
  26550. 20:46:41chatbot and virtual assistant. Large
  26551. 20:46:43language models power conversational
  26552. 20:46:45agents, chatbots, and virtual assistant
  26553. 20:46:47providing more interactive and humanlike
  26554. 20:46:49user interaction. The third one is
  26555. 20:46:51machine translation. Large language
  26556. 20:46:53models have been used for automatic
  26557. 20:46:55language translation enabling text
  26558. 20:46:57translation between different languages
  26559. 20:46:59with improved accuracy. The fourth one
  26560. 20:47:01is sentiment analysis. LLMs can analyze
  26561. 20:47:03and classify the sentiment or emotion
  26562. 20:47:05expressed in a piece of text which is
  26563. 20:47:08valuable for market research, brand
  26564. 20:47:10monitoring and social media analysis.
  26565. 20:47:12The fifth one is content recommendation.
  26566. 20:47:14These models can be employed to provide
  26567. 20:47:15personalized content recommendations
  26568. 20:47:17enhancing user experience and engagement
  26569. 20:47:20on platforms such as news website or the
  26570. 20:47:22streaming services. So these application
  26571. 20:47:24highlight the potential impact of large
  26572. 20:47:25language models in various domains for
  26573. 20:47:27improving language understanding
  26574. 20:47:29automation. So hello guys welcome to
  26575. 20:47:31this demo part of this video. So here
  26576. 20:47:35what I will do I will go to new then
  26577. 20:47:37Python 3 file
  26578. 20:47:40then here
  26579. 20:47:42I will give it the name called
  26580. 20:47:46exploratory
  26581. 20:47:54data
  26582. 20:47:56analysis. Basically we will so we have
  26583. 20:48:00one data set file of
  26584. 20:48:04roller coaster basically. So we will be
  26585. 20:48:07using that and you can download that
  26586. 20:48:09file from the description box below from
  26587. 20:48:11the below link driving link. Okay. So we
  26588. 20:48:14will be doing some small basic functions
  26589. 20:48:17using Python and later on we will uh
  26590. 20:48:20make some good charts. Okay. We will
  26591. 20:48:22remove duplicates and all we will do all
  26592. 20:48:24that
  26593. 20:48:26thing. We'll do data preparation. We'll
  26594. 20:48:28do feature engineering. Okay. And uh
  26595. 20:48:32we'll remove the duplicates. We'll check
  26596. 20:48:35for the duplicates. We'll make charts
  26597. 20:48:38like histogram, KD blocks, box plot and
  26598. 20:48:43like many more similar to that like heat
  26599. 20:48:46map. We will make scatter plot group by
  26600. 20:48:48comparison. Okay, we all do that, right?
  26601. 20:48:52So just stick with me and you can write
  26602. 20:48:56side along with me here this code. Okay.
  26603. 20:49:00Okay. So let's start with importing
  26604. 20:49:05pandas first.
  26605. 20:49:07I guess everyone know what is pandas and
  26606. 20:49:09numpy
  26607. 20:49:11fine
  26608. 20:49:13as np then I will import
  26609. 20:49:17numpy
  26610. 20:49:21as np why this is np and pd sorry my bad
  26611. 20:49:26pd is here because I don't want to write
  26612. 20:49:29again and again this pandas this numpy
  26613. 20:49:32so basically I can write this small
  26614. 20:49:34version okay Yeah. So if you guys don't
  26615. 20:49:38know what is panda. So panda is very
  26616. 20:49:40popular library for working with data.
  26617. 20:49:43Okay. It's goal is to be the most
  26618. 20:49:46powerful and flexible opensource tool.
  26619. 20:49:48Okay. So it has reached that goal. So
  26620. 20:49:51data frames are the center of pandas.
  26621. 20:49:54And what is data frame? A data frame is
  26622. 20:49:56structured like a table or a
  26623. 20:49:57spreadsheet. Okay. The rows and the
  26624. 20:49:59columns. Okay. Whereas numpy numpy is an
  26625. 20:50:02open-source Python library again that
  26626. 20:50:05facilitates uh you know efficient
  26627. 20:50:07numerical operation on large quantities
  26628. 20:50:09of data. Okay. So there are many
  26629. 20:50:12functions in numpy as well those we can
  26630. 20:50:17use in the pandas data frame. Got it. So
  26631. 20:50:21we'll import one more
  26632. 20:50:25dot piplot
  26633. 20:50:28dotp. Okay. as
  26634. 20:50:32plt. Okay, there is one more library
  26635. 20:50:34mattplot lip for plotting the graph and
  26636. 20:50:38and there is one more import
  26637. 20:50:42seabbon
  26638. 20:50:44as SNS. So what is se? Sebon is again
  26639. 20:50:47the Python data visualization library
  26640. 20:50:49based on Matt plot lip. Okay, it
  26641. 20:50:51provides a highle interface for drawing
  26642. 20:50:53attractive and informative statical you
  26643. 20:50:56know graphs. Okay. Then I will write
  26644. 20:50:59here plt dot style
  26645. 20:51:03dot use.
  26646. 20:51:06We'll write ggplot.
  26647. 20:51:08Got it. ggplot.
  26648. 20:51:12Fine.
  26649. 20:51:15So yeah,
  26650. 20:51:17let me run it. So what is ggplot? So
  26651. 20:51:20ggplot is an again this is an opensource
  26652. 20:51:23data visualization package for
  26653. 20:51:25historical programming. Okay. So yeah uh
  26654. 20:51:29you can say or a general scheme for data
  26655. 20:51:32visualization which breaks up graph into
  26656. 20:51:34you know semantic components such as
  26657. 20:51:36scales and layers. Okay. ML lip py
  26658. 20:51:42pyab
  26659. 20:51:46p. Okay. Yeah. So now
  26660. 20:51:50we will import our data set df. DF means
  26661. 20:51:53data frame. You can write any word of
  26662. 20:51:56your choice. Then pd again pandas dot
  26663. 20:52:00read
  26664. 20:52:01csv used for readings CSV file. Okay.
  26665. 20:52:05Excel files. Got it? Then here I will
  26666. 20:52:09write my
  26667. 20:52:11this coaster
  26668. 20:52:14dot CSV. Okay. I'm not writing any path
  26669. 20:52:17because my this uh data set is here
  26670. 20:52:22itself. Coaster Coaster CO this is
  26671. 20:52:26coaster.csv CSV. Okay. If you have your
  26672. 20:52:30data set in another location as in like
  26673. 20:52:32in C drive, D drive or whatever. Okay.
  26674. 20:52:34You can give that path.
  26675. 20:52:37Okay. Let me run it. Yeah. So now we
  26676. 20:52:41will do some data understanding. Okay.
  26677. 20:52:44We'll see data frame shape head and tail
  26678. 20:52:47data types and describe like small small
  26679. 20:52:50function. We will use DF dot shape.
  26680. 20:52:55Right? So we have 1087 rows and 56
  26681. 20:52:59columns in our data set. Okay. Then we
  26682. 20:53:04will see df.hat
  26683. 20:53:06five.
  26684. 20:53:10So df do.head means it will give me top
  26685. 20:53:13five rows of my data set. Okay. You can
  26686. 20:53:15see 1 2 3 4 5. Okay. Five rows and 56
  26687. 20:53:19columns. Here 56 column but we have 1087
  26688. 20:53:23rows. Okay. So we have coaster name,
  26689. 20:53:26length, speed, location, status, opening
  26690. 20:53:28date, type, this, this, this, this.
  26691. 20:53:30Okay, we will do
  26692. 20:53:34uh you know we'll make some graphs using
  26693. 20:53:36this these columns. Okay, and there is
  26694. 20:53:39one more df.tail
  26695. 20:53:42again last five rows. So df.tail gives
  26696. 20:53:45you the last last five rows. See 1086
  26697. 20:53:481085. Okay. And if you want to see
  26698. 20:53:53full data, it is here. Okay. 0 to 108.
  26699. 20:53:59Fine. Yeah. So we have one more df doc
  26700. 20:54:04columns to check
  26701. 20:54:07all the columns.
  26702. 20:54:09See coaster name, land, the speed,
  26703. 20:54:11location, status, opening date, type and
  26704. 20:54:14these all are my columns names. Fine. So
  26705. 20:54:18we have one more data types. Actually we
  26706. 20:54:22have like many data types but let me
  26707. 20:54:24show you some important ones or you can
  26708. 20:54:27say some the basics on one. Okay. So DF
  26709. 20:54:31dot D types. Dypes means data types.
  26710. 20:54:33Coaster name is object type length
  26711. 20:54:35object and inversion is float. Okay.
  26712. 20:54:39We'll check inversion.
  26713. 20:54:43Where is inversion? Yeah float type.
  26714. 20:54:45Okay,
  26715. 20:54:47then everything is object and eer
  26716. 20:54:50introduces int numeric one and latitude
  26717. 20:54:54float. Okay, so basically we have three
  26718. 20:54:58types float, int and object.
  26719. 20:55:01Okay, then what we will do? Let's just
  26720. 20:55:05quickly check this describe.
  26721. 20:55:11Okay. So what is the count of this
  26722. 20:55:14inversion 932?
  26723. 20:55:17Okay. It won't include any you know what
  26724. 20:55:22is it empty cells. Okay. It won't count
  26725. 20:55:24empty cells. Right. Mean of this
  26726. 20:55:27particular table then standard deviation
  26727. 20:55:29then mean minimum value 25% 70% 75% and
  26728. 20:55:33max. It will describe you all this.
  26729. 20:55:36Okay. And there is one more df.info info
  26730. 20:55:40to get the info. See coaster name 1087
  26731. 20:55:45value non null object then length 953
  26732. 20:55:49null. Okay. Like this.
  26733. 20:55:52Yeah. So now moving forward what we will
  26734. 20:55:55do? We will do data preparation. Okay.
  26735. 20:55:58What comes in data preparation like
  26736. 20:56:00dropping irrelevant columns and rows
  26737. 20:56:02which we don't want. Okay. Then second
  26738. 20:56:05thing is identifying duplicates columns.
  26739. 20:56:07Then third is renaming columns.
  26740. 20:56:10Then we'll do some feature creations,
  26741. 20:56:13right?
  26742. 20:56:15So if you want to drop a column, okay,
  26743. 20:56:18how you can drop?
  26744. 20:56:20So you have to write just df dot drop.
  26745. 20:56:23First I will give here stag data
  26746. 20:56:28repage.
  26747. 20:56:30Okay. Yeah. So how you can drop a
  26748. 20:56:33column? Okay. DF.
  26749. 20:56:36Then here I will write
  26750. 20:56:40opening
  26751. 20:56:43date=
  26752. 20:56:47to 1.
  26753. 20:56:50Okay.
  26754. 20:56:52Opening. Okay. X is wrong.
  26755. 20:56:58Instead of this
  26756. 20:57:01maybe what is the opening date? Okay. O
  26757. 20:57:03is capital here.
  26758. 20:57:08Instead of this you can use double
  26759. 20:57:11equals to
  26760. 20:57:14okay axis is not defined.
  26761. 20:57:18My bad. So what you have to do? You have
  26762. 20:57:21to give this here and yeah.
  26763. 20:57:26Okay. Next is not defined.
  26764. 20:57:29Yeah. Fine. Okay. So as you can see here
  26765. 20:57:35first I will show you this question name
  26766. 20:57:38length speed location status and opening
  26767. 20:57:41date is there fine.
  26768. 20:57:44So if you will go here question name
  26769. 20:57:48length speed location status nothing
  26770. 20:57:50opening date is there right so this is
  26771. 20:57:53how you can drop a table fine for a
  26772. 20:57:56while I'm making this as a comment maybe
  26773. 20:58:00in future
  26774. 20:58:02upcoming you know making graph I'll
  26775. 20:58:04leave this okay so yeah so I will write
  26776. 20:58:08here df equals to
  26777. 20:58:11df
  26778. 20:58:14poster name,
  26779. 20:58:17comma, location
  26780. 20:58:21then comma
  26781. 20:58:24status.
  26782. 20:58:27Then here I will write manufacturer.
  26783. 20:58:30Fine. Then again comma.
  26784. 20:58:34Then I will give here.
  26785. 20:58:38Okay. Year
  26786. 20:58:42introduced
  26787. 20:58:45then
  26788. 20:58:46comma
  26789. 20:58:48latitude.
  26790. 20:58:50Wait I will tell you why I'm doing this.
  26791. 20:58:52Then longitude
  26792. 20:58:55latitude longitude
  26793. 20:58:58then
  26794. 20:59:03type main.
  26795. 20:59:07Okay. Then
  26796. 20:59:10opening
  26797. 20:59:13date
  26798. 20:59:15clean.
  26799. 20:59:19Fine.
  26800. 20:59:21Then speed into
  26801. 20:59:25m/ hour r
  26802. 20:59:29speed
  26803. 20:59:31and comma what else then
  26804. 20:59:38height
  26805. 20:59:40foot
  26806. 20:59:42inversion
  26807. 20:59:45clean
  26808. 20:59:49geforce
  26809. 20:59:52clean dot copy. Yeah.
  26810. 20:59:58So here I will write okay
  26811. 21:00:02type in
  26812. 21:00:05okay one more mistake is here
  26813. 21:00:10fine.
  26814. 21:00:12So now what I will do
  26815. 21:00:15I will write here opening
  26816. 21:00:22date
  26817. 21:00:24clean equals to PD2
  26818. 21:00:29dot date time
  26819. 21:00:32DF
  26820. 21:00:36opening
  26821. 21:00:38date
  26822. 21:00:40Clean fine.
  26823. 21:00:44What happened?
  26824. 21:00:50So what does this pd do to date time do?
  26825. 21:00:54Okay. So it converts argument to data
  26826. 21:00:58time date time. Okay. So this function
  26827. 21:01:01you know you can say converts a scalar
  26828. 21:01:04array like or series or data frame
  26829. 21:01:06dictionary like to a pandas datetime
  26830. 21:01:08object. Okay.
  26831. 21:01:12So now let's rename the columns.
  26832. 21:01:15Okay.
  26833. 21:01:17Then to you know for better things DF
  26834. 21:01:23equals to DF dot rename
  26835. 21:01:26columns equals to
  26836. 21:01:32poster
  26837. 21:01:34name. Then
  26838. 21:01:37poster
  26839. 21:01:39name,
  26840. 21:01:45year
  26841. 21:01:48introduced then
  26842. 21:01:51here
  26843. 21:01:54introduced.
  26844. 21:01:55Okay.
  26845. 21:01:58Opening date clean to
  26846. 21:02:05open
  26847. 21:02:08date.
  26848. 21:02:10Okay. Then I will write speed
  26849. 21:02:16I then write caps speed.
  26850. 21:02:26Got it? Then
  26851. 21:02:30height
  26852. 21:02:31into foot
  26853. 21:02:34I will write it as
  26854. 21:02:41catch 50.
  26855. 21:02:44So why I'm writing this because this is
  26856. 21:02:46a very good practice as a professional
  26857. 21:02:48way or as a data analyst or as a
  26858. 21:02:52business analyst whatever you are
  26859. 21:02:53working on machine learning projects or
  26860. 21:02:55whatever this is a good practice.
  26861. 21:02:59Okay. So inverions
  26862. 21:03:02in versions
  26863. 21:03:06clean
  26864. 21:03:08then should be like
  26865. 21:03:11invers
  26866. 21:03:25fine.
  26867. 21:03:26Let me run it.
  26868. 21:03:29Okay. Now let's check the column.
  26869. 21:03:33Yeah. So now you can see our column name
  26870. 21:03:36is changed.
  26871. 21:03:38Fine.
  26872. 21:03:41So I will write this na
  26873. 21:03:46dot sum.
  26874. 21:03:51So what does this is na dot sum do? So
  26875. 21:03:54is na function returns a boolean value
  26876. 21:03:56of you know true if the value is n and
  26877. 21:03:59the false otherwise and the sum function
  26878. 21:04:01returns the sum of the true values which
  26879. 21:04:04equals to the number of n values in the
  26880. 21:04:06column. So here we have zero and n here
  26881. 21:04:10same the status 213
  26882. 21:04:13here 217 5
  26883. 21:04:16height foot is 916 okay so now what we
  26884. 21:04:19will do we will write here df dot
  26885. 21:04:25location
  26886. 21:04:27df dot duplicate
  26887. 21:04:31I told you in starting
  26888. 21:04:34we'll
  26889. 21:04:35duplicate
  26890. 21:04:37Okay.
  26891. 21:04:39So, question name, location, status,
  26892. 21:04:41man. Okay. It's duplicated now. Okay.
  26893. 21:04:47So, now what we will do? I will check
  26894. 21:04:50duplicates for the coaster name. So, you
  26895. 21:04:52can write df. Loc
  26896. 21:04:56then df
  26897. 21:04:58df dot
  26898. 21:05:01duplicated and subset equals to
  26899. 21:05:06poster
  26900. 21:05:08name then dot head. Okay. Head of five.
  26901. 21:05:14Now everyone know right what do
  26902. 21:05:18what this head does.
  26903. 21:05:21Okay. Yeah. So here you can see so these
  26904. 21:05:26are duplicate okay of the question name
  26905. 21:05:30right.
  26906. 21:05:32Why? because you can just check the
  26907. 21:05:33typing
  26908. 21:05:35and all. Fine. So now checking with an
  26909. 21:05:39example duplicate. Let's check
  26910. 21:05:42dot query
  26911. 21:05:47question name
  26912. 21:05:53crystal.
  26913. 21:05:56Now you will get each equals cyclone.
  26914. 21:06:00Okay. I took this crystal B cipher.
  26915. 21:06:03Fine.
  26916. 21:06:05Then run it. So now you can see 39 and
  26917. 21:06:0943 are the same.
  26918. 21:06:12Okay. Everything is same. So this one is
  26919. 21:06:14duplicate. So now what I will do? I will
  26920. 21:06:17write here. Just let me give some space.
  26921. 21:06:19Yeah. DF dot columns.
  26922. 21:06:24Then I will write here DF
  26923. 21:06:29dot location.
  26924. 21:06:32Then TF dot
  26925. 21:06:35duplicated
  26926. 21:06:39subset
  26927. 21:06:41question
  26928. 21:06:43name
  26929. 21:06:46then
  26930. 21:06:48location
  26931. 21:06:52then
  26932. 21:06:56opening
  26933. 21:06:58date.
  26934. 21:07:00Okay.
  26935. 21:07:02dot
  26936. 21:07:04d set
  26937. 21:07:07index
  26938. 21:07:09then drop that.
  26939. 21:07:12Okay.
  26940. 21:07:18So now what we will do we will do some
  26941. 21:07:19feature understanding.
  26942. 21:07:21Okay.
  26943. 21:07:24So now we will do some feature
  26944. 21:07:26understanding.
  26945. 21:07:32Okay.
  26946. 21:07:34So in this we will plot some feature
  26947. 21:07:37distribution like histogram KD box plot.
  26948. 21:07:39Okay. For that I will write here DF
  26949. 21:07:45year introduced.
  26950. 21:07:49Okay. Then I will write value
  26951. 21:07:54count.
  26952. 21:07:56So now what I will do? Let's create bar
  26953. 21:08:01chart. Okay.
  26954. 21:08:04Ax goals to der
  26955. 21:08:10introduced dot
  26956. 21:08:13value
  26957. 21:08:16counts. Okay. Then dot
  26958. 21:08:21add 10 max. I need
  26959. 21:08:27then dot plot kind equals to I need bar
  26960. 21:08:33then title is top
  26961. 21:08:38I write what what what should we give
  26962. 21:08:40top 10
  26963. 21:08:43years
  26964. 21:08:45coasters introduced.
  26965. 21:08:49Okay, then
  26966. 21:08:53I will write here ex dot set
  26967. 21:08:59X label. Then I will write here
  26968. 21:09:06introduced.
  26969. 21:09:09Okay. Then I will write here ex dot
  26970. 21:09:14set Y label.
  26971. 21:09:18than
  26972. 21:09:20count. Okay, let me run it. Okay,
  26973. 21:09:24unexpected character after line
  26974. 21:09:27continuation.
  26975. 21:09:28Okay,
  26976. 21:09:33we can remove this
  26977. 21:09:36unexpected intended.
  26978. 21:09:39Now let me run it. Yeah, so this is our
  26979. 21:09:42bar plot. Okay. So this is how you can
  26980. 21:09:46create bar plot using your data. Okay.
  26981. 21:09:49So a bar plot is you know one of the
  26982. 21:09:53most common types of graphics in this
  26983. 21:09:55data visualization or EDA. It shows the
  26984. 21:09:58relationship between a numerical and the
  26985. 21:10:00categorical variable. Okay. So each
  26986. 21:10:03entity of this categoric variable is
  26987. 21:10:07represented as a bar.
  26988. 21:10:10Got it? So now let's do some more. So
  26989. 21:10:15what uh let's make a stogram class.
  26990. 21:10:19Okay,
  26991. 21:10:21histogram.
  26992. 21:10:23So for that I will write
  26993. 21:10:26a ex = to df
  26994. 21:10:31speed
  26995. 21:10:33r then plot
  26996. 21:10:37kind equals to
  26997. 21:10:44Then comma I will write bins here. Okay.
  26998. 21:10:48Bins equals to 20
  26999. 21:10:53then
  27000. 21:10:55title
  27001. 21:10:58coaster
  27002. 21:10:59speed.
  27003. 21:11:01Okay. meter per hour then AX
  27004. 21:11:06dot set
  27005. 21:11:09X level
  27006. 21:11:12speed
  27007. 21:11:15okay forgot to give this
  27008. 21:11:18okay yeah so histogram displays
  27009. 21:11:22numerical data by grouping data into
  27010. 21:11:24bins of equal width so here each bin is
  27011. 21:11:28plotted as a bar whose height correspond
  27012. 21:11:30to how How many data points are in that
  27013. 21:11:32bin? So bins you can say are also
  27014. 21:11:35sometimes called the intervals or
  27015. 21:11:36classes or buckets. Basically
  27016. 21:11:39it's the same.
  27017. 21:11:42So here for this this much is the bin.
  27018. 21:11:44For this this much is a bin. This is
  27019. 21:11:46like that. Okay.
  27020. 21:11:50So now let's create KD plot. So for that
  27021. 21:11:55it's very simple. Ax = to DF.
  27022. 21:12:00Then I can write as speed in me of R
  27023. 21:12:04then plot
  27024. 21:12:08kinda
  27025. 21:12:13then title
  27026. 21:12:16I can write the coaster
  27027. 21:12:21speed
  27028. 21:12:23then ex dot set
  27029. 21:12:27X
  27030. 21:12:30speed
  27031. 21:12:33can.
  27032. 21:12:35Yeah. So, KD, what does KD means? A
  27033. 21:12:39kernel density estimate. This plot is a
  27034. 21:12:42method of, you know, visualization the
  27035. 21:12:45distribution of observation in a data
  27036. 21:12:47set analog to a stola. So, KD represent
  27037. 21:12:51the data using a continuous probability
  27038. 21:12:53density curve in one or more dimension.
  27039. 21:12:57Okay. In one or more dimension that
  27040. 21:12:59curve fine. So now let's do some feature
  27041. 21:13:03relationship. So in this we will make
  27042. 21:13:05scatter plot, heat map correlation and
  27043. 21:13:08pair plot. Okay. Or we can also do some
  27044. 21:13:10group by comparison. Fine. So now let's
  27045. 21:13:13first make
  27046. 21:13:16scatter plot. Okay. So, df dot plot
  27047. 21:13:22kind equals to
  27048. 21:13:25scatter
  27049. 21:13:27comma
  27050. 21:13:28x = to
  27051. 21:13:31speed
  27052. 21:13:32m/ hour
  27053. 21:13:35comma y = to
  27054. 21:13:39height in foot.
  27055. 21:13:42Fine.
  27056. 21:13:44Then comma let's give the title
  27057. 21:13:48equals to coaster speed
  27058. 21:13:52versus
  27059. 21:13:54height.
  27060. 21:13:56Fine then pl do
  27061. 21:14:02okay some error is there
  27062. 21:14:05speed meter per hour. Okay. Okay. S
  27063. 21:14:08should be capital and H should be
  27064. 21:14:12capital.
  27065. 21:14:14Yeah.
  27066. 21:14:16So this is scatter plot speed versus
  27067. 21:14:19height. This is a speed versus height.
  27068. 21:14:22Okay. If the I can see here height is
  27069. 21:14:26directly proportional to speed somewhat
  27070. 21:14:29because if you can see the 350 is the
  27071. 21:14:31speed less than
  27072. 21:14:33120 km/h. No it's not like that. Okay
  27073. 21:14:36fine.
  27074. 21:14:38So a scatter plot identifies the
  27075. 21:14:41possible relationship between change
  27076. 21:14:43observed in two different sets of
  27077. 21:14:44variable. Here the variables are height
  27078. 21:14:47and the speed. Okay. It provides a
  27079. 21:14:49visual and aesthetical uh you know means
  27080. 21:14:51to test the strength of relationship
  27081. 21:14:53between two variable. Fine. Now let's
  27082. 21:14:59okay now let's make one more to give you
  27083. 21:15:02better idea. Okay with legend.
  27084. 21:15:06Now I will wait SNS dot scatter
  27085. 21:15:11plot then x = to
  27086. 21:15:15speed
  27087. 21:15:18r
  27088. 21:15:22y = to
  27089. 21:15:27height into foot
  27090. 21:15:32whatever you can say then hue will be
  27091. 21:15:34there the year
  27092. 21:15:38introduced.
  27093. 21:15:41Okay. Then data equals to DF.
  27094. 21:15:47Then ex= to set title.
  27095. 21:15:53Set title. Then
  27096. 21:15:56coaster
  27097. 21:15:58speed versus
  27098. 21:16:01height.
  27099. 21:16:03then plt dot show.
  27100. 21:16:07Yeah.
  27101. 21:16:08So now you can see
  27102. 21:16:11so this color you know the light color I
  27103. 21:16:14will let me zoom it. Okay. So this color
  27104. 21:16:17we have some here which are introduced
  27105. 21:16:20in '90s. Okay. And this color which are
  27106. 21:16:23introduced in 1925 and these color which
  27107. 21:16:26are introduced in 2000. Okay. So you can
  27108. 21:16:30see like this as well.
  27109. 21:16:32So now let's make pair plot. Okay, it's
  27110. 21:16:37look amazing.
  27111. 21:16:40So let me write SNS dot
  27112. 21:16:44dot pair plot then TF comma
  27113. 21:16:51variable is equals to I need
  27114. 21:16:56introduce
  27115. 21:16:57then speed
  27116. 21:17:01meter per R
  27117. 21:17:04comma S should be capital height
  27118. 21:17:09foot. Just remember the spelling, okay?
  27119. 21:17:13It's case sensitive.
  27120. 21:17:15Inversion
  27121. 21:17:18inversions, comma,
  27122. 21:17:21inversions, comma, geforce.
  27123. 21:17:26Okay. Comma. U I will put what? What?
  27124. 21:17:30What? What? Okay. I will put type
  27125. 21:17:36main.
  27126. 21:17:37Fine then pl do
  27127. 21:17:42okay let me run it okay some error is
  27128. 21:17:46there
  27129. 21:17:48year introduced
  27130. 21:17:51y capital
  27131. 21:17:55let's say here so that's why I'm saying
  27132. 21:17:58just remember the proper spelling again
  27133. 21:18:01some
  27134. 21:18:04error in versions is version.
  27135. 21:18:11Now let me run it.
  27136. 21:18:14Okay. Again some
  27137. 21:18:21what inversions?
  27138. 21:18:23Okay. Let me check here while renaming
  27139. 21:18:28inversions only.
  27140. 21:18:31Okay. Fine.
  27141. 21:18:33in versions.
  27142. 21:18:37Let me paste in but nothing changes. Let
  27143. 21:18:41me check again.
  27144. 21:18:44Okay, some error is there.
  27145. 21:18:48Wait. So yeah, you can see it's running
  27146. 21:18:52fine.
  27147. 21:18:54So after this what I will do?
  27148. 21:18:58So what is the pair plot? So pair plot
  27149. 21:19:00function allows the user to get you know
  27150. 21:19:02then an axis grid via which each
  27151. 21:19:05numerical variable is stored in the data
  27152. 21:19:06is shared across the x and the y okay in
  27153. 21:19:09the structured column.
  27154. 21:19:12So this is how see you know type mean
  27155. 21:19:14wood other and steel this red one is
  27156. 21:19:17wood other or the blue one and the
  27157. 21:19:20purple are the steel one okay the
  27158. 21:19:24different different format okay now
  27159. 21:19:26let's create the last graph which is
  27160. 21:19:31heat map okay so let me write df
  27161. 21:19:35correlation equals to df
  27162. 21:19:39here introduced,
  27163. 21:19:46comma,
  27164. 21:19:48speed m/a
  27165. 21:19:55height
  27166. 21:19:57into foot
  27167. 21:19:59canions,
  27168. 21:20:09comma
  27169. 21:20:10G force
  27170. 21:20:15then
  27171. 21:20:16drop then coalition okay DFO
  27172. 21:20:29introduced I
  27173. 21:20:36capital yeah
  27174. 21:20:39so So this is correlation values of all
  27175. 21:20:41the things right. So now let's write SNS
  27176. 21:20:45dot
  27177. 21:20:47heat map
  27178. 21:20:50heat map TF C
  27179. 21:20:55not will be true.
  27180. 21:20:59Okay.
  27181. 21:21:01Now let me run it. Yeah. So to create
  27182. 21:21:05heat map in Python so you can use this.
  27183. 21:21:07Okay. C bond library for this heat map.
  27184. 21:21:11So this function takes a data frame you
  27185. 21:21:13know as a input and generates a heat map
  27186. 21:21:16type of things as the output. Okay. So
  27187. 21:21:20this is how you can perform EDA using
  27188. 21:21:22any data set or you can show your data
  27189. 21:21:26or insights with a beautiful
  27190. 21:21:28representation using graphs and all like
  27191. 21:21:31this. Fine. Web scraping is a powerful
  27192. 21:21:34technique that allows you to
  27193. 21:21:36automatically extract data from website.
  27194. 21:21:38Turning the vast amount of information
  27195. 21:21:40available online into something you can
  27196. 21:21:43easily analyze and use. Whether you are
  27197. 21:21:45gathering data for research, building a
  27198. 21:21:47data set for machine learning project,
  27199. 21:21:48or just curious about how websites work
  27200. 21:21:51behind the scenes. Web scraping is an
  27201. 21:21:53essential skills to have in your
  27202. 21:21:55toolkit. On the other hand, Python is
  27203. 21:21:57one of the most popular programming
  27204. 21:21:58languages for web scraping thanks to its
  27205. 21:22:01simplicity and the wealth of libraries
  27206. 21:22:03available. In this video, we will
  27207. 21:22:05explore how to use Python to scrape data
  27208. 21:22:07from website. And we will dive into
  27209. 21:22:09practical examples using Python
  27210. 21:22:10libraries like request and beautiful
  27211. 21:22:13soup to fetch and parse web content. But
  27212. 21:22:16it's not just about the code. Web
  27213. 21:22:17scripping comes with its own set of
  27214. 21:22:20challenges and ethical constitution. We
  27215. 21:22:22will talk about how to scrape
  27216. 21:22:24responsibly, respecting the rules set by
  27217. 21:22:26websites and ensuring that your scraping
  27218. 21:22:28activity don't negatively impact the
  27219. 21:22:31sites you are collecting data from. So
  27220. 21:22:33by the end of this video, you will have
  27221. 21:22:35a solid understanding of how to start
  27222. 21:22:37scraping data from the web using Python.
  27223. 21:22:40Whether you are new to programming or
  27224. 21:22:42looking to add web scraping to your
  27225. 21:22:43skill set, this video will give you the
  27226. 21:22:45knowledge and tools you need to get
  27227. 21:22:47started. So let's jump in and see how
  27228. 21:22:50Python can help you unlock the full
  27229. 21:22:51potential of the web. So without any
  27230. 21:22:53further ado, let's get started. So here
  27231. 21:22:55I am using this Google Collab for the
  27232. 21:22:57web scraping. Okay. You can use your own
  27233. 21:23:01like Jupyter notebook, Visual Code
  27234. 21:23:03Studio, any thing. Okay. So here I'll
  27235. 21:23:06write scraping
  27236. 21:23:09using Python.
  27237. 21:23:12Okay. Then here first you have to
  27238. 21:23:15install some libraries like
  27239. 21:23:19you know uh request
  27240. 21:23:22and you have to install beautiful. So
  27241. 21:23:29you have to install that you know pandas
  27242. 21:23:33because we will create one data frame
  27243. 21:23:35and we will save it then we will check
  27244. 21:23:38our data. Okay. And you can install some
  27245. 21:23:42basic Python library like numpy and all
  27246. 21:23:44that. Okay. So here first I will import
  27247. 21:23:51request.
  27248. 21:23:53Okay. Then I will write from PS4
  27249. 21:23:59import
  27250. 21:24:02beautiful
  27251. 21:24:06soap.
  27252. 21:24:09Okay.
  27253. 21:24:11So
  27254. 21:24:12then I will write import
  27255. 21:24:15pandas
  27256. 21:24:17as pay.
  27257. 21:24:20Fine. Then now what I will do? Okay
  27258. 21:24:24let's see what is this request and all
  27259. 21:24:27the so request is an HTTP client library
  27260. 21:24:30for the Python programming language. So
  27261. 21:24:32request is one of the most you know
  27262. 21:24:34downloaded Python libraries. Okay. It's
  27263. 21:24:37like more like over 2,000 not exactly
  27264. 21:24:412,000 sorry 200 or 300 million monthly
  27265. 21:24:44download. Okay. So what it does it maps
  27266. 21:24:46the HTTP protocol onto Python subject
  27267. 21:24:49oriented semantics. And here beautiful
  27268. 21:24:52soap. So
  27269. 21:24:54SOAP is a Python package for you know
  27270. 21:24:57parsing the HTML and XML documents
  27271. 21:24:59including those with you know malphone
  27272. 21:25:02markup. It creates a parse tree for
  27273. 21:25:04documents that can be you know used to
  27274. 21:25:06extract data HTML from HTML which is
  27275. 21:25:10useful for web scraping. Let me run it
  27276. 21:25:13here. Our second step will be
  27277. 21:25:17define the URL and the headers. Okay,
  27278. 21:25:21URL and the headers.
  27279. 21:25:26So URL okay from where you want to
  27280. 21:25:29extract your data. So I will write here
  27281. 21:25:32simply learn
  27282. 21:25:35then this okay let's open this PMP
  27283. 21:25:37certification
  27284. 21:25:39it yeah so here I will write this
  27285. 21:25:44and headers
  27286. 21:25:48equals to
  27287. 21:25:53so these are the headers okay so like in
  27288. 21:25:57which you know device or you're working
  27289. 21:26:00on which browser you are working on. So
  27290. 21:26:04these are for the headers. So now what
  27291. 21:26:06what I will do I will send a get request
  27292. 21:26:10to the simple page. Okay to this URL for
  27293. 21:26:14that you know right let me write this
  27294. 21:26:17sending
  27295. 21:26:19get request.
  27296. 21:26:23Okay, here write response
  27297. 21:26:26equals to
  27298. 21:26:28request
  27299. 21:26:30dot get
  27300. 21:26:33then URL
  27301. 21:26:36comma headers
  27302. 21:26:38equals to headers.
  27303. 21:26:41Okay, let me run it. Okay, working fine.
  27304. 21:26:45So now uh let's check if the request is
  27305. 21:26:49successful or not. For that if
  27306. 21:26:53response
  27307. 21:26:55dot status code plus equals to = 200
  27308. 21:27:03then
  27309. 21:27:05so equals to
  27310. 21:27:09false.
  27311. 21:27:11Okay. Then response
  27312. 21:27:17dot content
  27313. 21:27:23do
  27314. 21:27:26HTML
  27315. 21:27:29dot parser. Okay.
  27316. 21:27:34Here.
  27317. 21:27:35So here I will write course
  27318. 21:27:39titles. Why? because here I'm you know
  27319. 21:27:43initializing the list to store our
  27320. 21:27:46titles or whatever the things are okay
  27321. 21:27:49basically the data okay why I'm writing
  27322. 21:27:52course here because this is again course
  27323. 21:27:54page that's why nothing else okay here I
  27324. 21:27:57will give the empty list
  27325. 21:28:00so now we will find all course titles
  27326. 21:28:04based on the actual HTML structure what
  27327. 21:28:06is HTML structure
  27328. 21:28:08you have to go to the page right click
  27329. 21:28:10Click then inspect.
  27330. 21:28:14Okay. According to this HTML structure
  27331. 21:28:17means what is the class name? What is
  27332. 21:28:20the you know this PMP certification is
  27333. 21:28:22which heading? H1 heading, H2 heading,
  27334. 21:28:25which heading it is. Okay. So let's see
  27335. 21:28:30which heading it is. Okay. Okay. This
  27336. 21:28:34PMP certification training H1 heading.
  27337. 21:28:37Okay. Just remember H1 heading.
  27338. 21:28:41So here I will write for
  27339. 21:28:45quotes in soap
  27340. 21:28:48dot find
  27341. 21:28:52or
  27342. 21:28:55h1.
  27343. 21:28:57Okay.
  27344. 21:29:00Then title
  27345. 21:29:05custom course text
  27346. 21:29:09dot strip.
  27347. 21:29:12Okay.
  27348. 21:29:14Then here I'm write course
  27349. 21:29:19titles
  27350. 21:29:20dot append
  27351. 21:29:24title. Okay.
  27352. 21:29:27Now what I will do? I will check
  27353. 21:29:31if any data was extracted before it not.
  27354. 21:29:35Okay, I will write if course
  27355. 21:29:40titles.
  27356. 21:29:43Here I will create a data frame. DF
  27357. 21:29:46equals to PD dot data frame.
  27358. 21:29:53In this I will write course title. it
  27359. 21:29:56will be our you know uh that column
  27360. 21:29:59name. So for I will write it course
  27361. 21:30:01data. Okay.
  27362. 21:30:04Then
  27363. 21:30:06course
  27364. 21:30:09titles.
  27365. 21:30:12Okay.
  27366. 21:30:14Then here I will write df
  27367. 21:30:18dot to
  27368. 21:30:22csv.
  27369. 21:30:24Then uh just create simply learn dot
  27370. 21:30:28CSV. Okay. So our data will be saved in
  27371. 21:30:31this simply learn dot csv. Okay. CSV.
  27372. 21:30:35Fine. Here what I will do I will write
  27373. 21:30:39index equals to false.
  27374. 21:30:44Print
  27375. 21:30:46data.
  27376. 21:30:48Print data is saved in CSV.
  27377. 21:30:53Simply learn dot CSV.
  27378. 21:30:57Fine.
  27379. 21:30:59So here I will write
  27380. 21:31:03else
  27381. 21:31:07print
  27382. 21:31:09web page.
  27383. 21:31:13Okay.
  27384. 21:31:15write field
  27385. 21:31:17to
  27386. 21:31:21retrieve the web page. Okay, that's it.
  27387. 21:31:27Let me run it.
  27388. 21:31:30Okay, here if response storage to this
  27389. 21:31:34group
  27390. 21:31:36spawn object has no attribute
  27391. 21:31:39status code. Okay.
  27392. 21:31:46Okay. Title.
  27393. 21:31:52It is text.
  27394. 21:31:54Some minor spelling mistakes are there.
  27395. 21:31:59Okay. Dear.
  27396. 21:32:03Sorry. Sorry. My bad
  27397. 21:32:09again.
  27398. 21:32:12Okay. Sorry.
  27399. 21:32:16D will be capital, F will be capital.
  27400. 21:32:20Yeah. So now you can see data is saved
  27401. 21:32:22in simply dot CSV. So what I will do? I
  27402. 21:32:25will write data equals to everyone know
  27403. 21:32:28how to read CSV in Python. Read
  27404. 21:32:33CSV.
  27405. 21:32:35What was our file name? Simply learn
  27406. 21:32:39CSV.
  27407. 21:32:41Okay. Let me copy paste. Run it. Link
  27408. 21:32:45fine data.
  27409. 21:32:49Okay.
  27410. 21:32:51PMP certification training data is fed.
  27411. 21:32:54Okay. Here you can see H1 we gave and
  27412. 21:32:59that's why PMP certification training
  27413. 21:33:00came. Let's take what's in H2.
  27414. 21:33:06Okay. In H2 there is leading premier PMI
  27415. 21:33:10partner something is there. Okay. I will
  27416. 21:33:12change here
  27417. 21:33:14H1 to H2 then I will run it. Okay
  27418. 21:33:20fine.
  27419. 21:33:22See leading premier P P P P P P P P P P
  27420. 21:33:24P P P P P P P P P P PMI. Okay, let me
  27421. 21:33:26make it bigger. Yeah. So now you can see
  27422. 21:33:30course data we mentioned. So
  27423. 21:33:33if you will see this H1 didn't come why
  27424. 21:33:37because we have mentioned H2 that's why.
  27425. 21:33:40So these all are the H2.
  27426. 21:33:44Okay. So what is this? I don't know.
  27427. 21:33:46Okay. This is a time series graph. This
  27428. 21:33:49is Google collaping. Okay. So this is
  27429. 21:33:52how you can literally retrieve your data
  27430. 21:33:56from any website. Okay. Just remember
  27431. 21:33:59that some websites don't give you access
  27432. 21:34:02to their you know for their scrapping
  27433. 21:34:05just like Amazon don't give so you have
  27434. 21:34:06to use at that time API and this is not
  27435. 21:34:10ethical also to use someone's data okay
  27436. 21:34:15without asking or whatever you can
  27437. 21:34:16>> welcome to the neural network tutorial
  27438. 21:34:19my name is Richard Kersner I'm with the
  27439. 21:34:21simply learn team what's in it for you
  27440. 21:34:24well today we're going to cover what is
  27441. 21:34:25a neural network what can neural neural
  27442. 21:34:28networks do, how does a neural network
  27443. 21:34:31work, types of neural networks, and then
  27444. 21:34:33we're going to jump into a use case to
  27445. 21:34:35classify between the photos of dogs and
  27446. 21:34:38cats, and we'll do that on the KAS with
  27447. 21:34:40the TensorFlow in the back, but it's a
  27448. 21:34:42Python script. So, that's always my
  27449. 21:34:44favorite part is when we dive into the
  27450. 21:34:45actual script. So, what is a neural
  27451. 21:34:48network? So, hi guys. I heard you want
  27452. 21:34:51to know what a neural network is. Here
  27453. 21:34:52we have uh looks like he just went
  27454. 21:34:54shopping at a red tag sale. My robots's
  27455. 21:34:56back. So as a matter of fact, you have
  27456. 21:34:58been using neural network on a daily
  27457. 21:35:01basis. In today's world, it's just
  27458. 21:35:03amazing how much we use our new
  27459. 21:35:04technology. We're not even aware of it.
  27460. 21:35:06When you ask your mobile assistant to
  27461. 21:35:08perform a search for you, you know, like
  27462. 21:35:10saying you're Google or Siri or whoever
  27463. 21:35:12you use, Amazon Web, self-driving cars.
  27464. 21:35:15So that's the newest thing coming out.
  27465. 21:35:17They're just now trying to make those
  27466. 21:35:18legal in different states in the US and
  27467. 21:35:20around the world. Even in the UK, they
  27468. 21:35:22now have self-driving cars going up and
  27469. 21:35:24down the street. It's pretty amazing.
  27470. 21:35:25These are all neural network driven.
  27471. 21:35:27Computer games use it. A lot of computer
  27472. 21:35:29games are driven by neural networks in
  27473. 21:35:31the back end as part of the game system
  27474. 21:35:34and how it adjusts to the players. And
  27475. 21:35:36it's also used in processing the map
  27476. 21:35:38images on your phone. So every time you
  27477. 21:35:39do a navigation someplace and it opens
  27478. 21:35:42it up, they now use neural networks to
  27479. 21:35:44help you find the quickest way to get
  27480. 21:35:45there. Neural network. A neural network
  27481. 21:35:48is a system or hardware that is designed
  27482. 21:35:51to operate like a human brain. In
  27483. 21:35:53today's development, this is so
  27484. 21:35:55important to understand because we don't
  27485. 21:35:57have anything else to compare it to. I'm
  27486. 21:35:59sure someday in the future, the computer
  27487. 21:36:01will redefine or the neural network or
  27488. 21:36:03the AI artificial intelligence will
  27489. 21:36:05redefine what these mean. But as far as
  27490. 21:36:08we can today's world, in today's
  27491. 21:36:10commercial development, we have to
  27492. 21:36:11compare it to what humans do. So, it's
  27493. 21:36:13we want to compare and how it operates
  27494. 21:36:15to a human brain and how it solves
  27495. 21:36:17problems like a human does. What can a
  27496. 21:36:19neural network do? And really, we're
  27497. 21:36:22just going to dive in deeper to we just
  27498. 21:36:24covered and look at other examples. So,
  27499. 21:36:26what can a neural network do? Well,
  27500. 21:36:28let's list out the things neural
  27501. 21:36:30networks can do for you. Translate text.
  27502. 21:36:32Boy, we got Google Translate and
  27503. 21:36:34Microsoft has their own translate. They
  27504. 21:36:37have some really cool. They actually
  27505. 21:36:38have an earpiece. It's supposed to start
  27506. 21:36:39translating as you talk. What a cool
  27507. 21:36:42technology. What a cool time to live.
  27508. 21:36:44Identify faces. Can you imagine all the
  27509. 21:36:46uses for facial identification? In the
  27510. 21:36:48case of our uh sample or our code that
  27511. 21:36:51we're going to look at later, we'll be
  27512. 21:36:52identifying dogs and cats. So, not quite
  27513. 21:36:54as detailed as uh understanding whose
  27514. 21:36:56face belongs to who. I I'm waiting for
  27515. 21:36:58the Google glasses to come out so I can
  27516. 21:37:00see who's who and the identify faces as
  27517. 21:37:02I'm walking around. Have a little name
  27518. 21:37:03tag over them. Not out there yet, but
  27519. 21:37:05boy, we are close. We can identify the
  27520. 21:37:07faces and they have all kinds of
  27521. 21:37:08technologies to bring that information
  27522. 21:37:10back to us. Recognize speech goes along
  27523. 21:37:13with the translate text. So now as
  27524. 21:37:16you're talking into your assistant, it
  27525. 21:37:18can use that to do commands, turn lights
  27526. 21:37:20on, all kinds of things you can do with
  27527. 21:37:22recognizing speech. Read handwritten
  27528. 21:37:24text. They're starting to translate all
  27529. 21:37:26these old text documents that they've
  27530. 21:37:28had in storage instead of doing it
  27531. 21:37:30individually where somebody's going
  27532. 21:37:31through each text by themsel in a room.
  27533. 21:37:34Picture like an old Raiders of the Lost
  27534. 21:37:36Arc theme where he's in the back, you
  27535. 21:37:37know, archaeologist studying the text.
  27536. 21:37:39Now it's fed into a computer. They take
  27537. 21:37:41a picture. They even use neural networks
  27538. 21:37:43to take a scroll that is so messed up
  27539. 21:37:46that they can't undo the scroll and they
  27540. 21:37:48x-ray it and then they use that x-ray to
  27541. 21:37:51translate the text off of it without
  27542. 21:37:53ever opening the scroll. I mean just way
  27543. 21:37:56cool stuff they're starting to do with
  27544. 21:37:57all this. And of course control robots.
  27545. 21:38:00What would be a neural network without
  27546. 21:38:01bringing in the robots? And we have our
  27547. 21:38:03own favorite robot in the middle who
  27548. 21:38:05goes to our red tag cell and goes
  27549. 21:38:06shopping for us. So, you know, these are
  27550. 21:38:08just a few of the wonderful things that
  27551. 21:38:10neural networks are being applied to.
  27552. 21:38:12It's such an infant stage technology.
  27553. 21:38:15What a wonderful time to jump in. And
  27554. 21:38:17there are a lot of other things it goes
  27555. 21:38:18into. I mean, we could spend just
  27556. 21:38:20forever talking about all the different
  27557. 21:38:21applications from business to whatever
  27558. 21:38:23you can even imagine. They're now
  27559. 21:38:25applying neural networks to help us
  27560. 21:38:26understand. So, now we talked a little
  27561. 21:38:29bit about all the cool things you can do
  27562. 21:38:30with a neural network. Let's dive in and
  27563. 21:38:32say, how does a neural network work? So
  27564. 21:38:35now we've come far enough to understand
  27565. 21:38:36how neural network works. Let's go ahead
  27566. 21:38:39and walk through this in a nice
  27567. 21:38:40graphical representation. They usually
  27568. 21:38:42describe a neural network as having
  27569. 21:38:44different layers. And you'll see that
  27570. 21:38:46we've identified a green layer, an
  27571. 21:38:48orange layer, and a red layer. The green
  27572. 21:38:50layer is the input. So you have your
  27573. 21:38:52data coming in. It picks up the input
  27574. 21:38:54signals and passes them to the next
  27575. 21:38:56layer. The next layer does all kinds of
  27576. 21:38:58calculations and feature extraction.
  27577. 21:39:00It's called the hidden layer. A lot of
  27578. 21:39:02times there's more than one hidden
  27579. 21:39:04layer. We're only showing one in this uh
  27580. 21:39:06picture, but we'll show you how it looks
  27581. 21:39:07like in a more detail in a little bit.
  27582. 21:39:10And then finally, we have an output
  27583. 21:39:11layer. This layer delivers the final
  27584. 21:39:14result. So the only two things we see is
  27585. 21:39:16the input layer and the output layer.
  27586. 21:39:18Now let's make use of this neural
  27587. 21:39:20network and see how it works. Wonder how
  27588. 21:39:22traffic cameras identify vehicles
  27589. 21:39:24registration plate on the road to detect
  27590. 21:39:26speeding vehicles and those breaking the
  27591. 21:39:29law? They got me going through a red
  27592. 21:39:30light the other day. Well, last month.
  27593. 21:39:32That's like the horrible thing. They
  27594. 21:39:33send you this picture of you and all
  27595. 21:39:35your information because they pulled it
  27596. 21:39:36up off of your license plate and your
  27597. 21:39:38picture. I shouldn't have gone through
  27598. 21:39:39the red light. So, here we are and we
  27599. 21:39:41have an image of a car and you can see
  27600. 21:39:43the license plate on there. So, let's
  27601. 21:39:45consider the image of this vehicle and
  27602. 21:39:46find out what's on the number plate. The
  27603. 21:39:49picture itself is 28x 28 pixels and the
  27604. 21:39:52image is fed as an input to identify the
  27605. 21:39:54registration plate. Each neuron has a
  27606. 21:39:57number called activation that represents
  27607. 21:39:59the grayscale value of the corresponding
  27608. 21:40:01pixel range. And we range it from zero
  27609. 21:40:04to one. One for a white pixel and zero
  27610. 21:40:06for a black pixel. And you can see down
  27611. 21:40:08here we have an example where one of the
  27612. 21:40:09pixels is registered as like 082.
  27613. 21:40:12Meaning it's probably pretty dark. Each
  27614. 21:40:14neuron is lit up when its activation is
  27615. 21:40:16close to one. So as we get closer to
  27616. 21:40:18black on white, we can really start
  27617. 21:40:20seeing the details in there. And you can
  27618. 21:40:22see again the pixel shows us one up
  27619. 21:40:24there. It's like part of the car and so
  27620. 21:40:26it lights up. So pixels in the form of
  27621. 21:40:28arrays are fed to the input layer. And
  27622. 21:40:30so we see here the pixels of a car image
  27623. 21:40:32fed as an input. And you're going to see
  27624. 21:40:33that the input layer which is green is
  27625. 21:40:36one dimension while our image is
  27626. 21:40:38two-dimension. Now when we look at our
  27627. 21:40:40setup that we're programming in Python,
  27628. 21:40:41it has a cool feature that automatically
  27629. 21:40:43does the work for us. If you're working
  27630. 21:40:45with an older neural network pattern
  27631. 21:40:47package, you then convert each one of
  27632. 21:40:49those rows so it's all one array. So
  27633. 21:40:51you'd have like row one and then just
  27634. 21:40:53tack row two onto the end. You can
  27635. 21:40:54almost feed the image directly into some
  27636. 21:40:56of these neural networks. The key is
  27637. 21:40:58though is that if you're using a 28x 28
  27638. 21:41:01and you get a picture of this 30x30,
  27639. 21:41:03shrink the 30x30 down to fit the 28x 28.
  27640. 21:41:06So you can't increase the number of
  27641. 21:41:09input in this case green dots. It's very
  27642. 21:41:11important to remember when you work on
  27643. 21:41:12neural networks. And let's name the
  27644. 21:41:14inputs x1 x2 x3 respectively. So each
  27645. 21:41:17one of those represents one of the
  27646. 21:41:18pixels coming in. And the input layer
  27647. 21:41:20passes it to the hidden layer. And you
  27648. 21:41:22can see here we now have two hidden
  27649. 21:41:23layers in this image in the orange. And
  27650. 21:41:26each one of those pixels connects to
  27651. 21:41:28each one of those hidden layers. And the
  27652. 21:41:31interconnections are assigned weights at
  27653. 21:41:33random. So they get these random weights
  27654. 21:41:35that come through. If x1 lights up, then
  27655. 21:41:38it's going to be x1 times this weight
  27656. 21:41:40going into the hidden layer. And we sum
  27657. 21:41:42those weights. The weights are
  27658. 21:41:43multiplied with the input signal and a
  27659. 21:41:45bias is added to all of them. So as you
  27660. 21:41:47can see here we have X1 comes in and it
  27661. 21:41:49actually goes to all the different
  27662. 21:41:51hidden layer nodes or in this case uh
  27663. 21:41:53whatever you want to call them network
  27664. 21:41:55setup the orange dots and so you take
  27665. 21:41:57the value of X1 you multiply it by the
  27666. 21:42:00weight for the next hidden layer. So X1
  27667. 21:42:03goes to hidden layer 1 X1 goes to hidden
  27668. 21:42:06layer two X1 goes hidden layer 1 node
  27669. 21:42:09two hidden layer one node three and so
  27670. 21:42:11on. And the bias a lot of times they
  27671. 21:42:14just put the bias in as like another
  27672. 21:42:16green dot or another orange dot and they
  27673. 21:42:18give the bias a value one and then all
  27674. 21:42:21the weights go in from the bias into the
  27675. 21:42:23next node. So the bias can change. We
  27676. 21:42:26always just remember that you need to
  27677. 21:42:27have that bias in there. There's things
  27678. 21:42:29that can be done with it. Generally most
  27679. 21:42:31of packages out there control that for
  27680. 21:42:33you so you don't have to worry about
  27681. 21:42:34figuring out what the bias is. But if
  27682. 21:42:37you ever dive deep into neural networks,
  27683. 21:42:38you got to remember there's a bias or
  27684. 21:42:40the answer won't come out correctly. The
  27685. 21:42:42weighted sum of the input is fed as an
  27686. 21:42:44input to the activation function to
  27687. 21:42:46decide which nodes to fire. And for
  27688. 21:42:48feature extraction, as a signal flows
  27689. 21:42:50within the hidden layers, the weighted
  27690. 21:42:52sum of inputs is calculated and is fed
  27691. 21:42:54to the activation function in each layer
  27692. 21:42:56to decide which nodes to fire. So here's
  27693. 21:42:58our feature extraction of the number
  27694. 21:43:00plate. And you can see these are still
  27695. 21:43:01hidden nodes in the middle. And this
  27696. 21:43:03becomes important. We're going to take a
  27697. 21:43:05little detour here and look at the
  27698. 21:43:06activation function. So, we're going to
  27699. 21:43:08dive just a little bit into the math so
  27700. 21:43:10you can start to understand where some
  27701. 21:43:12of the games go on when you're playing
  27702. 21:43:13with neural networks in your
  27703. 21:43:15programming. So, let's look at the
  27704. 21:43:16different activation functions before we
  27705. 21:43:18move ahead. Here's our friendly red tag
  27706. 21:43:20shopping robot. And so, one is a sigmoid
  27707. 21:43:23function. And the sigmoid function which
  27708. 21:43:25is 1 over 1 + e to the minus x takes the
  27709. 21:43:28x value and you can see where it
  27710. 21:43:31generates almost a zero and almost a one
  27711. 21:43:34with a very small area in the middle
  27712. 21:43:35where it crosses over and we can use
  27713. 21:43:37that value to feed into another
  27714. 21:43:40function. So if it's really uncertain it
  27715. 21:43:42might have a 0.1 or 2 or 3 but for the
  27716. 21:43:45most part it's going to be really close
  27717. 21:43:46to one and really close to this case
  27718. 21:43:48zero zero to one the threshold function.
  27719. 21:43:50So if you don't want to worry about the
  27720. 21:43:52uncertainty in the middle, you just say,
  27721. 21:43:54"Oh, if x is greater than or equal to
  27722. 21:43:56zero, if not, then uh x is zero." So
  27723. 21:43:58it's either zero or one. Really
  27724. 21:44:00straightforward. There's no in between
  27725. 21:44:02in the middle. And then you have the
  27726. 21:44:04what they call the reel relu function.
  27727. 21:44:06And you can see here where it puts out
  27728. 21:44:08the value, but then it says, well, if
  27729. 21:44:10it's over one, it's going to be one. And
  27730. 21:44:13if it's uh less than zero, it's zero. So
  27731. 21:44:15it kind of just deadends it on those two
  27732. 21:44:17ends, but allows all the values in the
  27733. 21:44:19middle. And again, this like the sigmoid
  27734. 21:44:21function allows that information to go
  27735. 21:44:23to the next level. So it might be
  27736. 21:44:24important to know if it's a 0.1 or a
  27737. 21:44:26minus.1. The next hidden layer might
  27738. 21:44:29pick that up and say, "Oh, this piece of
  27739. 21:44:31information is uncertain or this value
  27740. 21:44:33has a very low certainty to it." And
  27741. 21:44:35then the hyperbolic tangent function.
  27742. 21:44:37And you can see here it's a 1 - e to the
  27743. 21:44:40-2x over 1 + e - 2x. And it's very much
  27744. 21:44:44along the same theme, a little bit
  27745. 21:44:46different in here in that it goes
  27746. 21:44:47between minus one and one. So you'll see
  27747. 21:44:49some of these it goes 0ero to one, but
  27748. 21:44:51this one goes minus one to one. And if
  27749. 21:44:53it's less than zero, it's, you know, it
  27750. 21:44:55doesn't fire and if it's over zero, it
  27751. 21:44:57fires. And it also still puts out a
  27752. 21:44:59value. So you still have a value you can
  27753. 21:45:01get off of that just like you can with
  27754. 21:45:03the sigmoid function and the relu
  27755. 21:45:04function. Very similar in use. And I
  27756. 21:45:07believe the originally used to be
  27757. 21:45:08everything was done in the sigmoid
  27758. 21:45:10function. That was the most uh commonly
  27759. 21:45:11used. And now they just kind of use more
  27760. 21:45:13the reloo function. The reason is one,
  27761. 21:45:15it processes faster because you already
  27762. 21:45:17have the value and you don't have to add
  27763. 21:45:20another compute the 1 / 1 + e to the
  27764. 21:45:22minus x for each hidden node and the
  27765. 21:45:25data coming off works pretty good as far
  27766. 21:45:27as putting it into the next level. If
  27767. 21:45:29you want to know just how close it is to
  27768. 21:45:30zero, how close is it not to
  27769. 21:45:32functioning, you know, is it minus.1
  27770. 21:45:34minus.2 usually they're float values.
  27771. 21:45:36You get like minus point minus.00138
  27772. 21:45:39or something. So, you know, important
  27773. 21:45:41information, but the Reu is most
  27774. 21:45:42commonly used these days as far as the
  27775. 21:45:44setup we're using. But you'll also see
  27776. 21:45:46the sigmoid function very commonly used
  27777. 21:45:48also. Now that you know what an
  27778. 21:45:50activation function is, let's get back
  27779. 21:45:53to the neural network. So, finally, the
  27780. 21:45:55model would predict the outcome of
  27781. 21:45:56applying a suitable activation function
  27782. 21:45:58to the output layer. So, we go in here,
  27783. 21:46:00we look at this, and we have the optical
  27784. 21:46:02character recognition OCR is used on the
  27785. 21:46:04images to convert it into a text in
  27786. 21:46:06order to identify what's written on the
  27787. 21:46:08plate. And as it comes out, you'll see
  27788. 21:46:10the red node. And the red node might
  27789. 21:46:12actually represent just the letter A. So
  27790. 21:46:14there's usually a lot of outputs when
  27791. 21:46:16you're doing text identification. We're
  27792. 21:46:18not going to show that on here, but you
  27793. 21:46:19might have it even in the order. It
  27794. 21:46:21might be what order the license plates
  27795. 21:46:23in. So you might have ABCDE E FG, you
  27796. 21:46:26know, all the alphabet plus the numbers.
  27797. 21:46:28And you might have the 1 2 3 4 5 6 7 8 9
  27798. 21:46:3210 places. So it's a very large array
  27799. 21:46:34that comes out. It's not a small amount
  27800. 21:46:36of uh, you know, we show three dots
  27801. 21:46:38coming in, eight hidden layer nodes, you
  27802. 21:46:40know, two sets of four. We just show one
  27803. 21:46:42red coming out. A lot of times this is
  27804. 21:46:44uh, you know, 28 * 28. If you did 30 *
  27805. 21:46:4730, that's, you know, 900 nodes. So 28
  27806. 21:46:50is a little bit less than that uh, just
  27807. 21:46:52on the input. And so you can imagine the
  27808. 21:46:54hidden layer is just as big. Each hidden
  27809. 21:46:57layer is just as big if not bigger. Then
  27810. 21:46:58the output is going to be there's so
  27811. 21:47:00many digits. You know, it's a lot.
  27812. 21:47:01There's it's a huge amount of input and
  27813. 21:47:03output. But we're only showing you just,
  27814. 21:47:04you know, it' be hard to show in one
  27815. 21:47:06picture. And so it comes up and this is
  27816. 21:47:07what it finally gets out in the output
  27817. 21:47:09as it identifies a number on the plate.
  27818. 21:47:11And in this case, we have 08-d3858.
  27819. 21:47:16Error in the output is back propagated
  27820. 21:47:18through the network and weights are
  27821. 21:47:20adjusted to minimize the error rate.
  27822. 21:47:22This is calculated by a cost function.
  27823. 21:47:24When we're training our data, this is
  27824. 21:47:27what's used and we'll look at that in
  27825. 21:47:28the code when we do the data training.
  27826. 21:47:30So, we have stuff we know the answer to
  27827. 21:47:33and then we put the information through
  27828. 21:47:35and it says yes, that was correct or no,
  27829. 21:47:38cuz remember we randomly set all the
  27830. 21:47:40weights to begin with. And if it's
  27831. 21:47:41wrong, we take that error. How far off
  27832. 21:47:43are you? You know, are you off by is it
  27833. 21:47:45if it was like minus one, you're just a
  27834. 21:47:47little bit off. If it's like minus 300
  27835. 21:47:50was your output, remember when we're
  27836. 21:47:51looking at those different options, you
  27837. 21:47:53know, hyperbolic or whatever, and we're
  27838. 21:47:54looking at the could doesn't have an
  27839. 21:47:57limit on top or bottom. it actually just
  27840. 21:48:00generates a number. So if it's way off,
  27841. 21:48:02you have to adjust those weights a lot.
  27842. 21:48:04But if it's pretty close, you might
  27843. 21:48:05adjust the weights just a little bit.
  27844. 21:48:07And you keep adjusting the weights until
  27845. 21:48:09they fit all the different training
  27846. 21:48:10models you put in. So you might have 500
  27847. 21:48:13training models and those weights will
  27848. 21:48:15adjust using the back propagation. It
  27849. 21:48:17sends the error backward. The output is
  27850. 21:48:19compared with the original result and
  27851. 21:48:21multiple iterations are done to get the
  27852. 21:48:23maximum accuracy. So, not only does it
  27853. 21:48:25look at each one, but it goes through it
  27854. 21:48:27and just keeps cycling through these the
  27855. 21:48:29data making small changes in the network
  27856. 21:48:31until it gets the right answers. With
  27857. 21:48:33every iteration, the weights at every
  27858. 21:48:35interconnection are adjusted based on
  27859. 21:48:38the error. We're not going to dive into
  27860. 21:48:40that math because it is a differential
  27861. 21:48:41equation and it gets a little
  27862. 21:48:43complicated, but I will talk a little
  27863. 21:48:45bit about some of the different options
  27864. 21:48:46they have when we look at the code. So,
  27865. 21:48:49we've explored a neural network. Let's
  27866. 21:48:51look at the different types of
  27867. 21:48:52artificial neural networks. And this is
  27868. 21:48:55like the biggest area growing is how
  27869. 21:48:57these all come together. Let's see the
  27870. 21:48:59different types of neural network. And
  27871. 21:49:02again, we're comparing this to human
  27872. 21:49:03learning. So here's a human brain. I
  27873. 21:49:06feel sorry for that poor guy. So we have
  27874. 21:49:08a feed for forward neural network.
  27875. 21:49:11Simplest form of a they call it a ann a
  27876. 21:49:14neural network. Data travels only in one
  27877. 21:49:17direction input to output. This is what
  27878. 21:49:20we just looked at. So as the data comes
  27879. 21:49:22in, all the weights are added, it goes
  27880. 21:49:24to the hidden layer, all the weights are
  27881. 21:49:26added, it goes to the next hidden layer,
  27882. 21:49:27all the weights are added, and it goes
  27883. 21:49:29to the output. The only time you use the
  27884. 21:49:31reverse propagation is to train it. So
  27885. 21:49:34when you actually use it, it's very
  27886. 21:49:35fast. When you're training it, it takes
  27887. 21:49:37a while because it has to iterate
  27888. 21:49:38through all your training data. And you
  27889. 21:49:40start getting into big data because you
  27890. 21:49:42can train these with a huge amount of
  27891. 21:49:44data. The more data you put in, the
  27892. 21:49:45better trained they get. The
  27893. 21:49:47applications vision and speech
  27894. 21:49:49recognition actually they're pretty much
  27895. 21:49:51everything we talked about a lot of
  27896. 21:49:52almost all of them use this form of
  27897. 21:49:55neural network at some level radio basis
  27898. 21:49:57function neural network this model
  27899. 21:50:00classifies a data point based on its
  27900. 21:50:02distance from a center point. What that
  27901. 21:50:05means is that you might not have
  27902. 21:50:07training data. So you want to group
  27903. 21:50:09things together and you create central
  27904. 21:50:11points and it looks for all the things
  27905. 21:50:13you know some of these things are just
  27906. 21:50:14like the other. If you've ever watched
  27907. 21:50:16the Sesame Street as a kid, that dates
  27908. 21:50:18me. So, it brings things together and
  27909. 21:50:20this is a great way if you don't have
  27910. 21:50:21the right training model, you can start
  27911. 21:50:23finding things that are connected you
  27912. 21:50:24might not have noticed before.
  27913. 21:50:25Applications power restoration systems.
  27914. 21:50:29They try to figure out what's connected
  27915. 21:50:30and then based on that they can fix the
  27916. 21:50:33problem if you have a huge power system.
  27917. 21:50:35and self-organizing neural network
  27918. 21:50:38vectors of random dimensions are input
  27919. 21:50:40to discrete map comprised of neurons. So
  27920. 21:50:43they basically find a way to draw they
  27921. 21:50:46call them they say dimensions or vectors
  27922. 21:50:48or planes because they actually chop the
  27923. 21:50:51data in one dimension, two dimension,
  27924. 21:50:53three dimension, four, five, six. They
  27925. 21:50:55keep adding dimensions and finding ways
  27926. 21:50:56to separate the data and connect
  27927. 21:50:58different data pieces together.
  27928. 21:51:00Applications used to recognize patterns
  27929. 21:51:02in data like in medical analysis. The
  27930. 21:51:04hidden layer saves its output to be used
  27931. 21:51:07for future prediction. Recurrent neural
  27932. 21:51:09networks. So the hidden layers remember
  27933. 21:51:11its output from last time and that
  27934. 21:51:13becomes part of its new input. Uh you
  27935. 21:51:16might use that especially in robotics or
  27936. 21:51:18flying a drone. You want to know what
  27937. 21:51:19your last change was and how fast it was
  27938. 21:51:22going to help predict what your next
  27939. 21:51:24change you need to make is to get to
  27940. 21:51:25where the drone wants to go.
  27941. 21:51:27Applications text to speech conversation
  27942. 21:51:30model. So, you know, I talked about
  27943. 21:51:31drones, but you know, just identifying
  27944. 21:51:33on Lexus or Google Assistant or any of
  27945. 21:51:36these, they're starting to add in I'd
  27946. 21:51:38like to play a song on my Pandora, and
  27947. 21:51:41I'd like it to be at volume 90%. So, you
  27948. 21:51:44now can add different things in there,
  27949. 21:51:45and it connects them together. The input
  27950. 21:51:47features are taken in batches like a
  27951. 21:51:50filter. This allows a network to
  27952. 21:51:51remember an image in parts. Convolution
  27953. 21:51:54neural network. today's world in photo
  27954. 21:51:57identification and taking apart photos
  27955. 21:51:59and trying to you know have you ever
  27956. 21:52:00seen that on Google where you have five
  27957. 21:52:02people together this is the kind of
  27958. 21:52:04thing separates all those people so then
  27959. 21:52:06it can do a face recognition on each
  27960. 21:52:08person applications used in signal and
  27961. 21:52:10image processing in this case I use
  27962. 21:52:12facial images or Google picture images
  27963. 21:52:14as one of the options modular neural
  27964. 21:52:17network it has a collection of different
  27965. 21:52:20neural networks working together to get
  27966. 21:52:22the output so wow we just went through
  27967. 21:52:24all these different types of neural
  27968. 21:52:26networks. And the final one is to put
  27969. 21:52:28multiple neural networks together. I
  27970. 21:52:30mentioned that a little bit when we
  27971. 21:52:32separated people in a larger photo and
  27972. 21:52:34individuals in the photo and then do the
  27973. 21:52:36facial recognition on each person. So
  27974. 21:52:38one network is used to separate them and
  27975. 21:52:40the next network is then used to figure
  27976. 21:52:43out who they are and do the facial
  27977. 21:52:44recognition. Applications still
  27978. 21:52:46undergoing research. This is a cutting
  27979. 21:52:48edge. you hear the term pipeline and
  27980. 21:52:51there's actual in Python code and in
  27981. 21:52:53almost all the different neural network
  27982. 21:52:55setups out there they now have a
  27983. 21:52:57pipeline feature usually and it just
  27984. 21:52:59means you take the data from one neural
  27985. 21:53:02network and maybe another neural network
  27986. 21:53:04or you put it into the next neural
  27987. 21:53:05network and then you take three or four
  27988. 21:53:07other neural networks and feed them into
  27989. 21:53:09another one. So how we connect the
  27990. 21:53:11neural networks is really just cutting
  27991. 21:53:13edge and it's so experimental. I mean
  27992. 21:53:15it's almost creative in its nature.
  27993. 21:53:17There's not really a science to it
  27994. 21:53:19because each specific domain has
  27995. 21:53:22different things it's looking at. So if
  27996. 21:53:23you're in the banking domain, it's going
  27997. 21:53:25to be different than the medical domain
  27998. 21:53:27than the automatic car domain. And
  27999. 21:53:29suddenly figuring out how those all fit
  28000. 21:53:31together is just a lot of fun and really
  28001. 21:53:33cool. So we have our types of artificial
  28002. 21:53:35neural network. We have our feed forward
  28003. 21:53:37neural network. We have a radial basis
  28004. 21:53:39function neural network. We have our
  28005. 21:53:40Cohen self-organizing neural network,
  28006. 21:53:43recurrent neural network, convolution
  28007. 21:53:46neural network, and modular neural
  28008. 21:53:48network where it brings them all
  28009. 21:53:49together. And u no the colors on the
  28010. 21:53:52brain do not match what your brain
  28011. 21:53:53actually does, but they do bring it out
  28012. 21:53:56that most of these were developed by
  28013. 21:53:58understanding how humans learn. And as
  28014. 21:54:00we understand more and more of how
  28015. 21:54:02humans learn, we can build something in
  28016. 21:54:05the computer industry to mimic that, to
  28017. 21:54:07reflect that. And that's how these were
  28018. 21:54:09developed. So exciting part, use case
  28019. 21:54:12problem statement. So this is where we
  28020. 21:54:13jump in. This is my favorite part. Let's
  28021. 21:54:15use the system to identify between a cat
  28022. 21:54:18and a dog. If you remember correctly, I
  28023. 21:54:19said we're going to do some Python code.
  28024. 21:54:21And you can see over here, my hair is
  28025. 21:54:24kind of sticking up over the computer,
  28026. 21:54:25cup of coffee on one side, and a little
  28027. 21:54:27bit of old school. A pencil and a pen on
  28028. 21:54:29the other side. Yeah, most people now
  28029. 21:54:31take notes. I love the stickies on the
  28030. 21:54:33computer. That's great. That's that is
  28031. 21:54:35my computer. I have sticky notes on my
  28032. 21:54:37computer in different colors. So, not
  28033. 21:54:39too far from uh today's programmer. So,
  28034. 21:54:41the problem is is we want to classify
  28035. 21:54:43photos of cats and dogs using a neural
  28036. 21:54:46network. And you can see over here we
  28037. 21:54:48have quite a variety of dogs in the
  28038. 21:54:50pictures and cats and you know just
  28039. 21:54:53sorting out it is a cat is pretty
  28040. 21:54:55amazing. And why would anybody want to
  28041. 21:54:56even know the difference between a cat
  28042. 21:54:58and a dog? Okay, you know why? Well, I
  28043. 21:55:00have a cat door. It'd be kind of fun
  28044. 21:55:02that instead of it identifying, instead
  28045. 21:55:05of having like a little collar with a
  28046. 21:55:06magnet on it, which is what my cat has,
  28047. 21:55:08the door would be able to see, oh,
  28048. 21:55:10that's the cat. That's our cat coming
  28049. 21:55:11in. Oh, that's the dog. We have a dog,
  28050. 21:55:13too. That's a dog I want to let in.
  28051. 21:55:15Maybe I don't want to let this other
  28052. 21:55:16animal in cuz it's a raccoon. So, you
  28053. 21:55:18can see where you could take this one
  28054. 21:55:19step further and actually apply this.
  28055. 21:55:21You could actually start a little
  28056. 21:55:23startup company idea, self-identifying
  28057. 21:55:25door. So, this use case will be
  28058. 21:55:28implemented on Python. I am actually in
  28059. 21:55:30Python 3.6. It's always nice to tell
  28060. 21:55:33people the version of Python because
  28061. 21:55:35that does affect sometimes which modules
  28062. 21:55:37you load and everything. And we're going
  28063. 21:55:38to start by importing the required
  28064. 21:55:40packages. I told you we're going to do
  28065. 21:55:42this in Kass. So we're going to import
  28066. 21:55:44from KAS models sequential from the Kass
  28067. 21:55:48layers conversion 2D or COV2D max
  28068. 21:55:51pooling 2D flatten and dense. And we'll
  28069. 21:55:55talk about what each one of these do in
  28070. 21:55:56just a second. But before we do that,
  28071. 21:55:58let's talk a little bit about the
  28072. 21:56:00environment we're going to work in. And
  28073. 21:56:01uh you know, in fact, let me go ahead
  28074. 21:56:03and open a uh the website, KASS's
  28075. 21:56:06website, so we can learn a little bit
  28076. 21:56:07more about KASS. So here we are on the
  28077. 21:56:09Kurass website, and it's uh ke.io.
  28078. 21:56:14That's the official website for Kurass.
  28079. 21:56:16And the first thing you'll notice is
  28080. 21:56:17that Kurass runs on top of either
  28081. 21:56:20TensorFlow, CNTK, and I think it's
  28082. 21:56:23pronounced Thano or Theo. What's
  28083. 21:56:25important on here is that TensorFlow and
  28084. 21:56:27the same is true for all these, but
  28085. 21:56:28TensorFlow is probably one of the most
  28086. 21:56:30widely used currently packages out there
  28087. 21:56:32with the KAS. And of course, you know,
  28088. 21:56:34tomorrow this is all going to change.
  28089. 21:56:36It's all going to disappear and they'll
  28090. 21:56:37have something new out there. So, make
  28091. 21:56:38sure when you're learning this code that
  28092. 21:56:40you understand what's going on and also
  28093. 21:56:42know the code. I mean, look, when you
  28094. 21:56:44look at the code, it's not as
  28095. 21:56:45complicated once you understand what's
  28096. 21:56:46going on. The code itself is pretty
  28097. 21:56:48straightforward. And the reason we like
  28098. 21:56:50KAS and the reason that people are
  28099. 21:56:52jumping on it right now, it's such a big
  28100. 21:56:54deal is if we come down here, let me
  28101. 21:56:56just scroll down a little bit. They talk
  28102. 21:56:58about user friendliness, modularity,
  28103. 21:57:00easy extensibility, work with Python.
  28104. 21:57:02Python's a big one because a lot of
  28105. 21:57:04people in data science now use Python,
  28106. 21:57:06although you can actually access Kass
  28107. 21:57:08other ways. Is if we continue down here
  28108. 21:57:10is layers. And this is where it gets
  28109. 21:57:12really cool. When we're working with
  28110. 21:57:14KASS, you just add layers on. Remember
  28111. 21:57:16those hidden layers we were talking
  28112. 21:57:18about? And we talked about the reelu
  28113. 21:57:21activation. You can see right here. Let
  28114. 21:57:22me just up that a little bit in size.
  28115. 21:57:24There we go. That's big. I can add in an
  28116. 21:57:27eelu layer. And then I can add in a
  28117. 21:57:29softmax layer in the next instance. We
  28118. 21:57:31didn't talk about softmax. So you can do
  28119. 21:57:33each layer separate. Now if I'm working
  28120. 21:57:35in some of the other kits I use, I take
  28121. 21:57:38that and I have one setup and then I
  28122. 21:57:40feed the output into the next one. This
  28123. 21:57:42one I can just add hidden layer after
  28124. 21:57:44hidden layer with the different
  28125. 21:57:45information in it which makes it very
  28126. 21:57:47powerful and very fast to spin up and
  28127. 21:57:50try different setups and see how they
  28128. 21:57:52work with the data you're working on.
  28129. 21:57:53And we'll dig a little bit deeper in
  28130. 21:57:54here. And a lot of this is very much the
  28131. 21:57:57same. So when we get to that part, I'll
  28132. 21:57:58point that out to you also. Now just a
  28133. 21:58:01quick side note, I'm using Anaconda with
  28134. 21:58:03Python in it. And I went ahead and
  28135. 21:58:05created my own package and I called it
  28136. 21:58:07the Kass Python 36 because I'm in Python
  28137. 21:58:0936. Anaconda is cool that You can create
  28138. 21:58:12different environments really easily. If
  28139. 21:58:13you're doing a lot of different
  28140. 21:58:14experimenting with these different
  28141. 21:58:16packages, probably want to create your
  28142. 21:58:17own environment in there. And the first
  28143. 21:58:19thing, as you can see right here,
  28144. 21:58:21there's a lot of dependencies. A lot of
  28145. 21:58:22these you should recognize by now if
  28146. 21:58:24you've done any of these videos. If not,
  28147. 21:58:26kudos for you for jumping in today. PIP,
  28148. 21:58:28install, numpy, sci, the scikitlearn,
  28149. 21:58:32pillow, and h5py
  28150. 21:58:34are both needed for the tensorflow and
  28151. 21:58:37then putting the kass on there. And then
  28152. 21:58:39you'll see here uh and pip is just a
  28153. 21:58:41standard installer that you use with
  28154. 21:58:42Python. You'll see here that we did pip
  28155. 21:58:44install TensorFlow since we're going to
  28156. 21:58:46do KAS on top of TensorFlow. And then
  28157. 21:58:48pip install and I went ahead and used
  28158. 21:58:50the GitHub. So git plusgit and you'll
  28159. 21:58:52see here github.com. This is one of
  28160. 21:58:55their releases, one of the most current
  28161. 21:58:56release on there that goes on top of
  28162. 21:58:58TensorFlow. And you can look up these
  28163. 21:58:59instructions pretty much anywhere. This
  28164. 21:59:01is for doing it on Anaconda. Certainly
  28165. 21:59:04you'd want to install these if you're
  28166. 21:59:05doing it in Iuntu server setup. you
  28167. 21:59:07you'd want to get I don't think you need
  28168. 21:59:09the H5 py and aru but you do need the
  28169. 21:59:11rest in there because they are
  28170. 21:59:12dependencies in there and it's pretty
  28171. 21:59:14straightforward and that's actually in
  28172. 21:59:15some of the instructions they have on
  28173. 21:59:17their website so you don't have to
  28174. 21:59:18necessarily go through this just
  28175. 21:59:19remember their website on there and then
  28176. 21:59:21when I'm under my uh Anaconda navigator
  28177. 21:59:24which I like you'll see where I have
  28178. 21:59:26environments and on the bottom I created
  28179. 21:59:28a new environment and I called it KAS
  28180. 21:59:30Python 36 just to separate everything
  28181. 21:59:32you can say I have Python 3.5 and Python
  28182. 21:59:3536 I used to have a bunch of other ones,
  28183. 21:59:37but it kind of cleaned house recently.
  28184. 21:59:39And of course, once I go in here, I can
  28185. 21:59:41launch my Jupyter Notebook, making sure
  28186. 21:59:43I'm using the right environment that I
  28187. 21:59:45just set up. This, of course, opens up
  28188. 21:59:47my um in this case, I'm using uh Google
  28189. 21:59:50Chrome. And in here, I could go and just
  28190. 21:59:52create a new document in here. And this
  28191. 21:59:54is all in your um browser window when
  28192. 21:59:56you use the Anaconda. Do you have to use
  28193. 21:59:58Anaconda and Jupyter Notebook? No. You
  28194. 22:00:01can use any kind of Python editor,
  28195. 22:00:03whatever setup you're comfortable with
  28196. 22:00:05and whatever you're doing in there. So,
  28197. 22:00:07let's go ahead and go in here and paste
  28198. 22:00:09the code in. And we're importing a
  28199. 22:00:11number of different settings in here. We
  28200. 22:00:14have import sequential. That's under the
  28201. 22:00:16models because that's the model we're
  28202. 22:00:17going to use as far as our neural
  28203. 22:00:19network. And then we have layers and we
  28204. 22:00:21have conversion 2D, max pooling 2D,
  28205. 22:00:24flatten dense. And you can actually just
  28206. 22:00:27kind of guess at what these do. We're
  28207. 22:00:29talking we're working in a 2D
  28208. 22:00:31photograph. And if you remember
  28209. 22:00:32correctly, I talked about how the actual
  28210. 22:00:35input layer is a single array. It's not
  28211. 22:00:37in two dimensions. It's one dimension.
  28212. 22:00:39All these do is these are tools to help
  28213. 22:00:41flatten the image. So, it takes a
  28214. 22:00:43two-dimensional image and then it
  28215. 22:00:44creates its own proper setup. You don't
  28216. 22:00:46have to worry about any of that. You
  28217. 22:00:48don't have to do anything special with
  28218. 22:00:49the photograph. You let the carass do
  28219. 22:00:51it. And we're going to run this. And
  28220. 22:00:52you'll see right here they have some
  28221. 22:00:53stuff that is going to be depreciated
  28222. 22:00:55and changed because that's what it does.
  28223. 22:00:56Everything's being changed as we go. You
  28224. 22:00:58don't have to worry about that too much.
  28225. 22:01:00If you have warnings, if you run it a
  28226. 22:01:01second time, the warning will disappear.
  28227. 22:01:03And this has just imported these
  28228. 22:01:04packages for us to use. Jupiter's nice
  28229. 22:01:07about this that you can do each thing
  28230. 22:01:08step by step. And I'll go ahead and also
  28231. 22:01:10zoom in there. A little control plus.
  28232. 22:01:13That's one of the nice things about
  28233. 22:01:14being in a browser environment. So, here
  28234. 22:01:17we are back. Another sip of coffee. If
  28235. 22:01:20you're familiar with my other videos,
  28236. 22:01:21you notice I'm always sipping coffee. I
  28237. 22:01:23always have a in my case latte next to
  28238. 22:01:24me, an espresso. So the next step is to
  28239. 22:01:26go ahead and initialize. We're going to
  28240. 22:01:28call it the CNN or classifier neural
  28241. 22:01:31network. And the reason we call it a
  28242. 22:01:33classifier is because it's going to
  28243. 22:01:34classify it between two things. It's
  28244. 22:01:36going to be cat or dog. So when you're
  28245. 22:01:38doing classification, you're picking
  28246. 22:01:40specific objects. You're specific. It's
  28247. 22:01:43a true or false. Yes, no. It is
  28248. 22:01:45something or it's not. So first thing
  28249. 22:01:48we're going to create our classifier and
  28250. 22:01:50it's going to equal sequential. So their
  28251. 22:01:51sequential setup is the classifier.
  28252. 22:01:54That's the actual model we're using.
  28253. 22:01:56That's the neural network. So we call it
  28254. 22:01:58a classifier. And uh the next step is to
  28255. 22:02:01add in our convolution. And let me just
  28256. 22:02:03do a uh let me shrink that down in size
  28257. 22:02:06so you can see the whole line. And let's
  28258. 22:02:07talk a little bit about what's going on
  28259. 22:02:09here. I have my classifier and I add
  28260. 22:02:11something. What am I adding? Well, I'm
  28261. 22:02:13adding my first layer. This first layer
  28262. 22:02:16we're adding in is probably the one that
  28263. 22:02:18takes the most work to make sure you
  28264. 22:02:20have it set correct. And the reason I
  28265. 22:02:22say that is this is your actual input.
  28266. 22:02:24And we're going to jump here to the part
  28267. 22:02:25that says input shape equals 64x 64x3.
  28268. 22:02:30What does that mean? Well, that means
  28269. 22:02:32that our pictures coming in. And there's
  28270. 22:02:35these pictures. Remember we had like the
  28271. 22:02:36picture of the car was 128x 128 pixels.
  28272. 22:02:40Well, this one is 64x 64 pixels. And
  28273. 22:02:43each pixel has three values. That's
  28274. 22:02:46where these numbers come from. And it is
  28275. 22:02:48so important that this matches. I
  28276. 22:02:50mentioned a little bit that if you have
  28277. 22:02:52like a larger picture, you have to
  28278. 22:02:53reformat it to fit this shape. If it
  28279. 22:02:56comes in as something larger, there's no
  28280. 22:02:58input notes. There's no input neural
  28281. 22:03:00network there that will handle that
  28282. 22:03:02extra space. So, you have to reshape
  28283. 22:03:03your data to fit in here. Now, the first
  28284. 22:03:06layer is the most important because
  28285. 22:03:07after that, KAS knows what your shape is
  28286. 22:03:10coming in here and it knows what's
  28287. 22:03:12coming out and so that really sets the
  28288. 22:03:15stage. Most important thing is that
  28289. 22:03:17input shape matches your data coming in.
  28290. 22:03:19And you'll get a lot of errors if it
  28291. 22:03:20doesn't. You'll go through there and
  28292. 22:03:22picture number 55 doesn't match it
  28293. 22:03:24correctly. And guess what it does? It
  28294. 22:03:26usually gives you an error. And then the
  28295. 22:03:27activation, if you remember, we talked
  28296. 22:03:29about the different activations on here.
  28297. 22:03:31We're using the reelu model. Like I
  28298. 22:03:34said, that is the most commonly used now
  28299. 22:03:36because one, it's fast. Doesn't have the
  28300. 22:03:39added calculations in it. It just says
  28301. 22:03:41here's the value coming out based on the
  28302. 22:03:43weights and the value going in. And um
  28303. 22:03:47from there, you know, it's uh if it's
  28304. 22:03:49over one, then it's good or over zero,
  28305. 22:03:51it's good. If it's under zero, then it's
  28306. 22:03:53considered not active. And then we have
  28307. 22:03:55this conversion 2D. What the heck is
  28308. 22:03:58conversion 2D? I'm not going to go into
  28309. 22:04:01too much detail in this because this has
  28310. 22:04:03a couple of things it's doing in here, a
  28311. 22:04:05little bit more in-depth than we're
  28312. 22:04:06ready to cover in this tutorial. But
  28313. 22:04:08this is used to convert from the photo
  28314. 22:04:11cuz we have 64x 64x3 and we're just
  28315. 22:04:14converting it to two-dimensional kind of
  28316. 22:04:16setup. So it's very aware that this is a
  28317. 22:04:18photograph and that different pieces are
  28318. 22:04:20next to each other. And then we're going
  28319. 22:04:21to add in uh a second convolutional
  28320. 22:04:24layer. That's what the cov stands for
  28321. 22:04:272D. So it's these are hidden layers. So
  28322. 22:04:29we have our input layer and our two
  28323. 22:04:31hidden layers and they are
  28324. 22:04:32two-dimensional because we're dealing
  28325. 22:04:34with a two-dimensional photograph. And
  28326. 22:04:36you'll see down here that on the last
  28327. 22:04:37one, we add a max pooling 2D and we put
  28328. 22:04:40a pool size equals 22. And so what this
  28329. 22:04:43is is that as you get to the end of
  28330. 22:04:45these layers, one of the things you
  28331. 22:04:47always want to think of is what they
  28332. 22:04:48call mapping and then reducing.
  28333. 22:04:50Wonderful terminology from the big data.
  28334. 22:04:52We're mapping this data through all
  28335. 22:04:53these layers. And now we want to reduce
  28336. 22:04:56it to only two sets. In this case, it's
  28337. 22:04:59already in two sets because it's a 2D
  28338. 22:05:01photograph. But we had, you know, two
  28339. 22:05:02dimensions by we actually have 64x 64
  28340. 22:05:05by3. So now we're just getting it down
  28341. 22:05:07to a 2x two. Just the two dimension
  28342. 22:05:10two-dimensional instead of having the
  28343. 22:05:11third dimension of colors. And we'll go
  28344. 22:05:13ahead and run these. We're not really
  28345. 22:05:15seeing anything in our run script
  28346. 22:05:17because we're just setting up. This is
  28347. 22:05:18all set up. And this is where you start
  28348. 22:05:20playing because maybe you'll add a
  28349. 22:05:21different layer in here to do something
  28350. 22:05:23else to see how it works and see what
  28351. 22:05:25your output is. That's what makes KAS so
  28352. 22:05:27nice is I can with just a couple flips
  28353. 22:05:29of code put in a whole new layer that
  28354. 22:05:31does a whole new processing and see
  28355. 22:05:33whether that improves my run or makes it
  28356. 22:05:35worse. And finally, we're going to do
  28357. 22:05:38the final setup, which is to flatten
  28358. 22:05:40classifier, add a flatten setup. And
  28359. 22:05:43then we're going to also add a layer, a
  28360. 22:05:44dense layer, and then we're going to add
  28361. 22:05:46in another dense layer. And then we're
  28362. 22:05:49going to build it. We're going to
  28363. 22:05:50compile this whole thing together. So,
  28364. 22:05:52let's flip over and see what that looks
  28365. 22:05:53like. And we've even numbered them for
  28366. 22:05:55you. So, we're going to do the
  28367. 22:05:56flattening. And flatten is exactly what
  28368. 22:05:58it sounds like. We've been working in a
  28369. 22:06:00two-dimensional array of picture, which
  28370. 22:06:03actually is in three dimensions because
  28371. 22:06:04of the pixels. The pixels have a whole
  28372. 22:06:06another dimension to it of three
  28373. 22:06:08different values. And we've kind of
  28374. 22:06:10resized those down to 2x two. But now
  28375. 22:06:12we're just going to flatten it. I don't
  28376. 22:06:14want to have multiple dimensions being
  28377. 22:06:16worked on by tensor and by kas. I want
  28378. 22:06:19just a single array. So, it's flattened
  28379. 22:06:21out. And then step four, full
  28380. 22:06:23connection. So we add in our final two
  28381. 22:06:26layers. And you could actually do all
  28382. 22:06:28kinds of things with this. You could
  28383. 22:06:29actually leave out this some of these
  28384. 22:06:31layers and play with them. You do need
  28385. 22:06:33to flatten it. That's very important.
  28386. 22:06:35Then we want to use the dents again.
  28387. 22:06:37We're taking this and we're taking
  28388. 22:06:40whatever came into it. So once we take
  28389. 22:06:42all those different the two dimensions
  28390. 22:06:43or three dimensions as they are and we
  28391. 22:06:45flatten it to one dimension. We want to
  28392. 22:06:47take that and we're going to pull it
  28393. 22:06:49into units of 128. They got that. You're
  28394. 22:06:52say where did they get 128 from? You
  28395. 22:06:53could actually play with that number and
  28396. 22:06:55get all kinds of weird results. But in
  28397. 22:06:56this case we took the 64 + 64 is 128.
  28398. 22:07:00You could probably even do this with 64
  28399. 22:07:02or 32. Usually you want to keep it in
  28400. 22:07:04the same multiple whatever the data
  28401. 22:07:05shape you're already using is in. And
  28402. 22:07:07we're using the activation the re lu
  28403. 22:07:10just like we did before. And then we
  28404. 22:07:11finally filter all that into a single
  28405. 22:07:15output. And it has how many units? One.
  28406. 22:07:17Why? Because we want to know whether
  28407. 22:07:19true or false. It's either a dog or a
  28408. 22:07:21cat. You could say one is dog, zero is
  28409. 22:07:24cat. Or maybe you're a cat lover and
  28410. 22:07:26it's one is cat and zero is dog. And if
  28411. 22:07:28you love both dogs and cats, you're
  28412. 22:07:31going to have to choose. And then we use
  28413. 22:07:33the sigmoid activation. If you remember
  28414. 22:07:35from before, we had the reel and there's
  28415. 22:07:38also the sigmoid. The sigmoid just makes
  28416. 22:07:40it clear it's yes or no. We don't want a
  28417. 22:07:42any kind of in between number coming
  28418. 22:07:44out. And we'll go ahead and run this.
  28419. 22:07:46And you'll see it's still all in setup.
  28420. 22:07:48And then finally, we want to go ahead
  28421. 22:07:49and compile. And let's put the compiling
  28422. 22:07:52our um classifier neural network. And
  28423. 22:07:54we're going to use the optimizer atom.
  28424. 22:07:56And I hinted at this just a little bit
  28425. 22:07:59before. Where does atom come in? Where
  28426. 22:08:01does an optimizer come in? Well, the
  28427. 22:08:03optimizer is the reverse propagation.
  28428. 22:08:06When we're training it, it goes all the
  28429. 22:08:08way through and says error and then how
  28430. 22:08:10does it readjust those weights. There
  28431. 22:08:12are a number of them. Atom is the most
  28432. 22:08:14commonly used and it works best on large
  28433. 22:08:17data. Most people stick with the atom
  28434. 22:08:20because when they're testing on smaller
  28435. 22:08:21data, see if their model is going to go
  28436. 22:08:23through and get all their errors out
  28437. 22:08:24before they run it on larger data sets.
  28438. 22:08:26They're going to run it on atom anyway,
  28439. 22:08:28so they just leave it on atom most
  28440. 22:08:30commonly used. But there are some other
  28441. 22:08:32ones out there. You should be aware of
  28442. 22:08:33that that you might try them if you're
  28443. 22:08:35stuck in a bind or you might blur that
  28444. 22:08:37in the future, but usually atom is just
  28445. 22:08:39fine on there. And then you have two
  28446. 22:08:40more settings. You have loss and
  28447. 22:08:42metrics. We're not going to dig too much
  28448. 22:08:44into loss or metrics. These are things
  28449. 22:08:46you really have to explore KAS because
  28450. 22:08:48there are so many choices. This is how
  28451. 22:08:50it computes the error. There's so many
  28452. 22:08:52different ways to on your back
  28453. 22:08:53propagation and your training. So we're
  28454. 22:08:55using the atom model, but you can
  28455. 22:08:57compute the error by um standard
  28456. 22:08:59deviation, standard deviation squared.
  28457. 22:09:02They use binary cross entropy. I'd have
  28458. 22:09:04to look that up to even know what that
  28459. 22:09:05is. There's so many of these. A lot of
  28460. 22:09:07times you just start with the ones that
  28461. 22:09:08look correct that are most commonly used
  28462. 22:09:11and then you have to go read the KAS
  28463. 22:09:12site and actually see what these
  28464. 22:09:14different losses and metrics and what
  28465. 22:09:16different options they have. So, we're
  28466. 22:09:18not going to get too much into them
  28467. 22:09:20other than to reference you over to the
  28468. 22:09:21KAS website to explore them deeper, but
  28469. 22:09:23we are going to go ahead and run them.
  28470. 22:09:25And now we've set up our classifier. So,
  28471. 22:09:27we have an object classifier. And if you
  28472. 22:09:29go back up here, you'll see that we've
  28473. 22:09:31added in step one. We added in our layer
  28474. 22:09:33for the input. We added a layer that
  28475. 22:09:36comes in there and uses the reelu for
  28476. 22:09:38activation. And then it pulls the data.
  28477. 22:09:41So this is even though these are two
  28478. 22:09:42layers, the actual neural network layer
  28479. 22:09:44is up here. And then it uses this to
  28480. 22:09:46pull the data into a 2x two. So into a
  28481. 22:09:49two-dimensional array from a
  28482. 22:09:50three-dimensional array with the colors.
  28483. 22:09:52Then we flatten it. So there's our adder
  28484. 22:09:54flatten. And then we add another dense
  28485. 22:09:56what they call dense layer. this dense
  28486. 22:09:58layer goes in there and it it downsizes
  28487. 22:10:00it to 128. It reduces it. So you can
  28488. 22:10:04look at this as uh we're mapping all
  28489. 22:10:06this data down the two-dimensional setup
  28490. 22:10:08and then we flatten it. So we map it to
  28491. 22:10:10a flatten map and then we take it and
  28492. 22:10:12reduce it down to 128 and we use the
  28493. 22:10:15reel again. And then finally we reduce
  28494. 22:10:18that down to just a single output and we
  28495. 22:10:20use a sigmoid to do that to figure out
  28496. 22:10:22whether it's yes, no, true, false, in
  28497. 22:10:25this case cat or dog. And then finally
  28498. 22:10:27once we put all these layers together we
  28499. 22:10:29compile them. That's what we've done
  28500. 22:10:30here and we've compiled them as far as
  28501. 22:10:33how it trains to use these settings for
  28502. 22:10:35the training back propagation. So if you
  28503. 22:10:37remember we talked about training our
  28504. 22:10:40setup and when we go into this you'll
  28505. 22:10:42see that we have two data sets. We have
  28506. 22:10:44one called the training set and the
  28507. 22:10:46testing set. And that's very standard in
  28508. 22:10:48any data processing is you need to have
  28509. 22:10:51that's pretty common in any data
  28510. 22:10:53processing is you need to have a certain
  28511. 22:10:55amount of data to train it and then you
  28512. 22:10:56got to know whether it works or not. Is
  28513. 22:10:58it any good and that's why you have a
  28514. 22:11:00separate set of data for testing it
  28515. 22:11:02where you already know the answer but
  28516. 22:11:03you don't want to use that as part of
  28517. 22:11:04the training set. So in here we jump
  28518. 22:11:07into part two fitting the classifier
  28519. 22:11:10neuron network to the images and then
  28520. 22:11:12from KAS let me just zoom in there. I
  28521. 22:11:15always love that about working with
  28522. 22:11:16Jupyter Notebooks. You can really see.
  28523. 22:11:18We're going to come in here. We do the
  28524. 22:11:19cross pre-processing an image. And we
  28525. 22:11:21import image data generator. It's so
  28526. 22:11:24nice of KAS. It's such a high-end
  28527. 22:11:26product right now going out. And since
  28528. 22:11:28images are so common, they already have
  28529. 22:11:30all this stuff to help us process the
  28530. 22:11:31data, which is great. And so, we come in
  28531. 22:11:33here, we do train data gen, and we're
  28532. 22:11:36going to create our object for helping
  28533. 22:11:38us train for reshaping the data so that
  28534. 22:11:40it's going to work with our setup. and
  28535. 22:11:43we use an image data generator and we're
  28536. 22:11:45going to rescale it. And you'll see here
  28537. 22:11:47we have one point which tells us it's a
  28538. 22:11:49float value on the rescale over 255.
  28539. 22:11:53Where does 255 come from? Well, that's
  28540. 22:11:55the scale in the colors of the pictures
  28541. 22:11:57we're using. They're value from 0 to
  28542. 22:11:59255. So, we want to divide it by 255 and
  28543. 22:12:02it'll generate a number between 0 and 1.
  28544. 22:12:05They have sheer range and zoom range.
  28545. 22:12:08Horizontal flip equals true. And this,
  28546. 22:12:10of course, has to do with if the photos
  28547. 22:12:12are different shapes and sizes. Like I
  28548. 22:12:14said, it's a wonderful package. You
  28549. 22:12:15really need to dig in deep to see all
  28550. 22:12:17the different options you have for
  28551. 22:12:19setting up your images. For right now
  28552. 22:12:21though, we're going to just stick with
  28553. 22:12:22some basic stuff here. And let me go
  28554. 22:12:23ahead and run this code. And again, it
  28555. 22:12:25doesn't really do anything because we're
  28556. 22:12:27still setting up the pre-processing.
  28557. 22:12:29Let's take a look at this next set of
  28558. 22:12:31code. And this one is just huge. We're
  28559. 22:12:33creating the training set. So the
  28560. 22:12:35training set is going to go in here and
  28561. 22:12:37it's going to use our train data gen we
  28562. 22:12:40just created flow from directory. It's
  28563. 22:12:42going to access in this case the path
  28564. 22:12:45data set training set. That's a folder.
  28565. 22:12:48So it's going to pull all the images out
  28566. 22:12:50of that folder. Now I'm actually running
  28567. 22:12:52this in the folder that the data sets
  28568. 22:12:54in. So if you're doing the same setup
  28569. 22:12:57and you load your data in there and
  28570. 22:12:59you're doing this, make sure wherever
  28571. 22:13:01your Jupyter notebook is saving things
  28572. 22:13:03to that you create this path or you can
  28573. 22:13:05do the complete path if you need to, you
  28574. 22:13:07know, C colon slash etc. And the target
  28575. 22:13:10size, the batch size and class mode is
  28576. 22:13:13binary. So the classes, we're switching
  28577. 22:13:15everything to a binary value. Batch
  28578. 22:13:17size. What the heck is batch size? Well,
  28579. 22:13:18that's how many pictures we're going to
  28580. 22:13:20batch through the training each time.
  28581. 22:13:22And the target size 64x 64. A little
  28582. 22:13:25confusing, but you can see right here
  28583. 22:13:26that this is just a general training and
  28584. 22:13:28you can go in there and look at all the
  28585. 22:13:30different settings for your training
  28586. 22:13:31set. And of course with different data,
  28587. 22:13:33we're doing pictures. There's all kinds
  28588. 22:13:35of different settings depending on what
  28589. 22:13:36you're working with. Let's go ahead and
  28590. 22:13:38run that and see what happens. And
  28591. 22:13:39you'll see that it found 800 images
  28592. 22:13:41belonging to one classes. So we have 800
  28593. 22:13:44images in the training set. And if we're
  28594. 22:13:47going to do this with uh the training
  28595. 22:13:50set, we also have to format the pictures
  28596. 22:13:52in the test set. Now, we're not actually
  28597. 22:13:55doing any predictions. We're not
  28598. 22:13:56actually programming the model yet. All
  28599. 22:14:00we're doing is preparing the data. So,
  28600. 22:14:01we're going to prepare a training set
  28601. 22:14:03and the test set. So, any changes we
  28602. 22:14:05make to the training set at this point
  28603. 22:14:07also have to be made to the test set.
  28604. 22:14:09So, we've done this thing. We've done a
  28605. 22:14:11train data generator. We've done our
  28606. 22:14:13training set. And then we also have
  28607. 22:14:15remember our test set of data. So I'm
  28608. 22:14:17going to do the same thing with that.
  28609. 22:14:18I'm going to create a test data gen and
  28610. 22:14:21we're going to do this image data
  28611. 22:14:22generator. We're going to rescale one
  28612. 22:14:24over 255. We don't need the other
  28613. 22:14:26settings, just the single setting for
  28614. 22:14:28the test data gen. And we're going to
  28615. 22:14:30create our test set. We're going to do
  28616. 22:14:31the same thing we did with the test set
  28617. 22:14:33except that we're pulling it from the
  28618. 22:14:34test set folder. And we'll run that. And
  28619. 22:14:37you'll see in our test set we found
  28620. 22:14:392,000 images. That's about right. We're
  28621. 22:14:42using 20% of the images as test and 80%
  28622. 22:14:45to train it. And then finally, we've set
  28623. 22:14:47up all our data. We've set up all our
  28624. 22:14:50layers, which is where all the work is
  28625. 22:14:52is cleaning up that data, making sure
  28626. 22:14:54it's going in there correctly. And we're
  28627. 22:14:55actually going to fit it. We're going to
  28628. 22:14:58train our data set. And let's see what
  28629. 22:15:00that looks like. And here we go. Let's
  28630. 22:15:02put the information in here. And let's
  28631. 22:15:03just take a quick look at what we're
  28632. 22:15:06looking at with our fit generator. We
  28633. 22:15:08have our classifier.fit
  28634. 22:15:09generator. That's our back propagation.
  28635. 22:15:12So the information goes through forward
  28636. 22:15:14with a picture and it says, "Oh, you're
  28637. 22:15:16either right or you're wrong." And then
  28638. 22:15:18the error goes backward and reprograms
  28639. 22:15:21all those weights. So we're training our
  28640. 22:15:23neural network. And of course, we're
  28641. 22:15:25using the training set. Remember, we
  28642. 22:15:27created the training set up here. And
  28643. 22:15:28then we're going steps per epic. So it's
  28644. 22:15:318,000 steps. Epic means that that's how
  28645. 22:15:34many times we go through all the
  28646. 22:15:36pictures. So we're going to rerun each
  28647. 22:15:37of the pictures. and we're going to go
  28648. 22:15:39through the whole data set 25 times, but
  28649. 22:15:41we're going to look at each picture
  28650. 22:15:43during each epic 8,000 times. So, we're
  28651. 22:15:46really programming the heck out of this
  28652. 22:15:47and going back over it. And then they
  28653. 22:15:50have validation data equals test set.
  28654. 22:15:52So, we have our training set and then
  28655. 22:15:54we're going to have our test set to
  28656. 22:15:56validate it. So, we're going to do this
  28657. 22:15:57all in one shot and we're going to look
  28658. 22:15:59at that and they're going to do 200
  28659. 22:16:00steps for each validation and we'll see
  28660. 22:16:02what that looks like in just a minute.
  28661. 22:16:04Let's go ahead and run our training
  28662. 22:16:05here. And we're going to fit our data.
  28663. 22:16:07And as it goes, it says epic one of 25.
  28664. 22:16:10You start realizing that this is going
  28665. 22:16:12to take a while. On my older computer,
  28666. 22:16:14it takes about 45 minutes. I have a dual
  28667. 22:16:18processor. You know, we're processing uh
  28668. 22:16:2110,000 photos. That's not a small amount
  28669. 22:16:23of photographs to process. So, if you're
  28670. 22:16:26on your laptop, you know, which I am,
  28671. 22:16:27it's going to take a while. So, let's go
  28672. 22:16:29ahead and uh go get our cup of coffee
  28673. 22:16:31and a sip and come back and see what
  28674. 22:16:33this looks like. So, I'm back. You
  28675. 22:16:35didn't know I was gone. That was
  28676. 22:16:36actually a lengthy pause there. I made a
  28677. 22:16:39couple changes. Let's discuss those
  28678. 22:16:41changes real quick and why I made them.
  28679. 22:16:43So, the first thing I'm going to do is
  28680. 22:16:44I'm going to go up here and insert a
  28681. 22:16:46cell above and let's paste the original
  28682. 22:16:48code back in there. And you'll see that
  28683. 22:16:50the original thing was steps per epic
  28684. 22:16:528,000, 25 epics, and validation steps
  28685. 22:16:552,000. And I changed these to 4,000
  28686. 22:16:58epics or 4,000 steps per epic, 10 epics,
  28687. 22:17:02and just 10 validation steps. And this
  28688. 22:17:05will cause problems if you're doing this
  28689. 22:17:07as a commercial release. But for demo
  28690. 22:17:09purposes, this should work. And if you
  28691. 22:17:11remember our steps per epic, that's how
  28692. 22:17:13many photos we're going to process. In
  28693. 22:17:15fact, let me go ahead and get my drawing
  28694. 22:17:16pen out. And uh let's just highlight
  28695. 22:17:18that right here. We have 8,000 pictures
  28696. 22:17:21we're going through. So for each epic,
  28697. 22:17:23I'm going to change this to 4,000. I'm
  28698. 22:17:24going to cut that in half. So, it's
  28699. 22:17:26going to randomly pick 4,000 pictures
  28700. 22:17:28each time it goes through an epic. And
  28701. 22:17:29the epic is how many processes. So, this
  28702. 22:17:32is 25. And I'm just going to cut that to
  28703. 22:17:3410. So, instead of doing 25 runs through
  28704. 22:17:378,000 photos each, which you can do the
  28705. 22:17:39math of 25 * 8,000, I'm only going to do
  28706. 22:17:4210 through 4,000. So, I'm going to run
  28707. 22:17:44this 40,000 times through the processes.
  28708. 22:17:47And the next thing I not you'll you'll
  28709. 22:17:49want to notice is that I also changed
  28710. 22:17:50the validation step. And this would
  28711. 22:17:52cause some major problems in releasing
  28712. 22:17:54cuz I dropped it all the way down to 10.
  28713. 22:17:56What the validation step does is it says
  28714. 22:17:58we have 2,000 photos in our training or
  28715. 22:18:01in our testing set and we're going to
  28716. 22:18:03use that for validation. Well, I'm only
  28717. 22:18:05going to use a random 10 of those to
  28718. 22:18:07validate. So, not really the best
  28719. 22:18:09settings, but let me show you why we did
  28720. 22:18:11that. Let's scroll down here just a
  28721. 22:18:13little bit and let's look at the output
  28722. 22:18:15here and see what that what's going on
  28723. 22:18:17there. So, I've got my drawing tool back
  28724. 22:18:19on, and you'll see here it lists a run.
  28725. 22:18:22So, each time it goes through an epic,
  28726. 22:18:24it's going to do 4,000 steps. And this
  28727. 22:18:26is where the 4,000 comes in. So, that's
  28728. 22:18:28where we have. We have epic one of 10,
  28729. 22:18:294,000 steps. So, it's randomly picking
  28730. 22:18:32half the pictures in the file and going
  28731. 22:18:33through them. And then we're going to
  28732. 22:18:34look at this number right here. That is
  28733. 22:18:36for the whole epic, and that's 24, 411
  28734. 22:18:40seconds. And if you remember correctly,
  28735. 22:18:42you divide that by 60, you get minutes.
  28736. 22:18:44If you divide that by 60, you get hours.
  28737. 22:18:47Or you can just divide the whole thing
  28738. 22:18:48by 60 * 60 which is 3600. If 3600 is an
  28739. 22:18:52hour, this is roughly 45 minutes right
  28740. 22:18:55here. And that's 45 minutes to process
  28741. 22:18:57half the pictures. So if I was doing all
  28742. 22:19:00the pictures, we're talking an hour and
  28743. 22:19:02a half per epic times 36 or no 25. They
  28744. 22:19:06had 25 up above 25. So that's roughly a
  28745. 22:19:09couple days. A couple days of
  28746. 22:19:11processing. Well, for this demo, we
  28747. 22:19:12don't want to do that. I don't want to
  28748. 22:19:14come back the next day. Plus, my
  28749. 22:19:15computer did a reboot in the middle of
  28750. 22:19:17the night. So, we look at this and we
  28751. 22:19:19say, "Okay, let's we're just testing
  28752. 22:19:20this out. My computer that I'm running
  28753. 22:19:22this on is a dual core processor. Uh,
  28754. 22:19:25runs 0.9 gigahertz per second. For a
  28755. 22:19:28laptop, you know, it's good about 4
  28756. 22:19:29years ago, but for running something
  28757. 22:19:31like this, it's probably a little slow.
  28758. 22:19:33So, we cut the times down. And the last
  28759. 22:19:34one was validation. We're only
  28760. 22:19:36validating it on a random 10 photos. And
  28761. 22:19:38this comes into effect because you're
  28762. 22:19:40going to see down here where we have
  28763. 22:19:42accuracy, value loss, value accuracy,
  28764. 22:19:46and loss. Those are very important
  28765. 22:19:48numbers to look at. So the 10 means I'm
  28766. 22:19:50only validating across 10 pictures. That
  28767. 22:19:53is where here we have value. This is ACC
  28768. 22:19:56is for accuracy. Value loss. We're not
  28769. 22:19:58going to worry about that too much. And
  28770. 22:19:59accuracy. Now accuracy is while it's
  28771. 22:20:02running, it's putting these two numbers
  28772. 22:20:04together. That's what accuracy is. And
  28773. 22:20:06value accuracy is at the end of the
  28774. 22:20:09epic. What's our accuracy into the epic?
  28775. 22:20:11What is it looking at? In this tutorial,
  28776. 22:20:12we're not going to go so deep, but these
  28777. 22:20:14numbers are really important when you
  28778. 22:20:16start talking about these two numbers
  28779. 22:20:19reflect bias. That is really important.
  28780. 22:20:22We just put that up there. And bias is a
  28781. 22:20:24little bit beyond this tutorial, but the
  28782. 22:20:26short of it is is if this accuracy,
  28783. 22:20:28which is being our validation per step
  28784. 22:20:30is going down and the value accuracy
  28785. 22:20:34continues to go up, that means there's a
  28786. 22:20:36bias. That means I'm memorizing the
  28787. 22:20:38photos I'm looking at. I'm not actually
  28788. 22:20:41looking for what makes a dog a dog, what
  28789. 22:20:44makes a cat a cat. I'm just memorizing
  28790. 22:20:46them. And so the more this discrepancy
  28791. 22:20:48grows, the bigger the bias is. And that
  28792. 22:20:50is really the beauty of the KAS neural
  28793. 22:20:54network. It has a lot of built-in
  28794. 22:20:55features like this that make that really
  28795. 22:20:57easy to track. So let's go ahead and
  28796. 22:20:59take a look at the next set of code. So
  28797. 22:21:01here we are into part three. We're going
  28798. 22:21:04to make a new prediction. And so we're
  28799. 22:21:06going to bring in a couple tools for
  28800. 22:21:07that. And then we have to process the
  28801. 22:21:09image coming in and find out whether
  28802. 22:21:11it's an actual dog or cat if we can
  28803. 22:21:13actually use this to identify it. And of
  28804. 22:21:15course the final step of part three is
  28805. 22:21:17to print prediction. We'll go ahead and
  28806. 22:21:19combine these. And of course you can see
  28807. 22:21:20me there adding more sticky notes to my
  28808. 22:21:22computer screen hidden behind the
  28809. 22:21:24screen. And you know last one was don't
  28810. 22:21:26forget to feed the cat and the dog.
  28811. 22:21:29So let's go and take a look at that and
  28812. 22:21:30see what that looks like in code and put
  28813. 22:21:32that in our Jupyter notebook. All right.
  28814. 22:21:34And let's paste that in here. And we'll
  28815. 22:21:36start by importing numpy as np. Numpy is
  28816. 22:21:40a very common package. I pretty much
  28817. 22:21:42import it on any Python project I'm
  28818. 22:21:44working on. Another one I use regularly
  28819. 22:21:46is pandas. They're just ways of
  28820. 22:21:47organizing the data. And then np is
  28821. 22:21:49usually the standard in most machine
  28822. 22:21:52learning tools as the return for the
  28823. 22:21:54data array. Although you know you use a
  28824. 22:21:56standard data array from Python. And we
  28825. 22:21:58have cross pre-processing import image.
  28826. 22:22:01This should all look familiar because
  28827. 22:22:02we're going to take a test image and
  28828. 22:22:04we're going to set that equal to in this
  28829. 22:22:06case cat or dog one as you can see over
  28830. 22:22:09here. And you know let me get my drawing
  28831. 22:22:11tool back on. So let's take a look at
  28832. 22:22:13this. We have our test image we're
  28833. 22:22:15loading and in here we have test image
  28834. 22:22:17one. And this one hasn't data hasn't
  28835. 22:22:19seen this one at all. So this is all
  28836. 22:22:20new. Oh, let me shrink the screen down.
  28837. 22:22:22Let me start that over. So here we have
  28838. 22:22:24my test image and we went ahead and the
  28839. 22:22:26cross processing has this nice image
  28840. 22:22:29setup. So we're going to load the image
  28841. 22:22:31and we're going to alter it to a 64x 64
  28842. 22:22:34print. So right off the bat, we're going
  28843. 22:22:36to cross is nice that way. It
  28844. 22:22:37automatically sets it up for us so we
  28845. 22:22:39don't have to redo all our images and
  28846. 22:22:41find a way to reset those. And then we
  28847. 22:22:42use also to set the image to an array.
  28848. 22:22:45So again, we're all in pre-processing
  28849. 22:22:47the data just like we pre-processed
  28850. 22:22:49before with our test information and our
  28851. 22:22:51training data. And then we use the
  28852. 22:22:53numpy. Here's our numpy that's uh from
  28853. 22:22:56our um right up here. Import numpy as in
  28854. 22:22:58p expand the dimensions test image axis
  28855. 22:23:01equal zero. So it puts it into a single
  28856. 22:23:03array. And then finally all that work
  28857. 22:23:07all that pre-processing and all we do is
  28858. 22:23:09we run the result. We click on here we
  28859. 22:23:10go result equals classifier predict test
  28860. 22:23:13image. And then we find out, well, what
  28861. 22:23:15is the test image? And let's just take a
  28862. 22:23:17quick look and just see what that is.
  28863. 22:23:18And you can see when I ran it, it comes
  28864. 22:23:20up dog. And if we look at those images,
  28865. 22:23:23there it is. Cat or dog. Image number
  28866. 22:23:25one. That looks like a nice floppy eared
  28867. 22:23:27lab. Friendly with his tongue hanging
  28868. 22:23:29out. It's either that or a very floppy
  28869. 22:23:31eared cat. I'm not sure which. But
  28870. 22:23:33according to our software, it says it's
  28871. 22:23:34a dog. And uh we have a second picture
  28872. 22:23:36over here. Let's just see what happens
  28873. 22:23:37when we run the second picture. We can
  28874. 22:23:39go up here and change this uh from dog
  28875. 22:23:40image one to two. We'll run that. and it
  28876. 22:23:43comes down here and says cat. You can
  28877. 22:23:45see me highlighting it down there as
  28878. 22:23:47cat. So, our process works. You're able
  28879. 22:23:49to label a dog a dog and a cat a cat
  28880. 22:23:51just from the pictures. There we go.
  28881. 22:23:53Cleared my drawing tool. And the last
  28882. 22:23:55thing I want you to notice when we come
  28883. 22:23:56back up here to when I ran it, you'll
  28884. 22:23:59see it has an accuracy of one and the
  28885. 22:24:02value accuracy of one. Well, the value
  28886. 22:24:05accuracy is the important one because
  28887. 22:24:07the value accuracy is what it actually
  28888. 22:24:09runs on the test data. Remember, I'm
  28889. 22:24:11only testing it on. and I'm only
  28890. 22:24:12validating it on a random 10 photos and
  28891. 22:24:14those 10 photos just happened to come up
  28892. 22:24:16one. Now, when they ran this on the
  28893. 22:24:18server, it actually came up about 86%.
  28894. 22:24:21This is why cutting these numbers down
  28895. 22:24:23so far for a commercial release is bad.
  28896. 22:24:26So, you want to make sure you're a
  28897. 22:24:27little careful of that when you're
  28898. 22:24:28testing your stuff that you change these
  28899. 22:24:30numbers back when you run it on a more
  28900. 22:24:31enterprise computer other than your old
  28901. 22:24:33laptop that you're just practicing on or
  28902. 22:24:35messing with. And we come down here and
  28903. 22:24:37again, you know, we had the validation
  28904. 22:24:38of cat. And so we have successfully
  28905. 22:24:40built a neural network that could
  28906. 22:24:42distinguish between photos of a cat and
  28907. 22:24:44a dog. Imagine all the other things you
  28908. 22:24:46could distinguish. Imagine all the
  28909. 22:24:48different industries you could dive into
  28910. 22:24:50with that. Just being able to understand
  28911. 22:24:51those two difference of pictures. What
  28912. 22:24:53about mosquitoes? Could you find the
  28913. 22:24:55mosquitoes that bite versus the
  28914. 22:24:56mosquitoes that are friendly? It turns
  28915. 22:24:58out the mosquitoes that bite us are only
  28916. 22:25:004% of the mosquito population, if even
  28917. 22:25:02that, maybe 2%. There's all kinds of
  28918. 22:25:04industries that use this and there's so
  28919. 22:25:06many industries that are just now
  28920. 22:25:08realizing how powerful these tools are.
  28921. 22:25:11Just in the photos alone, there is a
  28922. 22:25:13myriad of industries sprouting up. And I
  28923. 22:25:15said it before, I'll say it again. What
  28924. 22:25:17an exciting time to live in with these
  28925. 22:25:19tools and that we get to play with. So
  28926. 22:25:22key takeaways. Well, we covered what is
  28927. 22:25:24a neural network. We use all kinds of
  28928. 22:25:26processing the map images on your phone.
  28929. 22:25:29We talked about things that a neural
  28930. 22:25:30network can do. translate text, identify
  28931. 22:25:33faces all the way to control robots, you
  28932. 22:25:36know, lots of exciting things. How does
  28933. 22:25:38a neural network work? So, we discussed
  28934. 22:25:40that with the different layers going
  28935. 22:25:42from the picture to the input layer to
  28936. 22:25:44the hidden layers and their weights to
  28937. 22:25:46the final output layer. We also talked
  28938. 22:25:48about how it does the math and computing
  28939. 22:25:50the output as yes or no, categorically
  28940. 22:25:54true false. We discussed types of
  28941. 22:25:55artificial neural networks. A lot of
  28942. 22:25:57vocabulary there from the feed forward
  28943. 22:25:59neural network which is the most
  28944. 22:26:01commonly used. That's the one the neural
  28945. 22:26:03network we used is a feed forward neural
  28946. 22:26:04network that does backward propagation
  28947. 22:26:06to train. And there's a lot of other
  28948. 22:26:08ones out there. There's the radial
  28949. 22:26:10biases, the cohen self-organizing
  28950. 22:26:13recurrent neural network, convolution
  28951. 22:26:15neural network, modular neural network.
  28952. 22:26:17The big one was modular because it
  28953. 22:26:19incorporates pieces of all the other
  28954. 22:26:21ones. So that whatever you're working on
  28955. 22:26:23now is a huge conglomerate of multiple
  28956. 22:26:27networks. Just all cutting edge. All of
  28957. 22:26:29it's new. People even working on it
  28958. 22:26:31don't even know where it's going. Again,
  28959. 22:26:32very exciting times. And finally, we dug
  28960. 22:26:35through my favorite part. You can see
  28961. 22:26:37with my uh latte on one side, my old
  28962. 22:26:39school pens and pencil, and all my
  28963. 22:26:41sticky notes working away. That's not
  28964. 22:26:43actually me, by the way. You probably
  28965. 22:26:45guessed that. And we walked through and
  28966. 22:26:46actually did a cat and dog photo, a
  28967. 22:26:48simple cat and dog photo. And you could
  28968. 22:26:50see where some of the problems are in
  28969. 22:26:51processing large amounts of photographs
  28970. 22:26:53and data where that starts to become
  28971. 22:26:55going from a single machine on my laptop
  28972. 22:26:58with this, you know, lower amount of
  28973. 22:27:00resources all the way to big data. How
  28974. 22:27:02if you're processing hundreds and
  28975. 22:27:04thousands of these photos, this now
  28976. 22:27:06needs to be set up on an enterprise
  28977. 22:27:07machine or even on a cluster of
  28978. 22:27:09computers. Again, significantly past the
  28979. 22:27:11scope of this. The neat part about it
  28980. 22:27:13though is once you write this code, most
  28981. 22:27:15of this code, they now have tools that
  28982. 22:27:17you can almost take the same ideas, if
  28983. 22:27:19not the actual code, and push it right
  28984. 22:27:21onto a cluster computation. So really
  28985. 22:27:23cool times for this Python. My name is
  28986. 22:27:26Richard Kersner with the SimplyLearn
  28987. 22:27:28team. That's www.simplearn.com.
  28988. 22:27:30Get certified, get ahead. Although deep
  28989. 22:27:33learning is uh been around for a while,
  28990. 22:27:36it is just in its infant stages of
  28991. 22:27:38development as far as exploding on the
  28992. 22:27:40market. I mean it is right now they're
  28993. 22:27:42building robots with it. Deep learning
  28994. 22:27:44is used to train robots to perform human
  28995. 22:27:47tasks. Music composition. Deep neural
  28996. 22:27:49nets can be used to produce music by
  28997. 22:27:51making computers learn the patterns
  28998. 22:27:53involved in composing music. Image
  28999. 22:27:55colorization. Neural network recognizes
  29000. 22:27:57objects and uses information from the
  29001. 22:27:59images to color them. Machine
  29002. 22:28:01translation. Given a word, phrase or a
  29003. 22:28:03sentence in one language, neural
  29004. 22:28:04networks automatically translate them
  29005. 22:28:06into another language. Google Translate
  29006. 22:28:08is one such popular machine translator
  29007. 22:28:10you may have come across. And you'll
  29008. 22:28:13notice in here we didn't show any
  29009. 22:28:14examples of straight numbers like uh
  29010. 22:28:17projective cells in a business tracking
  29011. 22:28:20your favorite stock. You can certainly
  29012. 22:28:21do those with machine languages, but
  29013. 22:28:23this is the next level. Uh save that for
  29014. 22:28:25your regression models, your linear
  29015. 22:28:27regression where you're actually
  29016. 22:28:28processing and crunching just straight
  29017. 22:28:30numbers. With machine learning and deep
  29018. 22:28:32learning, we're going to a whole new
  29019. 22:28:33level as far as what we can figure out
  29020. 22:28:35on the computer. What's in it for you?
  29021. 22:28:37We're going to cover what is deep
  29022. 22:28:39learning. We're going to take a look at
  29023. 22:28:40the biological versus artificial
  29024. 22:28:42intelligence. What is neural network
  29025. 22:28:44activation function in your neural
  29026. 22:28:46network and the cost function and how do
  29027. 22:28:48neural networks work. How do neural
  29028. 22:28:50networks learn? So there's a little
  29029. 22:28:52you'll see a switch right there. We just
  29030. 22:28:54went from how are they working in the
  29031. 22:28:55math in the background to exactly how
  29032. 22:28:57are they learning. We'll be implementing
  29033. 22:28:58the neural network. We'll do a gradient
  29034. 22:29:00descent deep learning platforms and
  29035. 22:29:02we'll give an introduction to TensorFlow
  29036. 22:29:04and implementation in TensorFlow. That's
  29037. 22:29:06Google's platform that they open sourced
  29038. 22:29:08recently and it's probably one of the
  29039. 22:29:09most cutting edges in deep learning and
  29040. 22:29:12even it is still in the infant stage
  29041. 22:29:13which is one of the reasons they
  29042. 22:29:15released it to open source. What is deep
  29043. 22:29:17learning? Deep learning is a sub field
  29044. 22:29:19of machine learning that deals with
  29045. 22:29:20algorithms inspired by the structure and
  29046. 22:29:22function of the brain. And you can see
  29047. 22:29:24we have a nice picture here. We have
  29048. 22:29:25artificial intelligence which is kind of
  29049. 22:29:27the big bubble that encompasses all
  29050. 22:29:29these different things we're talking
  29051. 22:29:30about. This is ability of machine to
  29052. 22:29:32imitate intelligent human behavior. And
  29053. 22:29:34in there we have machine learning
  29054. 22:29:36application of AI that allows a system
  29055. 22:29:38to automatically learn and improve from
  29056. 22:29:41experience. And if you looked at any of
  29057. 22:29:42our other videos, you'll know that
  29058. 22:29:44machine learning covers a lot. So deep
  29059. 22:29:46learning is a subcategory of that. But
  29060. 22:29:48don't forget machine learning has all
  29061. 22:29:50kinds of other tools that people use to
  29062. 22:29:52do very basic uh descriptive and
  29063. 22:29:54predictive and postcriptive uh
  29064. 22:29:56analytics. And then you have deep
  29065. 22:29:58learning application of machine learning
  29066. 22:30:00that uses complex algorithms and deep
  29067. 22:30:03neural nets to train a model. Let's take
  29068. 22:30:05a look at the biological neuron versus
  29069. 22:30:07the artificial neuron. Now remember in
  29070. 22:30:09the human brain and and this is true for
  29071. 22:30:11most animals there are a lot of
  29072. 22:30:12different neurons going on. So this is
  29073. 22:30:14the very basic one. I mean there's
  29074. 22:30:16hundreds of different cells involved. So
  29075. 22:30:19when we talk about neural networks and
  29076. 22:30:20this is why I say it's in a very infant
  29077. 22:30:22stage. They're really basing it on uh
  29078. 22:30:24just the most basic thing that we're
  29079. 22:30:26able to figure out going on in the
  29080. 22:30:28neural networks. And you can see right
  29081. 22:30:30here we have dendrites fetch information
  29082. 22:30:32from an adjacent neurons and pass them
  29083. 22:30:34on as inputs. So you have your data
  29084. 22:30:36coming in and your data going out. Any
  29085. 22:30:39computer model should be looking at that
  29086. 22:30:40what's coming in what's going out. The
  29087. 22:30:42data is fed as an input to the neuron.
  29088. 22:30:44So we look at the artificial neuron. You
  29089. 22:30:47can see we have our inputs. They come in
  29090. 22:30:49each one is specially weighted into the
  29091. 22:30:51neuron and then the neuron has an
  29092. 22:30:52output. The cell nucleus processes the
  29093. 22:30:55information received from the dendrites
  29094. 22:30:57and the neuron processes the information
  29095. 22:30:59provided as inputs. Axons are the cables
  29096. 22:31:01over which the information is
  29097. 22:31:03transmitted and the information is
  29098. 22:31:05transferred over weighted channels. So
  29099. 22:31:06you can look at that uh I mentioned
  29100. 22:31:08weights briefly but you alter the data
  29101. 22:31:10coming in. So those weights are what
  29102. 22:31:12causes different information coming in
  29103. 22:31:15to be weighted differently and processed
  29104. 22:31:17differently. And the synapses receive
  29105. 22:31:19the information from the axons and
  29106. 22:31:20transmit it to the adjacent neurons.
  29107. 22:31:22That's in your biological model. And
  29108. 22:31:24then when we look at the artificial
  29109. 22:31:25neuron, the output is a final value
  29110. 22:31:27predicted by the artificial neuron. So
  29111. 22:31:29as we dig deeper into looking at the
  29112. 22:31:31theory behind the neural network and we
  29113. 22:31:34kind of flip back and forth between
  29114. 22:31:35these because there's two huge aspects
  29115. 22:31:38of it. One is from the outside. What are
  29116. 22:31:40you seeing and what's going on from the
  29117. 22:31:42inside so you can find to do what you
  29118. 22:31:44need to do and give the best results you
  29119. 22:31:46can. And we start off with what do we
  29120. 22:31:47feed? We feed an unlabeled image to a
  29121. 22:31:50machine which identifies it without any
  29122. 22:31:52human intervention. And so you can see
  29123. 22:31:54here we have a circle that comes in at
  29124. 22:31:56784 pixels and it comes in by 28x 28.
  29125. 22:31:59And you can see how it colors in the um
  29126. 22:32:01the circle on there. And we put a
  29127. 22:32:03triangle in. The triangle in also comes
  29128. 22:32:05in as 28x 28 and it has 784 pixels. So
  29129. 22:32:08you'll see between these two both of
  29130. 22:32:10them are 784 pixels. This machine is
  29131. 22:32:13intelligent enough to differentiate
  29132. 22:32:15between the various shapes. So that's
  29133. 22:32:17what we want to use our neural network
  29134. 22:32:18to do is to say hey this is a circle.
  29135. 22:32:20This is a triangle. That's more of a
  29136. 22:32:22categorical. You can also do a
  29137. 22:32:23regression model where you're actually
  29138. 22:32:25putting out float value or a numerical
  29139. 22:32:27value. We'll be looking at the true
  29140. 22:32:28false or the categorical model mostly
  29141. 22:32:30because that's where you usually start
  29142. 22:32:31at the different there is no real
  29143. 22:32:33difference when you as far as the way
  29144. 22:32:36the internal functioning goes when you
  29145. 22:32:38start flipping between them other than
  29146. 22:32:40well we'll talk about that in just a
  29147. 22:32:41minute. So you can actually go between
  29148. 22:32:42the two quite easily and the neural
  29149. 22:32:44network provides this capability. So
  29150. 22:32:46we're going to use this capability to
  29151. 22:32:48look between those two. One of the
  29152. 22:32:49things I want you to note in here is
  29153. 22:32:51that we're looking at 784 pixels. We're
  29154. 22:32:54looking at 784 inputs. That's very
  29155. 22:32:57different than stock with a high low or
  29156. 22:32:59last year's sales based on date or we're
  29157. 22:33:02looking at just a couple of numbers and
  29158. 22:33:04they're very clear. They're numbers.
  29159. 22:33:06They're very clear what they are, which
  29160. 22:33:07is something you'd put into a machine
  29161. 22:33:09learning linear regression model. This
  29162. 22:33:10is a step up from that in that we're
  29163. 22:33:12looking at complex patterns and how do
  29164. 22:33:14you figure those complex patterns out.
  29165. 22:33:16So, a neural network is a system modeled
  29166. 22:33:19on the human brain. And we looked at
  29167. 22:33:21that comparing the two. Let's go ahead
  29168. 22:33:22and look deeper into the neural network
  29169. 22:33:24itself. We have our inputs coming in. So
  29170. 22:33:26the inputs are fed to a neuron that
  29171. 22:33:28processes a data and gives us an output.
  29172. 22:33:30Input and output. This is the most basic
  29173. 22:33:33structure of a neural network known as a
  29174. 22:33:35perceptron. So if you see the term
  29175. 22:33:37perceptron, that's what we're talking
  29176. 22:33:38about. We're talking about this single
  29177. 22:33:40node that has inputs and an output.
  29178. 22:33:42However, neural networks are usually
  29179. 22:33:44much more complex. Let's start with
  29180. 22:33:47visualizing a neural network as a black
  29181. 22:33:49box. And I always love that symbol. It's
  29182. 22:33:51a black box. It's kind of magical. We
  29183. 22:33:53have our inputs coming in and we want
  29184. 22:33:55certain outputs. The box takes inputs,
  29185. 22:33:57processes them, and gives an output.
  29186. 22:34:00Let's have a look at what happens within
  29187. 22:34:02this box. And you can see me there in my
  29188. 22:34:04uh secret agent getup and I got my
  29189. 22:34:06hidden hood and everything. I guess I'm
  29190. 22:34:08part of the uh black skull or something
  29191. 22:34:10like that group. Uh so let's take a look
  29192. 22:34:12at what happens within this magic box.
  29193. 22:34:14And remember, we're skipping back and
  29194. 22:34:15forth between the theory of what's going
  29195. 22:34:17on in the box, which you have to know
  29196. 22:34:19how to fine-tune and how to build,
  29197. 22:34:21versus looking at it from the outside.
  29198. 22:34:23We're programming this box, and we have
  29199. 22:34:25an input and an output to the box as a
  29200. 22:34:27whole. Within the box exists a network
  29201. 22:34:29that is a core of deep learning. And you
  29202. 22:34:31can see here we're showing one layer and
  29203. 22:34:33we have our grid coming in. The network
  29204. 22:34:35consists of layers of neurons. Each
  29205. 22:34:38neuron is associated with a number
  29206. 22:34:40called the bias. And you can think of
  29207. 22:34:42the bias uh if you overly simplify this
  29208. 22:34:45and we're doing a linear regression
  29209. 22:34:46model. This is your y intercept in your
  29210. 22:34:49uklidian geometry. You have to have
  29211. 22:34:50something that offsets it. And so you
  29212. 22:34:52always have a bias in these cells.
  29213. 22:34:54Neurons of each layer transmit
  29214. 22:34:56information to neurons of the next layer
  29215. 22:34:58over channels. And so you can see each
  29216. 22:35:00of our layers going through from left to
  29217. 22:35:02right. These channels are associated
  29218. 22:35:04with numbers called weights. These
  29219. 22:35:06weights along with the biases determine
  29220. 22:35:08the information that is passed over from
  29221. 22:35:10the neuron to neuron. So just like the
  29222. 22:35:12bias is your y intercept in uklidian
  29223. 22:35:15geometry. You could look at the an one
  29224. 22:35:17weight. Remember this is very
  29225. 22:35:19complicated. So we're not looking at
  29226. 22:35:20just one weight. You could look at the
  29227. 22:35:21weight as your slope of the line. Or if
  29228. 22:35:23you're doing x= uh my y + c, it would be
  29229. 22:35:27the m value. Neurons of each layer
  29230. 22:35:29transmit information to neurons of the
  29231. 22:35:31next layer. And you can see here as they
  29232. 22:35:32light up going across into the final
  29233. 22:35:34layer. and then to the output. And in
  29234. 22:35:37this case, the output is going to be
  29235. 22:35:38either uh a square in this one or it
  29236. 22:35:41might light up the other one which is a
  29237. 22:35:42circle. The output layer emits a
  29238. 22:35:44predicted output. So in this case, we're
  29239. 22:35:46looking at a classification uh true
  29240. 22:35:49false. Is it a circle? Is it a triangle?
  29241. 22:35:51Is it a square? Let's now go deeper.
  29242. 22:35:53What happens within the neuron? So we're
  29243. 22:35:55going to dig deeper and start getting a
  29244. 22:35:57little bit closer to some of the math.
  29245. 22:35:58Don't worry, you don't have to be a
  29246. 22:36:00calculus expert and know your
  29247. 22:36:01differential equations. Even though this
  29248. 22:36:03is one giant differential equation, you
  29249. 22:36:06don't need to understand those to
  29250. 22:36:07understand what's going on. Within each
  29251. 22:36:09neuron, the following operations are
  29252. 22:36:11performed. The product of each input and
  29253. 22:36:13the weight of the channel it's passed
  29254. 22:36:15over is found. This is simply addition.
  29255. 22:36:17We're going to sum up the weight times
  29256. 22:36:19the output from the previous channel and
  29257. 22:36:21plus the bias. Sum of the weighted
  29258. 22:36:23products is computed. This is called the
  29259. 22:36:25weighted sum. Bias unique to the neuron
  29260. 22:36:27is added to the weighted sum. The final
  29261. 22:36:29sum is then subjected to the particular
  29262. 22:36:32function and we'll discuss those that
  29263. 22:36:34particular function. That part is really
  29264. 22:36:36important because those functions uh
  29265. 22:36:38have a huge impact on how well your
  29266. 22:36:40model performs under different
  29267. 22:36:42conditions. The final sum is then
  29268. 22:36:43subject to a particular function. This
  29269. 22:36:45is the activation function. So if you
  29270. 22:36:48ever hear the term activation function,
  29271. 22:36:50that's what we're talking about. What
  29272. 22:36:51activates this cell and what doesn't. As
  29273. 22:36:53we dig deeper into activation function,
  29274. 22:36:56an activation function takes the
  29275. 22:36:57weighted sum of the input as its input
  29276. 22:37:00adds a bias and provides an output. And
  29277. 22:37:03a lot of times you'll actually see one
  29278. 22:37:05formula for the sum of the weight the
  29279. 22:37:07weighted sum and the bias. You'll just
  29280. 22:37:09see that as a single line of everything
  29281. 22:37:10added together. And here we've broken it
  29282. 22:37:12apart because it makes it clear that
  29283. 22:37:14this bias is not computed the same as
  29284. 22:37:16the weighted sums. Here are the most
  29285. 22:37:18popular types of activation function.
  29286. 22:37:20And I always find these interesting
  29287. 22:37:22because at one point I was sitting at a
  29288. 22:37:24table with a gentleman who was finishing
  29289. 22:37:26his PhD. He was in his last year and he
  29290. 22:37:28said he went through all this stuff and
  29291. 22:37:30he ended up just trying the four
  29292. 22:37:32different activation functions on this
  29293. 22:37:34particular problem he was working on. So
  29294. 22:37:36knowing the math behind it doesn't
  29295. 22:37:38necessarily mean you're going to know it
  29296. 22:37:40right away. Uh so even somebody who
  29297. 22:37:42might have a PhD and be doing the
  29298. 22:37:43calculations on this comes back out of
  29299. 22:37:46it and ends up just trying the different
  29300. 22:37:48uh um activation functions to see what's
  29301. 22:37:50going to make a difference. And a lot of
  29302. 22:37:51times that's a final step. That's the
  29303. 22:37:53kind of thing where you built your whole
  29304. 22:37:54model. You've come back and you're like
  29305. 22:37:56wait a minute can I do a better deal
  29306. 22:37:58with a sigmoid function or the threshold
  29307. 22:38:00or the rectifier. Knowing what they're
  29308. 22:38:02doing is important so you can explain it
  29309. 22:38:03to somebody else. And again you probably
  29310. 22:38:05do this on a small set of data. If
  29311. 22:38:07you're working with big data, uh you
  29312. 22:38:08don't want to take down the full server
  29313. 22:38:10farm just to test out your three
  29314. 22:38:12different series. You take a small
  29315. 22:38:14portion of that data, test it, and then
  29316. 22:38:16you put it through to the big data. So
  29317. 22:38:17let's take a look at this. We have the
  29318. 22:38:19sigmoid function, and it's used for
  29319. 22:38:20models where we have to predict the
  29320. 22:38:22probability as an output. It exists
  29321. 22:38:24between zero and one. And you'll see
  29322. 22:38:26that's true of all of our activation
  29323. 22:38:28functions we're working with. Either the
  29324. 22:38:30cells on or off, it's true or false. And
  29325. 22:38:33there might be a little variation in
  29326. 22:38:34there which as an output could be used
  29327. 22:38:37to compute uncertainty in your solution.
  29328. 22:38:40So if you're getting a 7 with this
  29329. 22:38:42activation function, it might be well
  29330. 22:38:43I'm not sure if that's really a square
  29331. 22:38:45or I'm not sure that's really a
  29332. 22:38:46triangle. And that might be a flag for
  29333. 22:38:49it to be looked at by a human observer
  29334. 22:38:51at least in today's models where we're
  29335. 22:38:53at right now. And you can see here we
  29336. 22:38:54have the formula is simply equals 1 over
  29337. 22:38:571 + e the minus x where x is your value
  29338. 22:39:00coming in. and it's going to give you a
  29339. 22:39:02result that looks very similar to the
  29340. 22:39:03graph on there which is somewhere
  29341. 22:39:04between zero and one. Um, and right in
  29342. 22:39:07the middle you can see that there's a
  29343. 22:39:09huge uh kind of you can go through all
  29344. 22:39:11the different values and uncertainties
  29345. 22:39:13involved. So the sigmoid function is
  29346. 22:39:15probably the default on most of them. Uh
  29347. 22:39:17the next one is the threshold function.
  29348. 22:39:19It is a threshold-based activation
  29349. 22:39:20function. If x value is greater than a
  29350. 22:39:23certain value, the function is activated
  29351. 22:39:24and fired. Else not. Pretty
  29352. 22:39:26straightforward. Yes, no, true, false.
  29353. 22:39:29um I don't want to test for
  29354. 22:39:30improbabilities. I just want a straight
  29355. 22:39:32answer. I don't want to know if there's
  29356. 22:39:34a partial value on there. It either is
  29357. 22:39:36true or it's false. And the rectifier
  29358. 22:39:38function, it is the most widely used
  29359. 22:39:40activation function. I would debate
  29360. 22:39:42that. Um rectifier is pretty common one,
  29361. 22:39:45although I see that the sigmoid function
  29362. 22:39:46is used to be the basic one, but it's up
  29363. 22:39:48there. The rectifier function is very
  29364. 22:39:50commonly used. You get the output of X
  29365. 22:39:52if X is positive and zero otherwise. And
  29366. 22:39:54you can see here again just like um uh
  29367. 22:39:57it's either you kind of get a value
  29368. 22:39:59going up there. So max of x of zero. So
  29369. 22:40:02it's it's again it's like the threshold
  29370. 22:40:03function. Yes, no, true, false. Uh it's
  29371. 22:40:06either zero or it's uh some kind of
  29372. 22:40:08progressive value. And then we have the
  29373. 22:40:10rectifier function. I would argue with
  29374. 22:40:12this because the sigmoid function used
  29375. 22:40:13to be the most common one. But with the
  29376. 22:40:15rectifier function, it now says it is
  29377. 22:40:17the most commonly used or widely used
  29378. 22:40:19activation function and gives an output
  29379. 22:40:21of X if X is positive and zero
  29380. 22:40:24otherwise. This is kind of nice because
  29381. 22:40:26it now says absolutely not or it gives
  29382. 22:40:29you a value of probability. Now, when I
  29383. 22:40:31say a value of probability, be very
  29384. 22:40:33careful there. I'm not saying that it's
  29385. 22:40:34going to tell you this is 75% chance of
  29386. 22:40:36being a circle. I'm going to tell you
  29387. 22:40:38that it says, hey, if this says 0.1, it
  29388. 22:40:41probably needs to be looked at or 2 or
  29389. 22:40:433. It's going to depend on your data as
  29390. 22:40:45to what that value means. In general,
  29391. 22:40:48that just means it's flagging it that if
  29392. 22:40:49it's not a one, then chances are it
  29393. 22:40:51needs to be looked at by a person and
  29394. 22:40:53re-evaluated. And there's a hyperbolic
  29395. 22:40:55tangent function. This function is
  29396. 22:40:57similar to sigmoid function is bound to
  29397. 22:40:59a range of minus1 to 1. So you can see
  29398. 22:41:01there's our 1 - eus 2x and 1 plus over 1
  29399. 22:41:05+ eus 2x. Again, it's very similar to
  29400. 22:41:07the sigmoid function. The bonus of the
  29401. 22:41:10hyperbolic function is you have that
  29402. 22:41:12variable coming through the middle. So
  29403. 22:41:14again, you can look at it and you have a
  29404. 22:41:16little bit more weight as far as you can
  29405. 22:41:18process that down the line. That's a
  29406. 22:41:20little bit more advanced than than what
  29407. 22:41:21we're looking at right now. And a lot of
  29408. 22:41:22times it's not even necessary in a lot
  29409. 22:41:24of our different uh uses for these
  29410. 22:41:26activation functions. Now, we looked at
  29411. 22:41:28activation functions and I kind of said
  29412. 22:41:30those are a little bit like a black box
  29413. 22:41:32because even if you know all the math, a
  29414. 22:41:35lot of times you end up just playing
  29415. 22:41:36with them to find out what works. And it
  29416. 22:41:38also depends on what model you're
  29417. 22:41:39working with, whether you need a flat
  29418. 22:41:40yes, no, true, false, or you need to
  29419. 22:41:42have something in the middle that says,
  29420. 22:41:44hey, this isn't quite a one. You might
  29421. 22:41:46need to process this with the human
  29422. 22:41:47intervention. And you could look at
  29423. 22:41:49that. Uh, one example would be
  29424. 22:41:50self-driving cars. You don't want a car
  29425. 22:41:52to be yes, no, I'm going to go through
  29426. 22:41:54the the light. You want it to be like,
  29427. 22:41:56okay, if it's uh almost yes, maybe we
  29428. 22:41:58stop and have human intervention so we
  29429. 22:42:00don't get an accident. Cost function is
  29430. 22:42:02something you can really see and measure
  29431. 22:42:04and is very important. The cost value is
  29432. 22:42:07the difference between the neural net's
  29433. 22:42:08predicted output and the actual output
  29434. 22:42:11from a set of labeled training data. So
  29435. 22:42:13we have our group of data that's a
  29436. 22:42:15square circle and since we're looking at
  29437. 22:42:17geometrical shapes, we've had somebody
  29438. 22:42:19already labeled that data. They've
  29439. 22:42:21already said this is a triangle, this is
  29440. 22:42:23a square. And so if this is coming up
  29441. 22:42:25and it's giving us and it's saying a
  29442. 22:42:26square is a triangle and it's saying a
  29443. 22:42:28triangle is a circle, the output is
  29444. 22:42:30wrong. And so that output can then be
  29445. 22:42:33measured in the versus the actual output
  29446. 22:42:35and that's the cost. Uh you might also
  29447. 22:42:37hear this as error because that's the
  29448. 22:42:39error value being returned. How far off
  29449. 22:42:41is it? And what we're looking for is the
  29450. 22:42:43least cost or the least error value. And
  29451. 22:42:45it's obtained by making adjustments to
  29452. 22:42:47the weights and biases iteratively
  29453. 22:42:49throughout the training process. And
  29454. 22:42:51this this is called back propagation.
  29455. 22:42:54And we're going to look in that a little
  29456. 22:42:55deeper as we look into an example. It's
  29457. 22:42:57really hard to see when you're just
  29458. 22:42:58looking at arrows without actual numbers
  29459. 22:43:01and where that flow is coming from. But
  29460. 22:43:02you can look at this is here's our
  29461. 22:43:04inputs. They put out a prediction. The
  29462. 22:43:06prediction comes out and says, "Hey,
  29463. 22:43:08we've already labeled this data cuz
  29464. 22:43:09we're in training mode and the training
  29465. 22:43:11data is off. This is the cost. Can we
  29466. 22:43:13send that error or that cost back and
  29467. 22:43:16adjust those weights?" And we do it in
  29468. 22:43:18very small increments across large
  29469. 22:43:20amounts of data so that those weights
  29470. 22:43:23minimize that cost or that error. But
  29471. 22:43:25what happens within these neurons? So
  29472. 22:43:28let's look at a little example of this.
  29473. 22:43:29Kind of helps if you have some kind of
  29474. 22:43:31visual. Let's build a neural network
  29475. 22:43:32that predict bike prices based on a few
  29476. 22:43:35of its features. And we'll see here we
  29477. 22:43:37have our CC, our mileage, and our ABS.
  29478. 22:43:39And these are our three input layers.
  29479. 22:43:41And then we have the bike price and the
  29480. 22:43:43output layer. Now, it doesn't do us very
  29481. 22:43:45good to just uh pump it in from the
  29482. 22:43:47beginning and pump it out. And to be
  29483. 22:43:48honest, I would use a machine learning
  29484. 22:43:50linear regression model on this since
  29485. 22:43:52these are just straight numbers. But
  29486. 22:43:54because we want a simple example, we're
  29487. 22:43:57going to put this through and show you
  29488. 22:43:58as a neural network what that looks
  29489. 22:43:59like. And we got to put a hidden layer
  29490. 22:44:01in there. The hidden layer helps in
  29491. 22:44:02improving the output accuracy. And you
  29492. 22:44:05could look at this as a bunch of ores.
  29493. 22:44:08So it might say, hey, when we compare
  29494. 22:44:10these three values on the first hidden
  29495. 22:44:13layer neuron, we're looking at one set
  29496. 22:44:15of features and then we might weight
  29497. 22:44:17them in the second one. So these are a
  29498. 22:44:18bunch of different ores kind of how the
  29499. 22:44:20math comes out in behind the scenes. And
  29500. 22:44:22then they go out of course to the bike
  29501. 22:44:23trace or the output layer. And each of
  29502. 22:44:26the connections have a weight assigned
  29503. 22:44:27with it. And you'll see here we have a
  29504. 22:44:30mileage CC with the weight one and
  29505. 22:44:31weight two going into our first neuron.
  29506. 22:44:34And you'd also have your ABS going in
  29507. 22:44:35there. And so X1 * weight 1 + X2 *
  29508. 22:44:39weight 2 plus the bias of one. And step
  29509. 22:44:42two is our activation. The activation
  29510. 22:44:44function coming in there. When does this
  29511. 22:44:45fire? And the neuron takes a subset of
  29512. 22:44:48the inputs and processes it. And then we
  29513. 22:44:50go through and we do that with the um
  29514. 22:44:52second hidden layer neuron and the third
  29515. 22:44:54one and so on. So you process each layer
  29516. 22:44:56in order going forward. Now when I told
  29517. 22:44:59you this is in its infant stage, they
  29518. 22:45:01now have neurons that fire into the same
  29519. 22:45:04layer or back a layer so that you now
  29520. 22:45:06have a time series and there's all kinds
  29521. 22:45:08of wild things that they're
  29522. 22:45:09experimenting with on these layers. This
  29523. 22:45:11basic setup has been around since the
  29524. 22:45:13mid90s. It's only now because of our
  29525. 22:45:16technology that it's open to almost
  29526. 22:45:18everybody to play with it. And that's
  29527. 22:45:19why I say this is in an infant stage in
  29528. 22:45:21development is this basic math is here,
  29529. 22:45:24but what we can do with it is amazing.
  29530. 22:45:26And what they're actually doing with all
  29531. 22:45:27these different things is amazing. And
  29532. 22:45:28so we're just at the beginning of how to
  29533. 22:45:30use all these different tools and our
  29534. 22:45:32deep learning and our neural networks.
  29535. 22:45:34Uh and so once we have our hidden layer
  29536. 22:45:35computed, the information reaching the
  29537. 22:45:37neurons in the hidden layer is subjected
  29538. 22:45:39to the respective activation function.
  29539. 22:45:41And so each one of these fires an
  29540. 22:45:43activation output uh and then those are
  29541. 22:45:45each weighted to the final output layer.
  29542. 22:45:47So the processed information is now sent
  29543. 22:45:50to the output layer once again over
  29544. 22:45:52weighted channels. And you could look at
  29545. 22:45:54this as each one of these is um I always
  29546. 22:45:56look at this as like a group of people.
  29547. 22:45:58They're all looking at the bulletin
  29548. 22:45:59board and the first person says this is
  29549. 22:46:00what I project sales for the company and
  29550. 22:46:02the second person and the third and so
  29551. 22:46:04on. And then their perspectives are
  29552. 22:46:06weighted based on their expertise. So
  29553. 22:46:08your accountant might have a very high
  29554. 22:46:10weight where the um maybe your janitor
  29555. 22:46:12has a very low weight because their
  29556. 22:46:14expertise is not in accounting and then
  29557. 22:46:16that goes into the output layer and once
  29558. 22:46:18in the output layer it goes uh the
  29559. 22:46:20output which is the predicted value is
  29560. 22:46:22compared against the original value. So
  29561. 22:46:24now we have our output layer and since
  29562. 22:46:26we have like already a list of uh bikes
  29563. 22:46:29with their the different setups and what
  29564. 22:46:31their value is we can now generate an
  29565. 22:46:33error from this. The cost function
  29566. 22:46:35determines the error in prediction and
  29567. 22:46:37reports it back to the neural network.
  29568. 22:46:40So this is the cost. This is how far off
  29569. 22:46:42it is. This is your error coming back.
  29570. 22:46:44And as you can see, this is back
  29571. 22:46:46propagation going on. So now our error
  29572. 22:46:48is going in reverse because we know
  29573. 22:46:50we're not completely correct on this
  29574. 22:46:53particular channel. The weights are
  29575. 22:46:55adjusted in order to reduce the error.
  29576. 22:46:57So each time we go back, we are changing
  29577. 22:46:59those weights to reduce that error. and
  29578. 22:47:01we change them in small increments. You
  29579. 22:47:04don't want to fit one input. Remember,
  29580. 22:47:07you might have a data pool with a
  29581. 22:47:09terabyte of data. You don't want to
  29582. 22:47:11solve for the first set of data that
  29583. 22:47:12comes in and that be the main solution
  29584. 22:47:14because everything else will be off.
  29585. 22:47:16This is going to confuse you. That's
  29586. 22:47:17also called a bias. So, we have the bias
  29587. 22:47:20in the cell where we're adding a value,
  29588. 22:47:22the kind of like the y intercept, and we
  29589. 22:47:24have a bias of the whole neural network,
  29590. 22:47:27which means that it's weighted towards
  29591. 22:47:28one set of answers. So we want to make
  29592. 22:47:31small changes in these weights so we
  29593. 22:47:32don't create a bias and the weights are
  29594. 22:47:34adjusted in order to reduce the error or
  29595. 22:47:36the cost. The network is now trained
  29596. 22:47:38using the new weights. Once again the
  29597. 22:47:40cost is determined and back propagation
  29598. 22:47:42is continued until the cost cannot be
  29599. 22:47:44reduced any further. So let's go ahead
  29600. 22:47:46and plug in values and see how our
  29601. 22:47:48neural network works. So here we come in
  29602. 22:47:51here and initially our channels are
  29603. 22:47:52assigned with random weights. This is
  29604. 22:47:55important because if you assign them all
  29605. 22:47:57with the same weight, you might be able
  29606. 22:47:58to reproduce it. But it turns out that
  29607. 22:48:01if I put all my weights as one or all my
  29608. 22:48:03weights as zero, it takes longer to
  29609. 22:48:05train where if you have random weights,
  29610. 22:48:06they already have like a little bit of
  29611. 22:48:08adjustment and ores built in and that
  29612. 22:48:10will give us a better answer and train
  29613. 22:48:12faster. Our first neuron takes a value
  29614. 22:48:14of mileage and CC as inputs. So here
  29615. 22:48:17comes our computation whatever those
  29616. 22:48:19inputs are. And we do that again with
  29617. 22:48:21the second neuron with those values
  29618. 22:48:23coming in. You can see here we have
  29619. 22:48:24weight three and so on and then our
  29620. 22:48:26third neuron coming down and of course
  29621. 22:48:28our fourth neuron. So we're adding all
  29622. 22:48:30these different values coming in here in
  29623. 22:48:32our hidden layer. The process value from
  29624. 22:48:34each neuron is sent to the output layer
  29625. 22:48:36over weighted channels. So again here's
  29626. 22:48:38our weights coming in and we have N1,
  29627. 22:48:40N2, N3 and N4. Once again the values are
  29628. 22:48:43subjected to the activation function and
  29629. 22:48:45a single value is emitted as the output.
  29630. 22:48:47On comparing the predicted value to the
  29631. 22:48:49actual value, we clearly see that our
  29632. 22:48:51network requires training. So, here we
  29633. 22:48:53have it that our bike price uh we put
  29634. 22:48:55out, we thought it was worth 2,000 on
  29635. 22:48:57our random weights and the bike actually
  29636. 22:48:59was $4,000 on there. Guessing that's not
  29637. 22:49:02US dollars cuz that'd be a very
  29638. 22:49:04expensive bike. But maybe it is. There's
  29639. 22:49:05some $2,000 $4,000 bikes out there. The
  29640. 22:49:08cost function is calculated and back
  29641. 22:49:10propagation takes place. And this is
  29642. 22:49:12pretty simple. You can look at that as
  29643. 22:49:14our um we're subtracting one value from
  29644. 22:49:16the other. We square it and then we take
  29645. 22:49:18half of that and that is propagated back
  29646. 22:49:20up. And each layer generates its own
  29647. 22:49:23errors. Let's go back one because you
  29648. 22:49:24have your predicted Y and your actual Y.
  29649. 22:49:27That goes back to the first layer. And
  29650. 22:49:29then based on the value of the cost
  29651. 22:49:31function, certain weights are changed.
  29652. 22:49:33So when we look at the next layer, that
  29653. 22:49:35error is not the original 4,000 - 2,000
  29654. 22:49:39squar / 2. This error is based on the
  29655. 22:49:42error of each cell generated. How far
  29656. 22:49:45off is that cell as far as its weights.
  29657. 22:49:47We're not going to show you. It's
  29658. 22:49:48actually a very complicated differential
  29659. 22:49:50equation. And you can probably write it
  29660. 22:49:51out if you wanted to. You just write out
  29661. 22:49:53each formula that goes into the next
  29662. 22:49:55level and you add them all together and
  29663. 22:49:56you can write it out all the way
  29664. 22:49:57through. Computers make it so you don't
  29665. 22:49:59have to. And our neural network is
  29666. 22:50:01considered trained when the value for
  29667. 22:50:02the cost function is minimum. So when we
  29668. 22:50:05get our error way down as low as we can,
  29669. 22:50:07that's when our neural network is
  29670. 22:50:08trained. And there I mean just recently
  29671. 22:50:10they've come up with all kinds of
  29672. 22:50:12different means for measuring that
  29673. 22:50:14particular value. a little bit beyond
  29674. 22:50:16the scope of today's neural network, but
  29675. 22:50:18you can actually you can actually see
  29676. 22:50:19it. You know, how far do you do this
  29677. 22:50:21until the neural network doesn't need to
  29678. 22:50:23be trained anymore and you can overtrain
  29679. 22:50:25a neural network. Now, the tools that
  29680. 22:50:27we're looking at automatically let you
  29681. 22:50:29know when to stop, which is really nice.
  29682. 22:50:31And that is just like I said, we're at
  29683. 22:50:33the beginning stages in neural networks
  29684. 22:50:34and it's just really cool what they can
  29685. 22:50:36do now and how much of it's automated
  29686. 22:50:38and how much of it is experimental.
  29687. 22:50:40Right now, let's take a look at gradient
  29688. 22:50:41descent. But what approach do we take to
  29689. 22:50:44minimize the cost function? So here we
  29690. 22:50:46have nice error thing coming in. This is
  29691. 22:50:49our cost or our error. Uh let's start
  29692. 22:50:51with plotting the cost function against
  29693. 22:50:53the predicted value. And so you can see
  29694. 22:50:55they fed in multiple y's and these are
  29695. 22:50:58the errors coming in and the cost of
  29696. 22:51:00each of these inputs and changes going
  29697. 22:51:02on. Note we start at a random point on
  29698. 22:51:05the curve. So usually you put in you
  29699. 22:51:08know you pick up your data and you
  29700. 22:51:09randomly pick where to start in your
  29701. 22:51:11data. A lot of times you just run it
  29702. 22:51:12from the beginning because you're going
  29703. 22:51:13through so much data, it's not that big
  29704. 22:51:15of a deal. But you start with one point
  29705. 22:51:17going in. So your forward propagation
  29706. 22:51:19goes through. You're going to go ahead
  29707. 22:51:20and find your cost or your error. It
  29708. 22:51:22points that on the curve. And you can
  29709. 22:51:24see how we're plotting it right here.
  29710. 22:51:25Since the gradient at this point is
  29711. 22:51:27positive, we may move right. So we're
  29712. 22:51:29going to move a little bit to the right
  29713. 22:51:30on here. And this time the gradient is
  29714. 22:51:32negative. We move a little bit to the
  29715. 22:51:33left. Eventually we try out the point
  29716. 22:51:36where the gradient is zero. This is a
  29717. 22:51:38least value of cost function. You have
  29718. 22:51:41to be a little careful with this because
  29719. 22:51:43this particular I mean they make it look
  29720. 22:51:45nice and simple in this graph. Sometimes
  29721. 22:51:48these curves look like stair steps and
  29722. 22:51:50so there is global minimums and then
  29723. 22:51:53there is local there might be a local
  29724. 22:51:55point where the gradient is zero but
  29725. 22:51:57it's not the global one. Uh so it might
  29726. 22:51:59be way off to the left where it just
  29727. 22:52:01happens to step down a little bit and
  29728. 22:52:02you think you're in the right gradient.
  29729. 22:52:04And with that we have all the right
  29730. 22:52:05weights and we can say our network is
  29731. 22:52:07trained. So here we have um just some
  29732. 22:52:10major these are some of the big names
  29733. 22:52:12out there right now in development for
  29734. 22:52:14deep learning platforms. TensorFlow
  29735. 22:52:16which we'll actually do an example in in
  29736. 22:52:18a minute. Deep learning for J which is
  29737. 22:52:20in the Java platform. Uh so if you're a
  29738. 22:52:22Java programmer uh by the way is
  29739. 22:52:24TensorFlow is accessed most people are
  29740. 22:52:26using Python to access it but it is a
  29741. 22:52:28system that's kind of separate from a
  29742. 22:52:29lot of the programming languages which
  29743. 22:52:31makes it a lot more um flexible as far
  29744. 22:52:33as use. Deep learning forj is Java based
  29745. 22:52:36and then cross is just exploding right
  29746. 22:52:38now. And this is interesting. Cross is
  29747. 22:52:40uh working with TensorFlow. It actually
  29748. 22:52:42can sit on top of TensorFlow. And it can
  29749. 22:52:44also do its own thing. Uh so if you're
  29750. 22:52:47studying deep learning, you're getting
  29751. 22:52:48into it, you want to know the basics of
  29752. 22:52:50TensorFlow, but you also are going to
  29753. 22:52:52want to know the upper level of KAS
  29754. 22:52:53sitting on top of TensorFlow. We're just
  29755. 22:52:55looking at TensorFlow today though in
  29756. 22:52:57our example. And there's also Torch on
  29757. 22:52:59there. There's a bunch more that we
  29758. 22:53:00didn't list on here. Um even sklearn or
  29759. 22:53:03the uh side package in Python has a
  29760. 22:53:06neural network you can program a very
  29761. 22:53:08basic one and it is the same basic one
  29762. 22:53:10that you could do in TensorFlow if you
  29763. 22:53:12stripped everything out of it and then
  29764. 22:53:13TensorFlow has a lot of tools they've
  29765. 22:53:15added in and so has KAS but we're going
  29766. 22:53:17to be looking specifically at TensorFlow
  29767. 22:53:18in our example and TensorFlow is an
  29768. 22:53:21open- source tool used to define and run
  29769. 22:53:23computations on what they call tensors
  29770. 22:53:26very common language now so you more and
  29771. 22:53:28more we see the term tensor as being a
  29772. 22:53:30standard in the uh deep learning
  29773. 22:53:32language and this was originally
  29774. 22:53:34developed by Google. So let's dig a
  29775. 22:53:36little bit big in there. What are
  29776. 22:53:37tensors? Tensors are just another name
  29777. 22:53:40for arrays. So a tensor of dimension
  29778. 22:53:43five. You can see here we have ab kmq
  29779. 22:53:46whatever. So it's an array coming in.
  29780. 22:53:47And the tensor of dimension 54 more like
  29781. 22:53:50a picture. Very common to see that in a
  29782. 22:53:52picture. You can also see a tensor even
  29783. 22:53:54more detailed than a picture as we go to
  29784. 22:53:56the next one. Tensor of dimension 333.
  29785. 22:53:58This is 3D space. You might have a
  29786. 22:54:00picture that also has colors. That might
  29787. 22:54:02be the third dimension. You might have
  29788. 22:54:04four dimensions because you have both
  29789. 22:54:06your grid and your different color
  29790. 22:54:08channels and your zplot. You can see
  29791. 22:54:11where you can now process a very
  29792. 22:54:13highlevel set of data coming in whether
  29793. 22:54:16as an image or features. They could be
  29794. 22:54:18features that have nothing to do with
  29795. 22:54:19images. So there's a lot of stuff you
  29796. 22:54:21can do now with the tensors coming in.
  29797. 22:54:23Thus this where the term tensorflow
  29798. 22:54:25comes from. So we have um right now the
  29799. 22:54:27TensorFlow is the most popular library
  29800. 22:54:29in deep learning and I did mention KAS
  29801. 22:54:31now works with TensorFlow. So there's a
  29802. 22:54:33lot of stuff you can do between the two.
  29803. 22:54:35Uh it's an open-source software library
  29804. 22:54:37developed by Google. Uh so they hit a
  29805. 22:54:40roadblock and they realized hey this is
  29806. 22:54:43an infant stage technology. You know we
  29807. 22:54:45thought it was going to be the next
  29808. 22:54:46greatest thing and we were going to have
  29809. 22:54:48a hold on it but it's really infant as
  29810. 22:54:50far as how it's applied and what we can
  29811. 22:54:52do with it. Let's open source it so
  29812. 22:54:53everybody can work on it. uh let's take
  29813. 22:54:55it to the next level. And that's really
  29814. 22:54:56what open source does to a lot of these
  29815. 22:54:58uh packages when they release them. And
  29816. 22:55:00you can run on either a CPU or a GPU. So
  29817. 22:55:03when we look at the details, if you have
  29818. 22:55:05your graphic processing units, um what's
  29819. 22:55:08nice about those is they run a lot
  29820. 22:55:10faster. The downside is you have to play
  29821. 22:55:12with them a little bit to get them up
  29822. 22:55:13and running. And it's a hardware
  29823. 22:55:14upgrade. When we run it, I'll be running
  29824. 22:55:16it in the CPU mode. I have played with
  29825. 22:55:18it in my GPU on my personal computer.
  29826. 22:55:21you know, it does increase the
  29827. 22:55:22processing. Uh, but I did run into some
  29828. 22:55:24version problems with my Python and
  29829. 22:55:26stuff like that. And when I did finally
  29830. 22:55:28work it out, I went back to the CPU
  29831. 22:55:29because it didn't increase my speed
  29832. 22:55:31enough for what I was working on. But in
  29833. 22:55:32a larger group, you might be able put
  29834. 22:55:34that on. If you're working with a larger
  29835. 22:55:36stack of computers, you might want to
  29836. 22:55:37run it in the GPU. You can create a data
  29837. 22:55:39flow graphs that have nodes and edges.
  29838. 22:55:42So there's our edges coming in. We
  29839. 22:55:44didn't talk about edges, but that's very
  29840. 22:55:46up and cominging way of looking at your
  29841. 22:55:48analytical data is how do different
  29842. 22:55:50nodes connect? What do those edges look
  29843. 22:55:52like in between them? And it's used for
  29844. 22:55:54machine learning applications such as
  29845. 22:55:55neural networks. It is mostly a neural
  29846. 22:55:58network, but they have all kinds of
  29847. 22:56:00tools which sit on top of our basic
  29848. 22:56:02neural network. They have new stuff
  29849. 22:56:03evolving into the TensorFlow library.
  29850. 22:56:06So, it's very much uh just exploding.
  29851. 22:56:09great time to jump into TensorFlow
  29852. 22:56:10because there's all kinds of cool things
  29853. 22:56:12we're doing with it and all kinds of
  29854. 22:56:13cool applications you can now use uh
  29855. 22:56:15TensorFlow for. So let's take a look at
  29856. 22:56:17implementation in TensorFlow and we're
  29857. 22:56:20going to build a neural network to
  29858. 22:56:21identify handwritten digits using the uh
  29859. 22:56:24Mnest database or the MNIST database and
  29860. 22:56:28that stands for modified National
  29861. 22:56:30Institute of Standards and Technology
  29862. 22:56:32database. It is a collection of 70,000
  29863. 22:56:34handwritten digits and the digit labels
  29864. 22:56:37identify each of the digits from 0ero to
  29865. 22:56:39nine. This is a cool example because
  29866. 22:56:41it's simple enough that you could
  29867. 22:56:43actually run this through some basic
  29868. 22:56:45machine learning categorizing algorithms
  29869. 22:56:48and train them and you'll get about the
  29870. 22:56:50same answer because again it's it's
  29871. 22:56:51simple grid. The digits on the grid
  29872. 22:56:53don't have a huge amount of variation
  29873. 22:56:55like you would say an automated driving
  29874. 22:56:57car looking at the environment. So you
  29875. 22:56:59can still do this with a lot of your um
  29876. 22:57:02different linear models and stuff like
  29877. 22:57:03that. You can solve this and you'll get
  29878. 22:57:05about the same answer. When I ran a
  29879. 22:57:06comparison between TensorFlow and
  29880. 22:57:08between some basic uh regression models
  29881. 22:57:11or category models uh in machine
  29882. 22:57:13learning, they came up pretty even as
  29883. 22:57:15far as their output. Uh so this is kind
  29884. 22:57:17of where we start to see the complexity
  29885. 22:57:20of something coming in this case a
  29886. 22:57:22tensor you know or a grid of uh
  29887. 22:57:24information where the deep learning
  29888. 22:57:26model does as good as the regular models
  29889. 22:57:29and when you get past this kind of
  29890. 22:57:31complexity and features suddenly the
  29891. 22:57:34neural networks come up with better
  29892. 22:57:36answers better solutions and a better
  29893. 22:57:38build and that's why there's such a move
  29894. 22:57:40into neural networks is we live in a
  29895. 22:57:42complicated world and it's just really
  29896. 22:57:44cool we can do with this. So the
  29897. 22:57:45handwritten digits from the um NIST
  29898. 22:57:47database, they come in, the data set is
  29899. 22:57:50used to train the machine, a new image
  29900. 22:57:52of a digit is fed and the digit is
  29901. 22:57:55identified. Um and if you've looked at
  29902. 22:57:56any of our other machine learning tools
  29903. 22:57:59where we're doing training, uh where we
  29904. 22:58:00train our uh model to fit and then you
  29905. 22:58:03test it out, this should look pretty
  29906. 22:58:05familiar. Uh and there is some tools out
  29907. 22:58:07there for say untrained categorizing uh
  29908. 22:58:09where it's just looking for features
  29909. 22:58:11that fit together. So there are tools
  29910. 22:58:12that don't need that training. But this
  29911. 22:58:14is where uh when we talk about neural
  29912. 22:58:16networks, we do need to train them. And
  29913. 22:58:17this is what we're looking at.
  29914. 22:58:33So for this I'm going to use the
  29915. 22:58:35Anaconda Navigator just because it's a
  29916. 22:58:38very nice visual tool. You might be in
  29917. 22:58:40PyCharm or one of your other IDEs for
  29918. 22:58:43editing Python because we are looking at
  29919. 22:58:45Python TensorFlow. And under Anaconda,
  29920. 22:58:47we have the notebook, which is something
  29921. 22:58:49we use pretty regularly. And they have
  29922. 22:58:51the Jupyter Lab. The Jupyter Lab is the
  29923. 22:58:53Jupyter notebook, but with tabs and a
  29924. 22:58:55few new features. So, we'll be using the
  29925. 22:58:57Jupyter Lab today. And under the
  29926. 22:58:59environment, you'll want to go ahead and
  29927. 22:59:01and uh if you haven't yet, uh you'll see
  29928. 22:59:04that I have a number of different setups
  29929. 22:59:05in here. Right now I have the Python
  29930. 22:59:08version 36 and the TensorFlow. In this
  29931. 22:59:12case I have TensorFlow 1.12. If we
  29932. 22:59:14scroll down you can see that uh here we
  29933. 22:59:16go. TensorFlow and it's version 1.12.
  29934. 22:59:19And in here if you haven't yet you'll
  29935. 22:59:21need to install those and go in and just
  29936. 22:59:23open our terminal. And u if you've never
  29937. 22:59:26used the Anaconda or if you're in your
  29938. 22:59:29other thing you might have something
  29939. 22:59:30simple like pip. Is what I use for my
  29940. 22:59:33install. And you can simply do install
  29941. 22:59:36TensorFlow. And that should bring in the
  29942. 22:59:38most current version. Now, when I
  29943. 22:59:40installed this a few months ago, Python
  29944. 22:59:42version, I'm not going to run this
  29945. 22:59:44because I already have installed on
  29946. 22:59:45here. Python version 3.7, the newest one
  29947. 22:59:48out, still had a couple glitches with
  29948. 22:59:51the TensorFlow. I believe they've fixed
  29949. 22:59:53it as of writing of this, but um I'm
  29950. 22:59:55going to stick with 3.6 just so I don't
  29951. 22:59:56get any surprises on there. So, this is
  29952. 22:59:58Python version 36 with TensorFlow 1.12
  29953. 23:00:02on here. And if you haven't installed it
  29954. 23:00:04yet, you also want to install Numpy for
  29955. 23:00:06this example. That's Numbers Python or
  29956. 23:00:08uh nu py. You can just simply run an
  29957. 23:00:12install on there. Keep in mind if you're
  29958. 23:00:14in Anaconda uh and you've created one of
  29959. 23:00:16these environments specific to this,
  29960. 23:00:18keep withd.
  29961. 23:00:21If you're going to use pip, keep with
  29962. 23:00:22pip. Don't install one package with pip
  29963. 23:00:24and one under because that's how they
  29964. 23:00:26track those version numbers and how they
  29965. 23:00:28fit together and you can end up with a
  29966. 23:00:30problem. they don't pip doesn't see cond
  29967. 23:00:32and vice versa. Uh so just keep that in
  29968. 23:00:34mind when you're running your installs.
  29969. 23:00:35We'll go ahead and open up Jupyter Lab
  29970. 23:00:37and we're going to launch that. So
  29971. 23:00:39here's my Jupyter Lab. One of the really
  29972. 23:00:41cool features of Jupyter Lab is you have
  29973. 23:00:42tabs now. So you can open up multiple uh
  29974. 23:00:45notebooks. And this is nice cuz I have
  29975. 23:00:46my notes I'm working on and then our
  29976. 23:00:48actual window we're looking in. And
  29977. 23:00:50we'll go ahead and zoom in a little bit
  29978. 23:00:52here. There we go. So you have a nice u
  29979. 23:00:54hopefully easy to see fonts. And then
  29980. 23:00:56we'll go ahead and do a simple or get
  29981. 23:00:58our imports out of the way. Um, and so
  29982. 23:00:59we're going to import our TensorFlow as
  29983. 23:01:01TF. Uh, that's pretty much a standard
  29984. 23:01:03for TensorFlow, numpy, our numbers
  29985. 23:01:06Python as py, and we'll import our matt
  29986. 23:01:10plot library as plt. Again, these are
  29987. 23:01:13very common. So if you see TF or py or
  29988. 23:01:16plt, this is a standard that most people
  29989. 23:01:18use. Do you have to? No, you could just
  29990. 23:01:21do import numpy instead of doing as py.
  29991. 23:01:23And then from
  29992. 23:01:25tensorflow.acamples.tutorial
  29993. 23:01:28tutorials. This is always nice because
  29994. 23:01:30they actually include data set we're
  29995. 23:01:32going to play with. So, we're going to
  29996. 23:01:34import input data. So, there's our data
  29997. 23:01:37coming in. That's all we're doing is
  29998. 23:01:39telling it this is where it's coming
  29999. 23:01:40from. And if we're going to tell where
  30000. 23:01:41it's coming from, we need to go ahead
  30001. 23:01:42and create a variable with that
  30002. 23:01:44information in it. And we'll just call
  30003. 23:01:45this uh mnist
  30004. 23:01:48or minced. You know, I don't really know
  30005. 23:01:49how they pronounce that. I should
  30006. 23:01:50probably look that up. It's a very
  30007. 23:01:52common data set to use. And there's our
  30008. 23:01:54input data. And we're going to read data
  30009. 23:01:57sets. And this is um if you look at
  30010. 23:01:59this, we imported input data from our
  30011. 23:02:01TensorFlow. And so this is a TensorFlow
  30012. 23:02:05read statement for their tutorials. So
  30013. 23:02:07this isn't like some special Python
  30014. 23:02:09setup. This is just their setup. Makes
  30015. 23:02:11it easy to pull it in. So once we get
  30016. 23:02:13into their data sets, we need to go
  30017. 23:02:14ahead and tell it what kind of data set.
  30018. 23:02:16And again, this is what we brought in,
  30019. 23:02:18but it's going to be the nint data. And
  30020. 23:02:20this part is very important. one hot
  30021. 23:02:24equals true. This means that instead of
  30022. 23:02:27importing a value from 0 to 9, we
  30023. 23:02:31evaluate the data set. It's going to
  30024. 23:02:33bring it in as one hot. Whenever you see
  30025. 23:02:35one hot encoder, we're flattening that
  30026. 23:02:37out. And we have true false for zero,
  30027. 23:02:40true false for one, true false for two.
  30028. 23:02:42So our output, if you remember from our
  30029. 23:02:44output uh from the slide we did earlier,
  30030. 23:02:46uh in this case, I grabbed the one for
  30031. 23:02:48bike price. Doesn't really matter which
  30032. 23:02:50one we use. This has one output. So we
  30033. 23:02:52have our bike price on this. We're going
  30034. 23:02:54to have instead of one output, we're
  30035. 23:02:56going to have 10 outputs representing
  30036. 23:02:58each of the digits in there. And this
  30037. 23:03:00code really isn't going to show us
  30038. 23:03:02anything. It's good to see what we're
  30039. 23:03:03actually looking at. So um let's go
  30040. 23:03:05ahead and do a figure ax equals go into
  30041. 23:03:08our plot library subplots 10, 10. And
  30042. 23:03:13that is if you remember we talked about
  30043. 23:03:15tensor. Tensor being data coming in.
  30044. 23:03:17This is a 10 by 10 grid or 100 pixels on
  30045. 23:03:20there. And if we're going to display it,
  30046. 23:03:22uh let's go do K0
  30047. 23:03:24for I and range 10. Just a simple loop
  30048. 23:03:28through on the data. Let's do what is
  30049. 23:03:30it? Uh for J and range 10. And I
  30050. 23:03:33actually misqued that 10 uh 10 x 10 is
  30051. 23:03:36not the actual size of the pixels. Uh
  30052. 23:03:38the actual pixels are going to be um we
  30053. 23:03:40look at the shapes and we'll get into
  30054. 23:03:41that in just a second here. We'll take a
  30055. 23:03:42quick look at shape on there. Uh turns
  30056. 23:03:45out they're uh what are they? are, I
  30057. 23:03:47believe, 28x 28. Uh, so let's take a
  30058. 23:03:50look at that. And we're just going to
  30059. 23:03:51plot these. What are we looking at? What
  30060. 23:03:53are we working with? As a data
  30061. 23:03:54scientist, you should always be looking
  30062. 23:03:56back at your data and seeing what it
  30063. 23:03:59looks like and get that human
  30064. 23:04:00perspective because you just never know.
  30065. 23:04:02You know, the the computer may put
  30066. 23:04:04something out that looks makes no sense.
  30067. 23:04:06And at that point, you want to go back
  30068. 23:04:08and reevaluate what you did. Uh, so
  30069. 23:04:10we're going to go ahead and plot. We're
  30070. 23:04:11going to plot 10 digit, you know, 10 of
  30071. 23:04:13the digits by 10 of the digits. And
  30072. 23:04:14here's our ax. We'll create the J on our
  30073. 23:04:17subplots and we're going to do an image
  30074. 23:04:19show. We're going to look at the
  30075. 23:04:20training image for images of K. And then
  30076. 23:04:24we want to reshape this. We're going to
  30077. 23:04:25reshape this. And we're going to reshape
  30078. 23:04:27this 28x 28. That's how I knew I had it
  30079. 23:04:30wrong is cuz I looked down my notes. I
  30080. 23:04:31was like, oh no, that says 28. It's not
  30081. 23:04:3310 x 10. And I should know that already
  30082. 23:04:34cuz I've done enough messing with this
  30083. 23:04:36data set that I should have remembered.
  30084. 23:04:37Uh, but it's 28 x 28. And the aspect
  30085. 23:04:40we're going to do is auto. And this is
  30086. 23:04:41all, if you look at this, here's our
  30087. 23:04:43variable NIST. the NIST is coming from
  30088. 23:04:46data set. Uh so this is all part of the
  30089. 23:04:48TF TensorFlow learning or examples
  30090. 23:04:51tutorial in there. And then we'll go
  30091. 23:04:52ahead and do K plus equals 1. So we just
  30092. 23:04:56keep paging through our different um
  30093. 23:04:58images. And let's see what that looks
  30094. 23:04:59like. Let's go ahead and do a plot show.
  30095. 23:05:02Uh and we'll go ahead and run this so we
  30096. 23:05:03can take a look and see what we have
  30097. 23:05:04here. And so we have a nice plot here.
  30098. 23:05:06And you can just see that we have uh
  30099. 23:05:08some random numbers showing up in each
  30100. 23:05:09one of these little subplots. If you're
  30101. 23:05:11wanting a copy of this code, put a note
  30102. 23:05:14down in the YouTube video and let us
  30103. 23:05:16know or come visit us at
  30104. 23:05:17www.simplearn.com
  30105. 23:05:19and we'll send you out a copy of what
  30106. 23:05:20we're working on and get a copy of that
  30107. 23:05:22for your own setup. Uh so now we've
  30108. 23:05:24taken a look and we can just see we have
  30109. 23:05:26here's our pictures that are coming on.
  30110. 23:05:27We plotted them so we have an idea of
  30111. 23:05:29what we're looking at. Let's go ahead
  30112. 23:05:31and uh print. Let's look at the shape of
  30113. 23:05:33the features. Uh so when we have this we
  30114. 23:05:35have our nest train images and we'll do
  30115. 23:05:39the shape on there. Let's take a look
  30116. 23:05:41and just see what we're looking at uh as
  30117. 23:05:43far as uh our count and everything. And
  30118. 23:05:45so you can see here we have 55,000.
  30119. 23:05:49That's basically how many images we have
  30120. 23:05:50and this by 784. And in this data set
  30121. 23:05:53there's also our labels. So let's take a
  30122. 23:05:55look at that. We have our net train
  30123. 23:05:57labels shape. Let's take a look and see
  30124. 23:05:59what that looks like. Uh and there we
  30125. 23:06:01have 10 because there's 10 digits. So we
  30126. 23:06:03brought in that's our output we're
  30127. 23:06:04looking at. And so we have there we go
  30128. 23:06:0655,000. They match. They should match
  30129. 23:06:08because you should have equal numbers in
  30130. 23:06:10both of those. You know, here's our data
  30131. 23:06:12in and here's our answer. If you
  30132. 23:06:14remember, this is a bunch of zeros and
  30133. 23:06:16with one each each one will be 0001
  30134. 23:06:19would be what letter four or something
  30135. 23:06:21like that. So, let's take a look at what
  30136. 23:06:22our one hot encoding did for the first
  30137. 23:06:24observation. And this is when we're
  30138. 23:06:26exploring data, you really want to dig
  30139. 23:06:28in there and just see what the heck am I
  30140. 23:06:30looking at. So, we're going to look at
  30141. 23:06:31the labels. And this would be the first
  30142. 23:06:33label that comes up. And we'll go ahead
  30143. 23:06:35and run this. And we look at that. You
  30144. 23:06:37can see this is what I'm talking about.
  30145. 23:06:380 0 or 1 is 0 2 is 0 3 is 0 four is 0 5
  30146. 23:06:43is 0 6 is 0. 7 equals 1. So our very
  30147. 23:06:47first label is a seven. But our very
  30148. 23:06:49first label comes up that it's a seven.
  30149. 23:06:51And so we don't have like 0 through 9.
  30150. 23:06:54We have a bunch of zeros and just the
  30151. 23:06:56one to mark it as a seven on here. So
  30152. 23:06:58now we've kind of looked a quick look at
  30153. 23:07:00the data. And in here you might ask some
  30154. 23:07:02questions like what is 784? 24 * 24.
  30155. 23:07:05Remember that's the size of our grid on
  30156. 23:07:08there or our tensor coming in. So 784 is
  30157. 23:07:10a setup on there. And we've gone through
  30158. 23:07:12all this viewing the data. We'll go
  30159. 23:07:14ahead and start looking at our
  30160. 23:07:16tensorflow. So let's take our X
  30161. 23:07:18variable. This is going to be our
  30162. 23:07:19training set. We'll do a placeholder and
  30163. 23:07:22then we're going to have these come in
  30164. 23:07:24as float. Now if I remember correctly,
  30165. 23:07:26they're actually, you know, zero or one
  30166. 23:07:28for the values because they're either
  30167. 23:07:30but we have them coming in as a float
  30168. 23:07:31value. And we have a little bit of a
  30169. 23:07:33shape coming in here. And there's our
  30170. 23:07:34784. Uh so we let it know that this is
  30171. 23:07:37what's what our input is for our
  30172. 23:07:39TensorFlow. And this is our training
  30173. 23:07:41set. So we'll just put a label on there
  30174. 23:07:43to help us uh track that train set. And
  30175. 23:07:46then W. And with W, we'll go ahead and
  30176. 23:07:48do TF variables. And we'll do this as uh
  30177. 23:07:51zeros variables TF zeros. And we'll set
  30178. 23:07:54this as as 784 by 10. 10 being the
  30179. 23:07:58output. 784 being our number of
  30180. 23:08:01variables in and this is our weights.
  30181. 23:08:04Remember we have a bias in there too.
  30182. 23:08:06And I'll go back over this in just a
  30183. 23:08:08second as we see how that fits together
  30184. 23:08:11in our tensorflow. And we'll do this one
  30185. 23:08:13um with our variables again. We have 10.
  30186. 23:08:15So we're going to do the bias. We're
  30187. 23:08:17going do it the same kind of format and
  30188. 23:08:19setup on here. And so we'll do that as
  30189. 23:08:22as TF zeros of 10. So we'll just create
  30190. 23:08:25an array of 10 there. And this is our
  30191. 23:08:26bias. So with these three lines um and
  30192. 23:08:30there's actually they're coming out with
  30193. 23:08:32the eager execution which would bypass
  30194. 23:08:35some of what we're doing. But this is
  30195. 23:08:36important to understand is the first
  30196. 23:08:37thing you have to do with TensorFlow is
  30197. 23:08:39we have to allocate a space for the
  30198. 23:08:41variables and our TF placeholder and our
  30199. 23:08:44TF variable with our weights and our
  30200. 23:08:46biases. This actually hasn't done
  30201. 23:08:48anything yet. So all it is is
  30202. 23:08:49placeholders. That's why it's okay to
  30203. 23:08:51use zeros. Um you could have just as
  30204. 23:08:53easily used ones or anything else and it
  30205. 23:08:55wouldn't matter. The next stage is to go
  30206. 23:08:57ahead and set up some of the functions
  30207. 23:09:00going on. But before we do that, just
  30208. 23:09:02note that this hasn't done anything.
  30209. 23:09:03Even if I execute it, all it's done is
  30210. 23:09:05created placeholders until we do the
  30211. 23:09:07final initialization. And so we need to
  30212. 23:09:09go ahead and set up. We'll do y=
  30213. 23:09:11tf.n.oftmax.
  30214. 23:09:14And the code for this is tf.mmoxw.
  30215. 23:09:18And this is our uh sum. Let's just put a
  30216. 23:09:21note here so we can keep track of what's
  30217. 23:09:22going on. We're finding weighted sum of
  30218. 23:09:25inputs plus the bias. Uh so there's our
  30219. 23:09:29plus b the bias and then we need to go
  30220. 23:09:31ahead keep um let's do y underscore and
  30221. 23:09:34again another placeholder and this one
  30222. 23:09:36we'll set um it actually we'll put in as
  30223. 23:09:38tf placeholder on here tf placeholder
  30224. 23:09:41float none 10. There's our one hot
  30225. 23:09:43encoder going on there. So our 10 values
  30226. 23:09:46coming out and we'll do a cross entropy
  30227. 23:09:50on here and this is going to be minus tf
  30228. 23:09:52reduce sum and we'll do y here's our y
  30229. 23:09:55underscore which is remember we have
  30230. 23:09:57your y output and your actual output. Uh
  30231. 23:10:00so this will be our y underscore time
  30232. 23:10:02the tf log of y. And then finally um
  30233. 23:10:06before we do the actual initialization
  30234. 23:10:08of all our variables we'll set up our
  30235. 23:10:10train step. This equals our gradient
  30236. 23:10:12descent optimizer. Very important.
  30237. 23:10:14Remember we looked at that chart on our
  30238. 23:10:16um uh slides and so we've set up all
  30239. 23:10:18these formulas and here's our gradient
  30240. 23:10:20descent optimizer and as it keeps
  30241. 23:10:22looking it keeps looking for that zero
  30242. 23:10:24value. That's what we're doing with that
  30243. 23:10:26particular formula. So let's take a look
  30244. 23:10:28and see what we're doing here. We just
  30245. 23:10:30put together all of our pieces for
  30246. 23:10:32TensorFlow. And you know the devil's in
  30247. 23:10:34the details. We have here our training
  30248. 23:10:37set coming in. We have to put a
  30249. 23:10:39placeholder on there. We have our uh
  30250. 23:10:42variables with their weights. We have
  30251. 23:10:44our biases coming out and then we put in
  30252. 23:10:46our uh the weighted sum. So here's
  30253. 23:10:49summizing our summation here. Then we
  30254. 23:10:52have our y variable output. So there's
  30255. 23:10:54our y um how it works and then of course
  30256. 23:10:56the actual output on there. And then we
  30257. 23:10:58have our cross entropy coming in and
  30258. 23:11:01that's our minus tf.reduce sum the y *
  30259. 23:11:04the tf log of y. And then the training
  30260. 23:11:06step gradient descent optimizer and
  30261. 23:11:08we're using a 0.01 in this and we're
  30262. 23:11:10going to minimize cross entropy. So,
  30263. 23:11:12we're going to let it do all the work.
  30264. 23:11:14So, once we've set up all of these
  30265. 23:11:16different layers, we've allocated for
  30266. 23:11:18them, we need to go ahead and initialize
  30267. 23:11:20them. So, we're going to do an init tf
  30268. 23:11:22initialize, and it's going to be all
  30269. 23:11:24variables. Uh, one of the cool things is
  30270. 23:11:26they're in the process of doing away
  30271. 23:11:28with this. So, all these steps would be
  30272. 23:11:30bundled into one instead of having to
  30273. 23:11:32have placeholders. You initialize them
  30274. 23:11:33in the same process going on. And then
  30275. 23:11:36finally, everything in TensorFlow is
  30276. 23:11:38based on your session. Now, this is
  30277. 23:11:40changing that there's other options to
  30278. 23:11:42be able to run this, but we want to go
  30279. 23:11:43ahead and do uh session. There's our TF.
  30280. 23:11:46There's a TF session. And then we want
  30281. 23:11:48to go ahead and do session run. And what
  30282. 23:11:50are we going to run? Well, we did
  30283. 23:11:51initialization of all our variables. Uh
  30284. 23:11:54so, this is what we're running. And this
  30285. 23:11:55is we're actually once we do this, we
  30286. 23:11:57actually create our TensorFlow object.
  30287. 23:12:00So, this whole piece of code right here
  30288. 23:12:02is our TensorFlow object. We have our
  30289. 23:12:05input coming in with our weighted
  30290. 23:12:07variables coming in. Our soft max for
  30291. 23:12:10our metal going out. How does it add it
  30292. 23:12:12together for our y value? Uh and then we
  30293. 23:12:14have the actual uh float value coming
  30294. 23:12:17out. Checking on our all the way down.
  30295. 23:12:19So you can see all the different stages
  30296. 23:12:20going through that we're setting up. Um
  30297. 23:12:22and this is one of the reasons that a
  30298. 23:12:23lot of people like TensorFlow is because
  30299. 23:12:26you can designate all these different
  30300. 23:12:28pieces one step at a time. This is also
  30301. 23:12:30one of the reasons people don't like
  30302. 23:12:32TensorFlow is because you have to
  30303. 23:12:34designate all the different layers
  30304. 23:12:36coming down and there's a lot of steps
  30305. 23:12:38being made right now to minimize this to
  30306. 23:12:40make it either easier to automate it or
  30307. 23:12:43to allow you to do more complicated
  30308. 23:12:45things and all those steps are still at
  30309. 23:12:47play. So it's worth looking into the
  30310. 23:12:49more advanced version what's going on
  30311. 23:12:51with KAS on top of TensorFlow. It's also
  30312. 23:12:53important to understand what's going on
  30313. 23:12:55in these individual levels if you're
  30314. 23:12:56going to play with them. It's important
  30315. 23:12:58to understand, hey, what's going on with
  30316. 23:13:00the uh finding the weighted sum of the
  30317. 23:13:02inputs plus the bias because there's
  30318. 23:13:04other ways to do that. There's all kinds
  30319. 23:13:06of other tools in there now, but this is
  30320. 23:13:07the basic setup that you want to do on a
  30321. 23:13:09TensorFlow coming in. And we want to go
  30322. 23:13:11ahead and just run and admit our
  30323. 23:13:12TensorFlow. So, let's go ahead and do
  30324. 23:13:14that. Let's run this. We do get a
  30325. 23:13:16warning here because uh there's a move
  30326. 23:13:18to use global variables. This is one of
  30327. 23:13:20the changes they're making, but it as
  30328. 23:13:22far as this example, it's not going to
  30329. 23:13:23make a difference because we're doing
  30330. 23:13:25once we initialize it. This is
  30331. 23:13:26initializing our variables. And again,
  30332. 23:13:28these are only placeholders up here
  30333. 23:13:29until we initialize them. And I would
  30334. 23:13:31highly suggest put a note down there or
  30335. 23:13:33or go over to simplylearn.com and let
  30336. 23:13:35them know and have them email you a copy
  30337. 23:13:38of the code. So, you can actually play
  30338. 23:13:39with this code right here because this
  30339. 23:13:41is the body of what's going on in
  30340. 23:13:43TensorFlow. This is the build in neural
  30341. 23:13:46networks. And then once we've done that,
  30342. 23:13:48now comes kind of the fun part is we
  30343. 23:13:50need to go ahead and train it. Uh so
  30344. 23:13:52we've created our TensorFlow, we've
  30345. 23:13:53created our uh network and now we need
  30346. 23:13:56to go ahead and train it. Uh so let's
  30347. 23:13:57put together that training code and
  30348. 23:13:59let's just do uh for I in range u 0 to
  30349. 23:14:031,000. So we're just going to look at uh
  30350. 23:14:05the first 10,000 in our training. And
  30351. 23:14:07the way we pull that data from our mints
  30352. 23:14:10train next batch of 100. Uh so you look
  30353. 23:14:13at this. We're going to be doing groups
  30354. 23:14:14of 100 and then there's going to go
  30355. 23:14:16through a thousand of them. This is very
  30356. 23:14:18important that TensorFlow builds this
  30357. 23:14:21in. This is one of the downsides of
  30358. 23:14:23doing sklearn or one of the older
  30359. 23:14:25packages is they don't let you batch
  30360. 23:14:27groups in. Uh they wanted to have it all
  30361. 23:14:29up front and then you have to build your
  30362. 23:14:31own batch programs right now. Uh this
  30363. 23:14:33lets us go ahead and do that. And you
  30364. 23:14:34can see here we have batch x of s, batch
  30365. 23:14:36y of s. So there's our x and our y. You
  30366. 23:14:39could look at this as our training of X
  30367. 23:14:41and our train of Y or the data N and the
  30368. 23:14:45answer in. Uh and then we simply do our
  30369. 23:14:47session run. Uh so here's our session
  30370. 23:14:49that we've created. We're going to run
  30371. 23:14:50it and we want to do the train step. We
  30372. 23:14:53initialized our train step up here and
  30373. 23:14:55our TF. And so there's our train step
  30374. 23:14:58feed. It's a dictionary. Dictionary
  30375. 23:15:00coming in which we're going to create
  30376. 23:15:01right here. uh is x is our batch x of
  30377. 23:15:06our sample comma and our y underscore is
  30378. 23:15:09going to be our batch of our y sample.
  30379. 23:15:12Uh and so this goes through and we've
  30380. 23:15:14now hit the run button and we've trained
  30381. 23:15:16our session. We've trained this setup on
  30382. 23:15:19here. And once we've trained it, then we
  30383. 23:15:21need to go ahead and find out how good
  30384. 23:15:22our accuracy was and actually start
  30385. 23:15:24running some predictions through there.
  30386. 23:15:25Uh so we'll go ahead and create a a
  30387. 23:15:27correct prediction. And this is where
  30388. 23:15:29our tf.equal equal. We'll use our argmax
  30389. 23:15:31y of one and tf argmax of y of
  30390. 23:15:34underscore of one to help us get the
  30391. 23:15:35correct predictions on there. And then
  30392. 23:15:37we want to use that to feed into an
  30393. 23:15:39accuracy. And so our accuracy is going
  30394. 23:15:42to be tf reduce uh mean and we'll take
  30395. 23:15:45that and we'll do um a cast and this is
  30396. 23:15:48the correct prediction that we're
  30397. 23:15:49sending in there. And it is a uh float
  30398. 23:15:52value. Keep it simple. And let's go
  30399. 23:15:54ahead and print this out so we can see
  30400. 23:15:56what we're looking at. Uh so what are we
  30401. 23:15:57printing out? Uh we need to do a session
  30402. 23:15:59run. This session run is going to be on
  30403. 23:16:01the accuracy. Where did accuracy comes
  30404. 23:16:03from? This is our we're casting our TF
  30405. 23:16:05on there with the correct predictions on
  30406. 23:16:07that. So here's our accuracy feed in. So
  30407. 23:16:10it needs a dictionary for the data
  30408. 23:16:12coming in. We're going to create our
  30409. 23:16:13dictionary and x is going to be our nest
  30410. 23:16:16test images and y there is going to be
  30411. 23:16:21our nest.est
  30412. 23:16:23labels. Let me just double check and
  30413. 23:16:25make sure I have that typed in there
  30414. 23:16:26correctly. There we go. Oh, and let's go
  30415. 23:16:28ahead and run that and see what comes
  30416. 23:16:29up. And we end up with a N165
  30417. 23:16:33for our accuracy, which means our
  30418. 23:16:36trained neural network does a pretty
  30419. 23:16:38good job letting us know what these
  30420. 23:16:40different symbols are in guessing that a
  30421. 23:16:42seven and a three and a four, uh,
  30422. 23:16:44something that as humans we kind of take
  30423. 23:16:46for granted. I even have trouble reading
  30424. 23:16:48this. So, I don't know if I would be
  30425. 23:16:49able like that first one, I would sit
  30426. 23:16:51there for a long time figuring out
  30427. 23:16:52that's a seven versus a two. That could
  30428. 23:16:54have easily been a two to me. Sing seem
  30429. 23:16:56to do a pretty good job analyzing this
  30430. 23:16:58data. And this is used to analyze
  30431. 23:17:00something very complicated on these
  30432. 23:17:01images, very different than uh just a
  30433. 23:17:04straight value of uh cost of sales and
  30434. 23:17:07here's our return and our marketing. Uh
  30435. 23:17:09we can now create this nice neural
  30436. 23:17:10network that does all kinds of cool
  30437. 23:17:12things. Do you know friends that
  30438. 23:17:14according to the lending statistics the
  30439. 23:17:16demand for AI and ML specialist is
  30440. 23:17:18projected to surge by 40% between 2023
  30441. 23:17:22to 2027.
  30442. 23:17:24And on an average, an ML engineer is
  30443. 23:17:27expected to earn around 133 and $336 per
  30444. 23:17:31year. So if you are an aspiring ML
  30445. 23:17:34engineer and thinking about what
  30446. 23:17:36innovative projects you can show in your
  30447. 23:17:38portfolio, then your wait is over cuz in
  30448. 23:17:41this video I'll be covering eight
  30449. 23:17:43amazing ML projects that you can
  30450. 23:17:45showcase in your resume. So guys, let's
  30451. 23:17:48start first with a beginner level
  30452. 23:17:49project and the first project that we
  30453. 23:17:51are going to encounter that is home
  30454. 23:17:53value prediction. So guys, this project
  30455. 23:17:56aims to develop a predictive model to
  30456. 23:17:58estimate the value of residential
  30457. 23:17:59properties. The model will analyze
  30458. 23:18:02various features such as location,
  30459. 23:18:04square, footage, number of bedrooms and
  30460. 23:18:06bathrooms, age of the property and other
  30461. 23:18:09relevant factors. By leveraging
  30462. 23:18:11historical property data, the model will
  30463. 23:18:14be able to provide accurate home value
  30464. 23:18:15predictions which can be useful for real
  30465. 23:18:18estate agents, buyers and sellers. So
  30466. 23:18:21guys, the programming language that we
  30467. 23:18:23are going to use all over here will be
  30468. 23:18:25Python and machine learning libraries
  30469. 23:18:27that we will be using will be
  30470. 23:18:28scikitlearn, tensorflow, kas and for
  30471. 23:18:31data handling libraries we have pandas,
  30472. 23:18:34numpy and for visualization we have to
  30473. 23:18:37use mattplot and seabon. Now what will
  30474. 23:18:39be the approach for this one guys? So
  30475. 23:18:42guys the first one that we have a data
  30476. 23:18:44collection. So here what is going to
  30477. 23:18:46happen guys? So first you have to
  30478. 23:18:48collect the historical property data
  30479. 23:18:49from the sources like Zillow
  30480. 23:18:52retailer.com. You can also get database
  30481. 23:18:54from the public real estate databases
  30482. 23:18:56like Kaggle data sets where you have
  30483. 23:18:58Zillow home value prediction. Ensure
  30484. 23:19:00that the data set include features like
  30485. 23:19:02location where you have latitude,
  30486. 23:19:04longitude, square footage, number of
  30487. 23:19:06rooms, year built, property type and
  30488. 23:19:08previous sales. The next step that comes
  30489. 23:19:11is data cleaning. You have to handle the
  30490. 23:19:13missing values by using imputation
  30491. 23:19:15techniques or removing incomplete
  30492. 23:19:17records. Removing outliers that may skew
  30493. 23:19:20the model's prediction, normalize or
  30494. 23:19:22standardize the data to ensure
  30495. 23:19:24consistency. The third one that we have
  30496. 23:19:26is feature engineering. You have to
  30497. 23:19:28create new features such as proximity to
  30498. 23:19:30schools, crime rates and access to the
  30499. 23:19:32public transportation. Encode categorial
  30500. 23:19:35variables, example property type,
  30501. 23:19:37location using techniques like one hot
  30502. 23:19:38encoding. Generate interaction features
  30503. 23:19:41that capture relationship between
  30504. 23:19:42existing features. The fourth one that
  30505. 23:19:44we have is model selection. Use
  30506. 23:19:46regression models like linear
  30507. 23:19:48regression, random forest, gradient
  30508. 23:19:49boosting, neural networks. Experiment
  30509. 23:19:52with different models to identify the
  30510. 23:19:53best performing one. Now in the next
  30511. 23:19:56phase all you have to do guys is model
  30512. 23:19:58training and evaluation. Split the data
  30513. 23:20:00set into training and test sets. Train
  30514. 23:20:03the model on a training set and evaluate
  30515. 23:20:05their performance on the testing set
  30516. 23:20:07using metrics like RSM which means root
  30517. 23:20:10mean squared error. You can use cross
  30518. 23:20:12validation to ensure the model's
  30519. 23:20:14robustness and avoid overfitting. The
  30520. 23:20:17sixth one that we have all over here is
  30521. 23:20:19hyperparameter tuning. You can optimize
  30522. 23:20:21the model's hyperparameter using
  30523. 23:20:23techniques such as grid search or random
  30524. 23:20:25search to improve accuracy. And if
  30525. 23:20:27you're looking forward to deploy your
  30526. 23:20:29model, then you can develop a web
  30527. 23:20:31interface using flask or Django to allow
  30528. 23:20:33users to input property features and get
  30529. 23:20:35predictions. You can deploy the model on
  30530. 23:20:38the cloud platform like AWS for
  30531. 23:20:40scalability.
  30532. 23:20:41Now if we talk about the complexity
  30533. 23:20:43level of this, we all know that it is a
  30534. 23:20:45beginner level project. Now let us move
  30535. 23:20:48on to the one more set that is music
  30536. 23:20:50genre classification and generation. So
  30537. 23:20:53guys this is also one of the most
  30538. 23:20:55beginner level project. This project
  30539. 23:20:57aims to develop a system that can
  30540. 23:20:58classify music tracks into different
  30541. 23:21:00genres and generate new music
  30542. 23:21:02composition within specified genre. The
  30543. 23:21:04goal is to build a model that analyzes
  30544. 23:21:06audio features to categorize music and
  30545. 23:21:08uses deep learning techniques to create
  30546. 23:21:10new music. This project introduces
  30547. 23:21:13advanced concept of audio processing,
  30548. 23:21:15deep learning and generative models. So
  30549. 23:21:17guys, what will be used in this? So
  30550. 23:21:19we'll have programming language that
  30551. 23:21:21will be Python. For audio processing,
  30552. 23:21:23we'll be using librosa. For machine
  30553. 23:21:25learning libraries, we'll be using
  30554. 23:21:26tensorflow, kas, pytorch. For data
  30555. 23:21:29handling libraries, we'll be using
  30556. 23:21:30pandas, numpy. For visualization, we'll
  30557. 23:21:32be using mattplot, seabon. For the data
  30558. 23:21:35set guys, you can use gtzan music genre
  30559. 23:21:38data set or you can get it from free
  30560. 23:21:40music archive. So guys in the first
  30561. 23:21:42phase we are going to have data
  30562. 23:21:43collection. You can obtain data sets
  30563. 23:21:45containing music tracks and their
  30564. 23:21:47corresponding genre from the sources
  30565. 23:21:49from GTN music genre data set and the
  30566. 23:21:51free music archive. Ensure that a data
  30567. 23:21:54set includes diverse genre and
  30568. 23:21:55substantial number of tracks per genre.
  30569. 23:21:58Next we'll go for data prep-processing.
  30570. 23:22:00Use library librosa to load and
  30571. 23:22:02pre-process audio files including
  30572. 23:22:04feature extraction such as mil
  30573. 23:22:06frequency, septal coefficients, chroma
  30574. 23:22:08features and spectral contrast. You can
  30575. 23:22:11normalize the extracted features to
  30576. 23:22:13ensure consistent input for the given
  30577. 23:22:15models. Now if you talk about feature
  30578. 23:22:17engineering guys, you can extract
  30579. 23:22:19additional features from the audio files
  30580. 23:22:20such as tempo, beat, zero crossing rate
  30581. 23:22:23etc. Create a feature matrix that
  30582. 23:22:25represents the extracted audio features.
  30583. 23:22:27Then go for the model selection. Use
  30584. 23:22:30conventional neural network or recurrent
  30585. 23:22:32neural networks for the music genre
  30586. 23:22:33classification. Split the data set into
  30587. 23:22:36training and testing data set. Now if we
  30588. 23:22:38talk about model training and evaluation
  30589. 23:22:40guys then you can train the selected
  30590. 23:22:42classification model on the training
  30591. 23:22:43set. Evaluate the model's performance on
  30592. 23:22:46the testing set using metrics like
  30593. 23:22:47accuracy, precision, recall and F1
  30594. 23:22:50score. Use confusion matrices to
  30595. 23:22:52understand the classification
  30596. 23:22:53performance across different genres. Now
  30597. 23:22:55if we talk about model selection and
  30598. 23:22:57training for the music generation, what
  30599. 23:22:59you will do guys? You can use the
  30600. 23:23:01generative adversial networks or
  30601. 23:23:03recurrent neural networks such as LSTM,
  30602. 23:23:05long short-term memory for music
  30603. 23:23:07generation. Train the generative model
  30604. 23:23:09on the data set to create new music
  30605. 23:23:11sequences. Next, we have model training
  30606. 23:23:13and evaluation. You can train the
  30607. 23:23:15generative model on sequences of audio
  30608. 23:23:17features. You can evaluate the generated
  30609. 23:23:19music by listening tests and by
  30610. 23:23:21objective metrics like inception score
  30611. 23:23:23or fche audio distance. If I talk about
  30612. 23:23:26hyperparameter tuning guides, we can
  30613. 23:23:27optimize the model. You can use the
  30614. 23:23:29hyperparameters using techniques like
  30615. 23:23:31grid search or random search to improve
  30616. 23:23:33performance. If I talk about deployment
  30617. 23:23:35guys, you can deploy these models on
  30618. 23:23:37cloud platforms like AWS. Now let us
  30619. 23:23:40move on to our next project. So guys,
  30620. 23:23:43the complexity of this project is at the
  30621. 23:23:44beginner level. Now let us move to the
  30622. 23:23:47intermediate level projects. Next
  30623. 23:23:49project that we have all over here is
  30624. 23:23:50sentiment analysis of Twitter data. This
  30625. 23:23:53project aims to develop a sentiment
  30626. 23:23:55analysis model that can classify to its
  30627. 23:23:57side positive, negative or neutral. The
  30628. 23:24:00goal is to analyze public sentiment on
  30629. 23:24:02various topics or events using natural
  30630. 23:24:04language techniques. So guys, what will
  30631. 23:24:06be used all over here? So in this we
  30632. 23:24:09will have programming language like
  30633. 23:24:11Python. Okay. NLP libraries, NLTK spacy.
  30634. 23:24:15For machine learning libraries you can
  30635. 23:24:16use scikitlearn, tensorflow, kas. For
  30636. 23:24:19data handling libraries we have pandas,
  30637. 23:24:21numpy. For visualization we have
  30638. 23:24:23mattplot lilip seon and we can use the
  30639. 23:24:25API Twitter for data collection. Now how
  30640. 23:24:28you going to work on it guys? So guys if
  30641. 23:24:31I talk about the data collection use a
  30642. 23:24:33Twitter API to collect tweets based on
  30643. 23:24:35specific hashtags like keywords or
  30644. 23:24:37topics. Extract relevant fields like
  30645. 23:24:39tweet text, user information, timestamp
  30646. 23:24:42etc. Then if I talk about data
  30647. 23:24:43prep-processing guys, clean the tweet
  30648. 23:24:46text by removing special characters,
  30649. 23:24:48links, mentions, hashtags and stop
  30650. 23:24:50words. Tokenize the text and perform
  30651. 23:24:53limization or stemming to reduce the
  30652. 23:24:55words to their base form. Next, if you
  30653. 23:24:57talk about feature engineering guys, you
  30654. 23:24:59convert the clean text data into
  30655. 23:25:00numerical representation using TF, back
  30656. 23:25:03of words or word embedding. Now if I
  30657. 23:25:06talk about model selection guys, you can
  30658. 23:25:07choose a classification algorithm such
  30659. 23:25:09as logistic regression, n bias or LSDM.
  30660. 23:25:13Split the data set into training and
  30661. 23:25:14testing data set. Now if I talk about
  30662. 23:25:17model training and evaluation, then you
  30663. 23:25:19can train the selected model on the
  30664. 23:25:21training set. Evaluate the model's
  30665. 23:25:23performance on the testing set using
  30666. 23:25:24metrics like accuracy, precision,
  30667. 23:25:26recall, and fn score. Use cross
  30668. 23:25:29validation to ensure the model's
  30669. 23:25:31robustness. If I talk about
  30670. 23:25:32hyperparameter tuning guys, you can
  30671. 23:25:34optimize the model's hyperparameters
  30672. 23:25:36using grid search or random search to
  30673. 23:25:38improve the performance for deployment
  30674. 23:25:40which can be optional. You can deploy
  30675. 23:25:42your model on AWS for real-time
  30676. 23:25:44sentiment analysis. So if I talk about
  30677. 23:25:46the complexity level guys, its
  30678. 23:25:48complexity is intermediate. So guys, our
  30679. 23:25:51next project is customer segmentation
  30680. 23:25:53using K means clustering. This project
  30681. 23:25:56aims to segment customers into distinct
  30682. 23:25:58groups based on their purchasing
  30683. 23:25:59behavior and demographic information.
  30684. 23:26:01The objective is to understand customer
  30685. 23:26:03segments and tailor marketing strategies
  30686. 23:26:05accordingly. So guys, what programming
  30687. 23:26:08languages we'll be using? So basically
  30688. 23:26:10we'll be using Python. For machine
  30689. 23:26:12learning libraries, we will have
  30690. 23:26:13scikitlearn. For data handling
  30691. 23:26:15libraries, we'll have pandas, numpy. For
  30692. 23:26:17visualization libraries, we'll have
  30693. 23:26:18mattplot lab, seabon. And the data set
  30694. 23:26:21source will be e-commerce transaction
  30695. 23:26:23data. how we are going to work on this
  30696. 23:26:25one. For data collection, we can obtain
  30697. 23:26:27a data set of e-commerce transactions
  30698. 23:26:29that include customer demographics,
  30699. 23:26:30purchase history, and product
  30700. 23:26:32information. Next, we'll have data
  30701. 23:26:33prep-processing. For data
  30702. 23:26:35prep-processing, we are going to do the
  30703. 23:26:37cleaning of the data by handling the
  30704. 23:26:38missing values and outliers. Then, for
  30705. 23:26:41feature engineering, we are going to
  30706. 23:26:43create features like total purchase
  30707. 23:26:44amount, purchase frequency, and recency
  30708. 23:26:46of the purchases.
  30709. 23:26:48Then, we are going to proceed for the
  30710. 23:26:50model selection. You can use K means
  30711. 23:26:52clustering to segment the customers into
  30712. 23:26:54distinct groups. You can determine the
  30713. 23:26:56optimal number of clusters using methods
  30714. 23:26:58like album methods or silhou.
  30715. 23:27:01Now if I talk about model training and
  30716. 23:27:03evaluation, you can train the K means
  30717. 23:27:05model on the process data set. You can
  30718. 23:27:07evaluate the quality of clusters by
  30719. 23:27:09analyzing intracluster and intercluster
  30720. 23:27:11distances. Next we have the evaluation.
  30721. 23:27:14You can visualize the clusters using
  30722. 23:27:15techniques like PCA, principal component
  30723. 23:27:17analysis, TSN etc. Next we have the
  30724. 23:27:22hyperparameter tuning. Now now you can
  30725. 23:27:25tune this model and interpret the
  30726. 23:27:27characteristic of each segment. You can
  30727. 23:27:29develop a target marketing strategies
  30728. 23:27:30for each segment based on unique
  30729. 23:27:32behavior and preferences. Now deployment
  30730. 23:27:35is optional. You can develop a dashboard
  30731. 23:27:37using flask or Django to visualize
  30732. 23:27:39customer segments and track marketing
  30733. 23:27:41campaigns. So guys if I talk about the
  30734. 23:27:44complexity of this project. So this is
  30735. 23:27:47an intermediate level project. So guys
  30736. 23:27:49for data set you can use the Kaggle's
  30737. 23:27:51customer segmentation data set which is
  30738. 23:27:53available at the Kaggle's platform. Now
  30739. 23:27:56the third intermediate level project
  30740. 23:27:58that we have all over here is building a
  30741. 23:28:00chatbot with Rasa. This project aims to
  30742. 23:28:02build an intelligent chatbot using Rasa
  30743. 23:28:05framework. The chatbot will be capable
  30744. 23:28:06of understanding user queries and
  30745. 23:28:08providing appropriate responses making
  30746. 23:28:10it useful for customer support, personal
  30747. 23:28:13assistance or information retrieval.
  30748. 23:28:15What languages we are going to use? So
  30749. 23:28:17it will be Python based. We'll have the
  30750. 23:28:19NLP libraries like Rasa, NLTK, Spacey.
  30751. 23:28:22For machine learning libraries, we are
  30752. 23:28:24going to have scikitlearn, tensorflow,
  30753. 23:28:26kas, etc. For data handling libraries,
  30754. 23:28:28we are going to use pandas, numpy. So
  30755. 23:28:31guys, this was what we are going to do
  30756. 23:28:33it and how you can work on this one by
  30757. 23:28:35collecting the data, collect the
  30758. 23:28:37conversation data and FAQs from the
  30759. 23:28:39target domain, annotate the data to
  30760. 23:28:41create training examples for the
  30761. 23:28:43chatbot. Next comes is data
  30762. 23:28:45prep-processing. Clean the text data by
  30763. 23:28:47removing special characters and
  30764. 23:28:48normalizing the text. You can tokenize
  30765. 23:28:50and limitize the text to prepare for a
  30766. 23:28:53training. The third one we have the
  30767. 23:28:55models training. You can use Rasa's
  30768. 23:28:57NLU's component to train a model for
  30769. 23:28:59intent recognition and entity
  30770. 23:29:01extraction. You can define a dialog
  30771. 23:29:03management policies to handle different
  30772. 23:29:05conversation flows. Next guys, you can
  30773. 23:29:07perform the feature engineering and
  30774. 23:29:09integration. You can integrate the Ras
  30775. 23:29:11NLU and core components to build
  30776. 23:29:13complete chatbot. You can connect the
  30777. 23:29:15chatbot to messaging platform like
  30778. 23:29:17Facebook Messenger etc. For model
  30779. 23:29:19selection and testing what you can do
  30780. 23:29:21guys you can test this chatbot with
  30781. 23:29:23various inputs to ensure that it handles
  30782. 23:29:25the scenarios appropriately and you can
  30783. 23:29:28also select the right model using this.
  30784. 23:29:30Now if I talk about model training guys
  30785. 23:29:32what you have to do you have to collect
  30786. 23:29:34the user feedback and conversational
  30787. 23:29:36logs to continuously work on training
  30788. 23:29:38the model. Next, similarly you have to
  30789. 23:29:41retrain the model periodically with the
  30790. 23:29:43new data to see how it is working. So
  30791. 23:29:46that will be your evaluation. Now for
  30792. 23:29:48the hyperparameter tuning, what are you
  30793. 23:29:50going to do guys? You have to check in
  30794. 23:29:52those scenarios where it is able to tune
  30795. 23:29:54up with those scenarios where it can
  30796. 23:29:56handle the input appropriately. And next
  30797. 23:29:59is deployment. So guys, for deploying
  30798. 23:30:01it, you can use AWS. So guys, for a data
  30799. 23:30:04set, you can use Ras open source. So
  30800. 23:30:07that's a very good data set for you to
  30801. 23:30:09proceed. So guys, if you talk about
  30802. 23:30:11difficulty of this project, this is an
  30803. 23:30:13intermediate level project. Now let us
  30804. 23:30:15move on to the advanced level projects.
  30805. 23:30:17For advanced level projects, the first
  30806. 23:30:19one that comes up to my mind is movie
  30807. 23:30:21similarity from plot summaries. Now this
  30808. 23:30:24project aims to develop a system that
  30809. 23:30:26can find out recommended movies similar
  30810. 23:30:28to a given movie based on their plot
  30811. 23:30:31summaries. By analyzing the textual
  30812. 23:30:33content of the movie plot summaries, the
  30813. 23:30:35model will identify similarities and
  30814. 23:30:37suggests movies with similar themes,
  30815. 23:30:39story lines or genres. This project
  30816. 23:30:42introduces beginners to natural language
  30817. 23:30:44processing text similarly measures and
  30818. 23:30:46recommendation systems. What languages
  30819. 23:30:48we are going to use guys? We'll be using
  30820. 23:30:50Python NLP libraries like NLTK spacy
  30821. 23:30:53machine learning libraries like
  30822. 23:30:55scikitlearn data handling libraries like
  30823. 23:30:57pandas numpy visualization we can use
  30824. 23:31:00numpy and data set source will be IMDb
  30825. 23:31:02or kegel so guys this process is also
  30826. 23:31:05involving the data collection then you
  30827. 23:31:07have to go for data cleaning then
  30828. 23:31:09feature engineering next model selection
  30829. 23:31:12so similar process as I have discussed
  30830. 23:31:14in other projects so you have to also go
  30831. 23:31:16through the same one next what you have
  30832. 23:31:18to do device. Similarly, what you have
  30833. 23:31:21to do, you have to train the model, then
  30834. 23:31:23evaluate the model, then hypertune it
  30835. 23:31:25and finally proceed for the deployment.
  30836. 23:31:28So, this is overall process of this
  30837. 23:31:30project. Try to research on the website
  30838. 23:31:33a lot like how you can extract it. So,
  30839. 23:31:36guys, you can use kegel or towards data
  30840. 23:31:38science to research more about this
  30841. 23:31:40project. Now guys, if I talk about the
  30842. 23:31:42difficulty of this project and this is
  30843. 23:31:44an advanced level project. Now let us
  30844. 23:31:47move on to the next one that we have all
  30845. 23:31:50over here that is image segmentation
  30846. 23:31:52project for brain tumor prognosis. This
  30847. 23:31:55is a very very amazing project and
  30848. 23:31:57definitely you can put up on your
  30849. 23:31:58portfolio. Basically guys this project
  30850. 23:32:00aims to develop an image segmentation
  30851. 23:32:02model to identify and delinate brain
  30852. 23:32:05tumors from MRI scans. The goal is to
  30853. 23:32:07accurately segment the tumor regions
  30854. 23:32:09which can aid in prognosis treatment
  30855. 23:32:11planning surgical interventions. This
  30856. 23:32:13project introduces intermediate level
  30857. 23:32:15concepts of computer visions, deep
  30858. 23:32:17learning and medical image analysis. So
  30859. 23:32:19guys, what we'll be using all over here
  30860. 23:32:21for programming languages we can use
  30861. 23:32:23python for deep learning libraries we
  30862. 23:32:25can use tensorflow, kas, pytor. For
  30863. 23:32:27image processing libraries, we can use
  30864. 23:32:29opencv, scikit image. For data handling
  30865. 23:32:31libraries, we can use pandas, numpy. For
  30866. 23:32:34visualization, we can use mattplot,
  30867. 23:32:36cbond. Now if I talk about what is the
  30868. 23:32:39process of developing this project the
  30869. 23:32:41first step will be same data collection.
  30870. 23:32:44So next step you have to go for data
  30871. 23:32:46prep-processing. Third step you have to
  30872. 23:32:48do the model selection where you can use
  30873. 23:32:50CNN models for image segmentation task.
  30874. 23:32:53Then you go for model selection. Moving
  30875. 23:32:56ahead you're going to have the model
  30876. 23:32:58training and evaluation. You have to
  30877. 23:32:59split the data set into training and
  30878. 23:33:02validation and testing data sets. Next
  30879. 23:33:05proceed for the evaluation phase. Okay,
  30880. 23:33:07evaluate the model with certain metrics.
  30881. 23:33:09So here I can give you certain idea like
  30882. 23:33:11you can use dice coefficient,
  30883. 23:33:13intersection over union or accuracy.
  30884. 23:33:15Then go for hyperparameter tuning where
  30885. 23:33:17you have to optimize the model's
  30886. 23:33:19hyperparameters.
  30887. 23:33:20You can use grid search or random search
  30888. 23:33:22as we have discussed. And finally you
  30889. 23:33:24can deploy this model on AWS. Now guys
  30890. 23:33:27we have come to the final project. This
  30891. 23:33:30is also a very amazing project guys. So
  30892. 23:33:32guys the complexity level of this
  30893. 23:33:34project is advanced level.
  30894. 23:33:36Now let us move on to our final project
  30895. 23:33:39that is the impact of climate change on
  30896. 23:33:42birds. This is a very very amazing
  30897. 23:33:44project and definitely you can add it on
  30898. 23:33:46your resume. This project aims to
  30899. 23:33:48analyze the impact of climate change on
  30900. 23:33:50the bird population and migration
  30901. 23:33:52patterns by examining various climatic
  30902. 23:33:54factors and their correlation with bird
  30903. 23:33:56species data. The project seeks to
  30904. 23:33:58predict how climate change might affect
  30905. 23:34:00bird behavior and distribution. This
  30906. 23:34:02project will introduce you some advanced
  30907. 23:34:04level concepts like time series
  30908. 23:34:06analysis, environmental data modeling,
  30909. 23:34:08etc. So guys, what programming languages
  30910. 23:34:10we'll be using for data analysis? You
  30911. 23:34:12can see we'll have pandas, numpy. For
  30912. 23:34:14machine learning libraries, we are going
  30913. 23:34:15to have scikitlearn, tensorflow. For
  30914. 23:34:18visualization, we're going to have
  30915. 23:34:19mattplot lab, plotly. For geospatial, we
  30916. 23:34:22are going to have geopandas, folium. For
  30917. 23:34:24data source, we're going to have public
  30918. 23:34:26data sets on bird observation and
  30919. 23:34:27climate data sources from eird. Now what
  30920. 23:34:30will the process flow for this one guys?
  30921. 23:34:32First you have to proceed for data
  30922. 23:34:33collection. Gather bird observation from
  30923. 23:34:35the data like EIRD which provides
  30924. 23:34:37extensive record of bird sightings.
  30925. 23:34:40Okay. And for climate data you can
  30926. 23:34:42collect it from NOA including
  30927. 23:34:44temperature, precipitation and other
  30928. 23:34:46relevant climatic factors over the time.
  30929. 23:34:48And similar next process will be the
  30930. 23:34:50data prep-processing. Then you have to
  30931. 23:34:52proceed for feature engineering. Then
  30932. 23:34:54you have to go for model selection.
  30933. 23:34:56Okay. Moving ahead you have to go for
  30934. 23:34:58model training. then evaluation, then
  30935. 23:35:01hyperparameter tuning and finally you
  30936. 23:35:03have to deploy the model. So research
  30937. 23:35:05about this project, see what models you
  30938. 23:35:07are going to use. Suppose I can give you
  30939. 23:35:09a hint about this. You can use time
  30940. 23:35:10series analysis models like ARMA or ML
  30941. 23:35:14models. You can also use random forest
  30942. 23:35:16or gradient boosting for predicting
  30943. 23:35:18impact on the bird population. So guys,
  30944. 23:35:21use Google exhaustively to research
  30945. 23:35:23about this project. This is also a very
  30946. 23:35:25amazing project and it's going to give
  30947. 23:35:26you a lot of idea. Now if I talk about
  30948. 23:35:29the complexity level of this, it is an
  30949. 23:35:31advanced level project.
  30950. 23:35:33>> Welcome to deep learning interview
  30951. 23:35:34questions. My name is Richard Kersner
  30952. 23:35:37with the SimplyLearn team. That's
  30953. 23:35:39www.simplearn.com.
  30954. 23:35:41Get certified, get ahead. Today we're
  30955. 23:35:44going to help you prepare for interview
  30956. 23:35:45questions dealing with deep learning.
  30957. 23:35:47And we're going to go from the very
  30958. 23:35:48basics of neural networks and deep
  30959. 23:35:51learning into some of the more commonly
  30960. 23:35:54used models so you can have an
  30961. 23:35:56understanding of what kind of questions
  30962. 23:35:57are going to come up and what you need
  30963. 23:35:59to know in interview questions. We'll
  30964. 23:36:01start with a very general concept of
  30965. 23:36:03what is deep learning. This is where we
  30966. 23:36:05take large volumes of data in this case
  30967. 23:36:07on cats and dogs or whatever. A lot of
  30968. 23:36:09times you use um a training setup to
  30969. 23:36:12train your model. Remember it's kind of
  30970. 23:36:13like a magic black box going on there.
  30971. 23:36:16And then we use that to extract features
  30972. 23:36:17or extract information and in this case
  30973. 23:36:19classify the image of a cat and a dog.
  30974. 23:36:22So the primary takeaway we're talking
  30975. 23:36:23about deep learning is it learns from
  30976. 23:36:25large volumes of structured and even
  30977. 23:36:28unstructured data and uses complex
  30978. 23:36:30algorithms to train neural network. It
  30979. 23:36:32also performs complex operations to
  30980. 23:36:34extract hidden patterns and features.
  30981. 23:36:36And if we're going to discuss deep
  30982. 23:36:38learning in this very uh simplified
  30983. 23:36:40overview and we also have to go over
  30984. 23:36:42what is a neural network. This is a
  30985. 23:36:44common image you'll see of a drawing of
  30986. 23:36:46a forward propagation neural network and
  30987. 23:36:49it's it's a human brain inspired system
  30988. 23:36:51which replicate the way humans learn. So
  30989. 23:36:53this has inspired how our own neurons
  30990. 23:36:55and our brain fire but at a much
  30991. 23:36:57simplified level. Obviously it's not
  30992. 23:36:58ready to take over the human uh
  30993. 23:37:00population and and be our leader yet.
  30994. 23:37:02Not for many years. It's very much in
  30995. 23:37:04its infant stage. But it's inspired by
  30996. 23:37:06how our brains work. Um and they use a
  30997. 23:37:08lot of other inspirations. You can study
  30998. 23:37:10brains of moths and other animals that
  30999. 23:37:13they've used to figure out how to
  31000. 23:37:14improve these neural networks. The most
  31001. 23:37:16common one consists of three layers of
  31002. 23:37:18network and this is generally how you
  31003. 23:37:20view these networks is you have an
  31004. 23:37:21input, you have a hidden layer and an
  31005. 23:37:23output. And the neural network is uh
  31006. 23:37:26broken up into many pieces. But when we
  31007. 23:37:28focus just on the neural network, it's
  31008. 23:37:30always on the hidden layers that we're
  31009. 23:37:31making all the adjustments and figuring
  31010. 23:37:33out how to best set up those hidden
  31011. 23:37:36layers for their functions to both train
  31012. 23:37:39faster and to function better. When we
  31013. 23:37:41look at this, of course, we have our
  31014. 23:37:42input, hidden, and output. Each layer
  31015. 23:37:44contains neurons called as nodes perform
  31016. 23:37:47various operations. And you can see here
  31017. 23:37:49we have the list of the nodes. We have
  31018. 23:37:51both our input nodes and our output
  31019. 23:37:52nodes and then our hidden layer nodes.
  31020. 23:37:54And it's used in deep learning algorithm
  31021. 23:37:56like CNN, RNN, GN, etc. We'll address
  31022. 23:38:00some of these models a little closer, at
  31023. 23:38:01least the most common models as we go
  31024. 23:38:03down the list and we study the deep
  31025. 23:38:05learning and the neural network
  31026. 23:38:06framework. Let's start with what is a
  31027. 23:38:09multi-layer perceptron or MLP a lot of
  31028. 23:38:12time as they're referred to. And you'll
  31029. 23:38:14see these abbreviations. I'll be honest,
  31030. 23:38:16I have to write them down on a piece of
  31031. 23:38:18paper and go through them because I
  31032. 23:38:19never remember what they all mean even
  31033. 23:38:21though I play with them all the time.
  31034. 23:38:22What is a multi-layer perceptron? Well,
  31035. 23:38:24if you look at the image on the right,
  31036. 23:38:26it's very similar to what we just looked
  31037. 23:38:27at. You have your input layer, your
  31038. 23:38:29hidden layer, and your output layer. And
  31039. 23:38:31that's exactly what this is. It has the
  31040. 23:38:33same structure of a single layer
  31041. 23:38:34perceptron with one or more hidden
  31042. 23:38:36layers except the input layer, each node
  31043. 23:38:38in the other layers uses a nonlinear
  31044. 23:38:40activation function. What that means is
  31045. 23:38:43your input layer is your data coming in
  31046. 23:38:45and then your activation function is
  31047. 23:38:47based upon all those nodes and weights
  31048. 23:38:49being added together and then it has the
  31049. 23:38:51output. MLP uses supervised learning
  31050. 23:38:54method called back propagation for
  31051. 23:38:56training the model. Very key word there
  31052. 23:38:58is back propagation. Single layer
  31053. 23:39:00perceptron can classify only linear
  31054. 23:39:02separable classes with binary output 01.
  31055. 23:39:06But the MLP can classify nonlinear
  31056. 23:39:08classes. So let's break this down just a
  31057. 23:39:10little bit. The multi-layer perceptron
  31058. 23:39:12with an input layer and a hidden layer
  31059. 23:39:14and an output layer. As you see that it
  31060. 23:39:16comes in there, it has adds up all the
  31061. 23:39:18numbers and weights depending on how
  31062. 23:39:20your setup is. That then goes to the
  31063. 23:39:22next layer. That then goes to the next
  31064. 23:39:23hidden layer if you have multiple hidden
  31065. 23:39:25layers. And finally to the output layer.
  31066. 23:39:27The back propagation takes the error
  31067. 23:39:30that it sees. So whatever the output is,
  31068. 23:39:32it says, hey, this has an error to it.
  31069. 23:39:33It's wrong. And then sends that error
  31070. 23:39:35backwards from where it came from. And
  31071. 23:39:37there's a lot of different functions
  31072. 23:39:39used to uh train this based on that
  31073. 23:39:42error and how that error goes backwards
  31074. 23:39:44in the notes. Uh so forward is you get
  31075. 23:39:46your answers. Backward is for training.
  31076. 23:39:49You see this every day. Even my uh
  31077. 23:39:51Google Pixel phone has this. It they
  31078. 23:39:53train the neural network which takes a
  31079. 23:39:55lot more data to train than it does to
  31080. 23:39:57use. And then they load up that neural
  31081. 23:39:59network into in this case I have a Pixel
  31082. 23:40:012 which actually has a built-in neural
  31083. 23:40:03network for processing pictures. And so
  31084. 23:40:05it's just the forward propagation I use
  31085. 23:40:07when it processes my photos, but when
  31086. 23:40:10they were training it, you use the back
  31087. 23:40:11propagation to train it with the errors
  31088. 23:40:13they had. We'll be coming back to
  31089. 23:40:15different models that are used. For
  31090. 23:40:17right now though, multi-layer
  31091. 23:40:18perceptron, MLP, put that down as your
  31092. 23:40:21vocabulary word and of course back
  31093. 23:40:24propagation. What is data normalization
  31094. 23:40:26and why do we need it? This is so
  31095. 23:40:29important. We spend so much time in
  31096. 23:40:30normalizing our data and getting our
  31097. 23:40:32data clean and setting it up. Uh so we
  31098. 23:40:34talk about data there's a pre-processing
  31099. 23:40:36step to standardize the data. So
  31100. 23:40:38whatever we have coming in we don't want
  31101. 23:40:40it to be a uh you know one gigabyte file
  31102. 23:40:43here a 2 GBTE picture here and a 3
  31103. 23:40:46kilobyte text there. Even as a human I
  31104. 23:40:49can't process those all in the same
  31105. 23:40:50group. I have to reformat them in some
  31106. 23:40:52way that loops them together so they're
  31107. 23:40:54a standardized format. We use this uh
  31108. 23:40:56data normalization and and
  31109. 23:40:58pre-processing to reduce and eliminate
  31110. 23:41:01data redundancy. A lot of times the data
  31111. 23:41:04comes in and you end up with two of the
  31112. 23:41:06same images or uh uh the same
  31113. 23:41:08information in different formats. Then
  31114. 23:41:10we want to rescale values to fit into a
  31115. 23:41:12particular range for achieving better
  31116. 23:41:14convergence. What this means is with
  31117. 23:41:17most neural networks they form a bias.
  31118. 23:41:20We've seen this in recently in attacks
  31119. 23:41:22on neural networks where they light up
  31120. 23:41:24one pixel or one piece of the view and
  31121. 23:41:26it skews the whole answer. So suddenly u
  31122. 23:41:29because one pixel is really bright uh it
  31123. 23:41:32doesn't know what to do. Well when we
  31124. 23:41:34start rescaling it we put all the values
  31125. 23:41:36between say minus one and one and we
  31126. 23:41:38change them and refit them to those
  31127. 23:41:40values. It helps get rid of that bias
  31128. 23:41:42helps fix for some of those problems.
  31129. 23:41:44And then finally we restructure the data
  31130. 23:41:46and improve the integrity. We want to
  31131. 23:41:47make sure that we're not missing values
  31132. 23:41:49um or we don't have partial data coming
  31133. 23:41:51in. One way to look at this is uh bad
  31134. 23:41:54data in bad data out. And so you want
  31135. 23:41:57clean data in and you want good answers
  31136. 23:41:59coming out. One of the most basic models
  31137. 23:42:01used is a Boltzman machine. So let's
  31138. 23:42:04address what is a Boltzman machine. And
  31139. 23:42:06if you know we just did the MLP
  31140. 23:42:08multi-layer perceptron. So now we're
  31141. 23:42:10going to come into almost a simplified
  31142. 23:42:12version of that. And in this we have our
  31143. 23:42:14visible input layer and we have our
  31144. 23:42:16hidden layer. The Boltzman machines are
  31145. 23:42:18almost always shallow. They're usually
  31146. 23:42:19just two-layer neural nets that make
  31147. 23:42:21stochastic decisions whether a neuron
  31148. 23:42:24should be on or off. True or false? Yes.
  31149. 23:42:26No. First layer is a visible layer and
  31150. 23:42:28second layer is the hidden layer. Nodes
  31151. 23:42:30are connected to each other across
  31152. 23:42:32layers, but no two nodes of the same
  31153. 23:42:34layer are connected. Hence, it is also
  31154. 23:42:36known as restricted Boltzman machine.
  31155. 23:42:39Now that we've covered a basic MLP or
  31156. 23:42:41multi-layer perceptron, and we've gone
  31157. 23:42:44over the Boltzman machine, also known as
  31158. 23:42:46the restricted Boltzman machine, let's
  31159. 23:42:47talk a little bit about activation
  31160. 23:42:49formulas. And this is a huge topic that
  31161. 23:42:53can get really complicated but it also
  31162. 23:42:55is automated. So it's very simple. So
  31163. 23:42:57you have both a complicated and a simple
  31164. 23:42:59at the same time. So what is the role of
  31165. 23:43:02activation functions in a neural
  31166. 23:43:03network? Activation function decides
  31167. 23:43:05whether a neuron should be fired or not.
  31168. 23:43:08That's the most basic one and that
  31169. 23:43:10actually changes a little bit because
  31170. 23:43:11it's either whether fired or not in this
  31171. 23:43:13case activation function or what value
  31172. 23:43:16should come out when it's fired. But in
  31173. 23:43:17these models, we're looking at just the
  31174. 23:43:19boltsman restricted layers. So this is
  31175. 23:43:22what causes them to fire. Either they
  31176. 23:43:24don't or they do. It's a yes or no,
  31177. 23:43:26true, false, all or nothing. It accepts
  31178. 23:43:28the weighted sum of the inputs, the bias
  31179. 23:43:30as input to any activation function. So
  31180. 23:43:33whatever activation function is, it
  31181. 23:43:35needs to have the sum of the weights
  31182. 23:43:37times the input. So each input, if you
  31183. 23:43:39remember on that model, and let's just
  31184. 23:43:40go back to that model real quick. And
  31185. 23:43:42then you always have to add a bias. And
  31186. 23:43:44you can look at the bias if you remember
  31187. 23:43:46from your uklitian geometry. You draw a
  31188. 23:43:49straight line. Formula for that line has
  31189. 23:43:51a y-coordinate at the end. It's always
  31190. 23:43:54um cx plus m or something like that
  31191. 23:43:56where m is where it crosses the
  31192. 23:43:58ycoordinates. If you're doing a straight
  31193. 23:44:00line with these weights, it's very
  31194. 23:44:02similar, but a lot of times we just add
  31195. 23:44:04it in as its own weight. We take it as a
  31196. 23:44:07node of a one value coming in and then
  31197. 23:44:09we compute its new weight. And that's
  31198. 23:44:11how we compute that bias just like we
  31199. 23:44:12compute all the other weights coming in.
  31200. 23:44:14The node which gets fired depends on the
  31201. 23:44:16y value. And then we have a step
  31202. 23:44:18function. And the step function this is
  31203. 23:44:20where remember I said it's going to get
  31204. 23:44:21complicated and simple all at the same
  31205. 23:44:23time. We have a lot of different step
  31206. 23:44:25functions. We have the sigmoid function.
  31207. 23:44:28We have just a standard step function.
  31208. 23:44:30We have the ru is pronounced like ray
  31209. 23:44:33the ray of from the sun and lu like a
  31210. 23:44:35name. So ru function. And we have the
  31211. 23:44:37tangent h function. And if you look at
  31212. 23:44:39these, they all have something similar.
  31213. 23:44:41They all either force it to be um one
  31214. 23:44:44value or the other. They force it to be
  31215. 23:44:45in the case of the first three a zero or
  31216. 23:44:47one. And in the last one, it's either a
  31217. 23:44:50minus one or one. And you can easily
  31218. 23:44:51convert that to a 0, one, yes, no, true,
  31219. 23:44:54false. And on this, one of the most
  31220. 23:44:56common ones is the step function itself
  31221. 23:44:58because there is no middle value. There
  31222. 23:45:00is no um uh discrepancy that says, well,
  31223. 23:45:03I'm not quite sure. But as you get into
  31224. 23:45:05different models, probably the most
  31225. 23:45:07commonly used used to be the sigmoid was
  31226. 23:45:09most commonly used, but I see the relu
  31227. 23:45:11used more often. Really, depending on
  31228. 23:45:13what you're doing, you just have to play
  31229. 23:45:15with these and find out which one works
  31230. 23:45:16best depending on the data in your
  31231. 23:45:19output. The reason to have a non01
  31232. 23:45:22answer or something kind of in the
  31233. 23:45:24middle is when you're looking at this
  31234. 23:45:25and it's coming out, you can actually
  31235. 23:45:27process that middle ground as part of
  31236. 23:45:30the answer into another neural network.
  31237. 23:45:32So it might be that the relu function
  31238. 23:45:34says hey this is only a 6 not a one and
  31239. 23:45:38uh even though the one is what's going
  31240. 23:45:41into the next neural network or the next
  31241. 23:45:43hidden layer as an input the 6 value
  31242. 23:45:47might also be going in there to let you
  31243. 23:45:49know hey this is not a straight up one
  31244. 23:45:51or straight up zero it's someplace in
  31245. 23:45:53the middle this is a little uncertain
  31246. 23:45:54what's coming out here so it's a very
  31247. 23:45:56powerful tool in the basic neural
  31248. 23:45:58network you usually just use the step
  31249. 23:45:59function it's yes or no let's take a um
  31250. 23:46:02a big step back and take a kind of an
  31251. 23:46:05overview. The next function is what is a
  31252. 23:46:08cost function that we're going to cover.
  31253. 23:46:10This is so important because this is
  31254. 23:46:12your end result that you're going to do
  31255. 23:46:14over and over again and use to decide
  31256. 23:46:16whether the model is working or not,
  31257. 23:46:18whether you need to try a different step
  31258. 23:46:19function, whether you need to try a
  31259. 23:46:21different activation, whether you need
  31260. 23:46:22to try a fully different model used. Uh
  31261. 23:46:24so what is the cost function? Cost
  31262. 23:46:26function is a measure to evaluate how
  31263. 23:46:29good your model's performance is. It is
  31264. 23:46:31also referred as loss or error used to
  31265. 23:46:34compute the error of the output layer
  31266. 23:46:36during back propagation. There's our
  31267. 23:46:38back propagation where we're training
  31268. 23:46:39our model. That's one of our key words.
  31269. 23:46:42Mean squared error is an example of a
  31270. 23:46:44popular cost function. And so here we
  31271. 23:46:46have the cost function C = half of Y - Y
  31272. 23:46:50predicted. Um and then you square that.
  31273. 23:46:52So the first thing is um you know real
  31274. 23:46:54quick if you haven't done statistics
  31275. 23:46:56this is not a percentage. It's not a
  31276. 23:46:58percentage of how accurate it is. is
  31277. 23:47:00just a measurement of the error and we
  31278. 23:47:02take that error if we're training it and
  31279. 23:47:04we push that error backwards through the
  31280. 23:47:06neural network and we use that through
  31281. 23:47:08the different training functions
  31282. 23:47:10depending on what model you're using to
  31283. 23:47:12train the neural network. So when you
  31284. 23:47:14deploy the network you're usually done
  31285. 23:47:15training it because it takes a lot of
  31286. 23:47:17computational force to train it. Um this
  31287. 23:47:19is a very simple model and so you deploy
  31288. 23:47:21the train one. Uh but we want to know
  31289. 23:47:22how your error is and so how do we do
  31290. 23:47:24that? Well you split your data. part of
  31291. 23:47:26your data is for trading and part of
  31292. 23:47:28your data is for testing. And then we
  31293. 23:47:30can also test the error on there. So
  31294. 23:47:32it's very important. And then we're
  31295. 23:47:33going to go one more step on this. We
  31296. 23:47:36got to look at both the local and the
  31297. 23:47:38global setup. It might work great to
  31298. 23:47:40test your data on what you have on your
  31299. 23:47:42computer, but that's different than in
  31300. 23:47:44the field. So, when we're talking about
  31301. 23:47:46all these different tests and the error
  31302. 23:47:48test as far as your loss, you don't you
  31303. 23:47:50want to make sure that you're in a
  31304. 23:47:52closed environment when you do initial
  31305. 23:47:53testing, but you also want to open that
  31306. 23:47:55up and make sure you follow up with the
  31307. 23:47:56testing on the larger scale of data
  31308. 23:47:58because it will change. It might not fit
  31309. 23:47:59the larger scale. There might be
  31310. 23:48:01something in there in the way you
  31311. 23:48:02brought the data in specifically or the
  31312. 23:48:04data group you used or um any of those
  31313. 23:48:06could cause an error. So, it's very
  31314. 23:48:08important to remember that we're looking
  31315. 23:48:09at both the local and the global context
  31316. 23:48:11of our error. And just one other side
  31317. 23:48:14note on a lot of the newer models of
  31318. 23:48:16neural networks by comparing the error
  31319. 23:48:19we get on the data our training data
  31320. 23:48:21with a portion of the test data we can
  31321. 23:48:24actually figure out how good the model
  31322. 23:48:25is whether it's overfitted or not. We'll
  31323. 23:48:27go into that a little bit more as we go
  31324. 23:48:29into some of the different models. So we
  31325. 23:48:31have our output. We're able to um figure
  31326. 23:48:34out the error on it based on the square
  31327. 23:48:35means usually although there's other uh
  31328. 23:48:37functions used. So we want to talk about
  31329. 23:48:39what is gradient descent? Another
  31330. 23:48:41vocabulary word gradient descent is an
  31331. 23:48:44optimation algorithm to minimize the
  31332. 23:48:47cost function or to minimize the error.
  31333. 23:48:49Aim is to find the local or global
  31334. 23:48:51minima of a function. Determine the
  31335. 23:48:53direction the model should take to
  31336. 23:48:55reduce the error. So as we're looking at
  31337. 23:48:57this, we have our uh squared error that
  31338. 23:48:59we just figured out the co based on the
  31339. 23:49:01cost function. It says how bad is my
  31340. 23:49:03model fitting the data I just put
  31341. 23:49:05through it. And then we want to reduce
  31342. 23:49:07that error. So how do you figure out
  31343. 23:49:08what direction to do that in? Well, it
  31344. 23:49:10could be that you're looking at just
  31345. 23:49:12that line of that line of data coming
  31346. 23:49:14in. So that would be a local minima. We
  31347. 23:49:16want to know the error of that
  31348. 23:49:17particular setup coming in. And then you
  31349. 23:49:19have your global your global minima. We
  31350. 23:49:21want to minimize it based on the overall
  31351. 23:49:24data we're putting through it. And with
  31352. 23:49:25this we can figure out the global
  31353. 23:49:28minimum cost. We want to take all those
  31354. 23:49:30local minimum costs of each piece of
  31355. 23:49:33data coming in and figure out the global
  31356. 23:49:34one. How are we going to adjust this
  31357. 23:49:36model to fit all the data? We don't want
  31358. 23:49:38it to be biased just on three or four
  31359. 23:49:40lines of data coming in. We want it to
  31360. 23:49:42kind of extrapolate a general answer for
  31361. 23:49:45all the data coming in. But this of
  31362. 23:49:46course uh we mentioned it briefly about
  31363. 23:49:49back propagation. This is where really
  31364. 23:49:51comes in handy is training our model.
  31365. 23:49:53Neural network technique to minimize the
  31366. 23:49:56cost function helps to improve the
  31367. 23:49:58performance of the network. Back
  31368. 23:49:59propagates the error and updates the
  31369. 23:50:01weights to reduce the error. So as you
  31370. 23:50:03can see here is a very nice depiction of
  31371. 23:50:05a back propagation. We have our
  31372. 23:50:08predicted y coming out and then we have
  31373. 23:50:10since it's a training set we already
  31374. 23:50:12know the answer and the answer comes
  31375. 23:50:13back and based on case of the square
  31376. 23:50:16means was one of the functions we looked
  31377. 23:50:18at uh one of the activation functions
  31378. 23:50:20based on cost function that cost
  31379. 23:50:22function then depending on what you
  31380. 23:50:24choose for your back propagation method
  31381. 23:50:26and there's a number of them will change
  31382. 23:50:28the weights it will change the weight
  31383. 23:50:29going to each of one of those nodes in
  31384. 23:50:31the hidden layer and then based upon the
  31385. 23:50:34error that's still being carried back
  31386. 23:50:35it'll change the weights going to the
  31387. 23:50:37next hidden layer and then it computes
  31388. 23:50:39an error level on that and sends that
  31389. 23:50:41back up. And you're going to say, well,
  31390. 23:50:43if it computes the error into the first
  31391. 23:50:44hidden layer and fixes it, why would it
  31392. 23:50:46stop there? Well, remember, we don't
  31393. 23:50:48want to create a biased neural network.
  31394. 23:50:52So, we only make small adjustments on
  31395. 23:50:54these weights. We don't make a big
  31396. 23:50:56adjustment that changes everything right
  31397. 23:50:57off the bat. So, no matter how far back
  31398. 23:50:59you go, you're always going to have a
  31399. 23:51:00small amount of error, and that's still
  31400. 23:51:02going to continue to go all the way back
  31401. 23:51:03up the hidden layers. For right now,
  31402. 23:51:05focus on the back propagation is taking
  31403. 23:51:08that error and moving it backwards on
  31404. 23:51:11the neural network to change the weights
  31405. 23:51:13and help program it so that it'll have
  31406. 23:51:15the correct answers. So far, we've been
  31407. 23:51:17talking about forward propagation neural
  31408. 23:51:20networks. Everything goes forwards, goes
  31409. 23:51:22left to right. Uh but let's let's take a
  31410. 23:51:23little detour and let's see what is the
  31411. 23:51:25difference between a feed forward neural
  31412. 23:51:27network and a recurrent neural network.
  31413. 23:51:30Now, this is in the function, not when
  31414. 23:51:31we're training it using the back
  31415. 23:51:33propagation. So, you've got new
  31416. 23:51:34information coming in and you want to
  31417. 23:51:36get the answer and there's a couple
  31418. 23:51:37different networks out there and we want
  31419. 23:51:39to know we have a feed forward neural
  31420. 23:51:40network and we have a new uh vocabulary
  31421. 23:51:42term recurrent neural network. A feed
  31422. 23:51:45forward neural network signals travel in
  31423. 23:51:47one direction from input to output. No
  31424. 23:51:50feedback loops considers only the
  31425. 23:51:52current input cannot memorize previous
  31426. 23:51:54inputs. One example of one of these feed
  31427. 23:51:58forward neural networks. And we've
  31428. 23:51:59covered a number of them, but one of the
  31429. 23:52:00ones that has a big highlight nowadays
  31430. 23:52:02is the CNN, a convolutional neural
  31431. 23:52:05network. TensorFlow, the one put out by
  31432. 23:52:07Google is probably most known for their
  31433. 23:52:10CNN, where the information goes forward.
  31434. 23:52:12It uh first takes a picture, splits it
  31435. 23:52:15apart, goes through the individual
  31436. 23:52:16pixels on the picture, so it picks up a
  31437. 23:52:18different reading, then calculates based
  31438. 23:52:20on that, goes into a regular feed
  31439. 23:52:22forward neural network, and then gives
  31440. 23:52:24you a categorization on there. Now,
  31441. 23:52:26we're not covering the CNN today, but we
  31442. 23:52:28do have a video out that you can look up
  31443. 23:52:30on YouTube put out by SimplyLearn, the
  31444. 23:52:33convolutional neural network. wonderful
  31445. 23:52:35tutorial. Check that out and learn a lot
  31446. 23:52:37more about the convolutional neural
  31447. 23:52:38network. But you do need to know that
  31448. 23:52:40the CNN is a forward propagation neural
  31449. 23:52:43network only. So it's only moving in one
  31450. 23:52:45direction. So we want to look at a
  31451. 23:52:46recurrent neural network. Signals travel
  31452. 23:52:48in both directions making it a looped
  31453. 23:52:51network. Considers the current input
  31454. 23:52:53along with the previous received inputs
  31455. 23:52:55for generating the output of a layer.
  31456. 23:52:57Has the ability to memorize past data
  31457. 23:52:59due to its internal memory. And you can
  31458. 23:53:01see they have a nice uh image here. We
  31459. 23:53:03have our um input and for some reason
  31460. 23:53:06they always do the recurrent neural
  31461. 23:53:07network um in reverse from bottom up in
  31462. 23:53:10the images. It's kind of a standard
  31463. 23:53:11although I'm not sure why. Your X goes
  31464. 23:53:13into your hidden layer and your hidden
  31465. 23:53:16layer the answer for part of the answer
  31466. 23:53:17from that it generates feeds back into
  31467. 23:53:20the hidden layer. So now you have an
  31468. 23:53:22input of both X and part of the hidden
  31469. 23:53:24layer and then that feeds into your
  31470. 23:53:25output. Now if we go back to the forward
  31471. 23:53:28let me just go back a slide and we're
  31472. 23:53:30looking at uh our forward propagation
  31473. 23:53:32network. One of the tricks you can do to
  31474. 23:53:34use just a forward propagation network
  31475. 23:53:37is if you're in a what they call a time
  31476. 23:53:39sequence, that's a good uh term to
  31477. 23:53:41remember or a time series meaning that
  31478. 23:53:43it's sequential data. Each term comes
  31479. 23:53:45after the other. You can trick this by
  31480. 23:53:48creating your input nodes as with the
  31481. 23:53:51history. So if you know that uh you have
  31482. 23:53:53values one, five and seven going in and
  31483. 23:53:55you know what the output is from one
  31484. 23:53:57what those outputs are, you can expand
  31485. 23:53:59the input to include the history input.
  31486. 23:54:02That's one of the ways to trick a
  31487. 23:54:03forward propagation network into looking
  31488. 23:54:05at that. But when you do with a
  31489. 23:54:06recurrent neural network, you let the
  31490. 23:54:09hidden layer do that for you. It sends
  31491. 23:54:11that data and reprocesses it back into
  31492. 23:54:13itself. What are some of the
  31493. 23:54:14applications of recurrent neural
  31494. 23:54:16network? The RNN can be used for
  31495. 23:54:19sentiment analysis and text mining.
  31496. 23:54:21Getting up early in the morning is good
  31497. 23:54:23for health and it's a positive
  31498. 23:54:24sentiment. One of the catches you really
  31499. 23:54:26want to look at this when you're looking
  31500. 23:54:27at the language is that I could switch
  31501. 23:54:29this around and totally negate the
  31502. 23:54:32meaning of what I'm doing. So, it no
  31503. 23:54:33longer be positive. So, when you're
  31504. 23:54:35looking at a sentence, knowing the order
  31505. 23:54:37of the words is as important as the
  31506. 23:54:39meaning of the words. You can't just
  31507. 23:54:41count how many good words there are
  31508. 23:54:42versus bad words to get positive
  31509. 23:54:45sentiment. You know, have to know what
  31510. 23:54:46they're addressing. And there's lots of
  31511. 23:54:48other different uses. Uh, kids are
  31512. 23:54:50playing football or soccer as we call it
  31513. 23:54:52in the US. RN can help you caption an
  31514. 23:54:54image. So based on previous information
  31515. 23:54:56coming in, it refeeds that back in and
  31516. 23:54:58you have a image setter. And then time
  31517. 23:55:01series problems like predicting the
  31518. 23:55:03prices of stocks in a month or quarter
  31519. 23:55:05or sell of product can be solved using
  31520. 23:55:07an RNN. And this is a really good
  31521. 23:55:10example. You have whatever your stocks
  31522. 23:55:11were doing earlier this month will have
  31523. 23:55:13a huge effect of what they're doing
  31524. 23:55:15today if you're investing. So having an
  31525. 23:55:17RNN model, a recurrent neural network
  31526. 23:55:19feeding into itself what was happening
  31527. 23:55:21previously allows it to take that model
  31528. 23:55:23and program in that whole series without
  31529. 23:55:25having to put in the whole a month at a
  31530. 23:55:28time of data. You can only put in one
  31531. 23:55:30day at a time. But if you keep them in
  31532. 23:55:31order, it will look back and say, "Oh,
  31533. 23:55:33this because of what happened yesterday,
  31534. 23:55:34I need some information from that and
  31535. 23:55:35I'm going to use that to help predict
  31536. 23:55:37today's." And so on and so on. We're
  31537. 23:55:39going to go back to our activation
  31538. 23:55:40functions. Remember I told you uh ReLU
  31539. 23:55:43was one of the most common functions
  31540. 23:55:44used. Uh so let's talk a little bit more
  31541. 23:55:46about ReLU and also softmax. Softmax is
  31542. 23:55:49an activation function that generates
  31543. 23:55:51the output between zero and one. It
  31544. 23:55:54divides each output such that the total
  31545. 23:55:56sum of the outputs is equal to one. It
  31546. 23:55:58is often used in the output layers.
  31547. 23:56:00Softmax L of the N equals E to L the N
  31548. 23:56:03over the absolute value of E to the L.
  31549. 23:56:05So what does this function mean? I mean
  31550. 23:56:07what is actually going on here? So we
  31551. 23:56:09have our uh output nodes and our output
  31552. 23:56:12nodes are giving us uh let's say they
  31553. 23:56:13gave us 1.2.9 and point4. As a human
  31554. 23:56:17being I look at that and I say well the
  31555. 23:56:19greatest value is 1.2. So whatever
  31556. 23:56:21category that is if you have three
  31557. 23:56:23different categories maybe you're not
  31558. 23:56:24just doing if it's a cat or it's a dog
  31559. 23:56:27or u oh let's say it's a cow. We had
  31560. 23:56:29cats and dogs earlier. Why the cats and
  31561. 23:56:31dogs are hanging out with a cow. I don't
  31562. 23:56:33know. But we have a value and it might
  31563. 23:56:34say 1.2 2 is a cat, 0.9 is the dog, and
  31564. 23:56:38point4 is a cow. Uh, for some reason, it
  31565. 23:56:40thinks that there's a chance of it being
  31566. 23:56:41any one of these three items, and that's
  31567. 23:56:43how it comes out of the output layer.
  31568. 23:56:44Well, as a human, I can look at 1.2 and
  31569. 23:56:47say this is definitely what it is. It's
  31570. 23:56:48definitely a cat or whatever it is. Uh,
  31571. 23:56:50maybe it's looking at different kinds of
  31572. 23:56:51cars might be a better whether it's a
  31573. 23:56:53car, truck, or a motorcycle. Maybe
  31574. 23:56:55that'd be a better example. Well, from a
  31575. 23:56:57computer standpoint, that might be a
  31576. 23:56:59little confusing because they're just
  31577. 23:57:00numbers waving at us. And so with the
  31578. 23:57:02soft max, we want all those numbers to
  31579. 23:57:06always add up to one. So when I add
  31580. 23:57:08three numbers together, I want the final
  31581. 23:57:10output to be one on there. And so it
  31582. 23:57:12goes through this formula changes each
  31583. 23:57:13of these numbers. In this case, it
  31584. 23:57:15changes them to 46.34
  31585. 23:57:18and 2. They all add up to one. And
  31586. 23:57:20that's a lot easier to register because
  31587. 23:57:22it's very set. It's a set output. It's
  31588. 23:57:24never going to be more than one. It's
  31589. 23:57:26never going to be less than zero. And so
  31590. 23:57:27you can see here that there's probably a
  31591. 23:57:29pretty high chance that it's the first
  31592. 23:57:30one. So you're as a human being, we have
  31593. 23:57:32no problem knowing that. But this output
  31594. 23:57:34can then also go into say another input.
  31595. 23:57:37So it might be an automated car that's
  31596. 23:57:39picking up images and it says that image
  31597. 23:57:41in front of us is probably a big truck.
  31598. 23:57:43We should deal with it like it's a big
  31599. 23:57:45truck. It's probably not a motorcycle.
  31600. 23:57:46Um or whatever those categories are.
  31601. 23:57:48That's the softmax part of it. But now
  31602. 23:57:50we have the ru. Well, what where's the
  31603. 23:57:52ru coming from? Well, the ru is what's
  31604. 23:57:54generating the 1.2 and the 0.9 and the
  31605. 23:57:57point4. And so if you remember our relu
  31606. 23:57:59stands for rectified linear unit and is
  31607. 23:58:03the most widely used activation
  31608. 23:58:05function. We looked at a number of
  31609. 23:58:06different activation functions including
  31610. 23:58:08tangent h the step function. Remember I
  31611. 23:58:10said the step function is really used if
  31612. 23:58:12that's what your actual output is
  31613. 23:58:13because then you know it's a zero or
  31614. 23:58:15one. But the relu if you have that as
  31615. 23:58:17your output you now have a discrepancy
  31616. 23:58:19in there. And if that's going into
  31617. 23:58:21another neural network or another
  31618. 23:58:23process having that discrepancy is
  31619. 23:58:25really important. and it gives an output
  31620. 23:58:26of x if x is positive and zero
  31621. 23:58:29otherwise. So it says my x value is
  31622. 23:58:31going to be somewhere between zero or
  31623. 23:58:33one and then the uh usually unless it's
  31624. 23:58:36really uncertain the output's usually a
  31625. 23:58:38one or zero and then you have that
  31626. 23:58:40little piece of uncertainty there that
  31627. 23:58:41you can send forward to another network
  31628. 23:58:43or you can look at to know that there's
  31629. 23:58:45uncertainty involved and is often used
  31630. 23:58:47in the hidden layers. This is what's
  31631. 23:58:49coming out of the hidden layers into the
  31632. 23:58:50output layer usually or as we reference
  31633. 23:58:53the uh convolution neural network the
  31634. 23:58:56CNN you'd have to go to another video to
  31635. 23:58:58review the RLU is the most common used
  31636. 23:59:01for convolutional part of that network
  31637. 23:59:05has a bunch of little pieces that are
  31638. 23:59:06very simplified looking at all the
  31639. 23:59:08different images or different sections
  31640. 23:59:09of the map and the RLU works really good
  31641. 23:59:12for that like I said there's other
  31642. 23:59:13formulas used but that this is the most
  31643. 23:59:15common one and you'll see that in the
  31644. 23:59:17hidden layers going maybe between one
  31645. 23:59:18layer and the next layer. So just a
  31646. 23:59:20quick recap, we have our soft max, which
  31647. 23:59:23means that if you have uh numerous
  31648. 23:59:25categories, only one of them is going to
  31649. 23:59:27be picked, but you also want to have
  31650. 23:59:29some value attached to it, how well it
  31651. 23:59:31picked it, and you put that between 01.
  31652. 23:59:33So it's very uh standardized. So we have
  31653. 23:59:35our soft max. We looked at that. Let's
  31654. 23:59:37go back one. We looked at that here
  31655. 23:59:38where it transforms the numbers. And
  31656. 23:59:39then we have our ReLU function which
  31657. 23:59:41takes the information in the summation
  31658. 23:59:44and puts it between a zero and a one
  31659. 23:59:46where it's either clearly a zero or
  31660. 23:59:48depending on how confident our model is,
  31661. 23:59:52it'll go between the zero and one value.
  31662. 23:59:54What are hyperparameters? Oh, this is a
  31663. 23:59:57great interview question.
  31664. 23:59:58Hyperparameters. When you are doing
  31665. 24:00:00neural networks, this is what you're
  31666. 24:00:01playing with most of the time once
  31667. 24:00:03you've gotten the data formatted
  31668. 24:00:05correctly. A hyperparameter is a
  31669. 24:00:07parameter whose value is set before the
  31670. 24:00:09learning process begins. Determines how
  31671. 24:00:12a network is trained and the structure
  31672. 24:00:14of the network. This includes things
  31673. 24:00:15like the number of hidden units, how
  31674. 24:00:17many hidden layers are you going to have
  31675. 24:00:18and how many nodes in each layer.
  31676. 24:00:20Learning rate. Learning rate is usually
  31677. 24:00:23multiplied once you figured out the
  31678. 24:00:25error and how much you want to change
  31679. 24:00:26the weights. We talked about or I
  31680. 24:00:28mentioned it earlier just briefly. You
  31681. 24:00:29don't want to just make a huge change
  31682. 24:00:31otherwise you're going to have a biased
  31683. 24:00:32model. So you only take little
  31684. 24:00:34incremental changes and that's what the
  31685. 24:00:35learning rate is is those small
  31686. 24:00:37incremental changes. Epics, how many
  31687. 24:00:39times are you going to go through all
  31688. 24:00:41the data in your training set? So one
  31689. 24:00:43epic is one trip through all the data.
  31690. 24:00:45And there's a lot of other things
  31691. 24:00:46depending on which model you're working
  31692. 24:00:48with and which programming script you're
  31693. 24:00:50working with. Like the Python sklearn
  31694. 24:00:53package will have it slightly different
  31695. 24:00:55than say Google's TensorFlow package
  31696. 24:00:58which will be a little bit different
  31697. 24:00:59than the Spark machine learning package.
  31698. 24:01:01So these are just some examples of the
  31699. 24:01:03hyperparameters. And so you see in here
  31700. 24:01:05we have a nice image of our data coming
  31701. 24:01:07in and we train our model. Then we do a
  31702. 24:01:09comparison to see how good our model is.
  31703. 24:01:11And then we go back and we say, "Hey,
  31704. 24:01:13this this model's pretty good, but it's
  31705. 24:01:14biased." So then we send it back and we
  31706. 24:01:17change our hyperparameters to see if we
  31707. 24:01:19can get an unbiased model or we can have
  31708. 24:01:20a better prediction on it that matches
  31709. 24:01:22our data closer. What will happen if
  31710. 24:01:24learning rate is set too low or too
  31711. 24:01:27high? We have a nice couple graphs here.
  31712. 24:01:29We have one over here. It says a
  31713. 24:01:30learning rate set too low. And you can
  31714. 24:01:32see that it slowly works its way down
  31715. 24:01:34the curve. And on the right you can see
  31716. 24:01:36a learning rate set too high. It's just
  31717. 24:01:38bouncing back and forth. When your
  31718. 24:01:39learning rate is too low, that's what we
  31719. 24:01:41studied two slides ago. That's what the
  31720. 24:01:43learning rate was. Training of the model
  31721. 24:01:45will progress very slowly as we are
  31722. 24:01:47making very tiny updates to the weights.
  31723. 24:01:50We'll take many updates before reaching
  31724. 24:01:51the minimum point. So I just mentioned
  31725. 24:01:54epic going through all the data. You
  31726. 24:01:56might have to go through all the data a
  31727. 24:01:57thousand times instead of 500 times for
  31728. 24:01:59it to train. Learning rate too high
  31729. 24:02:01causes undesirable divergent behavior to
  31730. 24:02:04the loss function due to drastic updates
  31731. 24:02:06and weights. At times it may fail to
  31732. 24:02:08converge or even diverge. So if you have
  31733. 24:02:10your learning rate set too high and it's
  31734. 24:02:12training too quickly, maybe you'll get
  31735. 24:02:14lucky and it trains after one epic run,
  31736. 24:02:16but a lot of times it might never be
  31737. 24:02:18able to train because the weights are
  31738. 24:02:19changing too fast. They they flip back
  31739. 24:02:21and forth too easy. And you see down
  31740. 24:02:22here we've introduced uh two new terms
  31741. 24:02:25converge and diverge. Converge means
  31742. 24:02:29that our model has reached a point where
  31743. 24:02:32it's able to give a fairly good answer
  31744. 24:02:34for all the data we put in. All those
  31745. 24:02:36weights have adjusted and it's minimized
  31746. 24:02:38the error. Diverge means that the data
  31747. 24:02:40is so chaotic that it can never manage
  31748. 24:02:42to to train to that data. The data is
  31749. 24:02:45just too chaotic for it to train. So we
  31750. 24:02:46have two new words there. Converge and
  31751. 24:02:48diverge are important to know. Also what
  31752. 24:02:50is dropout and batch normalization?
  31753. 24:02:53Dropout is a technique of dropping out
  31754. 24:02:56hidden and visible units of a network
  31755. 24:02:58randomly to prevent overfitting of data.
  31756. 24:03:00It doubles the number of iterations
  31757. 24:03:02needed to converge the network. So here
  31758. 24:03:04we have our standard neural network and
  31759. 24:03:06then after applying dropout. Now it
  31760. 24:03:08doesn't mean we actually delete the
  31761. 24:03:10node. The node is still there and we're
  31762. 24:03:12still going to use that node. What it
  31763. 24:03:13means is that we're only going to work
  31764. 24:03:16with a few of the nodes. Um, a lot of
  31765. 24:03:18times I think the most common one right
  31766. 24:03:20now used is 20%. Uh, so you'll drop out
  31767. 24:03:2320% of the nodes when you do your
  31768. 24:03:25training, you reverse propagate your
  31769. 24:03:27data and then you'll randomly pick
  31770. 24:03:29another 20 nodes the next time you go
  31771. 24:03:31through an epic data training. So each
  31772. 24:03:32time you go through one epic, you will
  31773. 24:03:34randomly pick 20 of those nodes not to
  31774. 24:03:36not to mess with. And this allows for
  31775. 24:03:39less overfitting of the data. So by
  31776. 24:03:41randomly doing this you create some I
  31777. 24:03:43guess it just kind of pulls some nodes
  31778. 24:03:44off to the side and says we're going to
  31779. 24:03:45handle the data later on so we don't
  31780. 24:03:46overfit. Batch normalization is a
  31781. 24:03:49technique to improve the performance and
  31782. 24:03:51stability of neural network. The idea is
  31783. 24:03:54to normalize the inputs in every layer
  31784. 24:03:56so that they have mean output and
  31785. 24:03:57activation of zero and standard
  31786. 24:03:59deviation of one. This question covers a
  31787. 24:04:02lot of different things which is great.
  31788. 24:04:04It's a great uh interview question
  31789. 24:04:06because it pulls in that you have to
  31790. 24:04:07understand what the mean value is. So a
  31791. 24:04:10mean output activation of zero that
  31792. 24:04:12means our average activation is zero. So
  31793. 24:04:15when you normalize it remember usually
  31794. 24:04:16we're going between minus1 and one on a
  31795. 24:04:19lot of these. It's a very standard
  31796. 24:04:20setup. So you have to be very aware that
  31797. 24:04:22this is your mean output activation of
  31798. 24:04:24zero. And then we have our standard
  31799. 24:04:25deviation of one. So we want to keep our
  31800. 24:04:28error down to just a one value. The
  31801. 24:04:30benefits of this doing a batch
  31802. 24:04:32normalization is it provides
  31803. 24:04:34regularization. It trains faster, higher
  31804. 24:04:37learning rates and weights are easier to
  31805. 24:04:40initialize. What is the difference
  31806. 24:04:42between batch gradient descent and
  31807. 24:04:44stochastic gradient descent? Batch
  31808. 24:04:47gradient descent. Batch gradient
  31809. 24:04:49computes the gradient using the entire
  31810. 24:04:51data set. It takes time to converge
  31811. 24:04:53because the volume of data is huge and
  31812. 24:04:55weights update slowly. So you can look
  31813. 24:04:57at the batches. A lot of times if you're
  31814. 24:04:59using big data, batch the data in, but
  31815. 24:05:01you still go through a full epic. You
  31816. 24:05:03still go through all the data on there.
  31817. 24:05:05So bash gradient descent means you're
  31818. 24:05:07going to use it to fit all the data and
  31819. 24:05:08look for a convergence there. Stochastic
  31820. 24:05:11gradient descent. Stochastic gradient
  31821. 24:05:13computes the gradient using a single
  31822. 24:05:15sample. It converges much faster than
  31823. 24:05:17batch gradient because it updates weight
  31824. 24:05:19more frequently. Explain overfitting and
  31825. 24:05:21underfitting and how to combat them.
  31826. 24:05:23Overfitting happens when a model learns
  31827. 24:05:25the details and noise in the training
  31828. 24:05:27data to the degree that it adversely
  31829. 24:05:29impacts the execution of the model on
  31830. 24:05:31the new information. It is more likely
  31831. 24:05:33to occur with nonlinear models that have
  31832. 24:05:35more flexibility when learning a target
  31833. 24:05:37function. An example of this would be um
  31834. 24:05:39if you're looking at say cars and trucks
  31835. 24:05:42and motorcycles, it might only recognize
  31836. 24:05:45trucks that have a certain box-like
  31837. 24:05:47shape. It might not be able to notice a
  31838. 24:05:50flatbed truck unless it's only a
  31839. 24:05:52specific kind of flatbed truck or only
  31840. 24:05:54Ford trucks because that's what it saw
  31841. 24:05:55on the training set. This means that
  31842. 24:05:57your model performs great on your train
  31843. 24:05:59data and great on maybe a small test
  31844. 24:06:02amount of data, but when you go to use
  31845. 24:06:04it in the real world, it leaves out a
  31846. 24:06:06lot and start and is not very functional
  31847. 24:06:08outside of your small area, your slow
  31848. 24:06:11laboratory data coming in. Underfitting,
  31849. 24:06:13doing the opposite when you underfit
  31850. 24:06:14your data. Underfitting alludes to a
  31851. 24:06:16model that is neither well-trained on
  31852. 24:06:18training data nor can generalize to new
  31853. 24:06:21information. Usually happens when there
  31854. 24:06:22is less and improper data to train a
  31855. 24:06:25model. has a bur performance and
  31856. 24:06:27accuracy. So if you're using underfitted
  31857. 24:06:29data and you generate a model and you
  31858. 24:06:30distribute that in a commercial zone,
  31859. 24:06:32you'll have a lot of people unhappy with
  31860. 24:06:34you because it's not going to give them
  31861. 24:06:35very good answers. So we've explained
  31862. 24:06:37overfitting and underfitting. So now we
  31863. 24:06:39want to ask how to combat them.
  31864. 24:06:41Combating overfitting and underfitting,
  31865. 24:06:44resampling the data to estimate the
  31866. 24:06:46model accuracy, k-fold cross validation,
  31867. 24:06:49having a validation data set to evaluate
  31868. 24:06:51the model. So when we do the reampling,
  31869. 24:06:53we're randomly going to be picking out
  31870. 24:06:55data and we'll run it a few times to see
  31871. 24:06:57how that works depending on our random
  31872. 24:06:59data and how we sample the data to
  31873. 24:07:01generate our model and then we want to
  31874. 24:07:03go ahead and validate the data set by
  31875. 24:07:05having our training data and then
  31876. 24:07:07keeping some data on the side uh testing
  31877. 24:07:10data to validate it. How are weights
  31878. 24:07:12initialized in a network? Initializing
  31879. 24:07:14all weights to zero. All the weights are
  31880. 24:07:16set to zero. This makes your model
  31881. 24:07:18similar to a linear model. So if you
  31882. 24:07:20have linear data coming in, doing a
  31883. 24:07:22basic setup like that might work. All
  31884. 24:07:23the neurons in every layer perform the
  31885. 24:07:25same operation given the same output and
  31886. 24:07:28making the deep net useless. Right?
  31887. 24:07:30There's a key word. It's going to be
  31888. 24:07:31useless if you initialize everything to
  31889. 24:07:33zero. At that point be looking into some
  31890. 24:07:35other uh machine learning tools.
  31891. 24:07:37Initializing all weights randomly. Here
  31892. 24:07:39the weights are assigned randomly by
  31893. 24:07:41initializing them very close to zero. It
  31894. 24:07:43gives better accuracy to the model since
  31895. 24:07:45every neuron performs different
  31896. 24:07:46computations. And here we have the
  31897. 24:07:49weights are set randomly. We have our
  31898. 24:07:50input layer, the hidden layers and the
  31899. 24:07:52output layer. And W equals NP random
  31900. 24:07:54random N layer size L, layer size L
  31901. 24:07:57minus one. This is the most commonly
  31902. 24:07:59used is to randomly generate your
  31903. 24:08:01weights. What are the different layers
  31904. 24:08:03in CNN? Convolutional neural network.
  31905. 24:08:06First is the convolutional layer that
  31906. 24:08:08performs a convolutional operation. We
  31907. 24:08:11have our other video out if you want to
  31908. 24:08:13explore that more. and go into detail
  31909. 24:08:14exactly how the C the convolutional
  31910. 24:08:17layer works in the CNN as far as
  31911. 24:08:19creating a number of smaller uh picture
  31912. 24:08:21windows that go over the data. Uh the
  31913. 24:08:23second step is as a relu layer relu
  31914. 24:08:26brings nonlinearity to the network and
  31915. 24:08:28converts all the negative pixels to
  31916. 24:08:30zero. Output is rectified feature map.
  31917. 24:08:32So it goes into a mapping feature there.
  31918. 24:08:34Pooling layer pooling is a down sampling
  31919. 24:08:37operation that reduces the
  31920. 24:08:38dimensionality of the feature map. So we
  31921. 24:08:40have all our relu layer which is pulling
  31922. 24:08:42all these little maps out of our
  31923. 24:08:44convolutional layer. It's taking that
  31924. 24:08:46picture and little creating little tiny
  31925. 24:08:48neural networks to look at different
  31926. 24:08:49parts of the picture. Uh then we need to
  31927. 24:08:51pull it together and then finally the
  31928. 24:08:53fully connected layer. So we flatten our
  31929. 24:08:55pooling layer out and we have a fully
  31930. 24:08:57connected layer recognizes and
  31931. 24:08:59classifies the objects in the image. And
  31932. 24:09:01that's actually your forward propagation
  31933. 24:09:03reverse propagation training model
  31934. 24:09:05usually. I mean there's a number of
  31935. 24:09:06different models out there of course.
  31936. 24:09:08What is pooling in CNN and how does it
  31937. 24:09:11work? Pooling used to reduce the spatial
  31938. 24:09:13dimensions of a CNN performs down
  31939. 24:09:16sampling operation to reduce the
  31940. 24:09:18dimensionality. Creates a pulled feature
  31941. 24:09:20map by sliding a filter matrix over the
  31942. 24:09:22input matrix. I mentioned that briefly
  31943. 24:09:25on the previous slide. Um it's important
  31944. 24:09:27to know that you have if you see here
  31945. 24:09:29they have a rectified feature map. And
  31946. 24:09:31so each one of those colors like the
  31947. 24:09:33yellow color that might be one of the a
  31948. 24:09:35smaller little neural network using the
  31949. 24:09:37ReLU. You'll look at it'll just kind of
  31950. 24:09:39um go over the main picture and look at
  31951. 24:09:41all the different areas on the main
  31952. 24:09:42picture. So you might step one 2 3 four
  31953. 24:09:45spaces. Um and then you have another one
  31954. 24:09:47that's also looking at features and it
  31955. 24:09:49has a 2785. Each one of those is a map.
  31956. 24:09:52So it might be the first one might be a
  31957. 24:09:54map looking for cat ears and the second
  31958. 24:09:56one looking for human eyes. When it does
  31959. 24:09:58this, you then have this rectified
  31960. 24:10:00feature map looking at these different
  31961. 24:10:02features and the max pooling with a 2x
  31962. 24:10:04two filters and a stride of two. Stride
  31963. 24:10:06means instead of skipping every pixel,
  31964. 24:10:07you're going to go every two pixels. You
  31965. 24:10:09take the maximum values and you can see
  31966. 24:10:11over here we look at a pulled feature
  31967. 24:10:13map. One of the features says, hey, I
  31968. 24:10:15had a max value of eight. So somewhere
  31969. 24:10:17in here we saw a human eye labeled as
  31970. 24:10:20eight. Pretty high label. And maybe
  31971. 24:10:21seven was a human hand and maybe four
  31972. 24:10:23was cat whiskers or something that we
  31973. 24:10:25thought might be cat whiskers. Four is
  31974. 24:10:27kind of a low number in this particular
  31975. 24:10:29case compared to the other ones. So you
  31976. 24:10:30have your full pool feature map. You can
  31977. 24:10:32see the process here is we have our
  31978. 24:10:34stepping, we look for the max value and
  31979. 24:10:36then we create a poolled feature map of
  31980. 24:10:38the maxed values. How does a LSTM
  31981. 24:10:42network work? That's long shortterm
  31982. 24:10:44memory. So the first thing to know is
  31983. 24:10:46that an LSTMs are a special kind of
  31984. 24:10:48recurrent neural network capable of
  31985. 24:10:50learning long-term dependencies.
  31986. 24:10:52remembering information for long periods
  31987. 24:10:54of time is their default behavior. We
  31988. 24:10:57did look at the RNN briefly talked about
  31989. 24:10:59how the hidden layer feeds back into
  31990. 24:11:01itself. With the LSTM has a much more
  31991. 24:11:05complicated feedback and you can see
  31992. 24:11:06here we have the hidden layer of T minus
  31993. 24:11:09one and the hidden layer that's what the
  31994. 24:11:11H stands for hidden layer of T and the
  31995. 24:11:13formulas going in. As we can see here we
  31996. 24:11:15have the hidden layers we have T minus
  31997. 24:11:18one and then H of T where T stands for
  31998. 24:11:21time. So this is a series remember
  31999. 24:11:22working with series and we want to
  32000. 24:11:24remember the past and you can see you
  32001. 24:11:26have your ex your input of t and that
  32002. 24:11:28might be a frame in a video as a frame
  32003. 24:11:31comes in they usually use in this one
  32004. 24:11:33the tangent h activation formula but you
  32005. 24:11:36also see that it goes through a couple
  32006. 24:11:38other formulas the omega formula and so
  32007. 24:11:40when it combines these that then goes
  32008. 24:11:42into the next layer your next hidden
  32009. 24:11:44layer that then goes into the data
  32010. 24:11:46that's submitted to the next input so
  32011. 24:11:49you have your x of t + one. So when you
  32012. 24:11:51have that coming in, then you have your
  32013. 24:11:53H value that's coming forward from the
  32014. 24:11:55last process. And depending on how many
  32015. 24:11:57of these omega structures you put in
  32016. 24:11:59there depends on how long-term the
  32017. 24:12:01memory gets. So it's important to
  32018. 24:12:03remember this is more for your long-term
  32019. 24:12:05recurrent neural networks. The three
  32020. 24:12:07steps in an LSTM, step one decides what
  32021. 24:12:11to forget and what to remember. Step
  32022. 24:12:13two, selectively update cell state
  32023. 24:12:15values. So based on what we want to
  32024. 24:12:17remember and forget, we want to update
  32025. 24:12:18those cell values and then decides what
  32026. 24:12:20part of the current state make it to the
  32027. 24:12:22output. So now we have to also have an
  32028. 24:12:24output on there. What are vanishing and
  32029. 24:12:27exploding gradients? This is a great
  32030. 24:12:29question that affects all our neural
  32031. 24:12:30networks. While training an RNN, your
  32032. 24:12:33slope can become either too small or too
  32033. 24:12:35large and this makes the training
  32034. 24:12:37difficult. When the slope is too small,
  32035. 24:12:39the problem is known as vanishing
  32036. 24:12:41gradient. So our slope, we have our
  32037. 24:12:43change in x and our change in y. When
  32038. 24:12:45the slope decreases gradually to a very
  32039. 24:12:47small value, sometimes negative, and
  32040. 24:12:49makes training difficult. When the slope
  32041. 24:12:51tends to grow exponentially instead of
  32042. 24:12:53decaying, this problem is called
  32043. 24:12:55exploding gradient. The slope grows
  32044. 24:12:57exponentially. You can see a nice graph
  32045. 24:12:59of that here. Issues in gradient
  32046. 24:13:00problem, long training time, poor
  32047. 24:13:03performance, and low accuracy. What is
  32048. 24:13:05the difference between epic, batch, and
  32049. 24:13:07iteration in deep learning? Epic. An
  32050. 24:13:09epic represents one iteration over the
  32051. 24:13:12entire data set. So that's everything
  32052. 24:13:14you're going to go ahead and put into
  32053. 24:13:15that training model. Batch. We cannot
  32054. 24:13:17pass the entire data set into the neural
  32055. 24:13:19network at once. So we divide the data
  32056. 24:13:21set into a number of batches. And then
  32057. 24:13:24iteration. If we have 10,000 images as
  32058. 24:13:27data and a batch size of 200, then the
  32059. 24:13:30epic should run 10,000 times over 200.
  32060. 24:13:32So that means we have our total number
  32061. 24:13:34over the 200 equals 50 iterations. So in
  32062. 24:13:37each epic we're running over all the
  32063. 24:13:39data set, we're going to have 50
  32064. 24:13:41iterations. And each of those iterations
  32065. 24:13:43includes a batch of 200 images in this
  32066. 24:13:46case. Why TensorFlow is the most
  32067. 24:13:48preferred library in deep learning? Uh
  32068. 24:13:50well, first TensorFlow provides both C++
  32069. 24:13:52and Python APIs that makes it easier to
  32070. 24:13:55work on. Has a faster compilation time
  32071. 24:13:57than other deep learning libraries like
  32072. 24:13:59KAS and torch. TensorFlow supports both
  32073. 24:14:02CPUs and GPUs computing devices. So
  32074. 24:14:05right now TensorFlow is at the top of
  32075. 24:14:08the market because it's so easy to use
  32076. 24:14:10for both programmer side and for
  32077. 24:14:11hardware side and for the speed of
  32078. 24:14:13getting something up and running. What
  32079. 24:14:15do you mean by tensor in TensorFlow?
  32080. 24:14:17Tensor is a mathematical object
  32081. 24:14:19represented as arrays of higher
  32082. 24:14:20dimensions. These arrays of data with
  32083. 24:14:22different dimensions and ranks that are
  32084. 24:14:24fed as input to the neural network are
  32085. 24:14:26called tensors. And you can see here we
  32086. 24:14:28have a tensor of dimensions five, four.
  32087. 24:14:31So it's a two-dimensional tensor coming
  32088. 24:14:32in. Um, you can look at an image like
  32089. 24:14:34this that each one of those pixels is a
  32090. 24:14:37different value if it's a black and
  32091. 24:14:38white. So, it might be zero and ones and
  32092. 24:14:40then each one represents a black and
  32093. 24:14:41white image. In a color photo, you might
  32094. 24:14:43um either find a different value system
  32095. 24:14:46or you might have a tensor value that
  32096. 24:14:48has the xy coordinates as we see here
  32097. 24:14:50plus the colors. So, you might have
  32098. 24:14:51three more different dimensions for the
  32099. 24:14:54three different images, the red, the
  32100. 24:14:56blue, and the yellow coming in. And even
  32101. 24:14:58as you go from one layer or one tensor
  32102. 24:15:01to the next, these layers might change.
  32103. 24:15:03We might flatten them, might bring in
  32104. 24:15:04numerous. In the case of the convergence
  32105. 24:15:07neural network, we have all those
  32106. 24:15:08smaller different mappings of features
  32107. 24:15:10that come in. So each one of those
  32108. 24:15:12layers coming through is a tensor. If it
  32109. 24:15:14has multiple dimensions coming in and
  32110. 24:15:15weights attached to it, what are the
  32111. 24:15:17programming elements in TensorFlow?
  32112. 24:15:19Well, we have our constants. Constants
  32113. 24:15:21are parameters whose value does not
  32114. 24:15:22change. To define a constant, we use
  32115. 24:15:25tf.constant command. Example, A equals
  32116. 24:15:27TF.Constant 2.0 TF float 32. So it's a
  32117. 24:15:31tensor float value of 32. B equals TF
  32118. 24:15:34constant 3.0. Print AB. If we did a
  32119. 24:15:37print of AB, we'd have um TF.stant and
  32120. 24:15:40then of course uh B is that instance of
  32121. 24:15:42it. Variables. Variables allow us to add
  32122. 24:15:44new trainable parameters to graph. To
  32123. 24:15:47define a variable, we use TF.variable
  32124. 24:15:49command and initialize them before
  32125. 24:15:51running the graph in session. Example W
  32126. 24:15:53equals TF variable.3 DT type TF float 32
  32127. 24:15:57or B equals a TF variable minus 3, D
  32128. 24:15:59type float 32. Placeholders.
  32129. 24:16:01Placeholders allow us to feed data to a
  32130. 24:16:04TensorFlow model from outside a model.
  32131. 24:16:06It permits a value to be assigned later.
  32132. 24:16:08To define a placeholder, we use TF
  32133. 24:16:10placeholder command. Example A equals TF
  32134. 24:16:12placeholder b= a * 2 with the TF session
  32135. 24:16:16as SESS result equals session run B,
  32136. 24:16:20feed dictionary equals A3.0. 0 print
  32137. 24:16:22result. Uh so we have a nice example
  32138. 24:16:24there of a placeholder session. A
  32139. 24:16:26session is run to evaluate the nodes.
  32140. 24:16:28This is called as the tensorflow
  32141. 24:16:30runtime. So for example, you have a= tf
  32142. 24:16:33constant 2.0 b= tf constant 4.0 c= a
  32143. 24:16:36plus b. And at this point you'd go ahead
  32144. 24:16:38and create a session equals tf session.
  32145. 24:16:40And then you could evaluate the tensor C
  32146. 24:16:43print session run C. That would input C
  32147. 24:16:46as an input into your session. What do
  32148. 24:16:48you understand by a computational graph?
  32149. 24:16:51Everything in TensorFlow is based on
  32150. 24:16:52creating a computational graph. It has a
  32151. 24:16:55network of nodes where each node
  32152. 24:16:57performs an operation. Nodes represent
  32153. 24:16:59mathematical operation and edges
  32154. 24:17:01represent tensors. Since data flows in a
  32155. 24:17:04form of a graph, it is also called a
  32156. 24:17:06data flow graph. And we have a nice
  32157. 24:17:09visual of this graph or graphic image of
  32158. 24:17:11a computational graph. And you can see
  32159. 24:17:13here we have our input nodes, our add
  32160. 24:17:15multiply nodes and our multiply node at
  32161. 24:17:17the end. And then we have the edges
  32162. 24:17:19where the data flows. So we have from A
  32163. 24:17:22going to C, A going to D. You can see we
  32164. 24:17:24have a two flowing, a four flowing.
  32165. 24:17:26Explain generative adversarial network
  32166. 24:17:29along with an example. Suppose there is
  32167. 24:17:31a wine shop that purchases wine from
  32168. 24:17:33dealers which they will resell later. So
  32169. 24:17:35we have our dealer going to the wine,
  32170. 24:17:36our shop owner that then sells it for a
  32171. 24:17:38profit. But there are some malfactor
  32172. 24:17:40dealers who sell fake wine. In this
  32173. 24:17:42case, the shop owner should be able to
  32174. 24:17:44distinguish between fake and authentic
  32175. 24:17:46wine. The forger will try to different
  32176. 24:17:48techniques to sell fake wine and make
  32177. 24:17:50sure certain techniques go past the shop
  32178. 24:17:52owner's check. So, here's our forger
  32179. 24:17:54fake wine shop owner. The shop owner
  32180. 24:17:56would probably get some feedback from
  32181. 24:17:57the wine experts that some of the wine
  32182. 24:17:59is not original. The owner would have to
  32183. 24:18:01improve how he determines whether a wine
  32184. 24:18:03is fake or authentic. Goal of forger to
  32185. 24:18:05create wines that are indistinguishable
  32186. 24:18:07from the authentic ones. Goal of shop
  32187. 24:18:09owner to accurately tell if the wine is
  32188. 24:18:11real or not. There are two main
  32189. 24:18:12components of generative adversarial
  32190. 24:18:15network. And we refer to as a noise
  32191. 24:18:17vector coming in where we have our
  32192. 24:18:19forger who's going to generate fake wine
  32193. 24:18:21and then we have our real authentic wine
  32194. 24:18:24and of course our shop owner who has to
  32195. 24:18:25figure out whether it's real or fake.
  32196. 24:18:27The generator is a CNN that keeps
  32197. 24:18:30producing images that are closer in
  32198. 24:18:32appearance to the real images while the
  32199. 24:18:34discriminator tries to determine the
  32200. 24:18:35difference between real and fake images.
  32201. 24:18:38The ultimate aim is to make the
  32202. 24:18:39discriminator learn to identify real and
  32203. 24:18:41fake images. What is an autoenccoder?
  32204. 24:18:44The network is trained to reconstruct
  32205. 24:18:46its inputs. It is a neural network that
  32206. 24:18:48has three layers. Here the input neurons
  32207. 24:18:51are equal to the output neuron. The
  32208. 24:18:53network's target outside is same as the
  32209. 24:18:55input. It uses dimensionality reduction
  32210. 24:18:58to restructure the input. Input image
  32211. 24:19:01comes in. We have our Latin space
  32212. 24:19:03representation and then it goes back out
  32213. 24:19:05reconstructing the image. It works by
  32214. 24:19:07compressing the input to a Latin space
  32215. 24:19:09representation and then reconstructing
  32216. 24:19:11the output from this representation.
  32217. 24:19:13What is bagging and boosting? Bagging
  32218. 24:19:15and boosting are ensemble techniques
  32219. 24:19:17where the idea is to train multiple
  32220. 24:19:19models using the same learning algorithm
  32221. 24:19:21and then take a call. So we have in here
  32222. 24:19:23where we're bagging. We take a data set
  32223. 24:19:25and we split it. We're going to have our
  32224. 24:19:26training data and our test data. Very
  32225. 24:19:28standard thing to do. Then we're going
  32226. 24:19:29to randomly select data into the bags
  32227. 24:19:32and train your model separately. So we
  32228. 24:19:34might have bag one, model one, bag two,
  32229. 24:19:36model two, bag three, model 3, and so
  32230. 24:19:38on. In boosting the emphasis is to
  32231. 24:19:40select the data points which give wrong
  32232. 24:19:42output in order to improve the accuracy.
  32233. 24:19:45So in boosting we have our data set
  32234. 24:19:47again we split it to test data and train
  32235. 24:19:49data and we'll take a bag one and we'll
  32236. 24:19:51train the model. Data points with wrong
  32237. 24:19:53predictions then go into bag two and we
  32238. 24:19:55then train that model and repeat.
  32239. 24:19:57machine learning is which is a subset of
  32240. 24:20:00artificial intelligence, right? That's
  32241. 24:20:02uh basically
  32242. 24:20:04um machines learning from data
  32243. 24:20:08in order to uh make decisions
  32244. 24:20:11essentially. Um so this was a big
  32245. 24:20:14departure from the rules-based systems
  32246. 24:20:17at the time, right? That were explicitly
  32247. 24:20:19programmed to make decisions. So just
  32248. 24:20:22think of an example like a really big
  32249. 24:20:25kind of if this then that then that then
  32250. 24:20:27that and and else if this this this
  32251. 24:20:30right so bunch of rules that had to be
  32252. 24:20:33pre-programmed in order to um come out
  32253. 24:20:36with some final answer. Uh with machine
  32254. 24:20:39learning it's the exact opposite of
  32255. 24:20:40that. we're actually training something
  32256. 24:20:42from examples from existing data um in
  32257. 24:20:46order to predict something or um
  32258. 24:20:50make some type of decision. Uh and so
  32259. 24:20:52we're going to learn about the various
  32260. 24:20:54ways we can do machine learning. But if
  32261. 24:20:56you guys remember we um talked about
  32262. 24:20:59some of this like the differences and
  32263. 24:21:01the uh basically rules-based approaches
  32264. 24:21:05to learning from data approach. Um and
  32265. 24:21:09in included in that is going to be uh
  32266. 24:21:11complex unstructured data. So things
  32267. 24:21:14like images, text, audio. What handles
  32268. 24:21:16those really well is uh deep learning
  32269. 24:21:20which we will get to in the course after
  32270. 24:21:22this. But uh those are certainly in
  32271. 24:21:26there as learning from data even complex
  32272. 24:21:28data.
  32273. 24:21:31So we had this picture uh and I think
  32274. 24:21:33this is kind of around where we left off
  32275. 24:21:35last time was uh just distinguishing
  32276. 24:21:40between those three terms. We see
  32277. 24:21:41artificial intelligence, deep learning
  32278. 24:21:43and machine learning kind of used
  32279. 24:21:44interchangeably, but this is really how
  32280. 24:21:46they fit in. Artificial intelligence is
  32281. 24:21:48kind of a broad anything mimicking human
  32282. 24:21:51intelligence. Um which doesn't have to
  32283. 24:21:55be learning from data, but uh machine
  32284. 24:21:57learning is part of that. And then um
  32285. 24:22:00one way to accomplish machine learning
  32286. 24:22:02is to use neural nets which is the focus
  32287. 24:22:04of uh deep learning. Um and so deep
  32288. 24:22:08learning has been has found a lot of
  32289. 24:22:09success especially recently with uh
  32290. 24:22:12those complex data types like images,
  32291. 24:22:15speech, text, right? So deep learning
  32292. 24:22:18used all over the place. Even in um
  32293. 24:22:20modern like generative AI, we see deep
  32294. 24:22:23learning used quite a bit. Um it really
  32295. 24:22:26anything that's using neural nets is uh
  32296. 24:22:29going to be deep learning.
  32297. 24:22:32Um and again we'll focus on that later
  32298. 24:22:35but we're going to be mainly focused on
  32299. 24:22:37machine learning for this course.
  32300. 24:22:39Primarily machine learning that does not
  32301. 24:22:41use neural networks. Okay. So just
  32302. 24:22:43models that are not necessarily neural
  32303. 24:22:45networks
  32304. 24:22:48be our focus.
  32305. 24:22:51So in machine learning we had an example
  32306. 24:22:54of a game uh essentially um learning
  32307. 24:22:59what decisions to make uh based on the
  32308. 24:23:03uh kind of current um state of the
  32309. 24:23:06board. This could be a um you know
  32310. 24:23:09machine learning example that uh learns
  32311. 24:23:12from many previous examples. So a lot of
  32312. 24:23:15data around these games are used to
  32313. 24:23:18train these um kind of robots that can
  32314. 24:23:22play these games and play them at a very
  32315. 24:23:24high level. Um so there's been a lot of
  32316. 24:23:26successes actually in machine learning
  32317. 24:23:28and deep learning um around
  32318. 24:23:32uh playing games like chess or go
  32319. 24:23:36um using machine learning algorithms. So
  32320. 24:23:38pretty cool.
  32321. 24:23:41All right, so I think this is where we
  32322. 24:23:42ended. Last time we said there's a bunch
  32323. 24:23:43of different use cases for machine
  32324. 24:23:45learning. So um recommendation system is
  32325. 24:23:48going to be a big one and we will
  32326. 24:23:50actually study that uh in one of our
  32327. 24:23:53final lessons of this course. Um chat
  32328. 24:23:56bots like generative AI doing sentiment
  32329. 24:23:58analysis chat bots we'll study later but
  32330. 24:24:00those are certainly an application of
  32331. 24:24:02learning from data in order to uh
  32332. 24:24:06generate responses to text prompts
  32333. 24:24:08right. Um spam filtering that's a good
  32334. 24:24:11example like classifying an email as
  32335. 24:24:13spam or not spam. Um that that gets
  32336. 24:24:17trained from examples and uh learning
  32337. 24:24:20from data such as previous emails. Um
  32338. 24:24:24social media posts analysis is another
  32339. 24:24:26kind of text data um use case but you uh
  32340. 24:24:32can do a lot with that text like you can
  32341. 24:24:35predict the sentiment um you can predict
  32342. 24:24:38uh the category of what what the post is
  32343. 24:24:41talking about um those kind of things
  32344. 24:24:44all can be done with machine learning
  32345. 24:24:47>> and many other use cases not on this
  32346. 24:24:49list that we will uh cover
  32347. 24:24:51>> you know as we as we go further.
  32348. 24:25:02Okay, so this is where we kind of left
  32349. 24:25:04off. Um, so what's doing all the hard
  32350. 24:25:08work here is
  32351. 24:25:11>> uh machine learning algorithms. So these
  32352. 24:25:12are things that will um these are things
  32353. 24:25:16that will learn from the data. So they
  32354. 24:25:19are uh they they are basically um
  32355. 24:25:24algorithms or sets of rules that uh or
  32356. 24:25:28mathematical rules I should say not
  32357. 24:25:29formal rules like in the in the sense of
  32358. 24:25:31a rule system but mathematical um
  32359. 24:25:34formulas and mathematical uh rules
  32360. 24:25:37essentially that help us learn from the
  32361. 24:25:41data. So they correlate the data to some
  32362. 24:25:43type of outcome. So some type of
  32363. 24:25:46prediction uh whether that's going to be
  32364. 24:25:49as we will see whether that could be
  32365. 24:25:50like a number like we're predicting a
  32366. 24:25:53price or demand or sales
  32367. 24:25:56um or it could be a category like is
  32368. 24:25:58this transaction fraud or not fraud or
  32369. 24:26:01what's the probability that this is
  32370. 24:26:03fraud um so we have different kinds of
  32371. 24:26:07predictions we can make with machine
  32372. 24:26:08learning
  32373. 24:26:10um but uh we will study the kind of the
  32374. 24:26:14differences of those coming up. Um, but
  32375. 24:26:17machine learning algorithms are really
  32376. 24:26:19what power they're kind of the models,
  32377. 24:26:21right? They're the models that help
  32378. 24:26:23power uh machine learning to actually
  32379. 24:26:26learn from data.
  32380. 24:26:28So, we're going to spend a lot of time
  32381. 24:26:30in this course studying those algorithms
  32382. 24:26:33like the different models that we can
  32383. 24:26:34build and what their differences are,
  32384. 24:26:37what their strengths are, what their
  32385. 24:26:38weaknesses are. We'll we'll learn a lot
  32386. 24:26:40about those.
  32387. 24:26:44Okay. So I guess you can imagine like
  32388. 24:26:46everything is so data dependent, right?
  32389. 24:26:49Um we're learning from data. So uh it
  32390. 24:26:52makes sense that the quality of data
  32391. 24:26:55really really matters here in
  32392. 24:26:56determining how strong the model can be.
  32393. 24:26:59Um so you see this graph here charting
  32394. 24:27:03kind of the um high quality data um
  32395. 24:27:07versus just uh any old data but a decent
  32396. 24:27:11enough quantity of it. Um you can see
  32397. 24:27:14that performance and the performance is
  32398. 24:27:16measured by some evaluation metric. Um,
  32399. 24:27:20so think of it as uh something like an
  32400. 24:27:23accuracy. Like if we were predicting
  32401. 24:27:24fraud or not fraud, how accurate can our
  32402. 24:27:27model get at actually detecting fraud,
  32403. 24:27:30um, it gets better and better and better
  32404. 24:27:33the graph shows that the higher quality
  32405. 24:27:36of data that we have. So there's kind of
  32406. 24:27:39that there's a there's a saying in
  32407. 24:27:40machine learning um, called garbage in
  32408. 24:27:44garbage out. What that means is if you
  32409. 24:27:46have poor data, even the best model in
  32410. 24:27:48the world, poor data is not going to
  32411. 24:27:51result in having a good model that can
  32412. 24:27:53be accurate and perform well. Um, so it
  32413. 24:27:56needs to be high quality, meaning um
  32414. 24:28:00there needs to be a decent amount of it
  32415. 24:28:01and it needs to be labeled appropriately
  32416. 24:28:04as we will will talk about
  32417. 24:28:07um and it needs to not have any, you
  32418. 24:28:10know, significant outliers. it needs to
  32419. 24:28:12be clean, not have those missing values,
  32420. 24:28:15all of those things. Um, you can you
  32421. 24:28:18have a good chance at deriving good
  32422. 24:28:20predictions from higher quality data
  32423. 24:28:24as this kind of shows.
  32424. 24:28:30Okay.
  32425. 24:28:32So, one thing we're going to learn um as
  32426. 24:28:34we go along is
  32427. 24:28:36quantity matters as well. So, not only
  32428. 24:28:38quality, but a decent amount of it. And
  32429. 24:28:41um we're going to learn those kind of
  32430. 24:28:42rules of thumb like how much data do I
  32431. 24:28:45need for certain algorithms. Um one
  32432. 24:28:48thing that we will see is that uh the
  32433. 24:28:51the basic machine learning models that
  32434. 24:28:53we'll study don't need as much as a
  32435. 24:28:56neural network would. It you know neural
  32436. 24:28:59networks are going to require a lot more
  32437. 24:29:02um than a basic machine learning model
  32438. 24:29:05learning model. So uh that's something
  32439. 24:29:08we will see as we go along. But uh this
  32440. 24:29:10is something we'll talk about and
  32441. 24:29:12discuss with each model that we study is
  32442. 24:29:14kind of how much data do we actually
  32443. 24:29:16need to produce a high quality model.
  32444. 24:29:24Okay, any questions uh so far?
  32445. 24:29:34Okay, let's talk about the different
  32446. 24:29:37types of machine learning that we're
  32447. 24:29:39going to discuss. Pime, there's going to
  32448. 24:29:41be two primary ones that we will study
  32449. 24:29:43in this course and then a couple others
  32450. 24:29:46that'll be a little bit more advanced
  32451. 24:29:47that we won't get to but worth knowing
  32452. 24:29:50about. Um, so there's going to be four
  32453. 24:29:52total that we'll study or talk about and
  32454. 24:29:55they'll be on this list here, which is
  32455. 24:29:58um supervised learning and unsupervised
  32456. 24:30:00learning. Now, I'd say the majority of
  32457. 24:30:03our focus will probably be on supervised
  32458. 24:30:06learning, and we'll talk about what that
  32459. 24:30:08means, but we'll also cover unsupervised
  32460. 24:30:12learning as well. And so, we'll look at
  32461. 24:30:14the most popular techniques in each of
  32462. 24:30:16these types of machine learning.
  32463. 24:30:20Um,
  32464. 24:30:21and then we'll talk about these two, but
  32465. 24:30:24not really study them because they're
  32466. 24:30:25more advanced topics um that that will
  32467. 24:30:28be beyond the scope of what we'll do.
  32468. 24:30:30But uh these are going to be um
  32469. 24:30:34different styles of machine learning
  32470. 24:30:38that are going to be characterized by um
  32471. 24:30:42what kinds of predictions they make,
  32472. 24:30:43what kind of data they need and require.
  32473. 24:30:46Um and uh what kind of outcomes they're
  32474. 24:30:50actually producing. Um, so let's let's
  32475. 24:30:55get into each of these, but uh the the
  32476. 24:30:57one that we'll probably spend the
  32477. 24:30:59majority of our time on is going to be
  32478. 24:31:00supervised learning, but we will study
  32479. 24:31:03unsupervised learning as well. We'll
  32480. 24:31:05study both and we're going to talk about
  32481. 24:31:06we're going to define both of those um
  32482. 24:31:08coming up. And again, these will be a
  32483. 24:31:11little bit more advanced topics that we
  32484. 24:31:12won't spend too much time on.
  32485. 24:31:15Um, but but we'll discuss their
  32486. 24:31:17relevancy in machine learning um and
  32487. 24:31:21give a good definition to it.
  32488. 24:31:28Okay.
  32489. 24:31:32All right. Let's start with supervised
  32490. 24:31:34learning. Now this is going to be uh a
  32491. 24:31:38term that really refers to
  32492. 24:31:41using examples. So using labeled
  32493. 24:31:46examples. So here we say labeled data to
  32494. 24:31:50help our model train. In other words,
  32495. 24:31:54help our model be able to predict guided
  32496. 24:31:58by specific input output pairs. So
  32497. 24:32:01supervised really refers to the fact
  32498. 24:32:04that we have answers. We have examples,
  32499. 24:32:10we have answers with those and we use
  32500. 24:32:12that collection of data to build our
  32501. 24:32:15model off of so that we can predict
  32502. 24:32:19um those kinds of things like a price,
  32503. 24:32:23like a category, like a spam not spam.
  32504. 24:32:27in this in this slide like we would be
  32505. 24:32:29predicting if this shape is a square, a
  32506. 24:32:32triangle or a circle.
  32507. 24:32:34Um but but when we build a model for
  32508. 24:32:38that, we have data that has an answer
  32509. 24:32:42attached to it. Right? We've talked
  32510. 24:32:44about this before a little bit with
  32511. 24:32:45labels. So there's a guide there that
  32512. 24:32:50can guide us towards building our model.
  32513. 24:32:52there's an actual every every example
  32514. 24:32:55has an answer and that answer is really
  32515. 24:32:58critical to help build our model off of.
  32516. 24:33:01So, um that's it's almost like you have
  32517. 24:33:06um a you have a bunch of exercises
  32518. 24:33:11in let's say like a math textbook. You
  32519. 24:33:13have a bunch of exercises and you have
  32520. 24:33:15the answers and that way you can kind of
  32521. 24:33:18check your work. You think about model
  32522. 24:33:20training um that is the really a lot of
  32523. 24:33:24that process of model training as we are
  32524. 24:33:26going to discover is um basically
  32525. 24:33:29checking our work against these answers
  32526. 24:33:32in our data in our training data.
  32527. 24:33:35Okay. So supervised learning is any type
  32528. 24:33:38of machine learning that involves
  32529. 24:33:41learning from labeled data in order to
  32530. 24:33:44predict outcomes. Okay, predict outcomes
  32531. 24:33:47like now the the outcomes can be
  32532. 24:33:49numerical. They can be like a price,
  32533. 24:33:52temperature, demand, sales, revenue.
  32534. 24:33:56They can be numerical, but they can also
  32535. 24:33:58be categorical. So they can be like
  32536. 24:33:59spam, not spam, fraud, not fraud,
  32537. 24:34:01cancer, not cancer. Um, dog, cat,
  32538. 24:34:05giraffe, those kind of categories. Um,
  32539. 24:34:09we could predict those. It's some type
  32540. 24:34:11of outcome. Okay, some type of outcome.
  32541. 24:34:13The key is we're using labeled examples
  32542. 24:34:17to guide our model building. That's why
  32543. 24:34:20it's called supervised learning.
  32544. 24:34:23So we know in our data we know what the
  32545. 24:34:26inputs are. Of course, those are going
  32546. 24:34:28to be think of the inputs as like all of
  32547. 24:34:30our columns and then we have a special
  32548. 24:34:34label column that represents the output
  32549. 24:34:36we're trying to predict. So if you think
  32550. 24:34:37about that housing price data, the label
  32551. 24:34:40could be the price. And that's something
  32552. 24:34:42we would build a model to predict, but
  32553. 24:34:44we have answers for all of our examples
  32554. 24:34:47in our rows. We have answers to help
  32555. 24:34:50guide our model building.
  32556. 24:34:52They help tweak our model because we
  32557. 24:34:54know the answer ahead of time. So
  32558. 24:34:58they're they're really good examples to
  32559. 24:34:59build our model off of.
  32560. 24:35:03Okay. So that's that's supervised
  32561. 24:35:05learning.
  32562. 24:35:12Uh in this example is circle not in the
  32563. 24:35:14prediction because it's not part of the
  32564. 24:35:16test data even though it's in the
  32565. 24:35:18labeled data.
  32566. 24:35:20Um no it just not necessarily. It just
  32567. 24:35:23means that like we learn against all of
  32568. 24:35:27these examples that have these answers
  32569. 24:35:30and then when we observe new examples um
  32570. 24:35:33we can try to predict what those would
  32571. 24:35:35be based on what we've seen before. So I
  32572. 24:35:38if there was a you know it's just a
  32573. 24:35:40coincidence we only have two two
  32574. 24:35:42examples in our test data like we could
  32575. 24:35:44have a circle here in which case we
  32576. 24:35:46would predict circle
  32577. 24:35:48that's fine or at least we would hope
  32578. 24:35:50our model would predict circle right
  32579. 24:35:52that's what we're hoping may or may not
  32580. 24:35:54get it right
  32581. 24:35:56um but it's it's only not there because
  32582. 24:36:01we only like we're just assuming that we
  32583. 24:36:03only have two examples we're testing
  32584. 24:36:04against but in reality we would probably
  32585. 24:36:07do a lot more than two.
  32586. 24:36:09It's just it's just a coincidence
  32587. 24:36:10really.
  32588. 24:36:17In reality, we would test against a lot
  32589. 24:36:19more data. And we're actually going to
  32590. 24:36:21see why we would do that. Like why would
  32591. 24:36:24we train our model and then kind of use
  32592. 24:36:28additional data um to to evaluate it?
  32593. 24:36:32It's actually really important that we
  32594. 24:36:34do that step to get a sense of how good
  32595. 24:36:36our model is before we take it out in
  32596. 24:36:38the real world. So if we apply our model
  32597. 24:36:41that we build on our label data to
  32598. 24:36:45um this kind of set of test data that we
  32599. 24:36:49haven't been exposed to before. It helps
  32600. 24:36:52give us a sense of how good is our
  32601. 24:36:54model. So it's tested is usually used
  32602. 24:36:57for evaluation.
  32603. 24:37:00So that's something that's something
  32604. 24:37:01we'll study.
  32605. 24:37:04How do we train? Uh it depends on the
  32606. 24:37:07model. Um so training will be a sense uh
  32607. 24:37:11will be an algorithm that will um
  32608. 24:37:14basically update the model according to
  32609. 24:37:16the data. These labeled examples. Um
  32610. 24:37:19every model is going to be different in
  32611. 24:37:21exactly how it trains. So we're going to
  32612. 24:37:24we're going to talk about that when we
  32613. 24:37:25get to the individual models that we'll
  32614. 24:37:27study.
  32615. 24:37:29But uh loosely speaking, they're going
  32616. 24:37:31to use the data to adjust itself. Like
  32617. 24:37:35imagine adjust like tuning a bunch of
  32618. 24:37:37knobs. Um, like the best example I can
  32619. 24:37:40give you is we I think I did this one
  32620. 24:37:43last week where you have kind of a
  32621. 24:37:45function
  32622. 24:37:46that predicts the price and let's say it
  32623. 24:37:50has
  32624. 24:37:51um weights like weight one with feature
  32625. 24:37:54one, weight two with feature two,
  32626. 24:37:58weight three with feature three. So
  32627. 24:38:01imagine we had three input features and
  32628. 24:38:03we we built an answer according to that.
  32629. 24:38:05Essentially what we would do to train
  32630. 24:38:07the model is adjust these
  32631. 24:38:12um in order to get this correct based on
  32632. 24:38:15our our labeled examples.
  32633. 24:38:19Okay.
  32634. 24:38:21So that's something we're going to learn
  32635. 24:38:22about coming up shortly when we when we
  32636. 24:38:24actually dive into model. Every model is
  32637. 24:38:26going to be slightly different in how it
  32638. 24:38:27trains, but at a high level it's going
  32639. 24:38:29to use the training data with those
  32640. 24:38:33examples, right? the labeled examples to
  32641. 24:38:35help guide the formula essentially to
  32642. 24:38:38adjust to generate the proper kind of
  32643. 24:38:42model here.
  32644. 24:38:44The these things are going to be
  32645. 24:38:45adjusted according to the data
  32646. 24:38:48in order to produce the correct output.
  32647. 24:38:52So think about these as knobs that will
  32648. 24:38:54turn.
  32649. 24:38:59Okay.
  32650. 24:39:02Uh which type of machine learning is
  32651. 24:39:04used? Uh probably supervised um which is
  32652. 24:39:06what we're talking about now. So
  32653. 24:39:08probably supervised because most people
  32654. 24:39:10want to
  32655. 24:39:12um build some type of model to predict
  32656. 24:39:14something.
  32657. 24:39:16Uh so yeah, I'd say I'd say supervise.
  32658. 24:39:25Yes, we're are we are definitely going
  32659. 24:39:26to learn how to train. Yeah, we'll see.
  32660. 24:39:28We'll do the code. Um, I'll tell you
  32661. 24:39:31about how it's done. Yeah, we're
  32662. 24:39:33definitely going to learn it. But what I
  32663. 24:39:34was saying is it's kind of on a model
  32664. 24:39:36bymodel basis.
  32665. 24:39:38So, I want to wait till we get into the
  32666. 24:39:40individual models, then we'll talk about
  32667. 24:39:42how they're trained.
  32668. 24:39:44But yeah, we'll we'll learn how to do
  32669. 24:39:45that.
  32670. 24:39:52But yeah, supervisor is used all over
  32671. 24:39:54the place. Even even for uh generative
  32672. 24:39:57models, they use supervised learning
  32673. 24:40:00because um like an LLM
  32674. 24:40:03is going to use labeled examples in
  32675. 24:40:06order to train, right? In order to train
  32676. 24:40:09how to generate responses according to
  32677. 24:40:12prompts. Um it needs to learn against a
  32678. 24:40:15lot of text examples.
  32679. 24:40:18So that supervised learning is what um
  32680. 24:40:21results in that model,
  32681. 24:40:24right? Learning from those labeled
  32682. 24:40:26examples.
  32683. 24:40:36Okay.
  32684. 24:40:42It is yeah image image uh a lot of um
  32685. 24:40:47yeah a lot of image processing is
  32686. 24:40:48supervised like object detection. So the
  32687. 24:40:52YOLO model is an object detection model.
  32688. 24:40:54Yes. Um because it has to be trained
  32689. 24:40:57right. It has to be trained on uh it has
  32690. 24:41:00to be trained on images
  32691. 24:41:04with labels such as this is what object
  32692. 24:41:06is in this image. This is the box around
  32693. 24:41:09the object.
  32694. 24:41:11Um yes. So if if it's if it ever uses
  32695. 24:41:15label data to train and build the model,
  32696. 24:41:18it is supervised. So YOLO is definitely
  32697. 24:41:21supervised and we actually we will we
  32698. 24:41:24will cover the YOLO model later on in
  32699. 24:41:26our deep learning course. We talk about
  32700. 24:41:29object detection.
  32701. 24:41:31So we'll we'll study that.
  32702. 24:41:34But yeah, it's supervised
  32703. 24:41:45Okay. So on the slide we have some
  32704. 24:41:48common supervised learning algorithms
  32705. 24:41:50that are we will study. So all of these
  32706. 24:41:53we will study and understand what they
  32707. 24:41:55do and how they work but just giving you
  32708. 24:41:58some to name them. linear regression is
  32709. 24:42:00kind of the one I just drew out which is
  32710. 24:42:02the um this is the prototypical like
  32711. 24:42:06easiest to understand model that is kind
  32712. 24:42:08of the um exactly like this where we
  32713. 24:42:11have a weight times a feature
  32714. 24:42:14um a weight times a feature and then a
  32715. 24:42:17weight times a feature
  32716. 24:42:19and on and on and on. You can have as
  32717. 24:42:21many as you want.
  32718. 24:42:23um that is a linear regression. And so
  32719. 24:42:26that is um that's a supervised model
  32720. 24:42:29because we need this value here and we
  32721. 24:42:33need all of our inputs in order to um
  32722. 24:42:36actually train this model and generate
  32723. 24:42:38all those weights
  32724. 24:42:40um that that is uh that uses um labeled
  32725. 24:42:45examples to help tune all those knobs.
  32726. 24:42:48Um same with all these other models. So,
  32727. 24:42:49we're going to talk about decision
  32728. 24:42:50trees. We're going to talk about
  32729. 24:42:51logistic regression and and SVMs, which
  32730. 24:42:53are support vector machines. We'll talk
  32731. 24:42:56about all of those, but they're all
  32732. 24:42:57examples of supervised uh supervised
  32733. 24:42:59learning.
  32734. 24:43:01Okay, we'll talk about all of these.
  32735. 24:43:05They're all supervised because they all
  32736. 24:43:08require labeled examples in order to
  32737. 24:43:10train them and and then subsequently use
  32738. 24:43:13them. Okay.
  32739. 24:43:25Okay. So what are some use case
  32740. 24:43:26examples? So for for instance in uh
  32741. 24:43:29supervised learning we may be predicting
  32742. 24:43:30temperature based on yearly temperature
  32743. 24:43:33trends. So we would have that yearly
  32744. 24:43:36data as our um as our labeled examples
  32745. 24:43:40and those would supervise the learning
  32746. 24:43:42of a model that predicts temperature.
  32747. 24:43:44Um, same thing with predicting crop
  32748. 24:43:46yield based on um, seasonal crop quality
  32749. 24:43:50changes. So maybe we have a bunch of
  32750. 24:43:51features relating to crop quality. We
  32751. 24:43:54could predict crop yield. Um, we would
  32752. 24:43:58just need historical examples with those
  32753. 24:44:01labels, right? What the crop yield is
  32754. 24:44:03for each time period. Let's say we would
  32755. 24:44:07just need those uh, supervised examples
  32756. 24:44:09and we could easily build a model off of
  32757. 24:44:12it.
  32758. 24:44:13Um
  32759. 24:44:15uh this this last one sorting waste
  32760. 24:44:17based on known waste items and their
  32761. 24:44:19corresponding waste types. Um that's
  32762. 24:44:22kind of like spam. It's like filtering
  32763. 24:44:24basically like a spam filtering. Um so
  32764. 24:44:27think of it like the the shapes example.
  32765. 24:44:29We sorting things into squares, circles,
  32766. 24:44:32triangles. Um, same kind of idea here
  32767. 24:44:34where we have a bunch of examples on
  32768. 24:44:36what those um what those waste items
  32769. 24:44:40should uh should belong to, like what
  32770. 24:44:43wastist bins they would go to, for
  32771. 24:44:45example. Um, and those could be labeled
  32772. 24:44:50and therefore then we could um
  32773. 24:44:53understand what category of waste they
  32774. 24:44:55belong to.
  32775. 24:44:57Um, same thing with spam. Something is
  32776. 24:44:59fraud or not fraud. spam or not spam,
  32777. 24:45:02cancer or not cancer. All of those are
  32778. 24:45:03going to be supervised learning examples
  32779. 24:45:05because they're going to require in
  32780. 24:45:07order to train them, they're going to
  32781. 24:45:09require data that has those labels.
  32782. 24:45:12Okay? So, anything that has labels is
  32783. 24:45:15going to be supervised learning. So
  32784. 24:45:20again, this is where we will spend
  32785. 24:45:23probably the the majority of our time is
  32786. 24:45:26doing supervised learning problems, ones
  32787. 24:45:29that we have labeled data. We're
  32788. 24:45:31building a model and we're going to
  32789. 24:45:32predict those those uh labels
  32790. 24:45:35essentially.
  32791. 24:45:44Okay, before we go to unsupervised, any
  32792. 24:45:47questions about uh supervised.
  32793. 24:46:04Okay.
  32794. 24:46:07All right. So supervised requires labels
  32795. 24:46:11in order to have an example to go off of
  32796. 24:46:14to build your model. And that's because
  32797. 24:46:17you're predicting those kind of outcomes
  32798. 24:46:19like spam or not spam, cancer or not
  32799. 24:46:21cancer. Now unsupervised learning is
  32800. 24:46:25completely different. It's the opposite.
  32801. 24:46:28So unsupervised learning is where we do
  32802. 24:46:32not use labels whatsoever. So we're not
  32803. 24:46:35using any labels at all. So it's it it
  32804. 24:46:38can be completely unlabeled or even if
  32805. 24:46:40it's labeled, we're not using labels in
  32806. 24:46:42any way. But um we primarily would say
  32807. 24:46:45it's unlabeled data. We have no guidance
  32808. 24:46:48because we're not using the labels in
  32809. 24:46:50any way. We have no guidance to um
  32810. 24:46:54predict anything, but that's because
  32811. 24:46:56we're not really predicting anything in
  32812. 24:46:57unsupervised learning. Generally, what
  32813. 24:46:59we're doing is looking for some
  32814. 24:47:01structure or pattern.
  32815. 24:47:03Okay, with unsupervised learning, we're
  32816. 24:47:05looking for some structure or pattern.
  32817. 24:47:07So, um, one type of example that's very
  32818. 24:47:12very popular is going to be this second
  32819. 24:47:14one, which is, um, identification
  32820. 24:47:17identification of user groups based on
  32821. 24:47:20similarities or commonalities. Now, this
  32822. 24:47:22is going to be a problem basically known
  32823. 24:47:25as clustering
  32824. 24:47:29and it's a problem we will study quite a
  32825. 24:47:31bit. there's going to turn out to be
  32826. 24:47:33lots of different algorithms that can
  32827. 24:47:35accomplish clustering. So what
  32828. 24:47:37clustering attempts to do is basically
  32829. 24:47:39say um we have data that's like this and
  32830. 24:47:43then data over here and then data over
  32831. 24:47:46here. Let's just group these together.
  32832. 24:47:48So like this should be one group, this
  32833. 24:47:50should be one group and this should be
  32834. 24:47:51one group. And we can find those
  32835. 24:47:54structures and say okay this is group
  32836. 24:47:57one, this is group two and this is group
  32837. 24:48:00three.
  32838. 24:48:02one, two, three. And we can basically
  32839. 24:48:05build what we would call clusters of
  32840. 24:48:07data um based on how close together the
  32841. 24:48:11points are kind of located in these kind
  32842. 24:48:13of cluster zones like these boxes I've
  32843. 24:48:16drawn.
  32844. 24:48:18Okay. Now, that doesn't require any
  32845. 24:48:20label to do which is really fascinating.
  32846. 24:48:22So, unsupervised, you don't need any
  32847. 24:48:24label at all to accomplish the
  32848. 24:48:26algorithm. Um so, clustering is one good
  32849. 24:48:29example.
  32850. 24:48:30um finding outliers or anomalies is
  32851. 24:48:33another. So we don't necessarily have
  32852. 24:48:34any label of what is an outlier or what
  32853. 24:48:37is an anomaly. We are deriving that from
  32854. 24:48:40the features alone. There's no guidance.
  32855. 24:48:43There's no label um to doing like
  32856. 24:48:45outlier detection or anomaly detection.
  32857. 24:48:49Okay, so that's another good example.
  32858. 24:48:50One that's not listed on here um but is
  32859. 24:48:55also really important that we will study
  32860. 24:48:57is something known as dimensionality
  32861. 24:48:59reduction.
  32862. 24:49:01So dim reduction and what that what this
  32863. 24:49:05focuses on is basically compressing the
  32864. 24:49:08data set a bit. So we take our data and
  32865. 24:49:11basically compress it um so that but we
  32866. 24:49:15do it in such a way that we retain as
  32867. 24:49:17much information as we can. This is a
  32868. 24:49:20very like smart compression and what it
  32869. 24:49:23does is it lowers the dimension. Um
  32870. 24:49:26dimension think of the dimension as like
  32871. 24:49:29number of columns.
  32872. 24:49:32Number of columns.
  32873. 24:49:34So imagine we had 100 columns in a data
  32874. 24:49:37frame. What we could do is actually
  32875. 24:49:38reduce that down to 10. So like 10% of
  32876. 24:49:42that. So we reduce it down to 10. And um
  32877. 24:49:47but those 10 are it's not like we
  32878. 24:49:49chopped out um 90 other columns. We um
  32879. 24:49:54smartly kind of compressed all that
  32880. 24:49:56information into these 10 new columns um
  32881. 24:49:59that are compressed versions of the
  32882. 24:50:01hundred that we used to have. Um so
  32883. 24:50:04dimensionality reduction is is another
  32884. 24:50:07unsupervised technique. It requires no
  32885. 24:50:09guidance, no label to do, but is um a
  32886. 24:50:13really useful technique to reduce the
  32887. 24:50:15size of your data if you're doing things
  32888. 24:50:17with it. Um so this is another one that
  32889. 24:50:20we will we'll study how to do it and
  32890. 24:50:23basically more details behind it, what
  32891. 24:50:24the algorithms are.
  32892. 24:50:26Um we'll so probably those two in
  32893. 24:50:30unsupervised will spend the most amount
  32894. 24:50:31of time on clustering and dimensionality
  32895. 24:50:33reduction.
  32896. 24:50:42uh and supervised if some data is
  32897. 24:50:43present but we didn't label it means in
  32898. 24:50:46example we had circle triangle square in
  32899. 24:50:49the training data we add pentagon
  32900. 24:50:58but we didn't label that in that case
  32901. 24:51:04uh yeah so every um in supervised
  32902. 24:51:07learning, every row, think about it as
  32903. 24:51:10like every row in our data frame needs
  32904. 24:51:12to have a label
  32905. 24:51:14uh associated to it. It needs to have a
  32906. 24:51:16a column that represents the label.
  32907. 24:51:21So if we've never seen Pentagon before,
  32908. 24:51:23I can't use that as a label.
  32909. 24:51:29So, it has to the Pentagon has to exist
  32910. 24:51:32in the data if I'm going to be able to
  32911. 24:51:35predict it,
  32912. 24:51:41right? So, I can't predict, right? If
  32913. 24:51:43we've never seen it before, we have no
  32914. 24:51:45examples to go off. We have no guidance.
  32915. 24:51:47So, how could we predict that?
  32916. 24:51:50Right? We can't predict it.
  32917. 24:52:05if it's if it's in there. So if if we
  32918. 24:52:08have labels of Pentagon, let's say, then
  32919. 24:52:11yeah, we could predict Pentagon. We
  32920. 24:52:14could
  32921. 24:52:29remove. Remove what?
  32922. 24:52:36We wouldn't if it was talking about the
  32923. 24:52:37Pentagon, we wouldn't remove that. No,
  32924. 24:52:39let me go back to that page. We wouldn't
  32925. 24:52:42remove it. Um, it's just if it's not in
  32926. 24:52:46our labels, we're not going to be able
  32927. 24:52:47to predict it. So, Pentagon's a good
  32928. 24:52:50example here. Uh, Pentagon is not one of
  32929. 24:52:55our labels. So, it currently is not in
  32930. 24:52:58our data set as one of the labels. We
  32931. 24:53:00only have data that's either a triangle,
  32932. 24:53:02circle, or a square. We don't have
  32933. 24:53:05pentagon. So, I would never be able to
  32934. 24:53:08predict pentagon. I'll never be able to
  32935. 24:53:11do that if I haven't seen examples of it
  32936. 24:53:13before.
  32937. 24:53:15Okay. But let's say we had that in
  32938. 24:53:17there.
  32939. 24:53:19So we had Pentagon.
  32940. 24:53:24So if we had Pentagon, um we could have
  32941. 24:53:27an example of it in our labels
  32942. 24:53:37and then yeah, we it could be then we
  32943. 24:53:38could predict it.
  32944. 24:53:46Yeah. Yeah. The the don't get worried.
  32945. 24:53:48Don't worry about the test data. So the
  32946. 24:53:51test data is just saying here's a new
  32947. 24:53:53here's a shape. What is it? Okay, that's
  32948. 24:53:55a square. Here's a shape. What is it?
  32949. 24:53:57Okay, that's a triangle. And we could
  32950. 24:53:59have as many of those examples as we
  32951. 24:54:01want in our test data. So we could have
  32952. 24:54:03a circle and say, okay, what's this?
  32953. 24:54:06Should be circle,
  32954. 24:54:09right? The test data can be whatever it
  32955. 24:54:11whatever it wants. But yeah, if if we've
  32956. 24:54:13never seen Pentagon before, we're never
  32957. 24:54:15going to be able to predict it.
  32958. 24:54:25These are the the label data and labels
  32959. 24:54:29are basically the talking about the same
  32960. 24:54:31thing. The labels just mean what are the
  32961. 24:54:34categories that are present in our data.
  32962. 24:54:39So in this data we only have three
  32963. 24:54:41labels that are present.
  32964. 24:54:48So the labels is are relative to our
  32965. 24:54:51label data, right? It's saying
  32966. 24:54:54what labels,
  32967. 24:54:58excuse me, what labels uh do we have
  32968. 24:55:03in our data and we only have those three
  32969. 24:55:05circle, triangle, square. So so Pentagon
  32970. 24:55:08would not be part of those labels. We
  32971. 24:55:10couldn't predict it.
  32972. 24:55:24No. So unsupervised is not going to make
  32973. 24:55:27a prediction. That's the big difference
  32974. 24:55:29with unsupervised. They're not going to
  32975. 24:55:31make a prediction like this. Um so
  32976. 24:55:34unsupervised is not going to make a
  32977. 24:55:36prediction. It's going to do something
  32978. 24:55:37different like um basically say like
  32979. 24:55:40these guys are similar, these are
  32980. 24:55:42similar, these are similar, this is a
  32981. 24:55:45cluster, this is a cluster, this is a
  32982. 24:55:46cluster. It's not going to make a
  32983. 24:55:50prediction. That's what supervised
  32984. 24:55:52learning does.
  32985. 24:55:56Clustering, yes, which is unsupervised,
  32986. 24:55:59yes, clustering does not require any
  32987. 24:56:01labels. Unsupervised just means we don't
  32988. 24:56:03have any labels. We don't require any
  32989. 24:56:05labels.
  32990. 24:56:24So the other thing unsupervised might do
  32991. 24:56:26is it might say
  32992. 24:56:29and again without the labels it might
  32993. 24:56:31say that this is an outlier.
  32994. 24:56:36it might say that this guy is an outlier
  32995. 24:56:38because there's only there's only one of
  32996. 24:56:40those and they're not like the other. So
  32997. 24:56:42that that's something that um that's
  32998. 24:56:45something that uh unsupervised could do.
  32999. 24:56:54Um it it yeah and no. It kind of labels
  33000. 24:56:59a cluster in the sense that um it would
  33001. 24:57:03basically assign a number to it like
  33002. 24:57:05this is cluster one, this is cluster
  33003. 24:57:08two, this is cluster three.
  33004. 24:57:13It'll assign a number to it, but it's
  33005. 24:57:15not a very meaningful it doesn't assign
  33006. 24:57:17like a prediction label in the in the
  33007. 24:57:20traditional sense of a label.
  33008. 24:57:22It does provide like a numerical index
  33009. 24:57:24for the cluster to because what we want
  33010. 24:57:26to know is like okay this guy has the
  33011. 24:57:29cluster of one. This guy belongs to
  33012. 24:57:31cluster one. This guy belongs to cluster
  33013. 24:57:33one. This guy belongs to cluster two.
  33014. 24:57:35This guy belongs to cluster two. Does
  33015. 24:57:37that make sense? So there needs to be
  33016. 24:57:38some like index of what cluster you
  33017. 24:57:40belong to.
  33018. 24:57:43So it's kind of like a label but not in
  33019. 24:57:46the traditional like prediction sense.
  33020. 24:58:00Very
  33021. 24:58:15good. So again, unsupervised, no labels.
  33022. 24:58:19You're doing things like
  33023. 24:58:22identifying clusters,
  33024. 24:58:24um identifying outliers, doing
  33025. 24:58:27dimensionality reduction. These are all
  33026. 24:58:29like structure and pattern oriented
  33027. 24:58:32things. They're not predictions of a
  33028. 24:58:34label. Okay? They're not which is what
  33029. 24:58:37we would see in supervised learning.
  33030. 24:58:46Okay. So an example would be that we
  33031. 24:58:49take we put in the data um we can group
  33032. 24:58:53together uh data such as images into
  33033. 24:58:57categories based on similarities um
  33034. 24:59:00which would be like those clusters. So
  33035. 24:59:01there's no these would be groups that we
  33036. 24:59:04don't have any label on ahead of time
  33037. 24:59:06like we don't have we don't say that
  33038. 24:59:07this image should belong to this this
  33039. 24:59:09image should belong to this we derive
  33040. 24:59:12that from the characteristics of the
  33041. 24:59:14data. Um so think like a good example is
  33042. 24:59:18um customer groups. So we would identify
  33043. 24:59:21customers based on like okay do they
  33044. 24:59:24have similar spending levels? How many
  33045. 24:59:27days do they go shopping in a week? How
  33046. 24:59:29much money do they spend? And we can
  33047. 24:59:31kind of group together customers based
  33048. 24:59:33on similar qualities.
  33049. 24:59:36Clustering will find those groups that
  33050. 24:59:38should exist.
  33051. 24:59:41um it will discover those groups based
  33052. 24:59:43on um the similarities in the data, but
  33053. 24:59:47there's no labels that that say like
  33054. 24:59:49this person should be in this group,
  33055. 24:59:52this person should be in this ahead of
  33056. 24:59:54time. There's no labels of that. It gets
  33057. 24:59:57derived during the algorithm. It's
  33058. 24:59:59unsupervised,
  33059. 25:00:02right? There's no unsupervised really
  33060. 25:00:04literally means no guidance. There's no
  33061. 25:00:07guidance to doing it. We just derive
  33062. 25:00:09that from the structure of the data
  33063. 25:00:11which is the similarities.
  33064. 25:00:27Okay.
  33065. 25:00:33All right. So,
  33066. 25:00:36a couple more for you. So we had um
  33067. 25:00:39supervised which uses the labels. We
  33068. 25:00:44have unsupervised which uses no labels
  33069. 25:00:47looking for structure. And then we have
  33070. 25:00:49something that's kind of in between
  33071. 25:00:52which is um what is known as
  33072. 25:00:56semiupervised learning. And this is
  33073. 25:00:59where you use a combination of a little
  33074. 25:01:03bit of label data, but most of your data
  33075. 25:01:05is actually unlabeled data. Um, and you
  33076. 25:01:09try to get some use out of that label
  33077. 25:01:12data in order to um build a model out of
  33078. 25:01:17it. And so uh it uses the um it uses
  33079. 25:01:22that label data to um generally provide
  33080. 25:01:27some guidance on usually what happens
  33081. 25:01:29with semi-supervised learning is you use
  33082. 25:01:32your label data to kind of predict what
  33083. 25:01:35the label should be for the unlabelled
  33084. 25:01:38data and then you can go from there. So
  33085. 25:01:40you can create artificial labels on this
  33086. 25:01:42unlabeled data and then you can use all
  33087. 25:01:46of it once it's all been labeled kind of
  33088. 25:01:48like a supervised learning uh approach.
  33089. 25:01:51So but but this is semi-supervised
  33090. 25:01:53basically refers to the fact that you
  33091. 25:01:55start out with most of your data not
  33092. 25:01:58being labeled but you do have some
  33093. 25:02:01labeled examples and what you can do is
  33094. 25:02:03basically extrapolate those labels into
  33095. 25:02:06the unlabeled data set and then provide
  33096. 25:02:10some artificial labels and then now
  33097. 25:02:13everything has a label you can do
  33098. 25:02:14supervised learning.
  33099. 25:02:16Okay. So, it falls kind of between um
  33100. 25:02:20supervised and and unsupervised.
  33101. 25:02:23Uh and there so this is this is kind of
  33102. 25:02:26rare. Most of the time you're not going
  33103. 25:02:28to do that. You're actually just going
  33104. 25:02:30to um prefer to just start with all
  33105. 25:02:33label data. That's usually the preferred
  33106. 25:02:35approach. Most of the time you'll
  33107. 25:02:37actually just be doing supervised
  33108. 25:02:38learning, not really semi-supervised
  33109. 25:02:41learning. So, it's pretty rare, but um
  33110. 25:02:49it it could like if Yeah, it could if
  33111. 25:02:52the if we had a lot of examples of
  33112. 25:02:54Pentagon and we wanted and so they were
  33113. 25:02:56unlabeled and then we tried to guess
  33114. 25:02:58what kind of shape they were um and
  33115. 25:03:02provide an artificial label uh and then
  33116. 25:03:05um then use that whole data set to build
  33117. 25:03:07a model off of then yeah it could it
  33118. 25:03:10could fall into this category. Okay.
  33119. 25:03:18They Oh, going back to the question,
  33120. 25:03:20they still use some kind of label data
  33121. 25:03:21like age, gender. They use uh that's
  33122. 25:03:24those aren't those aren't really labels.
  33123. 25:03:26That's the features. So, yeah, they
  33124. 25:03:29still use the core features of the data.
  33125. 25:03:32They just don't have any like labels in
  33126. 25:03:34the traditional sense of a label. Like
  33127. 25:03:36you should think of a label as something
  33128. 25:03:39we are trying to predict.
  33129. 25:03:41So whether that's a price, whether
  33130. 25:03:43that's like a category like spam, not
  33131. 25:03:46spam, cancer, not cancer, it's something
  33132. 25:03:48we'd be interested in kind of
  33133. 25:03:49predicting. And so um in our data, we
  33134. 25:03:52would have an answer for every row. We'd
  33135. 25:03:55have one of our columns would be like
  33136. 25:03:56the the result like the outcome answer
  33137. 25:03:59that we're trying to predict. That's the
  33138. 25:04:01label.
  33139. 25:04:03So in unsupervised, we don't have any of
  33140. 25:04:05the labels.
  33141. 25:04:07We do have just the regular features
  33142. 25:04:09like gender, age, income, square
  33143. 25:04:13footage,
  33144. 25:04:17bedrooms, bathrooms, all those things.
  33145. 25:04:24Okay.
  33146. 25:04:28So, we have semi-supervised that falls
  33147. 25:04:30in between supervised. Now, the reason
  33148. 25:04:32it falls between is be is because
  33149. 25:04:35there's a decent amount of data that's
  33150. 25:04:37unlabeled. In fact, a majority of it
  33151. 25:04:39unlabeled. But what we can do is try to
  33152. 25:04:43label it. We can try to take what we
  33153. 25:04:45know from our existing labels and
  33154. 25:04:48predict an artificial label and then use
  33155. 25:04:51all that data together in kind of a
  33156. 25:04:54supervised fashion for a model down the
  33157. 25:04:56road.
  33158. 25:05:02So that's kind of what this picture uh
  33159. 25:05:04says is we can try to take um you know
  33160. 25:05:08maybe we try to infer some labels based
  33161. 25:05:10on we have some some labelled data here.
  33162. 25:05:14We have most of our data is unlabeled
  33163. 25:05:16and we try to supply some labels to it.
  33164. 25:05:20Um like maybe we have a babies category
  33165. 25:05:22of teens, a tween, uh you know youth and
  33166. 25:05:27um adults. Um and then we try so we we
  33167. 25:05:31take our our labels and we try to
  33168. 25:05:35extrapolate those into artificial labels
  33169. 25:05:37for this unlabelled data so that we can
  33170. 25:05:39use it now because then everything has a
  33171. 25:05:42label at this point and then we can just
  33172. 25:05:44go ahead and do supervised learning from
  33173. 25:05:46there.
  33174. 25:05:52So we can do supervised from there. What
  33175. 25:05:53we would prefer to do and what we'll do
  33176. 25:05:55in this course
  33177. 25:05:57um is just start with supervised. We'll
  33178. 25:06:00just start with the labels. We won't try
  33179. 25:06:02to derive artificial labels usually.
  33180. 25:06:04We'll just start with labels.
  33181. 25:06:15So one example in the real world is
  33182. 25:06:18something like Google photos which um
  33183. 25:06:22whenever you take a picture it can
  33184. 25:06:24provide uh uh labels based on previous
  33185. 25:06:28uh images in your library. So it can it
  33186. 25:06:32can produce tags or um labels on those.
  33187. 25:06:37Uh generally when you take that picture
  33188. 25:06:39it's kind of unlabeled unless you go in
  33189. 25:06:41and specifically provide some tags and
  33190. 25:06:44some labels. But um if you don't do that
  33191. 25:06:46it can still it can still uh make it can
  33192. 25:06:51artificially create one of those based
  33193. 25:06:53on the other label data that you already
  33194. 25:06:55have.
  33195. 25:06:57So that's um
  33196. 25:07:00that's an example.
  33197. 25:07:08Okay.
  33198. 25:07:12All right. Last one in terms of machine
  33199. 25:07:14learning. So we have supervised, we have
  33200. 25:07:18unsupervised.
  33201. 25:07:20Uh then we had semi-supervised which is
  33202. 25:07:22somewhere in between a mixture of having
  33203. 25:07:23some unlabelled data and label data. Um
  33204. 25:07:26now we're going to talk about
  33205. 25:07:27reinforcement learning which is
  33206. 25:07:29completely different. Um it's it's
  33207. 25:07:32completely different than the other
  33208. 25:07:33three. It's a type of machine learning
  33209. 25:07:36where we uh basically learn from
  33210. 25:07:39interaction with the environment. And
  33211. 25:07:41you might ask what are we learning? We
  33212. 25:07:44are learning what actions to take in the
  33213. 25:07:48environment. Um and the way we do that
  33214. 25:07:51is by reinforcing
  33215. 25:07:53positive actions that lead to a a
  33216. 25:07:56reward. Um, so that's where the word
  33217. 25:08:00reinforcement comes from is we we
  33218. 25:08:02basically uh imagine like a child
  33219. 25:08:05that's, you know, learning from trial
  33220. 25:08:07and error. Like they're trying to crawl,
  33221. 25:08:09they're trying to walk and they keep
  33222. 25:08:10falling down. um eventually they learn
  33223. 25:08:13how to do it through trial and error and
  33224. 25:08:15they might get a reward or they might um
  33225. 25:08:19reinforce some of those positive
  33226. 25:08:21movements that lead them to walk or
  33227. 25:08:24crawl um or they might learn from the
  33228. 25:08:28penalties, right? They might learn from
  33229. 25:08:31uh some type of feedback. So they might
  33230. 25:08:34learn from falling down like, "Oh, that
  33231. 25:08:35hurts. I should uh support myself a
  33232. 25:08:38little bit better, right?" Or be a
  33233. 25:08:39little more coordinated. Um
  33234. 25:08:42and so they they learn from those
  33235. 25:08:44actions and their interaction with the
  33236. 25:08:46environment. Um
  33237. 25:08:49uh so this is a complex um algorithm
  33238. 25:08:55essentially uh it's it deals a lot with
  33239. 25:08:59um again taking actions. Usually when
  33240. 25:09:02you take an action something changes in
  33241. 25:09:04the environment um then you kind of
  33242. 25:09:08observe some type of feedback. So, think
  33243. 25:09:10about like a a board game where you're
  33244. 25:09:13trying to figure out what move you
  33245. 25:09:15should make or another good example is
  33246. 25:09:17like with a robot um trying to navigate
  33247. 25:09:20a maze. So, like what route should it
  33248. 25:09:23take? Should it move forward? Should it
  33249. 25:09:24move backward? Should it move left or
  33250. 25:09:26right? Those are different actions it
  33251. 25:09:28can take. Also, like a self-driving car,
  33252. 25:09:30should it should it turn? Should it
  33253. 25:09:32speed up? Should it slow down? Those are
  33254. 25:09:34all good examples of things that have
  33255. 25:09:36been trained from reinforcement
  33256. 25:09:38learning.
  33257. 25:09:45Uh yeah. So real world examples would be
  33258. 25:09:48like in a board game uh a a reward would
  33259. 25:09:51be like if you win the game. Um or if
  33260. 25:09:55you like capture a piece like in
  33261. 25:09:57checkers or chess, that's a reward. A
  33262. 25:10:00penalty would be like if you lose the
  33263. 25:10:01game or lose one of your pieces, that
  33264. 25:10:03could be a a penalty.
  33265. 25:10:06um in a board game or sorry in like a a
  33266. 25:10:11robot navigation task, it could get
  33267. 25:10:14rewards for um moving in the right
  33268. 25:10:16direction
  33269. 25:10:18um towards the exit or like when it like
  33270. 25:10:21let's say you wanted to train a robot on
  33271. 25:10:22how to open the door and navigate a
  33272. 25:10:25room. Um you would penalize it for
  33273. 25:10:27bumping into the wall.
  33274. 25:10:29Um you would give it a reward for moving
  33275. 25:10:37usually oh like oh the algorithms
  33276. 25:10:40themselves usually it's like a a step
  33277. 25:10:42function um it's usually it's like a
  33278. 25:10:45discrete function that kind of is based
  33279. 25:10:47on the state so the reward it could be
  33280. 25:10:50like um like depending on the let's
  33281. 25:10:54let's go back to the board game example
  33282. 25:10:55like the reward could be like or even
  33283. 25:10:58the maze let's say like a navigating the
  33284. 25:11:01maze like getting to this let's say this
  33285. 25:11:02was the exit
  33286. 25:11:05And this was the entrance.
  33287. 25:11:09Then if they make it to here, they get a
  33288. 25:11:11numerical like if they make it to the
  33289. 25:11:13exit, they get a numerical reward of
  33290. 25:11:14like plus 100, let's say. So it's just a
  33291. 25:11:17number. And then if they uh like if they
  33292. 25:11:21bump if they go into here, like let's
  33293. 25:11:23say this is kind of like a death trap or
  33294. 25:11:25like a pit, this this would be like a
  33295. 25:11:27minus 100. So it could be like discrete
  33296. 25:11:30numerical values could be the reward if
  33297. 25:11:33they're moving in the right direction.
  33298. 25:11:34Like let's say we want to encourage
  33299. 25:11:36going this way then we could give
  33300. 25:11:37smaller intermediate rewards like this
  33301. 25:11:39should be a plus like if you move
  33302. 25:11:41forward this is a plus five this is a
  33303. 25:11:44plus 10 this is a plus 15 if you're
  33304. 25:11:47moving in the wrong direction away from
  33305. 25:11:49the exit. Um that would be like a minus5
  33306. 25:11:52or a minus 10. Does that make sense? So
  33307. 25:11:55they're they're numerical in nature and
  33308. 25:11:58what you're trying to do is collect the
  33309. 25:11:59most reward. You're trying to get the
  33310. 25:12:01largest reward you can through trial and
  33311. 25:12:05error. So you you try this out many many
  33312. 25:12:07many times. You basically simulate
  33313. 25:12:10running through this maze many many
  33314. 25:12:13times. And what dictates it what
  33315. 25:12:16dictates like where I should go is based
  33316. 25:12:19on what I've observed in the past. It's
  33317. 25:12:21almost like you're a child remembering
  33318. 25:12:22like, okay, what move should I make from
  33319. 25:12:24this space? Like if I'm here, if I'm
  33320. 25:12:27here, which way should I go? Should I go
  33321. 25:12:29down? Should I go right? Should I go
  33322. 25:12:31left? You kind of know that from
  33323. 25:12:33experience.
  33324. 25:12:35Does that make sense? Based on the
  33325. 25:12:36reward that I've seen in the past, like
  33326. 25:12:39when I've moved down, I've gotten a
  33327. 25:12:40higher reward than moving left or right.
  33328. 25:12:44Does that make sense? So, yeah, it's
  33329. 25:12:45it's a numerical value
  33330. 25:12:48as a reward.
  33331. 25:12:57Yeah, that's a great question. Um, how
  33332. 25:13:00does it differentiate rewards based on
  33333. 25:13:02gain and loss? I chess. So it's it's a
  33334. 25:13:05very comp complicated uh answer but
  33335. 25:13:07essentially every so in the chess board
  33336. 25:13:12you can think of the board as like every
  33337. 25:13:14every um
  33338. 25:13:17space is a state.
  33339. 25:13:21So I could be in this state I could be
  33340. 25:13:23in this state and then it's not not only
  33341. 25:13:26is every every uh space but where all
  33342. 25:13:29the other pieces are. So there's lots of
  33343. 25:13:31states that are possible.
  33344. 25:13:33Um, so
  33345. 25:13:36the way there's a way to quantify
  33346. 25:13:39essentially what's the value of taking a
  33347. 25:13:43certain action like moving my piece
  33348. 25:13:44left, moving it right, moving it up or
  33349. 25:13:47down um given the rest of the state. So
  33350. 25:13:51you're you're right, it may be
  33351. 25:13:52beneficial to sacrifice. Um, but we
  33352. 25:13:56would learn that through experience that
  33353. 25:13:58okay, the best move in this situation is
  33354. 25:14:00to sacrifice.
  33355. 25:14:02We would we would have to learn that
  33356. 25:14:04through trial and error many many many
  33357. 25:14:06times which is to say like okay if I'm
  33358. 25:14:08in this current state of the world right
  33359. 25:14:12all these pieces are distributed in this
  33360. 25:14:14way the best move for me right now in
  33361. 25:14:16the long run
  33362. 25:14:18to get the most reward in the long run
  33363. 25:14:21is to actually sacrifice my piece and
  33364. 25:14:22move it right move it into like a bad
  33365. 25:14:25position theoretically but we know from
  33366. 25:14:28experience that's actually the most
  33367. 25:14:29long-term reward is from that position
  33368. 25:14:32like moving it right may be the best
  33369. 25:14:35for me. So what you learn is how to take
  33370. 25:14:38actions
  33371. 25:14:40and actions are usually like move right,
  33372. 25:14:43move left, move up, move down. You think
  33373. 25:14:45about like a self-driving car though,
  33374. 25:14:47that's going to be like slow down, speed
  33375. 25:14:50up, turn your wheel 10 degrees. Um those
  33376. 25:14:55kind of actions.
  33377. 25:14:59So the the short answer is it's there's
  33378. 25:15:02a calculation there that you learn what
  33379. 25:15:06the long-term value of every state is
  33380. 25:15:10every unique state
  33381. 25:15:13and then you're trying to basically say
  33382. 25:15:15what action should I take from that
  33383. 25:15:17state
  33384. 25:15:19given that current state of the
  33385. 25:15:29Okay.
  33386. 25:15:31And I really I really like reinforcement
  33387. 25:15:34learning. It's actually probably my
  33388. 25:15:35favorite field of machine learning.
  33389. 25:15:38Unfortunately, we won't be covering it
  33390. 25:15:40um in our main uh course. We have
  33391. 25:15:43offered uh electives around
  33392. 25:15:45reinforcement learning in the past. So,
  33393. 25:15:47um stay tuned. Maybe when we get to the
  33394. 25:15:49end of this program, uh we'll offer an
  33395. 25:15:51elective on it and if enough people sign
  33396. 25:15:53up for it, we'll we'll run it. But, um
  33397. 25:15:57we it's not part of our we don't really
  33398. 25:15:59cover reinforcement learning as part of
  33399. 25:16:00our main topics. It's it is an advanced
  33400. 25:16:03uh more advanced topic than than what
  33401. 25:16:05we'll cover, but um I I really enjoy it.
  33402. 25:16:08I find it very fascinating.
  33403. 25:16:17Okay. So, all of this is kind of um
  33404. 25:16:20illustrating what I was saying, which is
  33405. 25:16:22um you think of like uh the thing that's
  33406. 25:16:25interacting in the environment like the
  33407. 25:16:27robot or the car or the human moving a
  33408. 25:16:30chest piece is known as the agent. It's
  33409. 25:16:34interacting with the environment by
  33410. 25:16:36taking actions which updates the state
  33411. 25:16:39um of of the environment. So that's
  33412. 25:16:41that's why you see this word state here.
  33413. 25:16:43This gets updated constantly every time
  33414. 25:16:45you take an action. Um ultimately what
  33415. 25:16:48reinforcement learning is trying to do
  33416. 25:16:49is learn the best action like what would
  33417. 25:16:52be the best action to take. Um
  33418. 25:16:56and the best action is is the one that
  33419. 25:16:58leads to the most long-term reward.
  33420. 25:17:01That's the best action. Um, so you have
  33421. 25:17:04to uh you have to learn what you know
  33422. 25:17:09what leads to a good reward by kind of
  33423. 25:17:12experiencing this over and over and over
  33424. 25:17:14through trial and error. So there's a
  33425. 25:17:16lot of um kind of simulation or letting
  33426. 25:17:19the robot try something a lot um in
  33427. 25:17:22order to kind of learn what's rewarding
  33428. 25:17:24and what's not. Think about it again
  33429. 25:17:26like I think a good example is like with
  33430. 25:17:27children, right? you kind of have to let
  33431. 25:17:29them try things until they learn on
  33432. 25:17:32their own what's what can they do and
  33433. 25:17:34what can they not do
  33434. 25:17:37what's the best actions right
  33435. 25:17:41so reinforcement learning has made its
  33436. 25:17:43way into other places so I I said like a
  33437. 25:17:46good example is self-driving cars or ro
  33438. 25:17:48robotics a lot of reinforce
  33439. 25:17:50reinforcement learning is used there one
  33440. 25:17:52place it's found its way into recently
  33441. 25:17:54is recommendation systems have kind of
  33442. 25:17:57merged with reinforcement learning
  33443. 25:17:58learning. Um, and this is because you
  33444. 25:18:03you can imagine there's kind of a
  33445. 25:18:05built-in reward for you clicking on a
  33446. 25:18:08video and kind of watching it.
  33447. 25:18:11Um, so that kind of reinforces that
  33448. 25:18:13recommendation and then uh that's where
  33449. 25:18:17um you can then kind of recommend a
  33450. 25:18:20similar thing and see if that's
  33451. 25:18:22rewarding and generates a click or
  33452. 25:18:25generates some view time or watch time
  33453. 25:18:27or whatever. Um so reinforcement
  33454. 25:18:30learning has found its way into a lot of
  33455. 25:18:32areas. Um recommendations being one of
  33456. 25:18:35them because it's just natural for the
  33457. 25:18:38idea of like what um should I recommend
  33458. 25:18:40next to generate the most reward. In
  33459. 25:18:43this case, the reward is kind of
  33460. 25:18:44correlated to did they click on it or
  33461. 25:18:46not or did they how long did they watch
  33462. 25:18:49watch for longer it's more rewarding
  33463. 25:18:52um those kind of things but uh place
  33464. 25:18:56places where reinforcement learning have
  33465. 25:18:57been used I said self-driving cars um
  33466. 25:19:01games so uh one of the most famous
  33467. 25:19:05examples if you want to look it up is
  33468. 25:19:06the um Alph Go this was in 2016 um the
  33469. 25:19:11Alph Go uh algorithm was a reinforcement
  33470. 25:19:15learning bot that beat um some of the
  33471. 25:19:18world's best Go players, which if you're
  33472. 25:19:21not familiar, Go is a um board game
  33473. 25:19:25that is a little bit more uh complex
  33474. 25:19:29than chess. It has more more uh it's a
  33475. 25:19:32larger board um more pieces to it. Um
  33476. 25:19:37but they there was a reinforcement
  33477. 25:19:39learning powered bot that actually um
  33478. 25:19:41learned how to play the game so
  33479. 25:19:43effective it could beat um world kind of
  33480. 25:19:46masters at the games was pretty amazing.
  33481. 25:19:49Um that's the alpha go and that was by
  33482. 25:19:51deep mind Google and deep mind in 2016.
  33483. 25:19:55That was pretty that was only 10 years
  33484. 25:19:57ago not that long.
  33485. 25:19:59Um so certain uh we said recommendation
  33486. 25:20:04uh even autocorrect um learning to
  33487. 25:20:07predict like what is the best correction
  33488. 25:20:10uh to generate a reward which would be
  33489. 25:20:12like you accept that correction or you
  33490. 25:20:14reject it would be a penalty. Um so
  33491. 25:20:17reinforced learning has been adapted to
  33492. 25:20:20these kind of problems very
  33493. 25:20:21successfully. Let's take a look at the
  33494. 25:20:24packages that we will use throughout. So
  33495. 25:20:27um of course we will rely on these three
  33496. 25:20:30which we've already relied on to do a
  33497. 25:20:33lot of things like numpy to do numerical
  33498. 25:20:36manipulations and calculations.
  33499. 25:20:39Uh mapplot lib to do any plotting and
  33500. 25:20:42not only map lib but maybe seabour as
  33501. 25:20:44well both of those to do plotting. Um,
  33502. 25:20:48pandas is a big one because
  33503. 25:20:51that's where all of our data is going to
  33504. 25:20:53be manipulated and prepped before it
  33505. 25:20:55goes into modeling.
  33506. 25:20:57So all of that stuff we learn from
  33507. 25:20:59pandis is definitely going to be applied
  33508. 25:21:01here in this course uh as we actually
  33509. 25:21:04build models. Um so of course like these
  33510. 25:21:07old ones that we've been working with
  33511. 25:21:09quite a bit um still going to be useful
  33512. 25:21:12here in the modeling stage. Um mainly
  33513. 25:21:16for different reasons though mostly to
  33514. 25:21:17get our data prepared to do some type of
  33515. 25:21:20modeling or maybe to visualize it before
  33516. 25:21:23we do modeling to get a sense of what it
  33517. 25:21:24looks like those kind of things.
  33518. 25:21:29Um,
  33519. 25:21:31sci-fi is sometimes useful for certain
  33520. 25:21:35uh um processing like in unsupervised
  33521. 25:21:38learning. We'll actually use scyp a
  33522. 25:21:39little bit to do dimensionality
  33523. 25:21:41reduction or help us do that. Um so
  33524. 25:21:44scypi will be used here and there and
  33525. 25:21:47we've seen it before with hypothesis
  33526. 25:21:49testing. We use scypi like the t test
  33527. 25:21:51and z test came from there. Um, some of
  33528. 25:21:54the unsupervised learning stuff will
  33529. 25:21:56come out of there, but the package we
  33530. 25:21:58will use by far the most in this course
  33531. 25:22:02is going to be Scikitlearn,
  33532. 25:22:05which is here. Um, and we've already
  33533. 25:22:09seen a little bit about scikitlearn in
  33534. 25:22:11terms of its pre-processing capability.
  33535. 25:22:14So, we use the uh minmax scaler and the
  33536. 25:22:18standard scaler from there from the
  33537. 25:22:20pre-processing module in scikitlearn.
  33538. 25:22:23but it has um many different models
  33539. 25:22:27built into it that we can use to help uh
  33540. 25:22:30do our training and predictions. Um so
  33541. 25:22:34it's a incredibly useful machine
  33542. 25:22:36learning library. It is the industry
  33543. 25:22:38standard machine learning library. Um if
  33544. 25:22:42you're going to do anything in machine
  33545. 25:22:43learning, it would be expected that you
  33546. 25:22:46know how to use scikitlearn.
  33547. 25:22:48Now what's really lucky about that is
  33548. 25:22:50that scikitlearn is a really easy
  33549. 25:22:53package to get used to. Nearly
  33550. 25:22:55everything we do in scikitlearn will
  33551. 25:22:57mostly follow the same pattern and so um
  33552. 25:23:00the code will be extremely simple. They
  33553. 25:23:02did a great job with that package of
  33554. 25:23:04making things really user friendly,
  33555. 25:23:06really simple. Um it's a really
  33556. 25:23:09fantastic package and we're going to get
  33557. 25:23:10a lot of practice with it uh as we go
  33558. 25:23:13along. Every model we build will
  33559. 25:23:14essentially be from scikitlearn
  33560. 25:23:17and not only like the models but um
  33561. 25:23:20doing the training doing the predictions
  33562. 25:23:22and then doing the evaluation will all
  33563. 25:23:24come from different uh scikitlearn u
  33564. 25:23:27modules. So that'll be really nice and
  33565. 25:23:30we'll get um good exposure to that
  33566. 25:23:33package throughout the course. So if
  33567. 25:23:35anything will come away from this course
  33568. 25:23:38as um psychit learn uh uh experts
  33569. 25:23:42that'll be very nice. So this is this
  33570. 25:23:45will be the new one for us psychitlearn
  33571. 25:23:47but we'll get a lot of practice with it.
  33572. 25:23:52Okay.
  33573. 25:23:56All right. So just to recap that lesson
  33574. 25:23:59before we move on to lesson three. Um we
  33575. 25:24:01talked about machine learning as
  33576. 25:24:02learning from data. um which is included
  33577. 25:24:06underneath the AI umbrella. But deep
  33578. 25:24:08learning is also included under machine
  33579. 25:24:10learning because it's still learning
  33580. 25:24:11from data but it's learning using neural
  33581. 25:24:14networks.
  33582. 25:24:15Um we talked about the four different
  33583. 25:24:17types of machine learning. We had
  33584. 25:24:18supervised, unsupervised,
  33585. 25:24:21semi-supervised and reinforcement. So
  33586. 25:24:24those are the the different types of
  33587. 25:24:25machine learning that are out there. Um
  33588. 25:24:28and then we talked about some of the pi
  33589. 25:24:30python packages uh that we will use the
  33590. 25:24:33main one being scikitlearn and of course
  33591. 25:24:35we'll use our older like pandas to
  33592. 25:24:37manipulate our data and get it uh pass
  33593. 25:24:39it into our model training etc.
  33594. 25:24:42But scikitlearn will be uh our go-to for
  33595. 25:24:46anything machine learning.
  33596. 25:24:50All right. So, some questions for you
  33597. 25:24:52guys, some checks.
  33598. 25:24:55So, let me know in the chat. What do you
  33599. 25:24:56guys think? Uh, which of the following
  33600. 25:24:59best describes machine learning?
  33601. 25:25:07Which choice do you think makes the best
  33602. 25:25:09is the best for this?
  33603. 25:25:52Very good. Very good. I see I see a lot
  33604. 25:25:54of choices for A and A would be the
  33605. 25:25:56correct choice. So machine learning is
  33606. 25:25:59definitely um a a subset of AI. that's
  33607. 25:26:04underneath that AI umbrella, but of
  33608. 25:26:05course we're learning from experience
  33609. 25:26:07and of course that experience is
  33610. 25:26:09recorded in the data um without being
  33611. 25:26:12explicitly programmed. Uh so it's the
  33612. 25:26:14exact opposite of BNC. We're definitely
  33613. 25:26:16not learning from rules and it's
  33614. 25:26:18definitely not just used for image and
  33615. 25:26:21speech recognition. It can be used for
  33616. 25:26:23many other things beyond those. So yeah,
  33617. 25:26:26A is the best choice there.
  33618. 25:26:29What do we say here?
  33619. 25:26:33Okay. What do you guys think about this?
  33620. 25:26:35Which example illustrates the use of
  33621. 25:26:37machine learning to enhance customer
  33622. 25:26:38experience in an ecommerce company?
  33623. 25:26:53In other words, what would be some what
  33624. 25:26:54would be some uh typical use cases of
  33625. 25:26:57machine learning?
  33626. 25:27:25Good. So I think uh C is going to be the
  33627. 25:27:28best answer here. Definitely C. So it's
  33628. 25:27:31using machine learning to do uh fraud
  33629. 25:27:34transactions. So so that would be a
  33630. 25:27:36prediction probably a supervised
  33631. 25:27:38learning right if if this is fraud or
  33632. 25:27:40not fraud. Um and then maybe some
  33633. 25:27:43customer behavior uh that might be
  33634. 25:27:46unsupervised. So maybe grouping together
  33635. 25:27:48customers uh clustering them based on
  33636. 25:27:50their data like their shopping behavior
  33637. 25:27:53and characteristics. Um that that might
  33638. 25:27:57be unsupervised but either way it's
  33639. 25:27:58machine learning.
  33640. 25:28:01Okay.
  33641. 25:28:06Okay. Final one. What distinguishes deep
  33642. 25:28:09learning from machine learning and
  33643. 25:28:11artificial intelligence? So what's
  33644. 25:28:12unique about deep learning?
  33645. 25:28:43Oh, very good. Yep. So, deep learning
  33646. 25:28:45uses neural networks as so you guys are
  33647. 25:28:49right on top of that. Neural deep
  33648. 25:28:50learning uses neural nets. That's what
  33649. 25:28:52makes it unique. So, machine learning
  33650. 25:28:55would be part A. Machine learning is
  33651. 25:28:57focused on learning from data.
  33652. 25:28:58underneath of that is learning from data
  33653. 25:29:01using neural networks which is what uh
  33654. 25:29:03deep learning is.
  33655. 25:29:07Very good.
  33656. 25:29:10All right, let's go to lesson three.
  33657. 25:29:14And lesson three has two notebooks.
  33658. 25:29:16We're going to be starting with 3.1.
  33659. 25:29:20So, you'll want to open up that
  33660. 25:29:21notebook. I'm going to go over to it
  33661. 25:29:23now. Give you a moment to open that up.
  33662. 25:29:33So, we're going to open the 3.1
  33663. 25:29:34notebook. Um, there's two of them. We'll
  33664. 25:29:37see how far if we can get into the
  33665. 25:29:39second one today. Probably will.
  33666. 25:29:42Um, but we're going to do the uh we're
  33667. 25:29:44going to start with 3.1 notebook. Do you
  33668. 25:29:46guys have this notebook? Should be in
  33669. 25:29:48your materials for for this course.
  33670. 25:29:53Let me give you a moment to open that
  33671. 25:29:54one.
  33672. 25:30:10Do you guys have it?
  33673. 25:30:27All right. So, we're going to start by
  33674. 25:30:30talking about uh supervised learning
  33675. 25:30:34um in our machine learning journey. So
  33676. 25:30:36remember, we're going to talk about uh
  33677. 25:30:38supervised and unsupervised after we do
  33678. 25:30:40supervised. Um and there's going to be a
  33679. 25:30:43lot to cover with supervised mainly
  33680. 25:30:45because um there are uh two different
  33681. 25:30:49types of problems we can tackle uh which
  33682. 25:30:52will be uh we'll talk about in a moment
  33683. 25:30:54predicting different kinds of values. Um
  33684. 25:30:57but let's talk about the kind of what
  33685. 25:30:59we're hoping to learn here which is um
  33686. 25:31:01talk about the different kinds of
  33687. 25:31:02problems that we'll study which are
  33688. 25:31:04these these categories of supervised
  33689. 25:31:06learning. Um those two categories are
  33690. 25:31:08going to be called classification and
  33691. 25:31:09regression. We'll talk about those and
  33692. 25:31:11their differences and then talk about
  33693. 25:31:13some applications and some uh example
  33694. 25:31:16algorithms
  33695. 25:31:18and that's just within this notebook. Um
  33696. 25:31:203.2 two we'll get into uh regression in
  33697. 25:31:24particular
  33698. 25:31:25um which will be uh very very
  33699. 25:31:28interesting. Okay. So that'll be our
  33700. 25:31:30first models that we'll build will be
  33701. 25:31:32over there in 3.2.
  33702. 25:31:36Okay. So if you guys remember um
  33703. 25:31:38supervised learning is where we learn
  33704. 25:31:40from labeled data. So we have input and
  33705. 25:31:43outputs in our in our data set. Um and
  33706. 25:31:47you so you train a model on this data
  33707. 25:31:50that includes input features and
  33708. 25:31:53corresponding outputs that are that are
  33709. 25:31:57the labels. Right? So um the goal is to
  33710. 25:32:01learn a relationship between the input
  33711. 25:32:04and the output. Of course that's what
  33712. 25:32:05any model is trying to do. Um, and what
  33713. 25:32:09this allows us to do is then take that
  33714. 25:32:12model and use it to make predictions on
  33715. 25:32:14never-beforeseen
  33716. 25:32:16uh data. Right? So then we have a
  33717. 25:32:19predictive model out of that that we can
  33718. 25:32:21use um going forward on new examples.
  33719. 25:32:25Um so
  33720. 25:32:27remember we will have in our data a
  33721. 25:32:30bunch of features which are columns and
  33722. 25:32:33then generally one of those columns will
  33723. 25:32:34be the label that we're trying to
  33724. 25:32:36predict.
  33725. 25:32:38And our model is going to try to learn
  33726. 25:32:40some type of relationship between those
  33727. 25:32:42inputs and the output label. So the
  33728. 25:32:46output label could be like fraud not
  33729. 25:32:47fraud, cancer not cancer, uh a price, a
  33730. 25:32:51temperature, those kind of things.
  33731. 25:32:55So let's talk about that. inside of um
  33732. 25:32:58supervised learning there are two
  33733. 25:33:00different types of learning that we can
  33734. 25:33:03do and they're really based on the label
  33735. 25:33:06or sometimes that label is known as the
  33736. 25:33:09target that we're trying to predict. Um
  33737. 25:33:12and depending on that type we get these
  33738. 25:33:15two different categories of learning or
  33739. 25:33:16two different types of learning. One is
  33740. 25:33:19known as regression. So that's generally
  33741. 25:33:22when we are predicting something that is
  33742. 25:33:25continuous or something that is a
  33743. 25:33:27numerical.
  33744. 25:33:30So numerical
  33745. 25:33:32numerical value. So think of price,
  33746. 25:33:35think of temperature, think of revenue.
  33747. 25:33:37We're trying to predict something like
  33748. 25:33:39that. Um versus something that is
  33749. 25:33:42categorical. So that the predicting
  33750. 25:33:45something categorical would be like
  33751. 25:33:46fraud, not fraud, spam, not spam. um
  33752. 25:33:49those are discrete categories and the
  33753. 25:33:53problem of predicting categories is is
  33754. 25:33:55known as classification because we're
  33755. 25:33:59trying to classify examples as belonging
  33756. 25:34:02to one category or another.
  33757. 25:34:06So we have these two main types of
  33758. 25:34:09supervised learning problems. we have
  33759. 25:34:11regression and we have classification
  33760. 25:34:13and they're going to be handled slightly
  33761. 25:34:16differently
  33762. 25:34:17um for many reasons that we're going to
  33763. 25:34:20uncover. Um one of the primary reasons
  33764. 25:34:23is that of course we're predicting
  33765. 25:34:26something that's continuous in the
  33766. 25:34:27regression case versus something
  33767. 25:34:28discrete. So the models have to be
  33768. 25:34:31slightly different to account for that.
  33769. 25:34:33Um but then a step beyond that is the
  33770. 25:34:37evaluation has to be different too. Um I
  33771. 25:34:40kind of alluded to this last week, but
  33772. 25:34:42when you're predicting a regression,
  33773. 25:34:43it's very very difficult to to get the
  33774. 25:34:46exact numerical answer. So um generally
  33775. 25:34:51we don't care about that. Um generally
  33776. 25:34:56we don't care about getting exactly uh
  33777. 25:34:59we don't care about getting it exactly
  33778. 25:35:00right.
  33779. 25:35:02um we just care about getting it um
  33780. 25:35:05we're just we care about getting it
  33781. 25:35:07nearby, getting it close enough. Um
  33782. 25:35:10whereas classification, we do care about
  33783. 25:35:13getting exactly right because it's a
  33784. 25:35:14discrete category. So we're going to be
  33785. 25:35:17able to evaluate that a little bit
  33786. 25:35:18differently to say did we get the answer
  33787. 25:35:20right or wrong. Regression is going to
  33788. 25:35:22be did we get close? Um because it's we
  33789. 25:35:25assume it's going to be nearly
  33790. 25:35:26impossible to predict a a continuous
  33791. 25:35:29number. Um, that's very hard to do.
  33792. 25:35:34Okay.
  33793. 25:35:36So, any questions on
  33794. 25:35:38uh that?
  33795. 25:35:44Any questions on those two differences?
  33796. 25:35:46Let me give you some examples. Maybe
  33797. 25:35:47it'll it'll help too.
  33798. 25:35:51So, again, the classification is going
  33799. 25:35:52to be predicting uh something that's
  33800. 25:35:55categorical. regression is going to be
  33801. 25:35:57predicting something that is continuous.
  33802. 25:36:04So think about trying to predict the
  33803. 25:36:06price of a house based on those other
  33804. 25:36:08features we talked about before like
  33805. 25:36:09square footage, bedrooms, bathrooms, all
  33806. 25:36:12those things we predict the price. That
  33807. 25:36:14would be a regression problem because
  33808. 25:36:16the price is a continuous value.
  33809. 25:36:19Let's take a look at an example here.
  33810. 25:36:22Um, imagine we were trying to uh predict
  33811. 25:36:26the temperature tomorrow. That's going
  33812. 25:36:29to be a regression problem, a a
  33813. 25:36:31supervised learning kind of regression
  33814. 25:36:33problem because we're trying to predict
  33815. 25:36:35a numerical temperature.
  33816. 25:36:39Okay? And versus a category like a
  33817. 25:36:43discrete category would be this would be
  33818. 25:36:45a classification. So this is a
  33819. 25:36:47regression on the left. This is a
  33820. 25:36:49classification
  33821. 25:36:52on the right. Classification
  33822. 25:36:56um because we are um predicting one of
  33823. 25:37:01two categories. Is it just hot or cold?
  33824. 25:37:03Now, we're not saying exactly where that
  33825. 25:37:05threshold is on what's hot or cold. That
  33826. 25:37:08would be a decision on on what we want
  33827. 25:37:10to what our discrete categories actually
  33828. 25:37:12mean.
  33829. 25:37:13But, um we only have two choices, hot or
  33830. 25:37:17cold.
  33831. 25:37:18versus predicting the entire temperature
  33832. 25:37:21which would be um a numerical prediction
  33833. 25:37:24of some exact number. Right? So that'd
  33834. 25:37:28be a regression and then on the right
  33835. 25:37:30would be a classification. Um now again
  33836. 25:37:34why is this so different? You can see
  33837. 25:37:35the types of predictions we're making
  33838. 25:37:37are completely different. One's a
  33839. 25:37:38number, one's a category. But again with
  33840. 25:37:40evaluation it's like if the if the true
  33841. 25:37:44answer in our labels was 84
  33842. 25:37:48and we predicted 83 that's a pretty good
  33843. 25:37:51result. That's still pretty close.
  33844. 25:37:53That's pretty close to this. So from an
  33845. 25:37:55evaluation perspective that's pretty
  33846. 25:37:57good. Um whereas like if I predicted
  33847. 25:38:00cold and it's actually hot that's that's
  33848. 25:38:02a wrong answer. So they're evaluated
  33849. 25:38:05slightly different.
  33850. 25:38:08Um, and that's something we're going to
  33851. 25:38:10see as we talk about evaluation of our
  33852. 25:38:13models once we build them is depending
  33853. 25:38:16on if it's classification regression,
  33854. 25:38:17there's going to be different ways of
  33855. 25:38:18evaluating them.
  33856. 25:38:22You can kind of see why it's very
  33857. 25:38:23difficult to say, okay, we got exactly
  33858. 25:38:2684 when it could be any number. Our
  33859. 25:38:30model is going to be predicting a
  33860. 25:38:31number. That's really hard to pin down
  33861. 25:38:34an exact floatingoint number. So, the
  33862. 25:38:36best we can do is kind of say, how close
  33863. 25:38:39did I get? Like, this would be a worse
  33864. 25:38:40answer. If I got something all the way
  33865. 25:38:42down here, that's a really long distance
  33866. 25:38:44to here. That's bad. That's a bad
  33867. 25:38:47prediction. But if I get something
  33868. 25:38:48really close, that's better, right?
  33869. 25:38:51That's a decent prediction because it's
  33870. 25:38:53pretty close,
  33871. 25:38:55right?
  33872. 25:38:57Of course, being perfect would be
  33873. 25:38:58getting exactly right, but that would be
  33874. 25:39:00nearly impossible to do.
  33875. 25:39:13Okay.
  33876. 25:39:18All right. Any questions on this?
  33877. 25:39:21Does it make sense on regression versus
  33878. 25:39:23classification? We're going to use those
  33879. 25:39:24words quite a bit as we go along. So
  33880. 25:39:27regression predicting that continuous
  33881. 25:39:29value classification predicting a
  33882. 25:39:31category
  33883. 25:39:34and they're going to be um different
  33884. 25:39:36models that do that
  33885. 25:39:41different models being used for
  33886. 25:39:42regression versus different models being
  33887. 25:39:44used for classification.
  33888. 25:39:52All right, let's talk about supervised
  33889. 25:39:55learning. uh applications here. So just
  33890. 25:39:58to name a few, we have HR operations.
  33891. 25:40:02Imagine your recruiter tasked with
  33892. 25:40:03finding the best candidates. Um so
  33893. 25:40:06supervised learning can help by um
  33894. 25:40:08rejecting or accepting candidates. Now
  33895. 25:40:10this is something that happens quite a
  33896. 25:40:12bit even today. Um and that it's kind of
  33897. 25:40:17like uh how recommendations happen like
  33898. 25:40:20this this resume should be um
  33899. 25:40:22recommended this should not um from a
  33900. 25:40:24whole pool of applications. Um so
  33901. 25:40:28there's those kind of use cases of of um
  33902. 25:40:32predicting a category that would be like
  33903. 25:40:34a classification. Should we should we
  33904. 25:40:36accept or reject the the candidate?
  33905. 25:40:39um finance. You see this all the time
  33906. 25:40:41with things like risk and loan
  33907. 25:40:44approvals.
  33908. 25:40:45Um you can uh predict the the the
  33909. 25:40:49category of like if the if the loan if
  33910. 25:40:52we should accept or reject the loan
  33911. 25:40:54application. Um you know that would be a
  33912. 25:40:57classification.
  33913. 25:40:59Um what's interesting about
  33914. 25:41:01classifications by the way so it says
  33915. 25:41:03here like we can predict the likelihood
  33916. 25:41:06of a of a loan being repaid.
  33917. 25:41:09um is a lot of classifications um we we
  33918. 25:41:13say that they predict a category but
  33919. 25:41:16under the hood they can actually predict
  33920. 25:41:18a probability and we turn that
  33921. 25:41:20probability into a category. So um you
  33922. 25:41:25know like we could say what's we could
  33923. 25:41:27say the likelihood of her loan being
  33924. 25:41:29repaid is very low. Let's say it's less
  33925. 25:41:31than 50% probability. Um then we could
  33926. 25:41:35label this as reject,
  33927. 25:41:38right? Right? We could label that as a
  33928. 25:41:39rejection. Um if it's greater than 50%.
  33929. 25:41:43Then we could label this as accept. So
  33930. 25:41:46we can set a threshold there
  33931. 25:41:49and say okay truly we're predicting a
  33932. 25:41:52prob like our model spits out a
  33933. 25:41:54probability but we turn that into a
  33934. 25:41:57category by saying should we accept if
  33935. 25:41:59it's less than 50% we should reject if
  33936. 25:42:02it's greater than we should accept.
  33937. 25:42:05Okay. So that's something we will see
  33938. 25:42:06with some of our classification models
  33939. 25:42:08is that they actually produce a
  33940. 25:42:10probability and we turn that probability
  33941. 25:42:12into a category label
  33942. 25:42:16um by by doing something simple like
  33943. 25:42:18this putting a threshold on it um for
  33944. 25:42:21the for the category.
  33945. 25:42:25So finances is used all over the place.
  33946. 25:42:27Not only just loans like fraud, we
  33947. 25:42:29talked about fraud, not fraud. That
  33948. 25:42:30would be a classification.
  33949. 25:42:32Um predicting sales revenue, that would
  33950. 25:42:36be a regression, right? What is the
  33951. 25:42:38revenue going to be in the next two
  33952. 25:42:40quarters? That's going to be a
  33953. 25:42:42regression problem.
  33954. 25:42:45Uh emails like spam, not spam, that's
  33955. 25:42:48going to be a classification.
  33956. 25:42:50um that's going to operate on the that's
  33957. 25:42:52going to take the text input and predict
  33958. 25:42:54if this email is a spam or a not spam.
  33959. 25:42:58That's going to be a uh supervised
  33960. 25:43:00learning problem, but it's going to be a
  33961. 25:43:02classification problem,
  33962. 25:43:05right? Uh manufacturing supervised
  33963. 25:43:08learning is used to inspect and uh
  33964. 25:43:11quality and classify products in
  33965. 25:43:12different grades. For example, a factory
  33966. 25:43:14might use a model to check for defects.
  33967. 25:43:16So this is actually something that
  33968. 25:43:17happens is you look at images of
  33969. 25:43:19products as they go through the assembly
  33970. 25:43:21line and you can take a look at those
  33971. 25:43:23images and predict if it's a high
  33972. 25:43:25quality, low quality, medium quality. Um
  33973. 25:43:28so they can be this is a classification,
  33974. 25:43:30right? They're going into different
  33975. 25:43:31categories of quality. Um so it's much
  33976. 25:43:35much like a manual kind of intervention
  33977. 25:43:37by some uh QA or quality control uh
  33978. 25:43:42specialist.
  33979. 25:43:44Okay. But that's a classification.
  33980. 25:43:51So in the maritime industry, supervised
  33981. 25:43:53learning can be used to predict current.
  33982. 25:43:55So current level
  33983. 25:43:57um and that can be used to forecast uh
  33984. 25:44:00supply and demand. Um so those would be
  33985. 25:44:03like regression models that are used to
  33986. 25:44:07predict um kind of like temperature but
  33987. 25:44:09in this case like title levels.
  33988. 25:44:14We talked about fraud already, so that's
  33989. 25:44:16there. Um, that would be a
  33990. 25:44:18classification.
  33991. 25:44:23Okay,
  33992. 25:44:26any questions on these uh examples?
  33993. 25:44:30Of course, there's many more. Um
  33994. 25:44:34recommendation is kind of like a
  33995. 25:44:36supervised learning problem uh where you
  33996. 25:44:39are
  33997. 25:44:41taking examples of things that people
  33998. 25:44:43have viewed in the past or or reviewed
  33999. 25:44:46in the past and using that to predict
  34000. 25:44:48what they would want to watch in the
  34001. 25:44:50future. Um so recommendation is
  34002. 25:44:54supervised learning. Um and it's like a
  34003. 25:44:58classification, you know, trying to
  34004. 25:45:00predict um uh certain number of
  34005. 25:45:03categories of of uh shows or movies that
  34006. 25:45:07you would want to watch. Um
  34007. 25:45:11and that's something that we will study
  34008. 25:45:13in the future. Recommend we'll we'll
  34009. 25:45:15have a whole lesson dedicated to
  34010. 25:45:16recommendation as well.
  34011. 25:45:21All right.
  34012. 25:45:23So when it comes down to the uh actual
  34013. 25:45:28models themselves, so there's going to
  34014. 25:45:30be lots of different models that we are
  34015. 25:45:31going to cover. Um and they are um going
  34016. 25:45:36to be different in their purpose and
  34017. 25:45:38kind of their uh what kinds of problems
  34018. 25:45:41they're used for. Um and uh their their
  34019. 25:45:46how they actually train is going to be
  34020. 25:45:48different. Um, but at a high level,
  34021. 25:45:51they're all trying to do the same thing,
  34022. 25:45:53which is learn some sort of relationship
  34023. 25:45:55between the input data and the and the
  34024. 25:45:57label, right? That's really what they're
  34025. 25:45:59trying to do because they're all
  34026. 25:46:00supervised. They're they have those
  34027. 25:46:02labels, trying to build some
  34028. 25:46:04relationship there. Um, they just do it
  34029. 25:46:07differently.
  34030. 25:46:09And what we're going to study is the
  34031. 25:46:11pros and cons of a lot of these models,
  34032. 25:46:13like when would I use one of them, when
  34033. 25:46:14would I use another. Um, so we'll try to
  34034. 25:46:17talk about that as we go along. Um, but
  34035. 25:46:20they're all trying to learn some
  34036. 25:46:23relationship between the input features
  34037. 25:46:25and the output, right? So you have to
  34038. 25:46:27keep that in mind. They're trying to
  34039. 25:46:29model that relationship. They just do it
  34040. 25:46:31in different ways. Okay? So as we go
  34041. 25:46:34along and learn about new models, um, we
  34042. 25:46:37will learn the details. will learn the
  34043. 25:46:38ins and outs um and those pros and cons,
  34044. 25:46:42but they're no matter what, they're all
  34045. 25:46:44trying to uh learn that relationship,
  34046. 25:46:48right? And be able to make predictions
  34047. 25:46:50on new data.
  34048. 25:46:53Okay,
  34049. 25:46:55so here's a list of models that we will
  34050. 25:46:58cover and work on throughout the uh the
  34051. 25:47:02sessions that we have. um we're not
  34052. 25:47:04going to do them all in one one sitting,
  34053. 25:47:07but um the first one that we're going to
  34054. 25:47:09start with and that we'll cover today is
  34055. 25:47:11going to be linear regression.
  34056. 25:47:14So we will cover linear regression and
  34057. 25:47:16then we'll cover the rest of these guys
  34058. 25:47:18mostly in the context of uh
  34059. 25:47:21classification.
  34060. 25:47:23So, um, what's interesting is some of
  34061. 25:47:26these guys can actually be used for both
  34062. 25:47:28regression and classification as long as
  34063. 25:47:30you make, um, certain adjustments to
  34064. 25:47:33them. They have variations that can be
  34065. 25:47:36used to do classification and regression
  34066. 25:47:38is very interesting. Um but we're going
  34067. 25:47:42to start with linear regression today
  34068. 25:47:45and then work our way through the rest
  34069. 25:47:47of these models when we do um we're
  34070. 25:47:49going to do a separate lesson four on
  34071. 25:47:51classification. So these all these guys
  34072. 25:47:53will come from lesson four.
  34073. 25:47:59Um and then uh we will do this guy in
  34074. 25:48:03lesson three in the 3.2 notebook. We'll
  34075. 25:48:06do all about linear regression.
  34076. 25:48:10Yeah, I so logistic regression is a
  34077. 25:48:12classification um which is kind of
  34078. 25:48:15strange that its name is regression but
  34079. 25:48:18it's doing a classification but the the
  34080. 25:48:20reason is that the logistic regression
  34081. 25:48:23um computes a probability. So it does a
  34082. 25:48:26regression to predict a number but that
  34083. 25:48:29number is actually a probability. So it
  34084. 25:48:31it produces a result that's between it
  34085. 25:48:34produces a probability that's between um
  34086. 25:48:38obviously uh zero and one.
  34087. 25:48:43So it uh and then we take that
  34088. 25:48:45probability and we turn it into a
  34089. 25:48:47category
  34090. 25:48:49like a spam not spam fraud not fraud.
  34091. 25:48:52Um but so so logistic regression is kind
  34092. 25:48:54of special. It's sort of like a
  34093. 25:48:57regression but it's predicting a very
  34094. 25:48:58specific type of value which is a
  34095. 25:49:00probability. So for for that reason it's
  34096. 25:49:03a classification uh algorithm primarily.
  34097. 25:49:12So we'll study that one in lesson four.
  34098. 25:49:15Uh but yeah, that's that's why it's
  34099. 25:49:17under that kind of umbrella of
  34100. 25:49:19classification is because it's it's
  34101. 25:49:21producing a probability as its main
  34102. 25:49:22output which we can then turn into a
  34103. 25:49:26category as long as we interpret that
  34104. 25:49:28probability as um in the right way uh
  34105. 25:49:32like the probability of spam,
  34106. 25:49:34probability of not spam.
  34107. 25:49:40Okay.
  34108. 25:49:44Okay. So, let me focus on um
  34109. 25:49:47let me focus on linear regression. I'm
  34110. 25:49:49not going to go through all of these
  34111. 25:49:50other use cases because we haven't
  34112. 25:49:52learned these models yet. Um so, I don't
  34113. 25:49:56think they're good. Uh I don't think
  34114. 25:49:59it's good to read about them yet until
  34115. 25:50:01we've covered them. So, once we cover
  34116. 25:50:04them in lesson four, I'll come back and
  34117. 25:50:06describe these examples to you guys and
  34118. 25:50:08we'll see why it makes sense. But I
  34119. 25:50:10think for a linear regression um which
  34120. 25:50:12is what we'll cover next, let me talk
  34121. 25:50:14about that example. So a prototypical
  34122. 25:50:16example would be like predicting the
  34123. 25:50:18house prices that we've seen in that
  34124. 25:50:20house price data set.
  34125. 25:50:22So um if we wanted to uh if we wanted to
  34126. 25:50:27predict um if we wanted to estimate the
  34127. 25:50:30market value of a house so the price
  34128. 25:50:35um we could do that by using the
  34129. 25:50:38features such as number of bedrooms,
  34130. 25:50:40square footage, location, age of the
  34131. 25:50:42property. Um and you know then when a
  34132. 25:50:46new when a new house comes on the market
  34133. 25:50:48we could estimate what the price should
  34134. 25:50:50be based on those features. So linear
  34135. 25:50:54regression is a good one to predict the
  34136. 25:50:56price like a housing price. Um and we'll
  34137. 25:50:59actually practice that in the next uh
  34138. 25:51:02notebook.
  34139. 25:51:05So we'll we'll uh and then all these
  34140. 25:51:08other now there's descriptions of these
  34141. 25:51:10other models but again we haven't
  34142. 25:51:11covered these guys yet. So I don't want
  34143. 25:51:13to really go through those until we get
  34144. 25:51:15to those models. So we get to those I'll
  34145. 25:51:17come back and mention the example.
  34146. 25:51:20Uh can K andN be used for clustering?
  34147. 25:51:23No. So um the clustering model is going
  34148. 25:51:27to be different. It's going to be uh K
  34149. 25:51:29means
  34150. 25:51:31K means that's the primary clustering
  34151. 25:51:33model. Not K nearest neighbors. K
  34152. 25:51:36nearest neighbors is used for uh it can
  34153. 25:51:39be used for regression. It can be used
  34154. 25:51:40for classification.
  34155. 25:51:44So we'll we'll talk about K andN which
  34156. 25:51:46is the K nearest neighbors in lesson
  34157. 25:51:48four.
  34158. 25:51:51It sounds really similar. Yeah, it
  34159. 25:51:54sounds really similar but K means is a
  34160. 25:51:56clustering algorithm that's that's
  34161. 25:51:57slightly different
  34162. 25:51:59different uh there's no labels used at
  34163. 25:52:02all. This K nearest neighbors is a is a
  34164. 25:52:06supervised learning algorithm. It uses
  34165. 25:52:08uh labels.
  34166. 25:52:18Good. Any any other questions so far?
  34167. 25:52:36Okay.
  34168. 25:52:38So that being said, let's move on to the
  34169. 25:52:413.2 notebook.
  34170. 25:52:44Let's move on to that which will be our
  34171. 25:52:47um first discussion around uh
  34172. 25:52:51regression. So going into supervised
  34173. 25:52:54learning and regression. Give you guys a
  34174. 25:52:56moment to pull up this notebook.
  34175. 25:52:59But yeah, you want to pull up the 3.2.
  34176. 25:53:01We'll do this one next. So we'll focus
  34177. 25:53:04in. And so our plan is to do regression
  34178. 25:53:06first and then we'll talk about
  34179. 25:53:08classification in lesson four
  34180. 25:53:15which we will cover all those other
  34181. 25:53:17models which you you could use for
  34182. 25:53:20classification uh on that list but then
  34183. 25:53:22we're going to talk about linear
  34184. 25:53:23regression uh first.
  34185. 25:53:33All right. So, we have a a big agenda.
  34186. 25:53:36This is a big notebook um to go through
  34187. 25:53:39a lot of material here surrounding
  34188. 25:53:42regression. So, we're we're going to
  34189. 25:53:44start with linear regression and see um
  34190. 25:53:47how we actually perform it, what that
  34191. 25:53:50model is doing. Um which we've kind of
  34192. 25:53:53seen the idea of it a little bit
  34193. 25:53:55already, so it should be somewhat
  34194. 25:53:56familiar. Um and then we'll talk about
  34195. 25:53:59how to adapt that linear regression idea
  34196. 25:54:02to um nonlinear what's called nonlinear
  34197. 25:54:06regression which is going to be using
  34198. 25:54:08like polomial
  34199. 25:54:10uh features. We'll talk about how to do
  34200. 25:54:11that. Um and then a big big big topic
  34201. 25:54:15for us is going to be evaluating the
  34202. 25:54:17model. So it'll be it'll be quite easy
  34203. 25:54:19to actually build it. building the model
  34204. 25:54:22will be really easy but evaluating and
  34205. 25:54:25interpreting that will be uh a lot of
  34206. 25:54:30interesting work there um because we
  34207. 25:54:33want to know what the performance of
  34208. 25:54:34that model is once we have it built
  34209. 25:54:36right we want to know how good of a
  34210. 25:54:38model is it is it worth using or do we
  34211. 25:54:41need to retrain it or get new data or
  34212. 25:54:43change the model up to talk about that
  34213. 25:54:46um how do you determine what to do based
  34214. 25:54:48on that performance
  34215. 25:54:50um and then we'll talk about here um a
  34216. 25:54:53couple things. We may not get to this
  34217. 25:54:55today, but regularization
  34218. 25:54:57which is used to boost the performance
  34219. 25:54:59uh in certain situations um whenever the
  34220. 25:55:03model is kind of uh performing um poorly
  34221. 25:55:07against test data even though it
  34222. 25:55:08performs pretty well on training data.
  34223. 25:55:11In that scenario, you can use offshoots
  34224. 25:55:14of linear regression that do some uh
  34225. 25:55:16what's called regularization. We'll talk
  34226. 25:55:18about that.
  34227. 25:55:20Um and then we'll talk about
  34228. 25:55:21hyperparameter tuning uh generally as a
  34229. 25:55:24strategy which is something you
  34230. 25:55:26generally do want to do when you're
  34231. 25:55:28training machine learning models. Um so
  34232. 25:55:31again these two we may not get to today
  34233. 25:55:35but um quite a quite a lot to get to be
  34234. 25:55:38prior to that mainly centered around
  34235. 25:55:40evaluation and building linear
  34236. 25:55:43regression.
  34237. 25:55:45Okay. So pretty cool. we'll get to our
  34238. 25:55:47first kind of model here. This linear
  34239. 25:55:49regression
  34240. 25:55:51to start with.
  34241. 25:55:55Okay,
  34242. 25:55:56so let's start with uh linear regression
  34243. 25:55:59here. Um, and really what linear
  34244. 25:56:04regression is attempting to do and I
  34245. 25:56:08want to show you this in this picture is
  34246. 25:56:11draw this line sometimes what is known
  34247. 25:56:14as the line of best fit. So this is our
  34248. 25:56:17model that kind of goes through the data
  34249. 25:56:20and it's generally a good predictor
  34250. 25:56:25um because if you give me um features uh
  34251. 25:56:30if you give me new features and let's
  34252. 25:56:33say they are let's say you give me a
  34253. 25:56:35feature that's right here.
  34254. 25:56:38So you say, okay, I have a feature
  34255. 25:56:39that's this value on the x- axis. Then I
  34256. 25:56:43know all I have to do is plug that into
  34257. 25:56:45my line equation, and I will generate a
  34258. 25:56:49a value that's like right here.
  34259. 25:56:52Okay, that's pretty that's on that line
  34260. 25:56:54at that input. And that's going to be my
  34261. 25:56:57prediction for what the output variable
  34262. 25:56:59should be. It's just going to be
  34263. 25:57:00something on that line. And what you can
  34264. 25:57:04see is this line is a decent estimate
  34265. 25:57:07for this data because it slices through
  34266. 25:57:10this pretty evenly. So it's a good guess
  34267. 25:57:13as to what the output should be given
  34268. 25:57:16any one of these inputs. It's a it's a
  34269. 25:57:18good estimator this line. And so our
  34270. 25:57:21goal building a linear regression is to
  34271. 25:57:23kind of build the equation of this line.
  34272. 25:57:26So we want this equation.
  34273. 25:57:30Equation of this line
  34274. 25:57:34is going to be our model.
  34275. 25:57:41Yes, it's going to look just like that.
  34276. 25:57:43MX plus B or yeah, MX plus C. It's going
  34277. 25:57:45to look exactly like that. uh except
  34278. 25:57:48that it's going to be more than just MX
  34279. 25:57:52because we have um generally more than
  34280. 25:57:56one feature. So you think of X as a
  34281. 25:57:57feature um it will be more than just MX.
  34282. 25:58:00It will generally be like uh it'll
  34283. 25:58:03generally look like this
  34284. 25:58:10and then plus maybe some bias here plus
  34285. 25:58:14uh an intercept. Yeah, it'll generally
  34286. 25:58:16look like that. So, yeah, you're exactly
  34287. 25:58:18right. MX plusb is the right idea.
  34288. 25:58:21Exactly right.
  34289. 25:58:24It'll generally look like that.
  34290. 25:58:28Nonlinear. It can be adapted to
  34291. 25:58:30nonlinear. Yeah. If we transform, we're
  34292. 25:58:33going to talk about that. If we
  34293. 25:58:34transform all of our features in a
  34294. 25:58:36nonlinear way, um we can apply linear
  34295. 25:58:39regression to it. Yes. And and that
  34296. 25:58:42would be a nonlinear regression. So yes,
  34297. 25:58:45we can do nonlinear things too.
  34298. 25:58:48We'll talk about that.
  34299. 25:58:58Okay. So linear regression again is the
  34300. 25:59:01art or science I should say not really
  34301. 25:59:04art but it is an exact science of
  34302. 25:59:07finding the equation of this line that
  34303. 25:59:09fits through this data. Um now why one
  34304. 25:59:12thing you should be thinking about is
  34305. 25:59:14why is this line a good predictor and
  34306. 25:59:18the argument is that if you take a look
  34307. 25:59:20at this distance from these blue points
  34308. 25:59:22so let's say these blue points are our
  34309. 25:59:24actual data points this line is going to
  34310. 25:59:28be found such that it minimizes this
  34311. 25:59:32distance
  34312. 25:59:34from the points to actually I should
  34313. 25:59:37draw it this way from the points to the
  34314. 25:59:39line.
  34315. 25:59:41So, we want this distance to be um
  34316. 25:59:44actually I should draw it that way, this
  34317. 25:59:47way. We want this distance to be kind of
  34318. 25:59:50at a minimum. So, it would be bad to
  34319. 25:59:52draw a line all the way out here because
  34320. 25:59:54then that's a lot of distance, right?
  34321. 25:59:56So, and that would be a lot of error um
  34322. 25:59:59contributed from not being able to
  34323. 26:00:01predict those points in our data set
  34324. 26:00:02very well. Um which is our training
  34325. 26:00:05data. That's why we have labels, right?
  34326. 26:00:07that that guide us in building this
  34327. 26:00:09line. Um so our goal is to build that
  34328. 26:00:12line especially so that this error or
  34329. 26:00:16this distance can be as minimum as
  34330. 26:00:20possible. Right? Which are all these
  34331. 26:00:22distances from these points to the line.
  34332. 26:00:25We want those to be as minimum as
  34333. 26:00:28possible. So our goal is to find this
  34334. 26:00:30equation.
  34335. 26:00:32So we're going to build a model that's
  34336. 26:00:34going to find this equation.
  34337. 26:00:39of the line
  34338. 26:00:41um such that our error
  34339. 26:00:46is minimal.
  34340. 26:00:50And what is the error? The error is the
  34341. 26:00:52distance
  34342. 26:00:56of our data points
  34343. 26:01:02to
  34344. 26:01:06to the line that we build. So
  34345. 26:01:08essentially what we'll do in order to
  34346. 26:01:10train this will be to adjust the
  34347. 26:01:13parameters or the or in that like I
  34348. 26:01:15think it's really good you brought up
  34349. 26:01:16the MX plus C. Basically the M and the C
  34350. 26:01:19will adjust. So we adjust those
  34351. 26:01:22accordingly to make this distance as
  34352. 26:01:25small as possible.
  34353. 26:01:28Okay to minimize that distance as much
  34354. 26:01:30as possible.
  34355. 26:01:39Okay.
  34356. 26:01:41So, um where is regression used? We've
  34357. 26:01:45already seen some examples. Here's some
  34358. 26:01:47more uh advertising like predicting
  34359. 26:01:50sales, predicting um oil and uh oil
  34360. 26:01:55production and demand. Those are like
  34361. 26:01:57forecast those are regression problems.
  34362. 26:02:00Um retail like demand forecasting for
  34363. 26:02:02inventory. Um healthc care predicting um
  34364. 26:02:07uh the levels of certain um uh blood
  34365. 26:02:12markers or you know something like that.
  34366. 26:02:14Um real estate predicting prices based
  34367. 26:02:18on those uh talked about like square
  34368. 26:02:20footage, bedrooms, bathrooms, those
  34369. 26:02:22things. So regression is used again
  34370. 26:02:24whenever we want to predict a number a
  34371. 26:02:26numerical output um that's a regression
  34372. 26:02:29problem.
  34373. 26:02:35So this kind of regression we're talking
  34374. 26:02:37about here is generally
  34375. 26:02:41um known as uh a when that equation is
  34376. 26:02:45linear that is known as a linear
  34377. 26:02:48regression. And so go back to that
  34378. 26:02:49picture when we have a when that
  34379. 26:02:52equation of the line that we find is a
  34380. 26:02:55linear equation meaning that it is
  34381. 26:02:58exactly the form I've been telling you.
  34382. 26:03:00So it's it's something like um weight
  34383. 26:03:04time feature
  34384. 26:03:06plus weight time feature
  34385. 26:03:09plus weight time feature
  34386. 26:03:13and then maybe some intercept um term
  34387. 26:03:17like some some bias term there. Um this
  34388. 26:03:21is a linear equation because all of the
  34389. 26:03:24features are to the single power. So
  34390. 26:03:27it's a linear power and this is a linear
  34391. 26:03:29combination of features with with those
  34392. 26:03:32different weights. So this is a linear
  34393. 26:03:35model
  34394. 26:03:36because it is uh it's what in math we
  34395. 26:03:40would call this a linear equation right
  34396. 26:03:43everything is to the first power. It
  34397. 26:03:45resembles mx plus b. It is a linear
  34398. 26:03:48equation or a linear model. Um so when
  34399. 26:03:53we talk about linear regression that is
  34400. 26:03:56a regression model so we're predicting
  34401. 26:03:58some continuous target that assumes we
  34402. 26:04:01are model our model is formed from this
  34403. 26:04:05kind of equation a linear equation.
  34404. 26:04:08So this is going to be our our model for
  34405. 26:04:12a linear
  34406. 26:04:14uh regression.
  34407. 26:04:17Okay.
  34408. 26:04:20And so when you when you train a linear
  34409. 26:04:23regression, your goal is to learn these
  34410. 26:04:25weights so that you can plug in um you
  34411. 26:04:30can plug in any one of your uh input
  34412. 26:04:32features and you um can generate a
  34413. 26:04:36prediction. You can which is going to be
  34414. 26:04:38something on that line, right? It's
  34415. 26:04:40going to be a value that's sitting here
  34416. 26:04:42on this line.
  34417. 26:04:44We put in all of our features and we end
  34418. 26:04:46up there somewhere on that line.
  34419. 26:04:50This output.
  34420. 26:04:54Okay.
  34421. 26:05:05Okay. Let me pause there. Any questions
  34422. 26:05:07on the linear model here or why it's
  34423. 26:05:12called linear regression?
  34424. 26:05:28Okay. And by the way in these notes um
  34425. 26:05:31this bullet point here where it says it
  34426. 26:05:32uses the least squares criterion to
  34427. 26:05:34estimate the coefficients that is
  34428. 26:05:36exactly what I said earlier with the
  34429. 26:05:38distance. So the distance is based on
  34430. 26:05:40the square
  34431. 26:05:42of this this quantity like how far away
  34432. 26:05:45you are from the line is based on this
  34433. 26:05:47square distance here and here and here
  34434. 26:05:51and here. So what we're trying to do is
  34435. 26:05:54find the least distance or least squares
  34436. 26:05:58which is that minimum distance. So
  34437. 26:06:00that's how we find all of these weights
  34438. 26:06:03is from minimize. We basically tune them
  34439. 26:06:06enough using our labels. So here's our
  34440. 26:06:09label which is the y. We basically plug
  34441. 26:06:12in our data and tune those enough to
  34442. 26:06:14minimize the error. It's it's a it's an
  34443. 26:06:16optimization problem,
  34444. 26:06:19right? We we're trying to find the
  34445. 26:06:21minimum of this quantity which is that
  34446. 26:06:26best fit line.
  34447. 26:06:40Okay.
  34448. 26:06:41So we have linear regression
  34449. 26:06:44um and we can do a simple linear
  34450. 26:06:49regression that only has one feature. So
  34451. 26:06:51if it only has one feature that's
  34452. 26:06:53exactly the so if there's only one input
  34453. 26:06:56feature sometimes that is known as um
  34454. 26:06:59simple regression or simple linear
  34455. 26:07:01regression and there's basically there's
  34456. 26:07:03only one feature. So one independent
  34457. 26:07:05variable is the feature.
  34458. 26:07:08There's only one feature. And so this
  34459. 26:07:11equation resembles the
  34460. 26:07:16exact equation that you guys just put in
  34461. 26:07:18there, which is um mx plus b,
  34462. 26:07:22right? It resembles exactly that. Um
  34463. 26:07:25we're just using different symbols for
  34464. 26:07:27those like beta beta 0 and beta 1. But
  34465. 26:07:30um basically exactly that simple line is
  34466. 26:07:34only one feature. So, and that's because
  34467. 26:07:37that line is going to um that line is
  34468. 26:07:42going to be generated uh according to
  34469. 26:07:44that equation. So, here's kind of what
  34470. 26:07:46it looks like.
  34471. 26:07:48This is the best fit line through all of
  34472. 26:07:50these blue dots. This is something we're
  34473. 26:07:52going to be able to build. We're going
  34474. 26:07:54to be able to build that equation um
  34475. 26:07:57pretty easily in scikitlearn.
  34476. 26:08:00So, we'll be able to find that um and it
  34477. 26:08:03won't be too hard. So this line will be
  34478. 26:08:06um y = beta 0 plus beta 1. So some
  34479. 26:08:11weight beta 1 times the only feature we
  34480. 26:08:15have x1.
  34481. 26:08:17Okay. So in this case um we would be
  34482. 26:08:21predicting sales. So sales would be the
  34483. 26:08:24value basically the label that we're
  34484. 26:08:26trying to predict and the feature that
  34485. 26:08:28we're putting in is uh I think it's the
  34486. 26:08:34number of TV expenses. Yep. TV expenses
  34487. 26:08:38which is on the x- axis. So there's one
  34488. 26:08:40feature which is um TV expense.
  34489. 26:08:48So um on this graph this would be this
  34490. 26:08:52would be our model.
  34491. 26:08:54Okay that would be our model. We only
  34492. 26:08:56have one feature and we have um these
  34493. 26:09:00two weights. We have an intercept B 0
  34494. 26:09:03and or beta 0 and then a one weight
  34495. 26:09:06which gets applied to that one feature
  34496. 26:09:08beta 1. And so our model would have
  34497. 26:09:11certain value for beta 0 and a certain
  34498. 26:09:13value for beta 1. That's what get that's
  34499. 26:09:16these guys get learned
  34500. 26:09:21learned during
  34501. 26:09:24model
  34502. 26:09:27training.
  34503. 26:09:33Okay. So those are what get learned
  34504. 26:09:36during our model training and they get
  34505. 26:09:38learned by a a a least what's called a
  34506. 26:09:41lease squares algorithm that is trying
  34507. 26:09:43to minimize that distance. It tries to
  34508. 26:09:45tweak beta 0 beta 1 to minimize this
  34509. 26:09:48distance of this line
  34510. 26:09:51um this line
  34511. 26:09:53to all of these points
  34512. 26:09:58trying to minimize this.
  34513. 26:10:02So imagine taking a line and kind of
  34514. 26:10:04moving it around and turning its its
  34515. 26:10:08slope, its angle um to try to find that
  34516. 26:10:11best fit,
  34517. 26:10:13which reduces that error the most.
  34518. 26:10:16Right? That's kind of what we're doing.
  34519. 26:10:40Uh can I explain? Yeah. So uh sales is
  34520. 26:10:44in dollars and and TV expense
  34521. 26:10:48um
  34522. 26:10:50uh
  34523. 26:10:52TV actually I think it's the other way
  34524. 26:10:54around. I think the sales is actually a
  34525. 26:10:56quantity. So I this is number of sales
  34526. 26:10:59that we have and TV expense is um I
  34527. 26:11:03think I think it's in dollars. So how
  34528. 26:11:05much money how much expense um did we
  34529. 26:11:08put into the into the product and then
  34530. 26:11:11this is how many sales did we have of
  34531. 26:11:13that product.
  34532. 26:11:17So I think it's the other way around
  34533. 26:11:24but what this what this graph is showing
  34534. 26:11:27is the blue points are our actual data
  34535. 26:11:31points. Okay. So so we have a collection
  34536. 26:11:34like we have a data frame that has so
  34537. 26:11:37imagine we had a data frame that has the
  34538. 26:11:39uh true values.
  34539. 26:11:42So it has the um TV expenses.
  34540. 26:11:46Um it has points that are like one. So
  34541. 26:11:50it has points that are like 120 and then
  34542. 26:11:53the sale sales could be like 700
  34543. 26:11:57700 units, let's say. And then it has um
  34544. 26:12:00so this is just our data set, right?
  34545. 26:12:01This would be like in a data frame that
  34546. 26:12:03we have. And then we had ones that were
  34547. 26:12:06um 50 and then this could be um this
  34548. 26:12:10could be 400 let's say and on and on and
  34549. 26:12:13on right so this is our data and this
  34550. 26:12:16data is plotted in the blue so these are
  34551. 26:12:19these blue points here
  34552. 26:12:21right so these are the blue points here
  34553. 26:12:24and the red points are is our model so
  34554. 26:12:27we built a linear regression model um
  34555. 26:12:31where we are putting in some values
  34556. 26:12:33we're putting in some fake x values
  34557. 26:12:36here and generating some predictions
  34558. 26:12:38which is this line,
  34559. 26:12:41this linear uh regression line, right?
  34560. 26:12:44And that line is derived from this data,
  34561. 26:12:48right? It gets learned from this
  34562. 26:12:50supervised uh examples.
  34563. 26:12:55Does that make sense?
  34564. 26:12:57That line is derived from the data. it's
  34565. 26:13:00actually um learned from like the line
  34566. 26:13:03of best fit is learned from that data
  34567. 26:13:10and the actual data is in the blue.
  34568. 26:13:14So you can see we're trying to build
  34569. 26:13:15this such that this distance is kind of
  34570. 26:13:17a minimum
  34571. 26:13:20so it's an optimal fit
  34572. 26:13:24to balance out these distances.
  34573. 26:13:34So it's just plotting. So it's just
  34574. 26:13:35building that relationship between the
  34575. 26:13:37input and output. Like when the when the
  34576. 26:13:39expenses are higher, um we seem to have
  34577. 26:13:42more sales.
  34578. 26:14:01Okay.
  34579. 26:14:10Uh what's perpendicular like the
  34580. 26:14:12distance? This should be this should be
  34581. 26:14:14perpendicular because it's a distance
  34582. 26:14:15here.
  34583. 26:14:19Is that what you mean? Like the distance
  34584. 26:14:20from the real points to the line? Yeah,
  34585. 26:14:22that should be perpendicular
  34586. 26:14:24because it's it's a it's a distance
  34587. 26:14:26formula.
  34588. 26:14:46Okay.
  34589. 26:14:50All right. So more generally now do do
  34590. 26:14:54we usually have one feature? No. So
  34591. 26:14:58generally we expand this to the more
  34592. 26:15:01general case where we have more than one
  34593. 26:15:04feature like what we see in the housing
  34594. 26:15:06data right where we could predict a
  34595. 26:15:08price but we have many different inputs
  34596. 26:15:10like bedrooms, bathrooms, square footage
  34597. 26:15:13etc.
  34598. 26:15:16So more broadly
  34599. 26:15:19instead of simple linear regression we
  34600. 26:15:22have what's known as multiple linear
  34601. 26:15:24linear regression which means we have
  34602. 26:15:26multiple variables or multiple features.
  34603. 26:15:29Um so this is exactly the equation I've
  34604. 26:15:32been talking about. Um so we just extend
  34605. 26:15:36that that one into many features. So
  34606. 26:15:39which is this case and then a intercept
  34607. 26:15:42term which is uh um there as sometimes
  34608. 26:15:47known as the bias. Um
  34609. 26:15:50but this is the intercept term to kind
  34610. 26:15:52of orient the line to start out in the
  34611. 26:15:54right place. Um and uh but this is the
  34612. 26:16:00um this is the equation that we would be
  34613. 26:16:03building the model. This is our model
  34614. 26:16:05essentially, right? This is the equation
  34615. 26:16:06we would be learning.
  34616. 26:16:09Intercept is like a constant. Yeah. So
  34617. 26:16:11if if all of the features were zero, um
  34618. 26:16:14this is what our our data would be. This
  34619. 26:16:16is what our result would be. If
  34620. 26:16:18basically if this was zero, this was
  34621. 26:16:19zero, this was zero, it would reduce to
  34622. 26:16:22this as the prediction. Yeah. It's like
  34623. 26:16:25a constant. Yes.
  34624. 26:16:33So in in geometry, the intercept's
  34625. 26:16:36actually really important because it it
  34626. 26:16:37orients where your line should start. So
  34627. 26:16:39it orients like so so these values are
  34628. 26:16:42kind of like the slope. They orient the
  34629. 26:16:44tilt of it. Like should it be tilted
  34630. 26:16:47like this or should it be more sloped?
  34631. 26:16:49But the intercept orients where it
  34632. 26:16:52should start like vertically like should
  34633. 26:16:54it start all the way up here? Should it
  34634. 26:16:56start more down here?
  34635. 26:16:58Um, that's what the intercept kind of
  34636. 26:17:01tells us.
  34637. 26:17:11Okay, so this is the situation. This is
  34638. 26:17:14going to be our linear regression model
  34639. 26:17:15that we will be building most of the
  34640. 26:17:17time because we will have again these
  34641. 26:17:19are all going to be features.
  34642. 26:17:22So this is some feature the X this is
  34643. 26:17:25some feature this is some feature
  34644. 26:17:29X1 etc. These are all features and what
  34645. 26:17:33gets learned during the training are
  34646. 26:17:35these coefficients. So all of these
  34647. 26:17:37coefficients including the beta 0ero um
  34648. 26:17:41will get learned. So these will get
  34649. 26:17:43learned
  34650. 26:17:45um from our data right they get learned
  34651. 26:17:48they will be trained from our data um in
  34652. 26:17:51order and and how do they get trained
  34653. 26:17:53it's from reducing that distance we try
  34654. 26:17:56to get that line of best fit by tweaking
  34655. 26:17:59those betas enough to uh until we reach
  34656. 26:18:02a minimum distance but there's there's
  34657. 26:18:04an algorithm behind that um that that
  34658. 26:18:07scikitlearn will run for us to find that
  34659. 26:18:10best fit Um, so we don't need to do that
  34660. 26:18:13manually, but that's that's the process
  34661. 26:18:15is basically tweaking those weights to
  34662. 26:18:18end up with that line of best fit. So in
  34663. 26:18:21higher dimensions, instead of a line,
  34664. 26:18:23you get more of what's called a plane
  34665. 26:18:26here. Um, which kind of looks like this.
  34666. 26:18:28So the best fit is actually this plane
  34667. 26:18:32where all um, it kind of dissects all
  34668. 26:18:34these points just like that um, in
  34669. 26:18:37higher dimensions. So this is uh instead
  34670. 26:18:40of a line you get this in in three
  34671. 26:18:42dimensions you get this plane like this
  34672. 26:18:45but it's still it's like a line of best
  34673. 26:18:47it's just a more general line of best
  34674. 26:18:49fit. It's still the same idea. Um we're
  34675. 26:18:52still trying to um come up with the best
  34676. 26:18:55coefficients to minimize that distance
  34677. 26:18:57from our from our points to the line.
  34678. 26:19:01Although in higher dimensions it's no
  34679. 26:19:03longer a line. It's more like a plane
  34680. 26:19:04like this. So you're trying to minimize
  34681. 26:19:06this distance from here down to the
  34682. 26:19:08plane
  34683. 26:19:09here up to the plane
  34684. 26:19:12in higher dimensions. So I want you to
  34685. 26:19:14keep in mind what we're trying to do
  34686. 26:19:17before we go into the code because the
  34687. 26:19:19code's going to make it seem really
  34688. 26:19:20really simple and that's because
  34689. 26:19:22scikitlearn is great and that's what it
  34690. 26:19:24does.
  34691. 26:19:26But we should realize that there's
  34692. 26:19:27something really complex going on which
  34693. 26:19:29is again finding the best value of these
  34694. 26:19:33weights
  34695. 26:19:35that minimizes the distance of this line
  34696. 26:19:39to the data points that we have. So
  34697. 26:19:42there's an algorithm there that will
  34698. 26:19:44keep trying to make adjustments to this
  34699. 26:19:47based on those distances. So it's going
  34700. 26:19:50to use those distances as a guide to
  34701. 26:19:52kind of tweak them to find the one that
  34702. 26:19:56results in the lowest amount of
  34703. 26:19:58distance. So we keep making tweaks, keep
  34704. 26:19:59making tweaks, keep making tweaks and
  34705. 26:20:02eventually we try to find we converge to
  34706. 26:20:05the set of weights that gives us that
  34707. 26:20:06best fitting line. Um and and there's an
  34708. 26:20:10algorithm there that occurs. Now luckily
  34709. 26:20:13that gets abstracted for us a bit behind
  34710. 26:20:16um scikitlearn
  34711. 26:20:18um finding that best fit. So there'll be
  34712. 26:20:21a function that we use in scikitlearn
  34713. 26:20:25when we build the model that will go
  34714. 26:20:26ahead and find the best weights for us
  34715. 26:20:30and that's then we now have our optimal
  34716. 26:20:33model right that then we can just plug
  34717. 26:20:35in different values of these features
  34718. 26:20:38and generate a prediction which is going
  34719. 26:20:41to be this uh result right so so that's
  34720. 26:20:44what we're ultimately trying to do is uh
  34721. 26:20:48train the model which will uh find all
  34722. 26:20:50those optimal weights and then uh we can
  34723. 26:20:54predict with it which would be plugging
  34724. 26:20:56in different feature values to to
  34725. 26:20:58generate a prediction.
  34726. 26:21:02Okay,
  34727. 26:21:03so let's see how that happens. It's
  34728. 26:21:05actually going to be super easy um with
  34729. 26:21:07scikitlearn.
  34730. 26:21:09So uh in this scenario we have um we're
  34731. 26:21:13going to import our pandas because we're
  34732. 26:21:15going to load our data from that. Um, so
  34733. 26:21:19of course we need some data to work
  34734. 26:21:20with. So we're going to load this uh
  34735. 26:21:22CSV.
  34736. 26:21:23Um, I
  34737. 26:21:26uh so I was not actually able to find
  34738. 26:21:30this CSV for this example, but I mean
  34739. 26:21:32that's okay because we'll do some we'll
  34740. 26:21:33do other examples where we'll work with
  34741. 26:21:35the data. If you happen to have it, um,
  34742. 26:21:38great. I didn't see it in in my files.
  34743. 26:21:42So just have to take the word for it
  34744. 26:21:43that these are the this is that TV and
  34745. 26:21:46sales columns here um from this data
  34746. 26:21:50set.
  34747. 26:21:52Okay. Um as an example. So um just to
  34748. 26:21:57see how it's fit um what we're going to
  34749. 26:22:01do and this is going to be a very
  34750. 26:22:04standard process for us for building a
  34751. 26:22:07model. These steps are going to be very
  34752. 26:22:09very standard for us which is going to
  34753. 26:22:11be first of all splitting the features
  34754. 26:22:15away from the label. That's the first
  34755. 26:22:18step that we always will take. So if you
  34756. 26:22:20take a look at this code, it's taking
  34757. 26:22:23all rows but only the first column.
  34758. 26:22:28Okay, so it's extracting all the
  34759. 26:22:30features from the data frame um which
  34760. 26:22:33happen to be which is just the first the
  34761. 26:22:35first column uh which is the TV uh
  34762. 26:22:39column right just that column there and
  34763. 26:22:42our target variable which is our label.
  34764. 26:22:46So our target variable aka the label um
  34765. 26:22:50is the second column, right? It's that
  34766. 26:22:54that sales column.
  34767. 26:22:57Um and so our first step here, let me
  34768. 26:23:01call that out here. First step is to
  34769. 26:23:05always split apart
  34770. 26:23:08features from labels.
  34771. 26:23:12Okay, so we put all those features into
  34772. 26:23:14a data frame called X and we have all of
  34773. 26:23:17our labels into technically a series but
  34774. 26:23:20uh sort of like a data frame, right? Um
  34775. 26:23:24called Y, which is just the um which is
  34776. 26:23:28just the uh uh labels. So that's just
  34777. 26:23:32the TV values. Um now you're going to
  34778. 26:23:35see why we do that. It's because we need
  34779. 26:23:39um our our features and labels split
  34780. 26:23:42apart to put them into the model
  34781. 26:23:44building function. It expects our
  34782. 26:23:48independent variables or our features to
  34783. 26:23:50be separated from our answers or our
  34784. 26:23:54labels that guide the model building.
  34785. 26:23:57That's the first thing you got to do is
  34786. 26:23:58separate those.
  34787. 26:24:01Okay, so this code will separate those
  34788. 26:24:03out into a capital X and a lowercase Y.
  34789. 26:24:06And that's actually pretty industry
  34790. 26:24:08standard notation. Whenever you split
  34791. 26:24:10apart all your features, usually you put
  34792. 26:24:13them into a data frame called capital X
  34793. 26:24:15and then you have a lowercase Y to
  34794. 26:24:18represent your labels. That's actually
  34795. 26:24:20pretty standard.
  34796. 26:24:23So it's pretty standard that um X
  34797. 26:24:26represents
  34798. 26:24:29features
  34799. 26:24:31and
  34800. 26:24:32Y represents labels
  34801. 26:24:37label column
  34802. 26:24:39whatever our label column is in this
  34803. 26:24:41case it is the sales because we're going
  34804. 26:24:44to be predicting sales
  34805. 26:24:49using the TV column the TV quant expense
  34806. 26:24:53quantity.
  34807. 26:25:00Yeah. So what it so the assignment is
  34808. 26:25:03that we are um the assignment is that we
  34809. 26:25:08are
  34810. 26:25:09uh we are um splitting apart our data.
  34811. 26:25:14So that when we first read in the data
  34812. 26:25:17um it is a data frame right that has two
  34813. 26:25:20columns TV and sales.
  34814. 26:25:24Oh perfect thank you Tim. I will I will
  34815. 26:25:28go ahead and so if we look at this data
  34816. 26:25:33it only has those two columns right it
  34817. 26:25:36only has those two columns. Okay. So
  34818. 26:25:39what we're doing with this is we are
  34819. 26:25:42splitting apart
  34820. 26:25:44our our independent variable our
  34821. 26:25:47features. So this this x will contain
  34822. 26:25:52our features
  34823. 26:25:56and y will contain
  34824. 26:25:59our label.
  34825. 26:26:03Does that make sense? We're splitting
  34826. 26:26:04this data apart. So, we're only grabbing
  34827. 26:26:06that first column here to be our
  34828. 26:26:09features. And then we're we're grabbing
  34829. 26:26:12the second column, which is the sales,
  34830. 26:26:13because we're going to predict the
  34831. 26:26:15sales. This is our label. We're going to
  34832. 26:26:17we're going to build a model to predict
  34833. 26:26:19the sales given the TV input, TV expense
  34834. 26:26:23input. So, the first thing we have to do
  34835. 26:26:26is split apart the features and the
  34836. 26:26:28label.
  34837. 26:26:31Okay, that's the first step we usually
  34838. 26:26:33will take. And the reason we have to do
  34839. 26:26:35that um just to reiterate, the reason we
  34840. 26:26:39have to do that is because our model
  34841. 26:26:42will expect our our data features to be
  34842. 26:26:45separate from the label. We will pass
  34843. 26:26:47those in separately.
  34844. 26:26:50X is TV. It's the first column
  34845. 26:26:54because we're using eyeling.
  34846. 26:27:09We are predicting the sales given the TV
  34847. 26:27:12expense value.
  34848. 26:27:17Yeah. which is why we split it into so
  34849. 26:27:20this is the second column right the
  34850. 26:27:22index one column
  34851. 26:27:28uh you just put in read CSV and pass in
  34852. 26:27:30the URL so you could so exactly the code
  34853. 26:27:34that was up earlier from Tim
  34854. 26:27:37um you just do this
  34855. 26:27:41and then data equals ed read CSV URL
  34856. 26:27:49So we split our data into X and Y here.
  34857. 26:27:53All right. Now, one other step that
  34858. 26:27:56we're going to take that's a very very
  34859. 26:27:58critical step and you're going to we're
  34860. 26:27:59going to see this step over and over and
  34861. 26:28:03over and over again. So splitting apart
  34862. 26:28:05into X and Y will become we'll do that
  34863. 26:28:08over and over and over and over again.
  34864. 26:28:10Not only that, but doing this next step,
  34865. 26:28:14which is what's called a train test
  34866. 26:28:17split. Now, let me show you what the
  34867. 26:28:19train test split does. It takes our data
  34868. 26:28:24and it's going to split apart our data
  34869. 26:28:27that we have, our X and our Y data. It's
  34870. 26:28:30going to split it apart into a
  34871. 26:28:32percentage that will be used to train
  34872. 26:28:34the data
  34873. 26:28:37and then a percentage that will be used
  34874. 26:28:39to test. Now, why would we want to do
  34875. 26:28:42that? It's mainly so we can do
  34876. 26:28:45evaluation. So, we build the model over
  34877. 26:28:48here and then we test it on data that
  34878. 26:28:51has not seen before. So, we reserve a
  34879. 26:28:55percentage of the data to be used for
  34880. 26:28:57test. Usually this this data is um
  34881. 26:29:01somewhere between uh 20 to 30%.
  34882. 26:29:06So somewhere between 20 to 30% of the
  34883. 26:29:09original data. So that means the
  34884. 26:29:12majority of it is used for training. So
  34885. 26:29:14the majority of the of that X and Y over
  34886. 26:29:17here is going to be between 70 to 80%.
  34887. 26:29:23will generally be used for for uh for
  34888. 26:29:27training. Okay. So somewhere between 20
  34889. 26:29:30to 30 the industry standard is some
  34890. 26:29:32anywhere in between there. Um a lot of
  34891. 26:29:35people like to use 30%, some people like
  34892. 26:29:37to use 20%. Um anything in that range is
  34893. 26:29:40acceptable. Um we will I think we
  34894. 26:29:43generally will favor like 30%.
  34895. 26:29:46um to be used for testing. But um the
  34896. 26:29:50the point is we don't we don't want to
  34897. 26:29:53mix those together. We want those to be
  34898. 26:29:55separated out so that we can have a fair
  34899. 26:30:00evaluation, right? We want to train our
  34900. 26:30:02data on this train our model on this
  34901. 26:30:05data and then see how well it performs
  34902. 26:30:08on this data that it has never seen
  34903. 26:30:10before.
  34904. 26:30:12Right? So in order to have data it's
  34905. 26:30:14never seen before, we're going to take
  34906. 26:30:16our x and our y and we're going to split
  34907. 26:30:17it using this function called train test
  34908. 26:30:21split that will do this kind of
  34909. 26:30:24splitting for us. Okay, so scikitlearn
  34910. 26:30:27has a function called train test split
  34911. 26:30:30that will go ahead and we're going to
  34912. 26:30:32pass our x and our y and we'll pass in a
  34913. 26:30:34percentage like 30% that we want to
  34914. 26:30:38split out into a test set and then the
  34915. 26:30:40remainder of that the 70% will be used
  34916. 26:30:44for training the model.
  34917. 26:30:47Okay.
  34918. 26:30:50So what we're going to get let me redraw
  34919. 26:30:53that. So what we're going to get out of
  34920. 26:30:55this for the train test split is we're
  34921. 26:30:57going to we're going to have an X and a
  34922. 26:30:59Y per
  34923. 26:31:02training and test. So we're going to get
  34924. 26:31:05now we're going to get an X train
  34925. 26:31:11and a Y train.
  34926. 26:31:15So we're going to get training features
  34927. 26:31:17and training labels. And then we're
  34928. 26:31:19going to get test features
  34929. 26:31:24to plug into our model and and test
  34930. 26:31:28answers or test labels
  34931. 26:31:32to do evaluation because what we should
  34932. 26:31:34be able to do is build the model over
  34933. 26:31:36here and then apply the model on this
  34934. 26:31:38data. Meaning we can take these features
  34935. 26:31:41and plug it into our model and then see
  34936. 26:31:44what answers we get and compare those
  34937. 26:31:47answers to this testing data. Right? We
  34938. 26:31:50should be able to do that to generate an
  34939. 26:31:52evaluation.
  34940. 26:31:56Okay. Now you may be wondering why do we
  34941. 26:31:59do any of that? What's the purpose of
  34942. 26:32:01that?
  34943. 26:32:03Evaluating it on this test data gives us
  34944. 26:32:06a good sense of will our model
  34945. 26:32:11generalize to new examples. Right? If it
  34946. 26:32:15performs pretty well on this data,
  34947. 26:32:18that's a good signal like when it's
  34948. 26:32:19performing pretty well on data it's
  34949. 26:32:21never seen before, that's a good
  34950. 26:32:24indicator that it's going to perform
  34951. 26:32:26pretty well when we use it on brand new
  34952. 26:32:28examples
  34953. 26:32:30um in the future.
  34954. 26:32:33Right. So that's a that's why we do this
  34955. 26:32:37evaluation on this data that it has not
  34956. 26:32:40seen before. It's going to see this
  34957. 26:32:43training data, right? We're going to
  34958. 26:32:44train the model on that data. But that
  34959. 26:32:47model will never be exposed to this test
  34960. 26:32:49data until we do the evaluation
  34961. 26:32:53and and generate some metrics to see how
  34962. 26:32:56good is this performing
  34963. 26:32:58and does it have a good chance of
  34964. 26:32:59generalizing to never before seen
  34965. 26:33:02examples which is what we want right
  34966. 26:33:04because we're going to use this model in
  34967. 26:33:06the real world. It's going to be being
  34968. 26:33:08used on new examples that it hasn't seen
  34969. 26:33:10before. We want it to perform well. So,
  34970. 26:33:13this is kind of our test, our
  34971. 26:33:15evaluation.
  34972. 26:33:20Okay. Any questions on the We're going
  34973. 26:33:23to do this in a moment. I'll show you
  34974. 26:33:24what it looks like in the code, but any
  34975. 26:33:27conceptually any questions on the train
  34976. 26:33:29test split idea. It's a very very
  34977. 26:33:32important idea that we um basically use
  34978. 26:33:37part of the data to train it and then
  34979. 26:33:38another part of it to evaluate. It's
  34980. 26:33:41very important we do that. By the way,
  34981. 26:33:43this has a term um this in machine
  34982. 26:33:46learning this is called cross
  34983. 26:33:50validation
  34984. 26:33:55because we are using one data set to
  34985. 26:33:58train the model and then we're cross
  34986. 26:34:00over we're crossing that over into
  34987. 26:34:03another data set to validate it which is
  34988. 26:34:06the uh the the testing that.
  34989. 26:34:13So this is called cross validation. Um
  34990. 26:34:15there's actually many ways to do cross
  34991. 26:34:17validation. That's something we'll
  34992. 26:34:18study. This is a very simple way of
  34993. 26:34:20doing cross validation. There's more
  34994. 26:34:21complex ways. You can take your data and
  34995. 26:34:24you can actually divide it into many
  34996. 26:34:26sections
  34997. 26:34:27and basically train it against most of
  34998. 26:34:29these and evaluate it against one at a
  34999. 26:34:32time and then rotate. So that's another
  35000. 26:34:35way to do cross validation. We're going
  35001. 26:34:36to study that. Um but this is the this
  35002. 26:34:40is the simplest way to do it here.
  35003. 26:34:48Okay.
  35004. 26:34:51So let me show you what you get when you
  35005. 26:34:52use train test split. So uh we're going
  35006. 26:34:55to import from sklearn.
  35007. 26:34:58We're uh from the model selection
  35008. 26:35:00module. Now we haven't used this before.
  35009. 26:35:03This is our first time using it. But
  35010. 26:35:04here's our model selection. We're going
  35011. 26:35:07to import this train test split function
  35012. 26:35:10and we're going to use it on our X and Y
  35013. 26:35:13and we're going to set a test size of
  35014. 26:35:1730% which is which is.3. So our test
  35015. 26:35:20size
  35016. 26:35:23is 30%.
  35017. 26:35:25Converted to decimal
  35018. 26:35:29right converted to.3 so that means we're
  35019. 26:35:32reserving 30% for that test set. Um you
  35020. 26:35:36can set a random state. Now that's
  35021. 26:35:37completely optional. Um the random state
  35022. 26:35:42is for reproducibility
  35023. 26:35:50because what the train test split is
  35024. 26:35:51going to do is it's actually going to
  35025. 26:35:53shuffle the data and then split it apart
  35026. 26:35:56into the 7030.
  35027. 26:35:58So um yes, the seed. Exactly. It's like
  35028. 26:36:02a seed. So it's it's saying like when
  35029. 26:36:04you do that shuffling every time I run
  35030. 26:36:06this notebook I'm going to get the same
  35031. 26:36:08result but it's going to be random the
  35032. 26:36:10first it's going to be random but I'm
  35033. 26:36:11going to be able to reproduce that
  35034. 26:36:13randomness with that random state. Yes,
  35035. 26:36:18it is like a seed.
  35036. 26:36:23Uh it's you can choose any number to be
  35037. 26:36:26your your um your random state. It 42
  35038. 26:36:29isn't important. You could choose zero.
  35039. 26:36:31You could choose one. Um, you could
  35040. 26:36:33choose any positive integer. Um, 42 is
  35041. 26:36:37kind of like the uh industry standard.
  35042. 26:36:41It's it's you'd have to look it up why
  35043. 26:36:43it is. Um, apparently 42 is a special
  35044. 26:36:47number. Um,
  35045. 26:36:50in kind of the history of development of
  35046. 26:36:52this stuff, there's nothing really
  35047. 26:36:54special about 42. You could choose a
  35048. 26:36:56random You could choose a random seed to
  35049. 26:36:58be uh zero. That's fine. It it doesn't
  35050. 26:37:01really it doesn't really matter.
  35051. 26:37:05Um you just want you can choose it to be
  35052. 26:37:08uh one, two, three. Um you can choose it
  35053. 26:37:11to be 15. You can choose it to be
  35054. 26:37:13anything you want it to be. It's really
  35055. 26:37:15so that your your shuffling is
  35056. 26:37:17consistent. Every time you run this
  35057. 26:37:19notebook, you get the same shuffle
  35058. 26:37:21result. So I'm always going to get the
  35059. 26:37:23same rows in these splits.
  35060. 26:37:29Hitch. There it is. I knew it was from
  35061. 26:37:31something.
  35062. 26:37:38Yeah. So 42 is kind of like a
  35063. 26:37:42it's it's just used ubiquitously
  35064. 26:37:46uh you know as kind of a um paying
  35065. 26:37:50tribute to the Hitchhiker's Guide to the
  35066. 26:37:52Galaxy, but it's no it's there's nothing
  35067. 26:37:54that special about 42. It doesn't it's
  35068. 26:37:56not going to change our result or
  35069. 26:37:58anything.
  35070. 26:37:59It's just so that this train set split
  35071. 26:38:02is going to shuffle our data and split
  35072. 26:38:04it apart into 7030.
  35073. 26:38:07You just want to set this to something
  35074. 26:38:08so that you get a cons every time we run
  35075. 26:38:11this notebook, we get a consistent
  35076. 26:38:13shuffle.
  35077. 26:38:15And so the data in these sets
  35078. 26:38:18are uh consistent. That's all.
  35079. 26:38:27Okay. But do you guys see how we pass in
  35080. 26:38:30our X and our Y and we generate four we
  35081. 26:38:33generate four different data uh
  35082. 26:38:36quantities here which is we generate
  35083. 26:38:38training features, test features,
  35084. 26:38:41training labels and test labels because
  35085. 26:38:43again we are generating these four
  35086. 26:38:47different we're generating data on these
  35087. 26:38:50two different sets a training set
  35088. 26:38:53and a test set. So we have training
  35089. 26:38:56features, training label,
  35090. 26:39:00and then test features, test label.
  35091. 26:39:04Okay, that's why it's so important to
  35092. 26:39:07split apart our data into the X and the
  35093. 26:39:09Y. We need those split apart in order
  35094. 26:39:12for this part to work.
  35095. 26:39:16So by the way, these two steps we will
  35096. 26:39:18always do for any model we build. We'll
  35097. 26:39:21generally do X and Y and then train test
  35098. 26:39:24split in order to generate the data that
  35099. 26:39:28we will use for building our model.
  35100. 26:39:35Okay. So this this data here is going to
  35101. 26:39:38be what we actually use to guide the
  35102. 26:39:40training of our model. So it's
  35103. 26:39:41definitely supervised, right? Linear
  35104. 26:39:43regression
  35105. 26:39:45um we we will use that
  35106. 26:39:55Okay, so we haven't built the model yet.
  35107. 26:39:57We're just getting our data split apart
  35108. 26:39:59and ready for the training. We haven't
  35109. 26:40:02actually built our model yet, right?
  35110. 26:40:04That'll be coming up uh in a moment. But
  35111. 26:40:07this is getting our data ready. We
  35112. 26:40:09started with our data frame. We split it
  35113. 26:40:11apart into uh an x and a y. And we split
  35114. 26:40:16that into a train test split. And um you
  35115. 26:40:22know then we can uh then we can go ahead
  35116. 26:40:25and um pass in to our model training
  35117. 26:40:28which we'll do in a moment.
  35118. 26:40:36Um you that's a good question. You could
  35119. 26:40:38run so what you could do is you could
  35120. 26:40:40run
  35121. 26:40:41um should we import numpy? Let's see.
  35122. 26:40:45We did. Okay. You could run the average
  35123. 26:40:49on the um you could check the MP mean on
  35124. 26:40:54the X train and see how it compares to
  35125. 26:40:58um
  35126. 26:41:00see how it compares to X.
  35127. 26:41:05So you could you could do that and see
  35128. 26:41:07what the average of this feature is um
  35129. 26:41:09compared to the average of the original.
  35130. 26:41:11They may not be perfect because we are
  35131. 26:41:13taking a reduced data set size. So I
  35132. 26:41:16don't think there's really any good
  35133. 26:41:18there's not like a one-sizefits-all
  35134. 26:41:20validation we can do because we're
  35135. 26:41:21taking a random shuffle and taking a
  35136. 26:41:24percent. We're taking 70% of the data
  35137. 26:41:26out. So we're not guaranteed to maintain
  35138. 26:41:28the same statistics. We can see if
  35139. 26:41:30they're close.
  35140. 26:41:32Um but does that make sense? Like we're
  35141. 26:41:34not guaranteed to get the same stats
  35142. 26:41:36because we're taking a slice of it.
  35143. 26:41:37We're taking 70%.
  35144. 26:41:40So it's not guaranteed to to to
  35145. 26:41:43be the same distribution really.
  35146. 26:41:50Delete that.
  35147. 26:41:53Uh is it good practice? Yes, it is.
  35148. 26:41:57It is. Uh 30% is the industry standard.
  35149. 26:42:00Anything between 20 to 30, so 0.2,
  35150. 26:42:030.25.3,
  35151. 26:42:05any of those are acceptable. It's really
  35152. 26:42:07up to you. Um I mostly see 30%.
  35153. 26:42:12Mo I think.3 is is a good good practice
  35154. 26:42:15to use for sure.
  35155. 26:42:17Um I did explain random state. Uh random
  35156. 26:42:20state is so that you get consistent
  35157. 26:42:23shuffling. Um you can set this to any
  35158. 26:42:26integer that you want it to be. It it
  35159. 26:42:28doesn't really matter. Um you can set it
  35160. 26:42:31to uh 100, you can set it to 10, you can
  35161. 26:42:34set it to 15. Um it just ensures because
  35162. 26:42:38what this split will do is it will
  35163. 26:42:40shuffle the data first. It'll shuffle
  35164. 26:42:42the rows and then um split it apart into
  35165. 26:42:45the into the train and test sets. So you
  35166. 26:42:49set the random state so that the next
  35167. 26:42:51time you run this you get the same
  35168. 26:42:53consistent shuffling. That's the only
  35169. 26:42:55that's the only thing it it helps you
  35170. 26:42:57with because it is randomized but when
  35171. 26:43:00you set a random state um it's so that
  35172. 26:43:03like if you run it again you'll get the
  35173. 26:43:05same shuffling.
  35174. 26:43:07You'll get the same the shuffling
  35175. 26:43:08matters because it it it uh dictates
  35176. 26:43:11what ends up in in these sets.
  35177. 26:43:19Okay.
  35178. 26:43:21All right. So let's see let's do let's
  35179. 26:43:24build the model. Um and let me show you
  35180. 26:43:28how easy this is going to be to build
  35181. 26:43:30the model. And this is really how it's
  35182. 26:43:31going to be for every single scikitlearn
  35183. 26:43:34model will basically look the exact same
  35184. 26:43:37for training it which is what's going to
  35185. 26:43:39make it really really nice. So the first
  35186. 26:43:41thing we have to do is import our model.
  35187. 26:43:45So from scikitlearn we're going to be
  35188. 26:43:46using a linear from the linear model
  35189. 26:43:49package or the linear model module I
  35190. 26:43:53should say within sklearn we're going to
  35191. 26:43:55be importing the linear regression
  35192. 26:43:58and we're going to create an instance of
  35193. 26:44:00the linear regression here.
  35194. 26:44:03Okay, so linear regression and look how
  35195. 26:44:07easy this is going to be. Nearly all
  35196. 26:44:12nearly all sklearn models use
  35197. 26:44:17ffit function to train.
  35198. 26:44:22So every one of them, no matter which
  35199. 26:44:25one we use, like the decision tree, like
  35200. 26:44:28the um logistic regression, any of those
  35201. 26:44:32like we use for classification that are
  35202. 26:44:33going to be coming up in lesson four,
  35203. 26:44:35they're all going to look the same in
  35204. 26:44:37terms of it's going to run.fit,
  35205. 26:44:40which is um scikitlearn's
  35206. 26:44:44uh generic function for training your
  35207. 26:44:46model. So this will execute the training
  35208. 26:44:50once we run this code. And what that
  35209. 26:44:53again the linear regression training is
  35210. 26:44:55going to do that least squares distance
  35211. 26:44:59procedure or algorithm to try to find
  35212. 26:45:02the right weights. It's trying to find
  35213. 26:45:05those weights that minimize that squared
  35214. 26:45:07distance uh from our line that it's
  35215. 26:45:10trying to build to the data.
  35216. 26:45:13And what I want you to notice is what we
  35217. 26:45:16put into the ffit. See how we put in the
  35218. 26:45:19training data where we put in the
  35219. 26:45:20training features and we put in the
  35220. 26:45:23training labels. Now this is supervised.
  35221. 26:45:27So of course we put in the labels,
  35222. 26:45:30right? Of course we put in these labels
  35223. 26:45:33here and of course we put in our
  35224. 26:45:35features here. So we're putting in all
  35225. 26:45:38of our examples from our training split
  35226. 26:45:43into this ffit which is going to train
  35227. 26:45:46the model uh so that we can we can use
  35228. 26:45:50it for prediction.
  35229. 26:45:52Okay, it's really fast. If I run this,
  35230. 26:45:55it's going to be pretty much instant.
  35231. 26:45:58Pretty much instantly it gets trained.
  35232. 26:46:00And you can see here we now have a
  35233. 26:46:02linear regression. you can see in this
  35234. 26:46:03little box. Um, and it and this
  35235. 26:46:06information says that it has been
  35236. 26:46:08fitted. So, it's now ready to be used.
  35237. 26:46:11Right? So, we now that's it. We've
  35238. 26:46:13trained our model. We try that's how
  35239. 26:46:15easy that was. We did ffit. Now, what we
  35240. 26:46:18should realize is there's a lot of work
  35241. 26:46:21going on behind the scenes of this ffit.
  35242. 26:46:24Okay. There's a lot of work being done
  35243. 26:46:26there to do the least squares algorithm
  35244. 26:46:30and find those weights and and create
  35245. 26:46:33that line of best fit. Right? So there
  35246. 26:46:36there's a lot of work being going on
  35247. 26:46:38there that's going on there behind the
  35248. 26:46:39scenes, but scikitlearn is abstracting
  35249. 26:46:42it away for us, right? And all we have
  35250. 26:46:44to do is fit when we're using this code.
  35251. 26:46:48Really easy. Really easy. Fit. And there
  35252. 26:46:52we go. We've trained our linear
  35253. 26:46:54regression model.
  35254. 26:46:58And by the way, if you want to see what
  35255. 26:47:01the coefficients are, you can actually
  35256. 26:47:03extract them if you do so if you take
  35257. 26:47:05your lin regression and you do um
  35258. 26:47:09coefficients like this.
  35259. 26:47:12COF with a with an underscore. So this
  35260. 26:47:16gives us the trained
  35261. 26:47:19weights
  35262. 26:47:21coefficients
  35263. 26:47:23also known as the coefficients right.
  35264. 26:47:26Um so if you run this you can see uh
  35265. 26:47:29right now we have this coefficient here
  35266. 26:47:33um which is the only coefficient we had
  35267. 26:47:35on our feature. So we only had one
  35268. 26:47:38feature coefficient there.
  35269. 26:47:51And we can take a look at our intercept
  35270. 26:47:57which is this.
  35271. 26:48:00So this gives us the train weights
  35272. 26:48:03and so we can look at the intercept we
  35273. 26:48:06can look at the the the coefficient. Um
  35274. 26:48:10so obviously if we have multiple
  35275. 26:48:12features our model has many features
  35276. 26:48:14it's going to have more values in that
  35277. 26:48:16coefficient but the intercept is just
  35278. 26:48:18the single value 7.23
  35279. 26:48:21and then the coefficient
  35280. 26:48:25is 0.046. So that's the weight that gets
  35281. 26:48:28learned.
  35282. 26:48:32Is there a size limit? No, not really.
  35283. 26:48:34There's no size limit. Um,
  35284. 26:48:37no. You can use as much data as you
  35285. 26:48:39want.
  35286. 26:48:41There's really no size limit other than
  35287. 26:48:43what like what you can fit in memory.
  35288. 26:48:48I'd say that's the only limit is
  35289. 26:48:49basically what the amount of data that
  35290. 26:48:51can fit in memory.
  35291. 26:49:00Okay.
  35292. 26:49:02All right. Were you guys able to run
  35293. 26:49:03this? Were you guys able to run the
  35294. 26:49:05linear regression ffit?
  35295. 26:49:09Okay, perfect.
  35296. 26:49:16Perfect. Do you Okay, great. Great.
  35297. 26:49:21So, we have a model and we can use it to
  35298. 26:49:24predict. Um, and so that's actually what
  35299. 26:49:27we're going to do next. If we go down
  35300. 26:49:29here, um we're going to have a function
  35301. 26:49:32that's going to um build a scatter plot
  35302. 26:49:35of our original test data.
  35303. 26:49:39Um so we're going to have our test data
  35304. 26:49:41here.
  35305. 26:49:43Um,
  35306. 26:49:45and we're going to then take our uh
  35307. 26:49:48we're going to take our training data
  35308. 26:49:51and plot we're going to use the uh this
  35309. 26:49:54data versus our sales predictions. So
  35310. 26:49:58you can see we're going to you this is
  35311. 26:50:00how by the way this is how you use the
  35312. 26:50:02scikitlearn model to predict. You have a
  35313. 26:50:05fit to train it and look at the function
  35314. 26:50:08you use to predict. It's literally just
  35315. 26:50:10called predict. That's how easy it is.
  35316. 26:50:13and you pass in your data, all your
  35317. 26:50:15features into this predict and it
  35318. 26:50:17generates a prediction for every row. So
  35319. 26:50:21every row in these features in this data
  35320. 26:50:23frame um will end up with a prediction
  35321. 26:50:27using our model. So what we're going to
  35322. 26:50:30do is plot our training uh features
  35323. 26:50:34against the predicted sales to see how
  35324. 26:50:38good of a fit that really was.
  35325. 26:50:41Okay. to see to see the regression fit.
  35326. 26:50:49Okay. And so there's the regression fit.
  35327. 26:50:52We have all of our test data here
  35328. 26:50:54plotted in the green. We have our blue,
  35329. 26:50:56which is our um we have our our blue,
  35330. 26:51:00which is our uh um training data line
  35331. 26:51:04that we built our model on. So that's a
  35332. 26:51:06pretty decent fit. Um and then our test
  35333. 26:51:10data is here. We just plotted in the
  35334. 26:51:12green scatter. But the thing I want you
  35335. 26:51:14to see is this prediction, right? We we
  35336. 26:51:17were able to generate some predictions
  35337. 26:51:19on that training um by running our
  35338. 26:51:23predict function with our model. Now
  35339. 26:51:24this model has been trained. So we've
  35340. 26:51:27already fit it and now we're using it to
  35341. 26:51:30predict, right? And so we're predicting
  35342. 26:51:32the sales and plotting that on the y
  35343. 26:51:34ais. So the sales are we're using the
  35344. 26:51:38predicted sales there which is our blue
  35345. 26:51:40line. So this is our line of best fit.
  35346. 26:51:44So this is our model prediction.
  35347. 26:51:52This is our model predictions. Right?
  35348. 26:51:56You can see it's a pretty decent uh
  35349. 26:51:58line, right? Pretty decent line of best
  35350. 26:52:00fit.
  35351. 26:52:02Of course, there's some error here. Like
  35352. 26:52:04there, you know, it's not perfect, but
  35353. 26:52:06it it does a decent job of being a best
  35354. 26:52:09fit line.
  35355. 26:52:23Okay.
  35356. 26:52:24So look how easy that was to
  35357. 26:52:27just to recap this to fit our model was
  35358. 26:52:30a linear regression.fit and of course
  35359. 26:52:32we're going to do more examples. So no
  35360. 26:52:35worries uh on that we're going to see
  35361. 26:52:37this many many many times throughout
  35362. 26:52:39this notebook. But we have linear
  35363. 26:52:42regression.fit to train it and then we
  35364. 26:52:45have linear regression.predict
  35365. 26:52:48to and we pass in our features and that
  35366. 26:52:50generates a predicted output.
  35367. 26:52:53Right? So what this is actually doing is
  35368. 26:52:57is computing this quantity.
  35369. 26:53:12We could do either.
  35370. 26:53:14We could do either. Um, so we could do,
  35371. 26:53:19so one thing we could do is plot uh, so
  35372. 26:53:22we could swap it out. We, we could do
  35373. 26:53:24either one. It doesn't, it's not a big
  35374. 26:53:26deal to do the training set. We could
  35375. 26:53:28do, so we could plot X test and then we
  35376. 26:53:31could plot linear regression X test.
  35377. 26:53:42So it's it's a similar line. Um it's
  35378. 26:53:47just different input features, but the
  35379. 26:53:48line is going to be the same. Just
  35380. 26:53:51different inputs,
  35381. 26:53:53but the coefficients are the same,
  35382. 26:53:55right? It's the same line. It's just we
  35383. 26:53:57generate different outputs.
  35384. 26:54:02So yeah, you could do either one.
  35385. 26:54:06This is This is honestly this is
  35386. 26:54:08probably better. I see what you're
  35387. 26:54:10saying. This is probably better because
  35388. 26:54:11this is the line of best fit through
  35389. 26:54:14this data. So that probably makes sense
  35390. 26:54:16to do to do predict on the test set.
  35391. 26:54:20Agreed on that. Probably makes about
  35392. 26:54:23most sense.
  35393. 26:54:30But you could do either one.
  35394. 26:54:42Yeah, I think that would be the most I
  35395. 26:54:44think that makes the most sense is for
  35396. 26:54:46it to be on the same one just to
  35397. 26:54:48validate. So like we could do we could
  35398. 26:54:50do training here and then train and
  35399. 26:54:52train just to see how that data lines
  35400. 26:54:55up. Really, what we're trying to do is
  35401. 26:54:58have our scattered data and then our
  35402. 26:55:00line of best fit on the same plot.
  35403. 26:55:03That's all we're trying to do, right?
  35404. 26:55:05So, yeah, I think I think they should be
  35405. 26:55:06the same.
  35406. 26:55:10I think that makes sense.
  35407. 26:55:14These values
  35408. 26:55:17or which values do you want to see?
  35409. 26:55:24Yeah, we could uh we could generate
  35410. 26:55:26those if we just do um let's go down
  35411. 26:55:29here. So the the line values
  35412. 26:55:34um are going to be uh the prediction. So
  35413. 26:55:38um the the
  35414. 26:55:40uh test
  35415. 26:55:43predictions
  35416. 26:55:45equals um
  35417. 26:55:50test predictions equals linear
  35418. 26:55:52regression.predict predict x test and
  35419. 26:55:54then we could uh we could print out our
  35420. 26:55:56test predictions.
  35421. 26:56:01Yeah. So we can see what those actual
  35422. 26:56:03values are on our uh on the test set.
  35423. 26:56:07Yeah.
  35424. 26:56:18Um we will do that. Yeah. So you thought
  35425. 26:56:20we were checking how well our data was
  35426. 26:56:21trained. We will do that. Yes, we
  35427. 26:56:23haven't learned how to evaluate this
  35428. 26:56:24yet. We're going to talk about that
  35429. 26:56:26coming up next. Yeah, we will do that.
  35430. 26:56:29We just haven't learned how to do proper
  35431. 26:56:31evaluation
  35432. 26:56:33of a regression model.
  35433. 26:56:36But yeah, it's something we're going to
  35434. 26:56:37talk about for sure
  35435. 26:56:39and see how to do in our code.
  35436. 26:56:46Okay.
  35437. 26:56:49All right. Any other uh questions on
  35438. 26:56:52this example?
  35439. 26:56:59Again, big takeaways
  35440. 26:57:02fit to train it and then predict to use
  35441. 26:57:06it.
  35442. 26:57:08Predict on the features to use the model
  35443. 26:57:11and make predictions with it.
  35444. 26:57:16So here is example. We we made all the
  35445. 26:57:18predictions. This these are all the
  35446. 26:57:19values that are on that line.
  35447. 26:57:22These are all our predictions. And
  35448. 26:57:23notice they this is a truly regression,
  35449. 26:57:25right? These are all floating point
  35450. 26:57:27values. Um so this is definitely a
  35451. 26:57:29regression, right?
  35452. 26:57:43Okay.
  35453. 26:57:50Uh, that's a good question. Um,
  35454. 26:57:54I'm not sure if there is
  35455. 26:57:58if there's like a verbose
  35456. 26:58:03there's not really no there's not really
  35457. 26:58:05a verbose. You can I mean you can look
  35458. 26:58:06at the source code if you really want to
  35459. 26:58:08see you can view the source code to see
  35460. 26:58:11um how it's done. I can tell you I mean
  35461. 26:58:14so generally linear regression is done
  35462. 26:58:17in two ways. Either you use a formula um
  35463. 26:58:21to to solve the optimization problem of
  35464. 26:58:24minimizing like this this uh distance
  35465. 26:58:27from the points to to the line. Um
  35466. 26:58:32or you use something called gradient
  35467. 26:58:33descent which is how a lot of these
  35468. 26:58:35things do it is they iterate through a
  35469. 26:58:39bunch of different iterations where they
  35470. 26:58:40update these weights according to um a
  35471. 26:58:44certain uh basically a gradient of the
  35472. 26:58:48the error function. The error function
  35473. 26:58:50in this case is the is the squared
  35474. 26:58:53distance from the line to the uh to to
  35475. 26:58:59the points.
  35476. 26:59:01So uh we can compute the gradient of
  35477. 26:59:03that and do um gradient descent. So if
  35478. 26:59:06you really want to look into it, I would
  35479. 26:59:08do some research on like linear
  35480. 26:59:10regression gradient descent.
  35481. 26:59:13Okay, linear regression gradient descent
  35482. 26:59:15to see how that's uh how that's being
  35483. 26:59:17done. Yeah, it it's it's a pretty simple
  35484. 26:59:21procedure. Um, again, you have the the
  35485. 26:59:25notion is that you want to minimize
  35486. 26:59:28minimize the loss or the error. Uh, in
  35487. 26:59:32this case, the loss is the square
  35488. 26:59:34distance. So, it's like um there's like
  35489. 26:59:38a it's a formula. It's like a sum of a
  35490. 26:59:41square distance from your prediction
  35491. 26:59:44um or your label sorry to your model
  35492. 26:59:48which is the beta 0 um plus beta 1 x1
  35493. 26:59:53plus beta 2 x2
  35494. 26:59:56etc like your model and then squared. So
  35495. 26:59:59this squared this is the squared
  35496. 27:00:01distance here and you're minimizing this
  35497. 27:00:04guy which is like a calculus problem.
  35498. 27:00:07You you find you basically find the this
  35499. 27:00:10is this is a I'm getting so far into the
  35500. 27:00:12weeds of this, but this is like a
  35501. 27:00:14parabola and you work your way No, no,
  35502. 27:00:17you're good. It's it's it's a good
  35503. 27:00:19question. Um you work your way down to
  35504. 27:00:22the minimum of it. Does that make sense?
  35505. 27:00:24Like you're working your way down here
  35506. 27:00:26and you do that through a descent
  35507. 27:00:28process, like a descent iteration.
  35508. 27:00:31Um
  35509. 27:00:33so
  35510. 27:00:35that's how these are found.
  35511. 27:00:37Um, but you don't see that happening in
  35512. 27:00:41the background. But if you look at the
  35513. 27:00:42source code, it I guarantee you it would
  35514. 27:00:44be it's either going to be this or
  35515. 27:00:46they're going to use the they're going
  35516. 27:00:48to use a a a matrix formula to basically
  35517. 27:00:51solve an equation um that involves this.
  35518. 27:00:57Basically, the derivative of this set
  35519. 27:00:59equal to zero and you find the minimum.
  35520. 27:01:02Either way, you're finding the minimum
  35521. 27:01:03of this.
  35522. 27:01:09Okay. But yeah, I don't think Psycharn
  35523. 27:01:12has like a uh maybe there's some type of
  35524. 27:01:15verbose flag you can look for.
  35525. 27:01:18I don't think they have that though. Not
  35526. 27:01:21that I've seen.
  35527. 27:01:31All right.
  35528. 27:01:33So I have uh an important um concept to
  35529. 27:01:37talk about next which is going to be uh
  35530. 27:01:40called overfitting and underfitting
  35531. 27:01:43um which is a really important concept
  35532. 27:01:45that's related to the training and test
  35533. 27:01:48data we just split apart to do
  35534. 27:01:51evaluation.
  35535. 27:01:53And um essentially the the issue with
  35536. 27:01:56machine learning is that it's not
  35537. 27:01:58perfect and it can struggle in different
  35538. 27:02:00ways. And the two ways that it primarily
  35539. 27:02:03struggles is going to be overfitting and
  35540. 27:02:05underfitting.
  35541. 27:02:06So overfitting is a situation where the
  35542. 27:02:11model basically memorizes the training
  35543. 27:02:14data so well that it's it fails to
  35544. 27:02:18generalize to new examples. So what we
  35545. 27:02:21see with overfitting is this exact sign
  35546. 27:02:25here where we have really good
  35547. 27:02:26performance on the training data. So
  35548. 27:02:28when so when we do that train test split
  35549. 27:02:30we see a really good accuracy or really
  35550. 27:02:34low error on the training data but it
  35551. 27:02:39does not perform anywhere near that on
  35552. 27:02:41that test data split. So what that means
  35553. 27:02:44is that the model is overfitting to the
  35554. 27:02:48training data. It's basically memorizing
  35555. 27:02:50it and it's not able to generalize very
  35556. 27:02:54well.
  35557. 27:02:56Now, why does that happen? It's usually
  35558. 27:02:58because the model is way too complex.
  35559. 27:03:01And that means generally you need to do
  35560. 27:03:04something to reduce the complexity.
  35561. 27:03:07Either you need to use a simpler model
  35562. 27:03:10or you need to use some type of
  35563. 27:03:12technique to mitigate overfitting. And
  35564. 27:03:15we're going to we're going to study some
  35565. 27:03:17of those techniques coming up in this
  35566. 27:03:18notebook. uh we might not get to it
  35567. 27:03:20today, but we're going to study
  35568. 27:03:22particularly what can we do to prevent
  35569. 27:03:24overfitting because overfitting is the
  35570. 27:03:26more common issue with machine learning
  35571. 27:03:28models. They tend to do so well at
  35572. 27:03:32learning from data that they pick up on
  35573. 27:03:34small details and patterns in the
  35574. 27:03:37training examples that they're exposed
  35575. 27:03:38to. They don't do a great job at
  35576. 27:03:41generalizing to new examples. They can
  35577. 27:03:43struggle with that.
  35578. 27:03:45So that's overfitting is struggling to
  35579. 27:03:48generalize to new examples, but you do
  35580. 27:03:50really well on your training data. So it
  35581. 27:03:53appears like you have a good model, but
  35582. 27:03:55it it's not able to go and make
  35583. 27:03:57predictions on test data very well,
  35584. 27:03:59which means we would not want to use
  35585. 27:04:01that model in the real world, right?
  35586. 27:04:03Because it's not able to generalize
  35587. 27:04:05outside of what it's already seen. And
  35588. 27:04:07that's not a good thing if we're trying
  35589. 27:04:08to use it for real world examples,
  35590. 27:04:10right?
  35591. 27:04:12So overfitting is a real issue. Um you
  35592. 27:04:16see it all the time. I've seen it many
  35593. 27:04:17many times in the real world, real
  35594. 27:04:19industry uh work that I've done.
  35595. 27:04:22Overfitting is a is a challenge for a
  35596. 27:04:25lot of machine learning models. And so
  35597. 27:04:26we need some techniques to overcome
  35598. 27:04:29overfitting. And we're going to study
  35599. 27:04:31some of those uh coming up shortly.
  35600. 27:04:35Um, one of the things that we can do,
  35601. 27:04:39one of the one of the things that we can
  35602. 27:04:40do to detect overfitting is exactly what
  35603. 27:04:43we just did, which is you split apart
  35604. 27:04:46your data into training and testing so
  35605. 27:04:48that you have a chance to do an
  35606. 27:04:50evaluation to see if you're even
  35607. 27:04:52overfitting in the first place. You want
  35608. 27:04:54to see that performance be consistent
  35609. 27:04:58from train to test, right? You want to
  35610. 27:05:00see consistency. What you don't want to
  35611. 27:05:02see is performance that drops off on the
  35612. 27:05:05test data. It's much worse. You don't
  35613. 27:05:08want to see that. That means that your
  35614. 27:05:10model is overfit uh to your training
  35615. 27:05:12data and it's not going to perform well
  35616. 27:05:14in the real world.
  35617. 27:05:17Okay. So, we're going to have a couple
  35618. 27:05:18ways to uh overcome that. Talk about
  35619. 27:05:22that. Um now, the opposite can actually
  35620. 27:05:25happen as well, which is called
  35621. 27:05:27underfitting. And underfitting
  35622. 27:05:30refers to the fact that a model is too
  35623. 27:05:33simple and it actually just performs
  35624. 27:05:37poorly across the board. So if we see
  35625. 27:05:40poor performance on the training and
  35626. 27:05:44testing data, that's a good signal that
  35627. 27:05:46the model's underfit and that means it's
  35628. 27:05:50too simple usually and you should try
  35629. 27:05:52using something more complex. Um, so the
  35630. 27:05:55best way to combat underfitting is to
  35631. 27:05:57use a more complex model. And as we go
  35632. 27:06:01through and learn about the models,
  35633. 27:06:03we're going to learn about which ones
  35634. 27:06:05are simple and which ones are complex.
  35635. 27:06:06So we're going to have a scale of kind
  35636. 27:06:09of complexity. And if you're
  35637. 27:06:11underfitting, you want to bump up to the
  35638. 27:06:13to a more complex model. If you're if
  35639. 27:06:16you're overfitting, one way of combating
  35640. 27:06:18that is to actually go down to something
  35641. 27:06:20more simple. Go the opposite way to
  35642. 27:06:22something simpler. So we need to learn
  35643. 27:06:24right now we've only learned linear
  35644. 27:06:26regression
  35645. 27:06:27but we will learn other models you know
  35646. 27:06:30in the future and we'll we'll talk about
  35647. 27:06:32uh their complexity and how they're
  35648. 27:06:34related to each other.
  35649. 27:06:36Okay, but these are two issues we see
  35650. 27:06:38just to draw that out again is if we
  35651. 27:06:41have a train test split where we have
  35652. 27:06:437030 split let's say and we perform
  35653. 27:06:46really well over here but we go to apply
  35654. 27:06:49that model over here and it fails its
  35655. 27:06:51accuracy drops off significantly more
  35656. 27:06:54error that's that's definitely
  35657. 27:06:56overfitting which is not good
  35658. 27:07:00right and then underfitting is just not
  35659. 27:07:02performing well in either case so even
  35660. 27:07:04on the training data itself your your
  35661. 27:07:06accuracy is not very good. So you're not
  35662. 27:07:09really learning effectively. You're
  35663. 27:07:11underfitting your model. So that's
  35664. 27:07:15that's um underfitting case.
  35665. 27:07:21Okay.
  35666. 27:07:26All right. Now the issue is that it can
  35667. 27:07:29be very difficult to balance these two
  35668. 27:07:32and get it correct. That's what makes
  35669. 27:07:33machine learning a little bit
  35670. 27:07:34challenging is getting this balance
  35671. 27:07:37correct of simplicity and complexity. So
  35672. 27:07:41you don't want to be overly complex that
  35673. 27:07:43you overfit, but you don't want to be
  35674. 27:07:45overly simple that you underfit and
  35675. 27:07:48you're not able to learn effectively. So
  35676. 27:07:51there's a bit of a tradeoff there. And
  35677. 27:07:53this trade-off is typically known in the
  35678. 27:07:55community as bias variance trade-off. Um
  35679. 27:07:58in which case uh it's basically like a
  35680. 27:08:01complexity simplicity trade-off. It's
  35681. 27:08:02another word for that. Um,
  35682. 27:08:06and so, uh, it's it's thought that, um,
  35683. 27:08:10if you, uh, if you have very, um, if you
  35684. 27:08:15have a situation where you're able to
  35685. 27:08:16fit the training data very well, you
  35686. 27:08:19risk not being able to generalize. In
  35687. 27:08:22other words, you risk overfitting, and
  35688. 27:08:24it's hard to um, it's hard to combat
  35689. 27:08:28that in a way. Um, and um, on the
  35690. 27:08:33reverse side, if you have something
  35691. 27:08:34really simple, um, you risk not learning
  35692. 27:08:38enough. Even if you're trying to combat
  35693. 27:08:40that overfitting, you risk not learning
  35694. 27:08:43enough and your model just doesn't
  35695. 27:08:45perform as well as it could. So, there's
  35696. 27:08:47a bit of a trade-off there of trying to
  35697. 27:08:49find the right balance between something
  35698. 27:08:51complex enough to learn, but something
  35699. 27:08:54not overly complex that it's going to
  35700. 27:08:58not generalize to new data. That's the
  35701. 27:09:01challenge. Um, like I said, we are going
  35702. 27:09:05to have techniques to overcome this. So
  35703. 27:09:08luckily there are things to basically
  35704. 27:09:10overcome this trade-off and um and help
  35705. 27:09:15us along the way so that we don't
  35706. 27:09:17overfit. They basically prevent
  35707. 27:09:18overfitting
  35708. 27:09:20um and allow us to use complex enough
  35709. 27:09:23models um that that won't be overfit.
  35710. 27:09:27This is in the um this was in our uh
  35711. 27:09:31lesson 3.2 notebook. So you want to pull
  35712. 27:09:34that one back up. We were working on
  35713. 27:09:35Monday.
  35714. 27:09:37Um, and just to recap this a little bit,
  35715. 27:09:40remember we were building a linear
  35716. 27:09:42regression, I wanted to recap some of
  35717. 27:09:44the steps we took there, um, that we
  35718. 27:09:48will be doing over and over again. And
  35719. 27:09:50really the same kind of steps, uh, that
  35720. 27:09:53we do here, we'll do in a lot of our
  35721. 27:09:55model building. Pretty much all of our
  35722. 27:09:57model building um, that we do, whether
  35723. 27:09:59it's regression or classification,
  35724. 27:10:01doesn't really matter. um we'll still be
  35725. 27:10:04doing a lot of these steps which are um
  35726. 27:10:07remember first we split apart our data
  35727. 27:10:09into kind of a features and a label
  35728. 27:10:13uh x and y and the reason that's
  35729. 27:10:16important is because um the model
  35730. 27:10:19training uses the features and the label
  35731. 27:10:23um to help train the model, right? They
  35732. 27:10:26use those separately. Um so we want to
  35733. 27:10:29split those apart whenever we can. And
  35734. 27:10:31so we have usually uh it's a good
  35735. 27:10:33practice to call your features capital X
  35736. 27:10:35and your labels lowercase Y. And what we
  35737. 27:10:39do with that is remember we immediately
  35738. 27:10:42split that into what we call the
  35739. 27:10:44training in a test set. And the picture
  35740. 27:10:47we had for that was something like this
  35741. 27:10:51where we had about 70% of the data
  35742. 27:10:55we used to train the model against and
  35743. 27:10:58then the other 30% of the data we use to
  35744. 27:11:01test the model against. Meaning that we
  35745. 27:11:04build a model over here and we apply it
  35746. 27:11:07to this set over here um to make
  35747. 27:11:10predictions. And then the that's where
  35748. 27:11:12the supervised learning really comes
  35749. 27:11:14into play, right? is on this test set.
  35750. 27:11:17We already have the answers. We already
  35751. 27:11:19have the label. And so we can apply our
  35752. 27:11:21model to this to the features over here.
  35753. 27:11:24Predict uh what the the label should be
  35754. 27:11:27and compare that. We can get a a metric,
  35755. 27:11:30right, that compares how close we are in
  35756. 27:11:33our prediction to the actual values. Um
  35757. 27:11:36and that was some of our performance
  35758. 27:11:38metrics. I'll recap some of those that
  35759. 27:11:40kind of measure that distance away from
  35760. 27:11:42our predictions to what the actual label
  35761. 27:11:45is. Um, but remember we had this train
  35762. 27:11:49test split function which helps us split
  35763. 27:11:52apart our features and our labels into
  35764. 27:11:55these uh four sets of data. So we have
  35765. 27:11:58our training features, our testing
  35766. 27:12:00features and then our training labels
  35767. 27:12:02and our testing labels. So we have all
  35768. 27:12:05of those and um really these two guys
  35769. 27:12:08are going to be used to train the model.
  35770. 27:12:10That's why they're called underscore
  35771. 27:12:12train. They're going to be used to train
  35772. 27:12:13that model. And then the then we're
  35773. 27:12:16going to predict on these set of
  35774. 27:12:18features and then com use those
  35775. 27:12:20predictions to compare to this set of
  35776. 27:12:23labels, right? That's on the test test
  35777. 27:12:25set. Um and you notice here our test
  35778. 27:12:28size is set to 30%. Um, that's a pretty
  35779. 27:12:32standard number. Anywhere between like
  35780. 27:12:3320 to 30% is pretty standard. Um, we'll
  35781. 27:12:37typically use.3, but it could be 02.
  35782. 27:12:40Anywhere in between is fine.
  35783. 27:12:44Okay, so we had that. Hopefully that uh
  35784. 27:12:47we remember that from Monday.
  35785. 27:12:49So we had a train and a test set. And
  35786. 27:12:51then building the model was actually
  35787. 27:12:53really really easy. Once you have those
  35788. 27:12:55train and test sets, um, we just import
  35789. 27:12:57our model object. So from uh scikitlearn
  35790. 27:13:00sklearn
  35791. 27:13:02um linear model uh module from that
  35792. 27:13:05package we import the linear regression
  35793. 27:13:08model and then we do um linear
  35794. 27:13:11regression.fit
  35795. 27:13:12and we pass in our features and our
  35796. 27:13:14labels and this is again this is where
  35797. 27:13:17that supervised learning is really
  35798. 27:13:18coming into play because we're passing
  35799. 27:13:21in these labels.
  35800. 27:13:23That's really what makes this work,
  35801. 27:13:24right? We need those labels to help
  35802. 27:13:26guide the model to make those updates.
  35803. 27:13:29If you guys remember, the model is
  35804. 27:13:32something that looks like this.
  35805. 27:13:37So, this was a bunch of different
  35806. 27:13:40coefficients
  35807. 27:13:41um times the features,
  35808. 27:13:44however many we have. Um, and so these
  35809. 27:13:49labels are really taking the place of
  35810. 27:13:51this and they're helping us um make the
  35811. 27:13:56correct updates to these to these
  35812. 27:13:58coefficients or sometimes we call them
  35813. 27:14:00weights. Um, these B 0, B1, B2. Um, we
  35814. 27:14:05find out what the optimal one is to get
  35815. 27:14:07the best fit, right? To get the line of
  35816. 27:14:09best fit. Um that's what the model
  35817. 27:14:13training when we call this ffit ffit
  35818. 27:14:15that's really what it's doing in the
  35819. 27:14:16background is finding all those
  35820. 27:14:18coefficients right to end up with the
  35821. 27:14:20line of best fit that has the lowest
  35822. 27:14:21amount of error.
  35823. 27:14:26Okay so hopefully that makes sense.
  35824. 27:14:28That's just a dofit fit um to train our
  35825. 27:14:31models. And that's really going to be um
  35826. 27:14:33the case for
  35827. 27:14:36uh pretty much every single model that
  35828. 27:14:39we uh train with scikitlearn. It's
  35829. 27:14:42pretty much going to be a fit. We pass
  35830. 27:14:43in our training uh features and our
  35831. 27:14:46training labels.
  35832. 27:14:49Okay, so we had that and this was the
  35833. 27:14:52visualization of that where we had our
  35834. 27:14:54test points kind of scattered and we see
  35835. 27:14:57our line of best fit is the one that
  35836. 27:14:59goes through there with that minimal
  35837. 27:15:01error. That's that's the whole goal.
  35838. 27:15:05Pretty decent predictor.
  35839. 27:15:10Okay. And then we talked about
  35840. 27:15:13overfitting underfitting. So just to
  35841. 27:15:15recap this overfitting is the concept of
  35842. 27:15:17our model basically memorizing our
  35843. 27:15:19training data. It performs really well
  35844. 27:15:21on that training set but it is not able
  35845. 27:15:24to generalize outside of that. So it
  35846. 27:15:27performs poorly on the test set or data
  35847. 27:15:30that it's never seen before. Um and
  35848. 27:15:33that's overfitting. So the reason that
  35849. 27:15:36it overfits is generally the model is
  35850. 27:15:38too complex and it needs to be um it
  35851. 27:15:42needs to be simplified a bit. And one of
  35852. 27:15:44the things we're going to do today is
  35853. 27:15:46see a couple of ways we can alter the
  35854. 27:15:48linear regression model um if we are
  35855. 27:15:51overfitting to prevent overfitting. Um
  35856. 27:15:55so there's going to be ways to handle
  35857. 27:15:57this. Um and so we're going to explore
  35858. 27:16:00some of those today.
  35859. 27:16:02uh underfitting is kind of the reverse
  35860. 27:16:04of that. Remember, it's where the model
  35861. 27:16:06is not learning enough. So, the
  35862. 27:16:07performance is poor even on the training
  35863. 27:16:09data. It's not good on the test data
  35864. 27:16:12either. Um that is a sign that the model
  35865. 27:16:16is probably too simple and maybe we
  35866. 27:16:18should use something more complex like
  35867. 27:16:20go from a linear regression maybe to use
  35868. 27:16:22a polomial regression. Um or maybe use
  35869. 27:16:25an entirely different model altogether.
  35870. 27:16:28um if we're underfitting, our
  35871. 27:16:29performance is poor, it's a good signal
  35872. 27:16:31we should try something else. Um
  35873. 27:16:36okay,
  35874. 27:16:37so we talked about those
  35875. 27:16:41and one of the things we also talked
  35876. 27:16:43about was evaluations. If you guys
  35877. 27:16:45remember, we had different metrics that
  35878. 27:16:47we could compute to get a gauge of how
  35879. 27:16:50good our model is actually performing.
  35880. 27:16:52Um one of those was MSE, which is this
  35881. 27:16:55mean squared error function. Um so we
  35882. 27:16:57did this example during class last time
  35883. 27:16:59on Monday um where we uh were able to
  35884. 27:17:04generate the mean squared error. That's
  35885. 27:17:07one of our metrics. And we can see what
  35886. 27:17:09the mean squared error is on the
  35887. 27:17:12training set and see what it is on the
  35888. 27:17:13test set by um just passing in our um
  35889. 27:17:17training predictions and our training
  35890. 27:17:18labels, our test predictions and our
  35891. 27:17:21test labels. pass those into this mean
  35892. 27:17:23squared error function and it computes
  35893. 27:17:25the MSE and that's that's a helpful
  35894. 27:17:27function from the scikitlearn metrics
  35895. 27:17:31um package um or module I should say and
  35896. 27:17:35we'll be using that quite a bit to do
  35897. 27:17:38you know evaluation of of especially of
  35898. 27:17:40regression right mean squared error is
  35899. 27:17:42pretty is probably the most common uh
  35900. 27:17:46performance metric we can have and if
  35901. 27:17:48you guys remember what it's really doing
  35902. 27:17:50is measuring these distances So mean
  35903. 27:17:52squared error is kind of like the
  35904. 27:17:54average distance away from our our
  35905. 27:17:56points to the actual um to the
  35906. 27:17:59predictions which the predictions are
  35907. 27:18:02all on this line. Um so it's like
  35908. 27:18:05measuring on average how how much error
  35909. 27:18:07do we have on average right? Um, and the
  35910. 27:18:10idea is the closer to zero the better.
  35911. 27:18:13Generally means that the distance away
  35912. 27:18:15from our prediction to our points is
  35913. 27:18:18pretty low. The closer to zero it is.
  35914. 27:18:20Um, which is pretty desirable.
  35915. 27:18:23So a low MSE is kind of what we're
  35916. 27:18:25looking for. Um, closer to zero the
  35917. 27:18:28better. And so um if one model has if
  35918. 27:18:32one model has um a low lower MSE than
  35919. 27:18:36another, it's it's a better performing
  35920. 27:18:38model, right? It has less error.
  35921. 27:18:42Okay. And then we also looked at the R r
  35922. 27:18:45squared or sometimes known as R2 um
  35923. 27:18:48score. Um this is another metric that we
  35924. 27:18:52could use that measures the the
  35925. 27:18:55variability
  35926. 27:18:56um of uh the predictions and if our
  35927. 27:19:01model is capturing that variability um
  35928. 27:19:03well um and so R squar is has a range of
  35929. 27:19:070 to one one is better that means the
  35930. 27:19:10model is capturing the the changes in in
  35931. 27:19:12the um output it um our predictions
  35932. 27:19:16follow along with those same changes um
  35933. 27:19:18so they're pretty close um so closer to
  35934. 27:19:22one would be a better score. So we have
  35935. 27:19:25those kind of metrics. So like on this
  35936. 27:19:27data um this would this would show that
  35937. 27:19:30this model was underfitting remember
  35938. 27:19:32because this
  35939. 27:19:34mean this MSE was bad and this MSE was
  35940. 27:19:38bad.
  35941. 27:19:40Um and what we should think of these in
  35942. 27:19:42the units of what our labels are. um
  35943. 27:19:47especially if we take the square root of
  35944. 27:19:49this the RMSSE that was another metric
  35945. 27:19:51we had um the square root of this is
  35946. 27:19:54actually in the exact units that we um
  35947. 27:19:58have for our labels. So uh in this
  35948. 27:20:00example this was the um this was the the
  35949. 27:20:04units or the sales versus the TV
  35950. 27:20:07products, right? Um and so this would
  35951. 27:20:11indicate that on average if we take the
  35952. 27:20:13square root of this um
  35953. 27:20:16and the square root of this um we have
  35954. 27:20:19uh
  35955. 27:20:20um we're on average about 11 sales units
  35956. 27:20:24off squared. So if we take the square
  35957. 27:20:26root of that um it's somewhere around 3
  35958. 27:20:27to four um somewhere in between three
  35959. 27:20:31and four units off. And this is as well.
  35960. 27:20:35Um, and because both of these are still
  35961. 27:20:37not close to zero. Um, this would be
  35962. 27:20:39under fit. And this shows that as well.
  35963. 27:20:42This isn't that close to one. It's
  35964. 27:20:44decent, but it's not um not that close
  35965. 27:20:47to one. So, we would say and performance
  35966. 27:20:49is poor on both training and test sets.
  35967. 27:20:52That's the key indicator of
  35968. 27:20:53underfitting. It's poor on both.
  35969. 27:21:01Yeah. Exactly. High MSE correlates to
  35970. 27:21:03underfitting. Yes. Yes. And it what's
  35971. 27:21:05key is it's high MSE on both on both the
  35972. 27:21:10training and the test sets.
  35973. 27:21:13If you have a high MSE on your test set
  35974. 27:21:15but a low MSE on your training set,
  35975. 27:21:17that's overfitting, right? Where it's
  35976. 27:21:20not generalizing from the training set
  35977. 27:21:22to the test data that it hasn't seen
  35978. 27:21:24before. That's overfitting. So the key
  35979. 27:21:27is high MSE on both sets.
  35980. 27:21:33All right. So we talked about that. Um
  35981. 27:21:35we did polomial regression last time. So
  35982. 27:21:38that was um doing
  35983. 27:21:42that was uh making a curved graph um by
  35984. 27:21:46transforming the features into polomial
  35985. 27:21:48features and then doing linear
  35986. 27:21:49regression with that. So you guys
  35987. 27:21:51remember from Monday we did this where
  35988. 27:21:54um we took our features and uh transform
  35989. 27:21:58them according to this polomial features
  35990. 27:22:00from scikitlearn. So we can go all the
  35991. 27:22:02way up to degree whatever degree we
  35992. 27:22:04want. So we put in four here, but
  35993. 27:22:06there's nothing special about four
  35994. 27:22:07really. This is just testing it out. Um
  35995. 27:22:10and we generate the the polomial
  35996. 27:22:12features and we can fit a linear
  35997. 27:22:14regression on those polomial features
  35998. 27:22:17and we get a slightly better model,
  35999. 27:22:20right? Um it fits the data a little bit
  36000. 27:22:23better than just a straight line. this
  36001. 27:22:25curved line with the polomial
  36002. 27:22:28features um performs a little bit better
  36003. 27:22:30and we could see that with the MSE right
  36004. 27:22:31we could evaluate the MSE of this um and
  36005. 27:22:35it would be lower
  36006. 27:22:37it would be lower than the curve line
  36007. 27:22:39and that's something we could do um we
  36008. 27:22:42would just have to pass in these test
  36009. 27:22:44predictions training predictions and
  36010. 27:22:45then the the test labels and training
  36011. 27:22:48labels and passes into the mean squared
  36012. 27:22:50error function and we could compute that
  36013. 27:22:51right wouldn't be hard to
  36014. 27:22:56All right. And then finally, where we
  36015. 27:22:58left off, um, you know, is on our
  36016. 27:23:01performance metrics. So, we talked about
  36017. 27:23:03mean squared error. That's that average
  36018. 27:23:05distance away from the labels to our
  36019. 27:23:08predictions. Um, and we take the square
  36020. 27:23:11root of that. It's it's basically
  36021. 27:23:13measuring the same thing, but it's the
  36022. 27:23:15square root of it is um more
  36023. 27:23:17interpretable because it's in the same
  36024. 27:23:19units as our label.
  36025. 27:23:21um mean absolute error is is the average
  36026. 27:23:25distance of the absolute value. So it's
  36027. 27:23:27not the squared distance formula like a
  36028. 27:23:29uklidian distance but it is a absolute
  36029. 27:23:32value. So it's a little bit um less
  36030. 27:23:34sensitive to outliers. They don't get
  36031. 27:23:36magnified as much. Um but it's not
  36032. 27:23:41typically used as much as a mean squared
  36033. 27:23:43error would be with regression. um we
  36034. 27:23:46talked about the last time because um
  36035. 27:23:48the distance formula or that distance is
  36036. 27:23:51actually what's used to train the model.
  36037. 27:23:53So it's a more natural um fit for a
  36038. 27:23:57performance metric for it.
  36039. 27:24:02All right. And then we had R square. We
  36040. 27:24:03just talked about that closer to zero
  36041. 27:24:05would be um worse. Closer to one would
  36042. 27:24:08be better. That means that the model
  36043. 27:24:10explains um all the variability in the
  36044. 27:24:13in the predictions. Uh it captures those
  36045. 27:24:16predictions um closely to the labels
  36046. 27:24:21um very well. So uh one would be better.
  36047. 27:24:26Closer to one would be better.
  36048. 27:24:29All right. So that's where we left off.
  36049. 27:24:31Um we're gonna pick up from there with
  36050. 27:24:33cross validation. Um, we've actually
  36051. 27:24:36already seen one method of cross
  36052. 27:24:38validation. So, we're going to study um
  36053. 27:24:40we're going to kind of recap that and
  36054. 27:24:41and then um talk about cross validation
  36055. 27:24:44in general um and look at some more
  36056. 27:24:48sophisticated techniques of it um coming
  36057. 27:24:51up next. But before I do that, any
  36058. 27:24:53questions about anything we've covered
  36059. 27:24:56um to this point in in the recap or
  36060. 27:24:59anything from Monday? Any questions on
  36061. 27:25:02that?
  36062. 27:25:04All right. So let's talk about uh cross
  36063. 27:25:07validation. Um now this term cross
  36064. 27:25:12validation refers to a technique that
  36065. 27:25:16evaluates performance. And what it does
  36066. 27:25:19is it divides our data into essentially
  36067. 27:25:23um training and test sets which we've
  36068. 27:25:25kind of already seen. And then we are
  36069. 27:25:27able to train a model on on the training
  36070. 27:25:30set, evaluate it on the test set, and
  36071. 27:25:32that's where that's where we get the
  36072. 27:25:34name cross validation because we're
  36073. 27:25:36crossing over our model from one batch
  36074. 27:25:38of data used to train it over to another
  36075. 27:25:41set of data used to validate those
  36076. 27:25:43predictions. Um, and there's actually
  36077. 27:25:47different ways to do cross validation.
  36078. 27:25:48So cross validation is a bit of an
  36079. 27:25:50umbrella term for multiple ways to do
  36080. 27:25:52that. We've already seen one way of
  36081. 27:25:54doing that um which I'm going to scroll
  36082. 27:25:56down to is um known as a hold out cross
  36083. 27:26:01validation. So that's um what we've been
  36084. 27:26:03doing so far. So this is just um
  36085. 27:26:06generating a train and a test set
  36086. 27:26:10train um split.
  36087. 27:26:13Um that's the that's what's known as the
  36088. 27:26:16hold out cross validation method. Um and
  36089. 27:26:19and this is exactly what we've been
  36090. 27:26:21doing so far, which is you split your
  36091. 27:26:24data into some type of split, usually
  36092. 27:26:267030,
  36093. 27:26:28um of a train and test
  36094. 27:26:31and then you um train your model on this
  36095. 27:26:34section of data and then apply it to
  36096. 27:26:37this to evaluate performance. Right? So
  36097. 27:26:39that's that's what's known as the hold
  36098. 27:26:41out method. Um it is uh you know
  36099. 27:26:46relatively simple. It's pretty fast to
  36100. 27:26:49do. Um, but there are more robust ways
  36101. 27:26:53to try to divide up our data a little
  36102. 27:26:56bit uh more evenly. Instead of just
  36103. 27:26:59having one split, we can actually do
  36104. 27:27:01many splits, which is the idea of um the
  36105. 27:27:04next kind of cross validation I'll
  36106. 27:27:06cover. But hold out method is one that
  36107. 27:27:09we've already studied. It's the most
  36108. 27:27:11basic type of cross validation you can
  36109. 27:27:13have. Um so hold out this is the most
  36110. 27:27:16basic
  36111. 27:27:18and we we've already been we've already
  36112. 27:27:21been uh working with this type. Okay.
  36113. 27:27:26So we've we've already seen hold out
  36114. 27:27:28method. Let me uh explain to you a more
  36115. 27:27:31sophisticated method a little bit more
  36116. 27:27:33advanced of a cross validation um which
  36117. 27:27:36is known as Kfold cross validation. So
  36118. 27:27:39this is um going to be a little bit more
  36119. 27:27:42advanced of a technique but this is the
  36120. 27:27:45idea of kfold is that you take your data
  36121. 27:27:47set
  36122. 27:27:49and you split it into k number of what
  36123. 27:27:54are called splits or folds. So you take
  36124. 27:27:57your data and you let's say it was let's
  36125. 27:27:59say k equals 5. So we have five splits
  36126. 27:28:02here.
  36127. 27:28:05Okay. So let's say k equals 5. we have
  36128. 27:28:08five splits. So what we're going to do
  36129. 27:28:13is we're going to we're going to train
  36130. 27:28:15our model on K minus one of those folds.
  36131. 27:28:19So if K was five, we had five splits.
  36132. 27:28:22We're going to take our model and train
  36133. 27:28:24it on four out of five of those uh
  36134. 27:28:28splits. So let's say it's these four.
  36135. 27:28:32We train it on these four.
  36136. 27:28:36Okay. And then what we do is the one
  36137. 27:28:39split that's left over we will we will
  36138. 27:28:43test our model against that split. So
  36139. 27:28:45we'll test here.
  36140. 27:28:50Okay. Now, this sounds very similar to
  36141. 27:28:52the hold out method where we're doing a
  36142. 27:28:54train test split, but it's a little bit
  36143. 27:28:56this kful cross validation is a little
  36144. 27:28:58bit more sophisticated because we repeat
  36145. 27:29:00this process that I just mentioned over
  36146. 27:29:03and over for all combinations of the
  36147. 27:29:06splits. So then what we'll do, this is
  36148. 27:29:08just one trial that we'll do it again,
  36149. 27:29:13but this time we will pick um four
  36150. 27:29:16different splits.
  36151. 27:29:18So, this time we might pick,
  36152. 27:29:21let me do blue. This time we might pick
  36153. 27:29:24this one, this one,
  36154. 27:29:27um,
  36155. 27:29:29this one,
  36156. 27:29:32and this one.
  36157. 27:29:35And then those four we will train our
  36158. 27:29:37data on. And then we will test against
  36159. 27:29:39this one. Okay? And we'll do we'll
  36160. 27:29:42repeat this
  36161. 27:29:45repeat for all combos of the folds.
  36162. 27:29:55Okay. So we'll repeat that. So
  36163. 27:29:58essentially what we're doing is rotating
  36164. 27:29:59through. Every time we rotate through
  36165. 27:30:02one of the folds is going to be left out
  36166. 27:30:03as a test set. Now this is a little bit
  36167. 27:30:07more robust than just a train test
  36168. 27:30:09split, right? because we are exposing
  36169. 27:30:12our model to more of the data in in
  36170. 27:30:15doing this, right? Because we're going
  36171. 27:30:17to split it evenly into five or 10
  36172. 27:30:20splits. Those are pretty common um
  36173. 27:30:22number of folds to use. 10 or five. Um
  36174. 27:30:25those are the ones I've most commonly
  36175. 27:30:27seen. Um but we're going to by rotating
  36176. 27:30:32through which folds are being used for
  36177. 27:30:33training, which ones being left out. um
  36178. 27:30:36we are exposing our our model to more of
  36179. 27:30:39the data this way than just doing a
  36180. 27:30:41single train test split. Right? So now
  36181. 27:30:44what do we do with with the results is
  36182. 27:30:47every time we do this we we generate um
  36183. 27:30:50an MSE let's say or some type of
  36184. 27:30:52performance metric. So let's say we
  36185. 27:30:54generate an MSE from this guy,
  36186. 27:30:57we generate an MSE from this version and
  36187. 27:31:00we generate an MSE for all combos.
  36188. 27:31:04each combo we generate MSE and then what
  36189. 27:31:07we do is we average
  36190. 27:31:10the metrics
  36191. 27:31:13or the in this case uh if we use MSE we
  36192. 27:31:16would average those together. So every
  36193. 27:31:19time we do a fold combination and we
  36194. 27:31:21keep four of them for training, one for
  36195. 27:31:23test and we rotate through all those
  36196. 27:31:25combinations, we are going to generate
  36197. 27:31:27an MSE for every combination
  36198. 27:31:30then we're just going to average those
  36199. 27:31:32MSE's to get a final. So the final MSE
  36200. 27:31:36of cross val of this kffold.
  36201. 27:31:40So the final metric
  36202. 27:31:43is just the average of the uh
  36203. 27:31:46performance on all of the fold
  36204. 27:31:48combinations. Okay. So our final MSE, we
  36205. 27:31:52just average all those MSE's from all of
  36206. 27:31:54our combinations.
  36207. 27:31:56Okay.
  36208. 27:31:58Now, what's the advantage to doing this?
  36209. 27:32:01It's way more robust of a estimate of
  36210. 27:32:04the of the performance of the model
  36211. 27:32:06because we're exposing it to all
  36212. 27:32:09basically all of our data, right? We're
  36213. 27:32:11getting a sense of how it performs
  36214. 27:32:12across all those different folds. Um
  36215. 27:32:16rather than just doing a single train
  36216. 27:32:18test split, which is a bit it's basic,
  36217. 27:32:20it works, but it's a bit basic. Um so
  36218. 27:32:23this is more robust estimate of the
  36219. 27:32:26performance.
  36220. 27:32:28Now, what's the drawback to doing this
  36221. 27:32:30is that it's more intensive. So, if you
  36222. 27:32:32have a lot of data, this is going to be
  36223. 27:32:34pretty expensive to do because you're
  36224. 27:32:36going to have to especially you have a
  36225. 27:32:37high number of folds, right? You're
  36226. 27:32:39going to have to divide your data into k
  36227. 27:32:41number of folds and you're going to have
  36228. 27:32:43to do this over and over again. Um, and
  36229. 27:32:45if it's a large data set, it might take
  36230. 27:32:47your model a long time to train. It's
  36231. 27:32:49going to be a little bit more uh
  36232. 27:32:52computationally intense than if we just
  36233. 27:32:55did a train test split.
  36234. 27:32:57Okay, we just did a single like 7030
  36235. 27:32:59split. We only do that once. We only
  36236. 27:33:02train the model once, right? We train it
  36237. 27:33:04on the 70, apply it to the 30% test data
  36238. 27:33:08and evaluate performance that way. Um,
  36239. 27:33:11so we're only really using the model and
  36240. 27:33:14training the model once, but in this
  36241. 27:33:16kfold, we're going to do it um, you
  36242. 27:33:19know, k number of times essentially
  36243. 27:33:23or I should say one for every
  36244. 27:33:24combination that we have to work through
  36245. 27:33:27of of all the folds.
  36246. 27:33:32Okay.
  36247. 27:33:35All right. Does that make sense? Any any
  36248. 27:33:38questions on Kfold cross validation? So
  36249. 27:33:41K K K K K K K K K K K K K K K K K K K K
  36250. 27:33:42K K K K K K K K K K K K K K K K K K K K
  36251. 27:33:42K is an important uh number here. It
  36252. 27:33:45it's how many folds, how many splits do
  36253. 27:33:48you have? A typical value for K is going
  36254. 27:33:50to be somewhere like five or 10.
  36255. 27:33:54So 10 folds or five folds. Those are
  36256. 27:33:57pretty pretty standard
  36257. 27:34:01from what from what I've seen.
  36258. 27:34:05But does the does the concept make sense
  36259. 27:34:07or is there any questions on it on in
  36260. 27:34:09terms of um you're always going to leave
  36261. 27:34:11one fold out. You're going to split it
  36262. 27:34:13up into K number of folds. Always leave
  36263. 27:34:15one out. Train on the rest of it.
  36264. 27:34:18Evaluate on that one that gets left out
  36265. 27:34:19and then rotate those through. And
  36266. 27:34:22you're going to do that for every
  36267. 27:34:23combination and average all those
  36268. 27:34:25metrics.
  36269. 27:34:32And by the way, there's going to be an
  36270. 27:34:33easy function in scikitlearn that will
  36271. 27:34:36do this for us. So managing all these
  36272. 27:34:38combinations will be really easy. It's
  36273. 27:34:41actually just built into scikitlearn. So
  36274. 27:34:43we don't have to um we don't have to do
  36275. 27:34:46this all by hand. Okay, this will be in
  36276. 27:34:48scikitlearn. It'll handle doing all
  36277. 27:34:50these combinations of folds for us and
  36278. 27:34:53computing the average metric will be
  36279. 27:34:55really easy. So um
  36280. 27:34:59we don't have to worry about that. We're
  36281. 27:35:00going to see an example of this coming
  36282. 27:35:02up shortly.
  36283. 27:35:04All right, of kfold cross validation.
  36284. 27:35:08But this is a this is a really widely
  36285. 27:35:10used technique. And again like the
  36286. 27:35:12purpose you may be wondering like what's
  36287. 27:35:13the purpose ultimately of doing this?
  36288. 27:35:15It's to get a sense of if our model is
  36289. 27:35:18going to perform well on new data.
  36290. 27:35:20That's really what we want to know. like
  36291. 27:35:22is the model going to perform well when
  36292. 27:35:23I start to use it on new data that it's
  36293. 27:35:26never seen before and this kffold is a
  36294. 27:35:30decent indicator of that because we are
  36295. 27:35:34varying which data it sees across many
  36296. 27:35:37different folds right so it's a it's
  36297. 27:35:40kind of a good um proxy to exposing it
  36298. 27:35:44to different kinds of data each time and
  36299. 27:35:46seeing how it performs
  36300. 27:35:49right all right because we're working
  36301. 27:35:50our way through each one of the folds
  36302. 27:35:51there's always going to be one fold left
  36303. 27:35:53out. We're going to change which fold
  36304. 27:35:55gets left out each time. And um that's
  36305. 27:35:58sort of mimicking the idea of we're
  36306. 27:36:00going to apply our model to new data and
  36307. 27:36:02see how it performs. And it's it's new
  36308. 27:36:05data every fold.
  36309. 27:36:25um how we know which model is best suits
  36310. 27:36:29for which scenario because we have Yeah,
  36311. 27:36:32that's a good question. Um, so my we're
  36312. 27:36:36going to learn this as we go along
  36313. 27:36:38because we haven't covered all the
  36314. 27:36:39models yet, but generally the best
  36315. 27:36:43advice I can give on that is
  36316. 27:36:46you you generally want to start as
  36317. 27:36:49simple as you can get and then if it's
  36318. 27:36:51not performing well then work your way
  36319. 27:36:53up to something more complex.
  36320. 27:36:56So we are going to have models that are
  36321. 27:36:58simpler. We're going to have models that
  36322. 27:37:00are more complex. The rule of thumb is
  36323. 27:37:02to start with the most simple model that
  36324. 27:37:04works.
  36325. 27:37:07So you're usually going to have the same
  36326. 27:37:10ones that you're going to try in the
  36327. 27:37:12beginning. And linear regression is a
  36328. 27:37:14very simple model. It's usually the
  36329. 27:37:16first one you want to try for regression
  36330. 27:37:18because it's the simplest.
  36331. 27:37:20Um, and for classification, we're going
  36332. 27:37:22to have a similar like logistic
  36333. 27:37:24regression is the simplest kind of
  36334. 27:37:26classification model we could have. So
  36335. 27:37:29usually want to start with that and then
  36336. 27:37:31if it underfits like if we see it's
  36337. 27:37:34producing a lot of error then we work
  36338. 27:37:37our way up to a more sophisticated
  36339. 27:37:39model.
  36340. 27:37:41So um that's the way we that's the way
  36341. 27:37:45it should usually go is simple to
  36342. 27:37:47complex it based on their performance.
  36343. 27:37:50So we evaluate it and then we can repeat
  36344. 27:37:52the process. If it's not performing well
  36345. 27:37:53we can try something different that's
  36346. 27:37:55more complex if it's underfitting.
  36347. 27:38:05Uh this is a good question. Does a model
  36348. 27:38:06reset after training each k minus one
  36349. 27:38:09fold? Um yeah, it's essentially like a
  36350. 27:38:12blank model every time uh every fold. So
  36351. 27:38:15um we imagine like you have a brand you
  36352. 27:38:19have a fresh model every um k minus one
  36353. 27:38:22combination. Yes.
  36354. 27:38:30And the reason the reason it has to be
  36355. 27:38:32that way is because you don't want the
  36356. 27:38:35other folds influencing the model that
  36357. 27:38:39like on on the next combination. You
  36358. 27:38:42don't want the previous combination to
  36359. 27:38:43influence the results on the next one,
  36360. 27:38:45right? Um you want it to be a fresh
  36361. 27:38:48evaluation on every combination of
  36362. 27:38:51folds.
  36363. 27:39:10Okay.
  36364. 27:39:12All right. So, let me describe to you a
  36365. 27:39:15variation on what we just um talked
  36366. 27:39:18about with the K-fold. So, there's
  36367. 27:39:20another cross validation known as
  36368. 27:39:22stratified K-fold. And um this is the
  36369. 27:39:26same exact procedure as kfold except
  36370. 27:39:29that when we this is used for
  36371. 27:39:31classification.
  36372. 27:39:32Um so when we do classification
  36373. 27:39:35uh we want to make sure that the
  36374. 27:39:37different categories are going to be um
  36375. 27:39:40split amongst those folds in a
  36376. 27:39:43proportional way. So we don't what we
  36377. 27:39:45don't want to happen is um when we split
  36378. 27:39:48apart the data. So, let's say we have
  36379. 27:39:50let's say we're predicting um spam not
  36380. 27:39:53spam. What we don't want to have happen
  36381. 27:39:56when we do our splits is we don't want
  36382. 27:39:58to have all of the spams end up in one
  36383. 27:40:01fold and then every other fold has no
  36384. 27:40:04spam, no spam, no spam, no spam, right?
  36385. 27:40:08That's not very good. Um because if we
  36386. 27:40:11if we train against all these guys, we
  36387. 27:40:13have no shot at predicting spam when
  36388. 27:40:16they've never seen spam before. So
  36389. 27:40:18stratify kayfold is is used in
  36390. 27:40:20classification
  36391. 27:40:23and it's to um it's to make our splits
  36392. 27:40:27ensure that they have basically a
  36393. 27:40:29balanced number of categories for each
  36394. 27:40:32split. Um so that we don't end up with
  36395. 27:40:34certain splits with way more spams than
  36396. 27:40:37not spams. Um so we we do what's called
  36397. 27:40:40stratifying where we make sure the
  36398. 27:40:42proportions are balanced across each uh
  36399. 27:40:45split. So this is only really useful in
  36400. 27:40:47classification, not really necessary in
  36401. 27:40:50regression because we're predicting a
  36402. 27:40:51value. But if we were predicting a
  36403. 27:40:53category,
  36404. 27:40:55like in classification like fraud, not
  36405. 27:40:58fraud, we don't want to do the split and
  36406. 27:41:00have every single fraud example um by
  36407. 27:41:03bad luck in our shuffling and split end
  36408. 27:41:05up in one split and every other um every
  36409. 27:41:09other split has no examples of fraud.
  36410. 27:41:11Right? So we want to stratify this to
  36411. 27:41:13spread out those um frauds against all
  36412. 27:41:16the other splits. Um so uh again um
  36413. 27:41:21scikitlearn will take care of that for
  36414. 27:41:23you. Um but if you're doing
  36415. 27:41:24classification and you have an
  36416. 27:41:26imbalanced data set um you you really
  36417. 27:41:29want to make sure you stratify k-fold.
  36418. 27:41:32um imbalanced meaning that you have a a
  36419. 27:41:36um different number. Like if you're
  36420. 27:41:38doing fraud, not fraud, you have way
  36421. 27:41:39more not frauds than frauds. Um where
  36422. 27:41:42where that category is imbalanced,
  36423. 27:41:45you want to make sure it's balanced
  36424. 27:41:46across all your splits.
  36425. 27:41:49Um so this is this is useful in
  36426. 27:41:52classification only, not really
  36427. 27:41:53regression, which is what we're talking
  36428. 27:41:54about right now. Um but it's just a
  36429. 27:41:57variation on this that ensures when we
  36430. 27:41:59do those folds um the data is
  36431. 27:42:02distributed evenly amongst those folds
  36432. 27:42:04as much as we can. The labels are I
  36433. 27:42:06should say.
  36434. 27:42:08Okay. So that's stratified kfold. It's
  36435. 27:42:12the same same procedure once we have our
  36436. 27:42:14splits. It's the same where we do k
  36437. 27:42:16minus one of them. We train test on that
  36438. 27:42:18last fold um and then rotate through all
  36439. 27:42:22the folds and and average all the
  36440. 27:42:24metrics. the same exact procedure. It's
  36441. 27:42:26just the splitting itself um is going to
  36442. 27:42:29be balanced in a stratified kfold.
  36443. 27:42:35Okay. So, hold out we've already talked
  36444. 27:42:36about um is just doing a single train
  36445. 27:42:39test split. We've talked about that. One
  36446. 27:42:42more variation that is a bit of an
  36447. 27:42:44extreme version of K-fold. So it's
  36448. 27:42:46actually the same process as Kfold, but
  36449. 27:42:48it's an extreme version is if you set K
  36450. 27:42:52equal to the number of data points. So
  36451. 27:42:54you basically are um this is a really
  36452. 27:42:57really extreme kfold where you um
  36453. 27:43:00basically are training on all the data.
  36454. 27:43:03Um so you're training on all the data
  36455. 27:43:06except one point and then you test
  36456. 27:43:10against that one point. Um now why would
  36457. 27:43:13you ever do this? Um it's mainly so for
  36458. 27:43:16this reason here. It's to um maximize
  36459. 27:43:20the amount of training data that your
  36460. 27:43:22model gets exposed to because instead of
  36461. 27:43:24just doing instead of just doing five
  36462. 27:43:26splits
  36463. 27:43:28um which would be like
  36464. 27:43:31you know these four folds are going to
  36465. 27:43:33be used and then we um test against one
  36466. 27:43:36fold. um we're essentially going to use
  36467. 27:43:3999% of the data, right? One point is
  36468. 27:43:42going to be left out. 99% of the data
  36469. 27:43:45gets used to train. Um and then we're
  36470. 27:43:47always going to leave out one point. And
  36471. 27:43:49and the issue is we're actually going to
  36472. 27:43:51do that over and over and over again and
  36473. 27:43:53rotate that one point to cover the whole
  36474. 27:43:55data set. So, we're going to train on
  36475. 27:43:5899%, leave one that one point out,
  36476. 27:44:01and then rotate through every
  36477. 27:44:03combination of points until we've left
  36478. 27:44:05out every single point, and then average
  36479. 27:44:08all those together. Um, so this is a
  36480. 27:44:10this is an extreme kfold. Again, the
  36481. 27:44:13number of folds is actually equal to the
  36482. 27:44:15number of data points in this case. So,
  36483. 27:44:16we have every point is its own fold and
  36484. 27:44:19we train on everything but one. Test on
  36485. 27:44:23that one. This gets you the maximum size
  36486. 27:44:26of your training data because you're
  36487. 27:44:28basically gonna have every point but one
  36488. 27:44:30used in the training.
  36489. 27:44:32This gets you the maximum size. However,
  36490. 27:44:34it gets you the maximum uh expense
  36491. 27:44:38especially for large data sets. This is
  36492. 27:44:40going to be usually you're not going to
  36493. 27:44:42use this um especially for large data
  36494. 27:44:44sets because it's just too extreme. It's
  36495. 27:44:48going to take you a really long time to
  36496. 27:44:49work through every single point being
  36497. 27:44:52left out. um it's just going to take a
  36498. 27:44:55while to do.
  36499. 27:44:57So, for that reason, the leave one out
  36500. 27:45:00um that that's why it's called leave one
  36501. 27:45:02out because it's you're leaving one out
  36502. 27:45:04every single time. Um is rarely used. I
  36503. 27:45:08I have don't really see it used that
  36504. 27:45:09often, but it is an extreme version of
  36505. 27:45:12kful cross validation.
  36506. 27:45:16Okay. But rarely ever actually used. I
  36507. 27:45:19think the the ones that get used the
  36508. 27:45:20most are definitely the hold out method
  36509. 27:45:22with just a regular train test split. Um
  36510. 27:45:25and then uh the other one that gets used
  36511. 27:45:27quite a bit is is kfold
  36512. 27:45:30or stratified kfold if you're if you're
  36513. 27:45:32doing classification,
  36514. 27:45:34but certainly kfold in the in a
  36515. 27:45:36regression case.
  36516. 27:45:38Okay.
  36517. 27:45:43All right. Um we're going to do an
  36518. 27:45:45example with these guys. So we'll do
  36519. 27:45:47that next.
  36520. 27:45:48um with with the different cross
  36521. 27:45:50validation techniques. Um but any
  36522. 27:45:54questions on what they are doing
  36523. 27:45:57conceptually before we actually do the
  36524. 27:45:59code example.
  36525. 27:46:16Okay.
  36526. 27:46:18Very good.
  36527. 27:46:25All right. So, let's see some examples.
  36528. 27:46:27Um, let's go into our code and build a
  36529. 27:46:32model and do the different cross
  36530. 27:46:34validation techniques on it. Um, you're
  36531. 27:46:37going to see it's actually going to be
  36532. 27:46:38really easy to do and we it sounds
  36533. 27:46:40complex like doing the kfold and leaving
  36534. 27:46:43one out and testing. It sounds kind of
  36535. 27:46:45complex, but I promise you scikitlearn
  36536. 27:46:47makes it really easy to do. Um,
  36537. 27:46:51and so, uh, we won't need to do too much
  36538. 27:46:54besides just use the right, uh, tools
  36539. 27:46:57from scikitlearn. Uh, so we're going to
  36540. 27:46:59we're going to see that. Um, so here we
  36541. 27:47:01have some imports. The, um, primary, uh,
  36542. 27:47:05thing that's a little bit new for us is
  36543. 27:47:07going to be these, um, different kinds
  36544. 27:47:09of cross validation techniques. So we
  36545. 27:47:11have our kfold, we have our stratified
  36546. 27:47:13kfold, leave one out. Um, which are
  36547. 27:47:15those different cross validation
  36548. 27:47:17techniques. Um, these are going to be
  36549. 27:47:19used in combination with this cross val
  36550. 27:47:24score which is going to keep track of
  36551. 27:47:27the different um metrics and then
  36552. 27:47:29average them
  36553. 27:47:31uh while we do one of these um cross
  36554. 27:47:35validation techniques. So this guy gets
  36555. 27:47:38used in combination with one of these to
  36556. 27:47:42um as as we're going to see in the code
  36557. 27:47:44uh to average those metrics um doing the
  36558. 27:47:48different folds, right? Perform doing
  36559. 27:47:49performance against the different folds.
  36560. 27:47:52Okay. And then of course we need a model
  36561. 27:47:55using a linear regression. That's that's
  36562. 27:47:57the one we've studied so far. Um and
  36563. 27:48:00then we have just a regular metrics. If
  36564. 27:48:02we want to compute those um using maybe
  36565. 27:48:05just hold out, right? And hold out um
  36566. 27:48:08which which is just a regular train test
  36567. 27:48:10split um we could use these guys to
  36568. 27:48:12evaluate performance.
  36569. 27:48:14But in a more sophisticated kfold style
  36570. 27:48:17of cross validation, we're going to use
  36571. 27:48:19this to evaluate the the performance.
  36572. 27:48:25Okay, let's see.
  36573. 27:48:28So, we're going to be working with this
  36574. 27:48:30housing with ocean proximity data. Um,
  36575. 27:48:33you guys should have this one. Uh,
  36576. 27:48:37so you guys should have this one. So, if
  36577. 27:48:40you want to follow along and run it
  36578. 27:48:41yourself, um, you can load that one in.
  36579. 27:48:46Um, I want to make sure that I have it.
  36580. 27:48:51Let me pull that one in. So, it should
  36581. 27:48:52be this guy.
  36582. 27:49:03I'm going to load that in so I can make
  36583. 27:49:05sure I run it with you guys.
  36584. 27:49:08Um,
  36585. 27:49:17let me run this.
  36586. 27:49:21Do you guys have that data?
  36587. 27:49:24the housing with ocean proximity.
  36588. 27:49:28It's another it's another housing data
  36589. 27:49:30set. Um
  36590. 27:49:33but it it's a little bit different than
  36591. 27:49:35the ones we've seen before. It has a a
  36592. 27:49:38special feature for how close it is to
  36593. 27:49:40the ocean at different locations.
  36594. 27:49:50So it looks kind of like this. If we
  36595. 27:49:52load it in and do our head, which is
  36596. 27:49:54usually what we do, right? We can see um
  36597. 27:49:57we can see that it's got these features.
  36598. 27:50:00So it's got uh uh bedrooms, total rooms,
  36599. 27:50:05um it's got uh median age. Now this is
  36600. 27:50:08this is looks a little strange for total
  36601. 27:50:10rooms and um uh bedrooms and population
  36602. 27:50:15etc. But it's um
  36603. 27:50:19it's it's got those uh it's got those
  36604. 27:50:22because it's representing an entire
  36605. 27:50:24neighborhood. So it's an entire
  36606. 27:50:27neighborhood. And we're looking at this
  36607. 27:50:29um this is actually going to be our
  36608. 27:50:30label is this median house value for the
  36609. 27:50:32entire neighborhood. So what's that
  36610. 27:50:34median value uh in the neighborhood? And
  36611. 27:50:38this is the total number of bedrooms,
  36612. 27:50:40total number of rooms, um population,
  36613. 27:50:43households. So, how many houses are
  36614. 27:50:46there? Um, median income. And of course,
  36615. 27:50:49these are scaled. So, these are um
  36616. 27:50:51likely times, you know, uh thousands. Um
  36617. 27:50:58but um that's our data. We could
  36618. 27:51:02describe it.
  36619. 27:51:07So we can see the average age, average
  36620. 27:51:10median age. Um which sounds a little um
  36621. 27:51:13weird, but that's it's because again
  36622. 27:51:15this is the median of data within a
  36623. 27:51:18neighborhood. Um so the average of those
  36624. 27:51:22is about 28 or 29. Um we have
  36625. 27:51:27u
  36626. 27:51:31total bedrooms. The we can look at the
  36627. 27:51:33men. There's some data that only has
  36628. 27:51:36one. So, it's likely only one house in
  36629. 27:51:38there. Um, which is what this
  36630. 27:51:41represents. There's only one house. So,
  36631. 27:51:43there there is some neighborhood that
  36632. 27:51:45only has one house. Um, and we see the
  36633. 27:51:48median um we see the minimum uh median
  36634. 27:51:53house values there. And then the maximum
  36635. 27:51:55down here um is a pretty big number.
  36636. 27:52:016,000 households is the largest that we
  36637. 27:52:03have in any any one of these
  36638. 27:52:05neighborhoods.
  36639. 27:52:07Okay. So, just a little bit of
  36640. 27:52:09description of the data.
  36641. 27:52:25Okay. So, then we can run.info. So, this
  36642. 27:52:28is um let me ask you guys, were you able
  36643. 27:52:30to load this? Were you able to run this?
  36644. 27:52:36If you're following along, were you able
  36645. 27:52:38to
  36646. 27:52:39load it and take a look at
  36647. 27:52:49Okay, great. Great.
  36648. 27:52:52Okay, so we're able to load that and
  36649. 27:52:55then look at head. Perfect. Um
  36650. 27:53:00Okay.
  36651. 27:53:02Um and then we run describe which gives
  36652. 27:53:05us that uh usual kind of statistical
  36653. 27:53:07description. Uh so we can see some
  36654. 27:53:10interesting stats about those.
  36655. 27:53:18What do you guys notice about the info?
  36656. 27:53:20Anything interesting that we see from
  36657. 27:53:22there?
  36658. 27:53:36Is there any missing data
  36659. 27:53:42any features that have missing data? Can
  36660. 27:53:44we see
  36661. 27:53:47object? Yeah, object type usually is
  36662. 27:53:49string. If it's an object type, that
  36663. 27:53:51usually means string. Python when we
  36664. 27:53:54read it into pandas it usually is just a
  36665. 27:53:56string.
  36666. 27:53:59So that that makes sense like we have
  36667. 27:54:00mostly numerical features but then we
  36668. 27:54:03have a this ocean proximity which is a
  36669. 27:54:05string.
  36670. 27:54:14Yeah. Total bedrooms has nles. That's
  36671. 27:54:16right. Because you can see here this
  36672. 27:54:18does not equal the number of uh rows
  36673. 27:54:21that we have. So this is the number of
  36674. 27:54:23rows which about 20,000 rows. That's a
  36675. 27:54:24good size data set, right? 20,000 rows.
  36676. 27:54:27That's decent. Um we're definitely
  36677. 27:54:30missing some data here for sure. Um we
  36678. 27:54:34could count how much we're missing
  36679. 27:54:35exactly by running this is NATO sum. Um
  36680. 27:54:40and so we see that total bedrooms is
  36681. 27:54:42missing about 200 uh 200 rows are
  36682. 27:54:46missing total bedroom uh value.
  36683. 27:54:52Okay. And then one thing I wanted to
  36684. 27:54:54look at is yes, this is a string. So
  36685. 27:54:56what remember what we can do with those?
  36686. 27:54:58That's a categorical.
  36687. 27:55:01So ocean proximity
  36688. 27:55:06is a categorical
  36689. 27:55:09string
  36690. 27:55:11feature.
  36691. 27:55:13So we can take a look at its value
  36692. 27:55:15counts, which is usually a good idea to
  36693. 27:55:17take a look and see what possible values
  36694. 27:55:20that feature could be. So if we look at
  36695. 27:55:23our
  36696. 27:55:25um what are we calling this? Housing
  36697. 27:55:27data.
  36698. 27:55:31housing data
  36699. 27:55:34ocean
  36700. 27:55:36proximity
  36701. 27:55:42value counts.
  36702. 27:55:49So, here's the different types that that
  36703. 27:55:51one can be. So, there's some
  36704. 27:55:53neighborhoods that are less than 1 hour
  36705. 27:55:54from the ocean. There's some that are
  36706. 27:55:56inland. There's some that are near the
  36707. 27:55:58ocean. There's some that are near a bay.
  36708. 27:56:01There's even five of them that are on an
  36709. 27:56:03island. So, these are the different
  36710. 27:56:05values of the ocean proximity. So,
  36711. 27:56:08remember, you can always do that. If you
  36712. 27:56:09see a string feature, you can always
  36713. 27:56:12take a look at what its um categories
  36714. 27:56:14are. And it looks like most things are
  36715. 27:56:17less than 1 hour from the ocean, but
  36716. 27:56:18it's kind of evenly distributed here. Um
  36717. 27:56:22otherwise
  36718. 27:56:25very few islands.
  36719. 27:56:33But as you can imagine like this feature
  36720. 27:56:34is probably going to be important for
  36721. 27:56:36determining um what the value is, right?
  36722. 27:56:40Probably going to be important.
  36723. 27:56:47Okay. So, um, we need to deal with these
  36724. 27:56:51NLES. If we're going to build a model,
  36725. 27:56:53right? So, um, this is all of our
  36726. 27:56:55typical data prep. If we want to build a
  36727. 27:56:57model, we're going to have to deal with
  36728. 27:56:58these NLES. What do you guys think we
  36729. 27:57:00should do with the NLES? What would you
  36730. 27:57:03what do you think for total bedrooms?
  36731. 27:57:05What do you think is a good strategy to
  36732. 27:57:06do? Keep in mind, we have 20,000 points,
  36733. 27:57:1220,000 rows I should say, and about 200
  36734. 27:57:15of them are null.
  36735. 27:57:22Right. So about 200 are null. Um so what
  36736. 27:57:27do you what do you guys think would be
  36737. 27:57:28like a good strategy to deal with those
  36738. 27:57:29NLES in that case?
  36739. 27:57:51average. We can't ignore it because we
  36740. 27:57:55can't ignore that column.
  36741. 27:57:58We can't ignore the whole column. So,
  36742. 27:58:00something needs to go there.
  36743. 27:58:09Probably don't want to make it zero.
  36744. 27:58:11I think average is a decent average is a
  36745. 27:58:14decent idea. Probably don't want to make
  36746. 27:58:16it zero because um that would indicate
  36747. 27:58:19that there's no bedrooms and yet we
  36748. 27:58:21still have a bunch of total rooms. So it
  36749. 27:58:24probably doesn't make sense to do zero.
  36750. 27:58:30Average, I think average could be a
  36751. 27:58:32decent one.
  36752. 27:58:34Now in this example, what we're actually
  36753. 27:58:37going to do is we're
  36754. 27:58:41rows.
  36755. 27:58:43We're actually going to drop the rows al
  36756. 27:58:45together. Now, why are we doing that?
  36757. 27:58:46It's because we have so much data and
  36758. 27:58:50only 200 of them are null.
  36759. 27:58:54Okay, only 200 of them are null. So,
  36760. 27:58:56we're actually just going to drop the
  36761. 27:58:57rows. Now, that's a choice.
  36762. 27:59:02Um, that's a choice, right? Is that we
  36763. 27:59:05could fill in with the average like you
  36764. 27:59:07guys are suggesting. What we're actually
  36765. 27:59:09going to do is just drop the rows. It it
  36766. 27:59:11makes up less. It makes up about 1% of
  36767. 27:59:15the whole data. So it's not that much of
  36768. 27:59:18it is missing. We can drop those rows.
  36769. 27:59:21So that's actually what we're going to
  36770. 27:59:22do here is we remove all the roles with
  36771. 27:59:26the NLES by doing drop NA. So this just
  36772. 27:59:28drops them. So those rows are cut out.
  36773. 27:59:31Um, it's arguable that we could replace
  36774. 27:59:37it's arguable that we could just replace
  36775. 27:59:38it with something and I think you guys
  36776. 27:59:40have good thoughts which is the average
  36777. 27:59:42a default
  36778. 27:59:45um assume total bedrooms. We could we
  36779. 27:59:47could try that. Yeah.
  36780. 27:59:51Assign a value based on comparable home
  36781. 27:59:53value. Yes, you could do that too.
  36782. 27:59:54That's a good strategy is to look at the
  36783. 27:59:57other rows that are similar to it and
  36784. 27:59:59fill in a value. That's absolutely fair.
  36785. 28:00:02Um, in this example, we're actually just
  36786. 28:00:04going to drop those rows,
  36787. 28:00:09but I think that's totally um totally
  36788. 28:00:11valid.
  36789. 28:00:13This is a choice.
  36790. 28:00:16We could fill NA with different values
  36791. 28:00:22such as the average
  36792. 28:00:26total bedrooms
  36793. 28:00:29um derive a value etc. So we could
  36794. 28:00:33derive something which I think Brent you
  36795. 28:00:35have a good suggestion that's a good
  36796. 28:00:37suggestion. Um we could derive something
  36797. 28:00:40like that uh and fill in the blank and
  36798. 28:00:42that's I think that's totally valid. Um,
  36799. 28:00:44we could take the average of the um
  36800. 28:00:48bedrooms. Uh, I meant total rooms here.
  36801. 28:00:51Sorry, total rooms. Um, we could fill in
  36802. 28:00:54we could fill it in with the total rooms
  36803. 28:00:56for that category um or for that row.
  36804. 28:01:00Um, many options. In this case, we're
  36805. 28:01:03actually just going to drop those rows
  36806. 28:01:05because they make up such a small
  36807. 28:01:07percentage relative to the 20,000 rows
  36808. 28:01:10that we have. It's about 1%. Right? 200
  36809. 28:01:14rows is about 1% of 20,000.
  36810. 28:01:17So, we're just going to drop them. But
  36811. 28:01:19that's a choice. We don't have to drop
  36812. 28:01:21them. We could fill in with something.
  36813. 28:01:24Um, and if we did that, we would use
  36814. 28:01:26fill NA rather than drop NA, right?
  36815. 28:01:42Uh after dropping the rows, how many? So
  36816. 28:01:45it's just so after we drop the rows, um
  36817. 28:01:48after we drop the rows, it's just going
  36818. 28:01:50to be we still have all our other rows
  36819. 28:01:53are intact, right? So if we look at this
  36820. 28:01:54now,
  36821. 28:02:03we now have um slightly uh slightly less
  36822. 28:02:07entries.
  36823. 28:02:09So now we have this this many um rather
  36824. 28:02:12than rather than this many,
  36825. 28:02:16right? We dropped those 200
  36826. 28:02:25But they're all filled in. Yeah, they're
  36827. 28:02:27So all the other columns are still
  36828. 28:02:28filled in. We're just we're we're
  36829. 28:02:30cutting out the whole row. So if you
  36830. 28:02:32think about our data set, um we have all
  36831. 28:02:35these rows and all these columns. What
  36832. 28:02:38we're doing is like if there's a null
  36833. 28:02:40here, we're just we're just getting rid
  36834. 28:02:42of that whole row, right? And so we
  36835. 28:02:45still have all the other rows intact.
  36836. 28:02:54Uh, we can drop them because we have a
  36837. 28:02:56good sample size. Yes,
  36838. 28:03:02that's exactly right, Ronald. Yep, we
  36839. 28:03:04can drop them because we have we have
  36840. 28:03:0520,000 rows and only 200 are missing
  36841. 28:03:08values. So, that's totally fine.
  36842. 28:03:15Uh, drop a removes all rows that has any
  36843. 28:03:17null. Yes, that's true. It it will go
  36844. 28:03:20ahead and just drop any row where
  36845. 28:03:22there's any null, no matter what column
  36846. 28:03:24it's in. Yes,
  36847. 28:03:35index. Yeah, the index is not getting
  36848. 28:03:37reset. Um, that's true. So, um, what we
  36849. 28:03:43what you can always do is you can reset
  36850. 28:03:45the index. So, um, if you want to, it's
  36851. 28:03:48optional. We we're not really going to
  36852. 28:03:50use the index for anything that
  36853. 28:03:52important, right? But what we could do
  36854. 28:03:54is, uh, reset index.
  36855. 28:03:59Uh,
  36856. 28:04:01we could do that, right? Which will
  36857. 28:04:02reset it.
  36858. 28:04:17So now now it gets reset.
  36859. 28:04:36But um let me actually I don't I don't
  36860. 28:04:39really want to do that. I'm going to
  36861. 28:04:41reset this.
  36862. 28:04:48Um,
  36863. 28:04:57yeah, we could do that.
  36864. 28:05:18Okay.
  36865. 28:05:20So now importantly there should be uh no
  36866. 28:05:23missing data of this of this new one
  36867. 28:05:25where we've dropped NAS. Right. So now
  36868. 28:05:27this is good. If you now the reason we
  36869. 28:05:29had to do this is because if we try to
  36870. 28:05:31build a linear regression and we have
  36871. 28:05:33NLES in there. Um the the issue is like
  36872. 28:05:37how do you build a model where you have
  36873. 28:05:40something like this
  36874. 28:05:48and these are null? Like what do how do
  36875. 28:05:51you multiply a number by a null?
  36876. 28:05:54Um we can't really do that, right?
  36877. 28:05:58we can't really do that. So, um,
  36878. 28:06:07so therefore, uh, we need to get rid of
  36879. 28:06:10NLES like the the null is not really
  36880. 28:06:12going to work in there. So, uh, we need
  36881. 28:06:15to get rid of them for linear regression
  36882. 28:06:17to to really have a chance to work,
  36883. 28:06:18right? To train it and be able to use
  36884. 28:06:20it.
  36885. 28:06:22You got to get rid of those nles.
  36886. 28:06:30All right,
  36887. 28:06:33any questions so far? So, we haven't
  36888. 28:06:34done any modeling yet. We're doing some
  36889. 28:06:35We're doing some data preparation before
  36890. 28:06:37we get to the modeling. And we haven't
  36891. 28:06:39done any cross validation yet. We
  36892. 28:06:41haven't set that up. We're just doing
  36893. 28:06:42our data preparation before we get to
  36894. 28:06:44the modeling. Right? So, we've dropped
  36895. 28:06:47some NAS. We've checked it. Um, we're
  36896. 28:06:50going to do one more prep step, which is
  36897. 28:06:52to um change that ocean proximity
  36898. 28:06:57feature into something numerical because
  36899. 28:06:58again, how do you build a model where
  36900. 28:07:00you're inserting a string into those
  36901. 28:07:03like beta 1, beta 2, beta 3 times of
  36902. 28:07:05features? You can't really do that when
  36903. 28:07:07it's a string. Um, so what we're going
  36904. 28:07:10to do, I'm going to get rid of this
  36905. 28:07:12because I don't think we really need
  36906. 28:07:13that. um is we are going to uh run this
  36907. 28:07:18get dummies function which is our um our
  36908. 28:07:22get dummies function is our usual one to
  36909. 28:07:26uh our git dummies one is our usual one
  36910. 28:07:29to um
  36911. 28:07:32uh get our one hot encoding.
  36912. 28:07:34So this is our uh one hot encoding here.
  36913. 28:07:42We now are going to have data that's
  36914. 28:07:45like this, right? So we have ocean. So
  36915. 28:07:48So by the way, this prefix
  36916. 28:07:50um this prefix is OP, which which is
  36917. 28:07:54short for ocean proximity, right? So we
  36918. 28:07:56have ocean proximity uh less than 1 hour
  36919. 28:07:59from the ocean, ocean proximity inland,
  36920. 28:08:02ocean proximity island, near bay, near
  36921. 28:08:04ocean. So these first five rows are near
  36922. 28:08:07the bay. Um so they have a one there and
  36923. 28:08:10a zero in the other spots. So this is
  36924. 28:08:12good. This one hot encodes that feature
  36925. 28:08:15into these numerical uh values,
  36926. 28:08:19right?
  36927. 28:08:21Were you guys able to run that one? they
  36928. 28:08:24get dummies.
  36929. 28:08:31So the reason that Yeah, that's a great
  36930. 28:08:33question. How did it go ocean proximity?
  36931. 28:08:35It's because um that is the only uh
  36932. 28:08:38string feature we have. That's the only
  36933. 28:08:41one we have. So it it's going to look
  36934. 28:08:43for any non-numericals and one hot
  36935. 28:08:45encode those however many however many
  36936. 28:08:47there are. So whatever objects we have
  36937. 28:08:51which are strings, it's going to
  36938. 28:08:52automatically oneh hot encode those.
  36939. 28:09:03Yeah, we could have Right. We could have
  36940. 28:09:05went here and did Right. We could have
  36941. 28:09:07done ocean
  36942. 28:09:11proximity,
  36943. 28:09:13but we only have one of those features.
  36944. 28:09:17So it's just going to do that to the
  36945. 28:09:18whole data frame
  36946. 28:09:20uh on that one feature. So what we're
  36947. 28:09:23going to do is um go ahead and split it
  36948. 28:09:26into an x and a y um which the x is
  36949. 28:09:30always what includes our features. The y
  36950. 28:09:34is what we are trying to predict which
  36951. 28:09:35is the label. Now, um, in order to
  36952. 28:09:40separate those out, what we're going to
  36953. 28:09:41do is assign X to be the variable that
  36954. 28:09:44is, um, our data frame minus this median
  36955. 28:09:48house value column. So what this is
  36956. 28:09:50doing is um uh it's not permanently
  36957. 28:09:55dropping because we're not uh dropping
  36958. 28:09:57it in place but it is returning us a
  36959. 28:10:01copy of the data frame with the median
  36960. 28:10:04house value column left out right it's
  36961. 28:10:06dropped. So this is this is uh something
  36962. 28:10:09we want to do because that will the rest
  36963. 28:10:11of it will contain our features right.
  36964. 28:10:14So, um this will temporarily or I should
  36965. 28:10:18say return a copy of the DF with um
  36966. 28:10:24median house value
  36967. 28:10:28dropped,
  36968. 28:10:30right? Median house value dropped. Um so
  36969. 28:10:33we go ahead and drop that one. Uh now
  36970. 28:10:37remember it's not permanent. It's just
  36971. 28:10:38giving us uh the remainder of it which
  36972. 28:10:41is this housing data. dropping this and
  36973. 28:10:44it's assigning that to X and then we're
  36974. 28:10:46taking the actual median house value
  36975. 28:10:49column from the original data and
  36976. 28:10:52assigning that to Y. So this is going to
  36977. 28:10:53be our labels,
  36978. 28:10:57right? So this is what we are trying to
  36979. 28:11:02predict.
  36980. 28:11:05Okay, so that is our Y and that's always
  36981. 28:11:07how it is. X is our features, Y is our
  36982. 28:11:10labels. Um hopefully that makes sense.
  36983. 28:11:12What this is doing is this is going to
  36984. 28:11:15get rid of that label column and
  36985. 28:11:17everything else will be our features and
  36986. 28:11:19then this will get rid of this will just
  36987. 28:11:22assign the label column to Y.
  36988. 28:11:26All right. And then what we can do is
  36989. 28:11:28pass X and Y into our train test split
  36990. 28:11:32function and this will generate the hold
  36991. 28:11:34out set. So if we want to do the hold
  36992. 28:11:36out cross validation this is how we
  36993. 28:11:39would do it is we would split the data
  36994. 28:11:41into X train X test Y train Y test um
  36995. 28:11:45using train test split. So this is what
  36996. 28:11:48we did last time. This would be this
  36997. 28:11:51would be for hold out cross validation
  36998. 28:11:58right where we are uh uh just have that
  36999. 28:12:03one one set for testing one set for uh
  37000. 28:12:06one set for training one test one set
  37001. 28:12:08for testing I should say right so this
  37002. 28:12:11is pretty standard train test split um
  37003. 28:12:14we pass in that x we pass in the y we
  37004. 28:12:16use a 30% test size which pretty
  37005. 28:12:18standard
  37006. 28:12:19and random state so that we get the
  37007. 28:12:21consistent shuffling if we were to run
  37008. 28:12:23this multiple times. Um we we get that
  37009. 28:12:26uh consistent randomization.
  37010. 28:12:30Okay,
  37011. 28:12:31so we have that and so now our X train
  37012. 28:12:36is a percentage um of the data frame of
  37013. 28:12:39the 20,000 uh rows and the X test is uh
  37014. 28:12:4430% of that. So it's only about 6,000
  37015. 28:12:46rows, which is what um the shape of that
  37016. 28:12:48is.
  37017. 28:12:52Yeah. X. So X is our features. So we're
  37018. 28:12:55we're putting all of our data in that is
  37019. 28:12:57our features into X. And so the the um
  37020. 28:13:01most efficient way of doing that is um
  37021. 28:13:05the most efficient way of doing that is
  37022. 28:13:06to
  37023. 28:13:08uh just take our data and drop the
  37024. 28:13:12median house value column because that's
  37025. 28:13:13our label column. So we just remove
  37026. 28:13:16that. The rest of the data is our
  37027. 28:13:18features. So that's what that's what
  37028. 28:13:20this X is, right? It's all of our
  37029. 28:13:22feature data. All of our columns that is
  37030. 28:13:24not the label column essentially is what
  37031. 28:13:27that's doing. And then Y is our label
  37032. 28:13:30column from our original data,
  37033. 28:13:34right? Y is our label column. And so
  37034. 28:13:38this this will um contain all of our
  37035. 28:13:41labels which is the median house value.
  37036. 28:13:44X X contains every column but the one
  37037. 28:13:47we're going to so we we ultimately
  37038. 28:13:50decide that but X contains um X is
  37039. 28:13:54everything that is not our dependent
  37040. 28:13:57variable which is what we're predicting.
  37041. 28:14:00So we're removing what we are trying to
  37042. 28:14:02predict from X. X should be everything
  37043. 28:14:04else. That's always how it's going to
  37044. 28:14:06be. X is X is always going to be all of
  37045. 28:14:09those independent variables that we're
  37046. 28:14:12using to predict the median house value.
  37047. 28:14:15So we are going to predict the median
  37048. 28:14:18house value. We need to remove it from
  37049. 28:14:20X.
  37050. 28:14:22So we're we're taking everything but
  37051. 28:14:24that column.
  37052. 28:14:29So it's the whole data frame. It's the
  37053. 28:14:32whole data frame minus this one column
  37054. 28:14:35with just the dependent variable. Right.
  37055. 28:14:39Exactly right. Removing the dependent
  37056. 28:14:41variable and keeping all the
  37057. 28:14:42independence. That's exactly right.
  37058. 28:14:44Exactly right. So think about it in
  37059. 28:14:47terms of the model. Let's go back to the
  37060. 28:14:49features. Right. Think about it in terms
  37061. 28:14:50of the model. We are trying to predict
  37062. 28:14:53this this value. We're building a model
  37063. 28:14:57to try to predict this. So we are going
  37064. 28:15:00to make sure x is everything but this
  37065. 28:15:04right. So this is actually just y.
  37066. 28:15:07That's our label. That's our dependent
  37067. 28:15:09variable. Right? That's y. Everything
  37068. 28:15:12else is belongs to x. Everything else
  37069. 28:15:15belongs to x including all of these.
  37070. 28:15:22Right? We choose this one to be y
  37071. 28:15:24because we're building a model to
  37072. 28:15:26predict that. That's our label.
  37073. 28:15:29All right. So, we have our we use X and
  37074. 28:15:32Y to do our train test split. So, we
  37075. 28:15:34have our our training features and our
  37076. 28:15:37test features and then our training
  37077. 28:15:38label and test labels here. Um, pretty
  37078. 28:15:42standard there.
  37079. 28:15:44Um, okay. So, this is what's new is if
  37080. 28:15:47we want to do k-fold uh validation, what
  37081. 28:15:50we're going to do is create a kfold
  37082. 28:15:52object. So, we have this kfold from
  37083. 28:15:54scikitlearn that we already imported. we
  37084. 28:15:57are going to create a kfold um where we
  37085. 28:16:02are going to specify how many folds we
  37086. 28:16:04want. So that is the in uh inslits
  37087. 28:16:07parameter as this says um this is going
  37088. 28:16:10to be uh uh in this case we're going to
  37089. 28:16:14do 10 folds. That's pretty standard. So
  37090. 28:16:16I think the typical number of folds that
  37091. 28:16:18I've seen and I've worked with in my in
  37092. 28:16:20my career is usually five or 10.
  37093. 28:16:24Five or 10 folds is the standard.
  37094. 28:16:29Okay. So, we're doing 10 folds in this
  37095. 28:16:31case and we're setting a random state
  37096. 28:16:33because we're going to do shuffling. So,
  37097. 28:16:35in order to produce those folds, we're
  37098. 28:16:37going to shuffle the data first and then
  37099. 28:16:39split it into five folds, right? So,
  37100. 28:16:42this this kffold object is going to
  37101. 28:16:44manage creating these splits for us,
  37102. 28:16:48right? These even splits. I know I I
  37103. 28:16:50didn't draw it even, but um it's going
  37104. 28:16:53to manage these five folds for us and
  37105. 28:16:55it's going to shuffle the data and
  37106. 28:16:57assign them to these different folds and
  37107. 28:16:59we're and then what we're going to do is
  37108. 28:17:01use those to do our training.
  37109. 28:17:04We're going to execute the cross
  37110. 28:17:05validation using this kfold object.
  37111. 28:17:09Okay, so we create the kfold
  37112. 28:17:12um we initialize our model as well. So,
  37113. 28:17:15of course, in order to train something
  37114. 28:17:18uh in the K-folds, we're going to need a
  37115. 28:17:20model. In this case, we're using linear
  37116. 28:17:22regression, right? Which is which is the
  37117. 28:17:24model we've been studying so far. So,
  37118. 28:17:26you have a linear regression. Um now,
  37119. 28:17:29look how easy it's going to be in order
  37120. 28:17:31to execute cross validation. All we need
  37121. 28:17:34to do is um all we need to do is create
  37122. 28:17:38a cross file score function
  37123. 28:17:42um or I should say use the cross file
  37124. 28:17:44score function from scikitlearn. So we
  37125. 28:17:46use that with the model we want to
  37126. 28:17:48train. So our model goes first. So
  37127. 28:17:51that's the linear regression object.
  37128. 28:17:54Then our data. So our extra our features
  37129. 28:17:57and our label for our training.
  37130. 28:18:00And then um let me skip over this for a
  37131. 28:18:03second. I'll explain what this is in a
  37132. 28:18:05second. Um but then we are using uh the
  37133. 28:18:10cross validation technique is our
  37134. 28:18:12K-fold. So this is where our K-fold
  37135. 28:18:14object goes in the CV parameter which is
  37136. 28:18:17cross validation. So what cross
  37137. 28:18:20validation strategy are you using? We're
  37138. 28:18:21using Kfold and the K-fold we're using
  37139. 28:18:24is this one we defined up here KF. So
  37140. 28:18:27we're putting that right here for this.
  37141. 28:18:29And then um in jobs um allows us to
  37142. 28:18:34parallelize this. So if we set it to
  37143. 28:18:36negative one that's the that that's the
  37144. 28:18:38default um it will do it will actually
  37145. 28:18:41train across the different combinations
  37146. 28:18:43in parallel um which speeds it up. So
  37147. 28:18:46you want to you want to keep this to
  37148. 28:18:48negative one if you can. So um now let
  37149. 28:18:52me describe the scoring. So what this
  37150. 28:18:55means is we put in our metric here. Um
  37151. 28:19:00and so you can put mean absolute error,
  37152. 28:19:03you can put in mean squared error. Um
  37153. 28:19:06those are the two that we can use. And
  37154. 28:19:09um the reason we it has a negative in
  37155. 28:19:11front of it is because we want to find
  37156. 28:19:15the one that has the lowest score.
  37157. 28:19:19That's going to be our best model is the
  37158. 28:19:21one that has the lowest score. So, we
  37159. 28:19:24take the absolute value.
  37160. 28:19:26I'm sorry. We take the abs the the the
  37161. 28:19:29metric and we take the negative of it.
  37162. 28:19:32Um because the highest scoring one is
  37163. 28:19:36going to be the closest to zero. Um so
  37164. 28:19:39it's just a we use the we use the
  37165. 28:19:41negative of the of the metric. Um
  37166. 28:19:44because on the number line like the the
  37167. 28:19:47highest um scoring one should be the
  37168. 28:19:50least um or I should say the maximum
  37169. 28:19:53negative that we can get. That's going
  37170. 28:19:55to be closest to zero. So if here's
  37171. 28:19:57zero, this will be like -1 is better
  37172. 28:20:00than -10. Right? So something that
  37173. 28:20:04scores um the maximum negative uh
  37174. 28:20:08absolute error would be closest to zero.
  37175. 28:20:12And something that has more is going to
  37176. 28:20:14be on this side.
  37177. 28:20:17So this is only the reason we need this
  37178. 28:20:19is only just to keep track of the scores
  37179. 28:20:22of each individual um fold. Okay.
  37180. 28:20:27So the one so the reason we can do that
  37181. 28:20:29is at the end we can kind of see which
  37182. 28:20:31which combination performed the best. um
  37183. 28:20:35it's going to be the one that has the
  37184. 28:20:37highest uh highest value of the negative
  37185. 28:20:41which is closest to zero.
  37186. 28:20:48That's just a convention.
  37187. 28:20:51Yeah, it's just because um it's because
  37188. 28:20:54the cross validation is looking to
  37189. 28:20:56maximize the metric. So whatever has the
  37190. 28:20:58best score
  37191. 28:21:00um whatever has the best score is
  37192. 28:21:03considered the best uh performance. Um
  37193. 28:21:05but we are using uh something where
  37194. 28:21:09lower is better. So we we take the
  37195. 28:21:11negative and like the the highest
  37196. 28:21:13negative would be closest to zero,
  37197. 28:21:17right? The highest negative is going to
  37198. 28:21:18be closest to zero.
  37199. 28:21:22So that so it's it's just because like
  37200. 28:21:25we want the lower score to be the best.
  37201. 28:21:29The lowest score should be the best.
  37202. 28:21:32So we take the negative of it. Um and so
  37203. 28:21:36something that is more negative is going
  37204. 28:21:38to be worse. Yeah, that's the reason.
  37205. 28:21:43So something that's down this way is
  37206. 28:21:45going to be worse.
  37207. 28:21:47Okay. So it runs this
  37208. 28:21:51and what you can see is if we actually
  37209. 28:21:53print this out, if we print out our
  37210. 28:21:55k-fold scores, what we should get is 10
  37211. 28:21:57different scores.
  37212. 28:22:01And you can see um we have 10 different
  37213. 28:22:04uh scores here, which are all negative
  37214. 28:22:07because we're taking the negative of the
  37215. 28:22:08absolute of the mean absolute error. Um
  37216. 28:22:12so what we would be looking for here is
  37217. 28:22:17um we want to take the average of these
  37218. 28:22:20scores but take the absolute value of
  37219. 28:22:22them to get the best performance. So
  37220. 28:22:24this is capturing like this is the score
  37221. 28:22:26on the first fold combination. This is
  37222. 28:22:29the score on the second fold
  37223. 28:22:30combination. This is the score on the
  37224. 28:22:32third fold combination and on and on and
  37225. 28:22:35on. And these are the absolute errors.
  37226. 28:22:38Okay, these are the absolute errors. Um,
  37227. 28:22:42so if we take a look at computing the uh
  37228. 28:22:45average, which by the way, we don't need
  37229. 28:22:47this import because we're using the
  37230. 28:22:49numpy average. So that's fine. Um, we
  37231. 28:22:52can take the absolute value of those um
  37232. 28:22:55and take a look at the average MSE
  37233. 28:23:00or sorry MAE. Now I want you to think
  37234. 28:23:04about this this uh average performance.
  37235. 28:23:07So this is our performance right here on
  37236. 28:23:09the cross validation.
  37237. 28:23:11This is our average
  37238. 28:23:13M AE across all of our fold
  37239. 28:23:16combinations. So that's a that's an
  37240. 28:23:18indicator of our performance, right? Um
  37241. 28:23:21for the cross validation.
  37242. 28:23:24Now what are the units of our original
  37243. 28:23:29uh the original median value? They're
  37244. 28:23:33already in the thousands, right? So if
  37245. 28:23:36we go to that feature, they're already
  37246. 28:23:40in these hundreds of thousands. So this
  37247. 28:23:43is not a very good error. It's it's kind
  37248. 28:23:48of high, right? Because it's in this is
  37249. 28:23:5049,000.
  37250. 28:23:52Um that's that's how far away we are in
  37251. 28:23:55absolute value on average is 49,000 um
  37252. 28:23:59dollars on the median value. That's not
  37253. 28:24:02very good. So this score
  37254. 28:24:06this score is
  37255. 28:24:09um not very good. So this model is not
  37256. 28:24:12performing that well and we can see that
  37257. 28:24:14by comparing this error to our actual uh
  37258. 28:24:17data. So this is right around 50,000
  37259. 28:24:22and our median uh house values are in
  37260. 28:24:25the hundreds of thousands. So on average
  37261. 28:24:28we're 50,000 off when we make a
  37262. 28:24:31prediction. That's a significant amount,
  37263. 28:24:34right? That's a significant amount on
  37264. 28:24:36average um when our when our data is in
  37265. 28:24:39about the hundreds of thousands here.
  37266. 28:24:43So we are um we have a significant
  37267. 28:24:46amount of error 50,000 relative to the h
  37268. 28:24:49to our units that our our data is in.
  37269. 28:24:52Right? Um so this score is not very
  37270. 28:24:54good. Um
  37271. 28:24:57and so we see that from the cross
  37272. 28:24:59validation. So look how easy the cross
  37273. 28:25:00valid is. Again we just do cross file
  37274. 28:25:03score. We put in our model. We put in
  37275. 28:25:04our data. We put in our cross validation
  37276. 28:25:08uh strategy here which is kfold. And we
  37277. 28:25:10can generate these metrics across all
  37278. 28:25:13the fold combinations. So it's this
  37279. 28:25:15function is taking care of rotating
  37280. 28:25:17those and doing every combo with just
  37281. 28:25:19the 10 different combinations here of
  37282. 28:25:22the of the folds.
  37283. 28:25:2510 different instances where you have
  37284. 28:25:26you know 10 different folds are the ones
  37285. 28:25:28that are left out for evaluation.
  37286. 28:25:30Um so it's managing that for us using
  37287. 28:25:33this data right using this training data
  37288. 28:25:35here. Um and we uh we generate these um
  37289. 28:25:42generate these scores.
  37290. 28:25:46Okay. So that's kf fold. It's not hard
  37291. 28:25:48to do. All you have to do is um just use
  37292. 28:25:52a cross file score. And we could change
  37293. 28:25:54this to mean squared error. That's you
  37294. 28:25:57know we could do that too. That'd be
  37295. 28:25:58pretty easy. Um, so that'd be no issue.
  37296. 28:26:03We just happen to be using the absolute
  37297. 28:26:05error here. Of course, we could use
  37298. 28:26:07squared error.
  37299. 28:26:09Were you guys able to get this to run?
  37300. 28:26:12K-fold scores.
  37301. 28:26:14It produces an array of 10 10 different
  37302. 28:26:17scores, which should make sense because
  37303. 28:26:19those are these are the um we're
  37304. 28:26:21splitting our data into 10 different
  37305. 28:26:23folds,
  37306. 28:26:25right?
  37307. 28:26:2710 different folds. than leaving one out
  37308. 28:26:29to do our evaluation on. So the one that
  37309. 28:26:31gets left out every time is what's
  37310. 28:26:33producing these scores. So it's 10
  37311. 28:26:35different ones get left out when we
  37312. 28:26:37rotate through all the combinations.
  37313. 28:26:42And so we average these scores
  37314. 28:26:51and we get this amount. We get about
  37315. 28:26:5350,000 in error on average.
  37316. 28:27:02Um, what do you think would be what do
  37317. 28:27:04you think would be acceptable? So, if
  37318. 28:27:05our if we're predicting the price, like
  37319. 28:27:07if we're a real estate agent and we're
  37320. 28:27:09predicting these prices and they
  37321. 28:27:12typically are
  37322. 28:27:14Yeah, close to zero would be great.
  37323. 28:27:16That'd be fantastic. Closer to zero
  37324. 28:27:18would be better. The average is um
  37325. 28:27:20206,000.
  37326. 28:27:23So 50,000 is a decent percentage of
  37327. 28:27:26that. Um so you know you can compute it
  37328. 28:27:29as a percentage right. So 50,000 is a
  37329. 28:27:32decent percentage of that. Um probably
  37330. 28:27:35you want this to be less than 20,000
  37331. 28:27:38would be about 10% error. 20,000
  37332. 28:27:43right? So maybe like 30,000 somewhere in
  37333. 28:27:48there.
  37334. 28:27:52Yeah. 10% would be 5% error. 10,000
  37335. 28:27:55would be 5% error. That's true. That's
  37336. 28:27:56true. So that would be that would be
  37337. 28:27:59much better. So being closer to zero,
  37338. 28:28:00like the smaller the better, of course.
  37339. 28:28:03Of course. Um but yeah, I would say an
  37340. 28:28:06acceptable percentage of error is
  37341. 28:28:08probably 20%.
  37342. 28:28:11Probably 20%, which would be um like
  37343. 28:28:1440,000 or less would probably be
  37344. 28:28:16acceptable.
  37345. 28:28:19Usually when we usually when you build
  37346. 28:28:21models um 80% accuracy is usually uh
  37347. 28:28:26considered decent.
  37348. 28:28:29Usually considered decent
  37349. 28:28:3180%. So I'd say 40,000 or less would be
  37350. 28:28:35kind of ideal.
  37351. 28:28:40Does that make sense
  37352. 28:28:42to answer the question?
  37353. 28:28:45That's a good question. What value is
  37354. 28:28:46acceptable? I think probably less than
  37355. 28:28:4940,000 would be ideal. That's right
  37356. 28:28:52around 20% error.
  37357. 28:28:56All right, so that's K-fold. Um let's do
  37358. 28:28:59just a regular hold out now. So this is
  37359. 28:29:01just using our training and test data.
  37360. 28:29:03Um doing model.fit and calculating an
  37361. 28:29:06MSE on the test data. So this is this is
  37362. 28:29:09just the um hold out strategy here where
  37363. 28:29:12we just have um this is less robust but
  37364. 28:29:15it's a lot quicker to do and easier to
  37365. 28:29:17set up. Right? So um this is using the
  37366. 28:29:22hold out strategy. So just a regular
  37367. 28:29:26um train test split.
  37368. 28:29:30Are we going to rebuild the model? No,
  37369. 28:29:32not necessarily. There's some things we
  37370. 28:29:34could do most likely. And like one thing
  37371. 28:29:38we did not do was scale our features.
  37372. 28:29:41Remember I said that's a pretty
  37373. 28:29:42important thing to do is to scale our
  37374. 28:29:45features. We did not do that. So that
  37375. 28:29:47would be an enhancement to this that
  37376. 28:29:49we're going to So I I actually do think
  37377. 28:29:50we'll do that later. Yes. So I think we
  37378. 28:29:53will actually do that now that I'm
  37379. 28:29:55thinking about it. Yes. One of the
  37380. 28:29:57things we can do is scale these features
  37381. 28:29:59using like a minmax scaler or a standard
  37382. 28:30:01scaler. that's actually going to help us
  37383. 28:30:03um that's going to help us do better
  37384. 28:30:06predictions.
  37385. 28:30:10So that that's one thing we could do. Um
  37386. 28:30:13but yeah, we will we'll try to see if we
  37387. 28:30:16can get better.
  37388. 28:30:18It should help it. Yeah, usually you
  37389. 28:30:20want to scale you want to scale the
  37390. 28:30:22data. That's something we didn't do in
  37391. 28:30:24our preparation step. We did a lot of
  37392. 28:30:26the things we should do. We removed nles
  37393. 28:30:27and we did one hot encoding to the
  37394. 28:30:30proximity feature like this one. Um
  37395. 28:30:33those are good to do but we didn't scale
  37396. 28:30:36any of these other we didn't scale any
  37397. 28:30:38of the features right we didn't scale
  37398. 28:30:40any of them. Um it you it will have an
  37399. 28:30:44effect. It usually when we scale it
  37400. 28:30:46it'll be a better model.
  37401. 28:30:49It'll it'll learn a little bit better if
  37402. 28:30:51we can scale the data. Um so that way
  37403. 28:30:55like these
  37404. 28:30:58um like ages aren't you know drastically
  37405. 28:31:01different than like in scale than total
  37406. 28:31:03bedrooms or income
  37407. 28:31:06uh those kind of things. So we usually
  37408. 28:31:08want these to be in a similar scale
  37409. 28:31:10range.
  37410. 28:31:13So we'll we will I think we'll scale
  37411. 28:31:15them coming up in a bit and it should
  37412. 28:31:18help the model.
  37413. 28:31:21We've talked about that before, right?
  37414. 28:31:22Scaling usually is a good idea to do
  37415. 28:31:25when you're prepping your data for
  37416. 28:31:26modeling.
  37417. 28:31:29No, you want to you want to scale your
  37418. 28:31:31test data as well. You're going to do
  37419. 28:31:33both. You're going to scale your
  37420. 28:31:34training data. You're going to scale it.
  37421. 28:31:36So that's actually a good point you
  37422. 28:31:38bring up is any transformations you do
  37423. 28:31:40on your training to build your model,
  37424. 28:31:43you should also do on your test set so
  37425. 28:31:45you get an applesto apples comparison.
  37426. 28:31:48You should always do the same
  37427. 28:31:49transformations.
  37428. 28:31:51Yes. Would scaling data impact K? Yeah,
  37429. 28:31:54it could. It could make it better. It
  37430. 28:31:56could uh Yeah, it should impact it. We
  37431. 28:31:58should get a better model. So, when we
  37432. 28:32:00do the different folds, we'll get
  37433. 28:32:01different we'll get better scores. Yeah,
  37434. 28:32:04it it will impact
  37435. 28:32:08uh yeah, if they're so that's a good
  37436. 28:32:11point. If they're going to use our
  37437. 28:32:12model, then yes, they have to scale the
  37438. 28:32:14data as well. If they're going to if we
  37439. 28:32:16build the model on the assumption that
  37440. 28:32:17the input is scaled, then yes, they have
  37441. 28:32:21to also scale their data when they're
  37442. 28:32:22using it with our model. That's true.
  37443. 28:32:29I mean, not really. I'll show you why.
  37444. 28:32:32There's something that's actually going
  37445. 28:32:33to make it easier um that that will
  37446. 28:32:36automate doing the scaling for them. So,
  37447. 28:32:38they don't they don't have to do the
  37448. 28:32:40scaling manually. it'll just it'll
  37449. 28:32:42happen automatically when they use the
  37450. 28:32:44model. I'm going to show you something
  37451. 28:32:45that's going to automate that which is
  37452. 28:32:47going to be called a pipeline.
  37453. 28:32:49So that part will be automated and they
  37454. 28:32:52won't have to do that. So it won't be
  37455. 28:32:53heavy on the user. No, in theory it is,
  37456. 28:32:57but
  37457. 28:32:58has a really helpful tool to make it
  37458. 28:33:00easy to do that. So I'm going to I'm
  37459. 28:33:02going to show us that um later on in the
  37460. 28:33:04notebook.
  37461. 28:33:07No, the data data is not for a single
  37462. 28:33:09house. It's for like a neighborhood. So
  37463. 28:33:12there's a certain number of households
  37464. 28:33:14in the neighborhood. And this is the
  37465. 28:33:16we're predicting the median house value
  37466. 28:33:18of that neighborhood.
  37467. 28:33:21Yeah. So there's a there's certain
  37468. 28:33:22number of households. There's there's
  37469. 28:33:23like an a median income, a population,
  37470. 28:33:26certain number of people that live
  37471. 28:33:28there. Um proximity generally of where
  37472. 28:33:32that location is. It also has a latitude
  37473. 28:33:34and longitude.
  37474. 28:33:38So,
  37475. 28:33:41and a median age in that neighborhood.
  37476. 28:33:43So, yeah, it's not just a single house.
  37477. 28:33:54Okay, let's go back to this was the hold
  37478. 28:33:58out strategy. So, this is a lot simpler.
  37479. 28:34:00This is just model.fit, right? This is
  37480. 28:34:02just model.fit on the training uh data.
  37481. 28:34:05And then we um can predict on the test
  37482. 28:34:07features and generate test predictions.
  37483. 28:34:10And then we can compute our error on
  37484. 28:34:12those um we can compute our error
  37485. 28:34:15amongst the test predictions and our
  37486. 28:34:16test uh label. So that's our useful mean
  37487. 28:34:21squared error function, right? To to
  37488. 28:34:23compute the MSE. Um let's see what the
  37489. 28:34:27MSE is. So MSE is right here.
  37490. 28:34:33Um now what we could do is we can take
  37491. 28:34:36the MSE
  37492. 28:34:38and we can take the square root of it.
  37493. 28:34:39So let's actually do that. Let's um do
  37494. 28:34:42MP. Square root of the
  37495. 28:34:45um test
  37496. 28:34:48MSE
  37497. 28:34:51and we get um 67 we get 67,000.
  37498. 28:34:56So that's pretty high on this. So when
  37499. 28:34:59we just now look at the difference of
  37500. 28:35:01that, right? When we just do a train
  37501. 28:35:03test split,
  37502. 28:35:05um
  37503. 28:35:08when we just do a train test split, we
  37504. 28:35:10get a worse score because it's not as
  37505. 28:35:12it's not as robust, right? We're not
  37506. 28:35:14showing that to many of the other uh
  37507. 28:35:17folds. So, we get a lot more error this
  37508. 28:35:20way on the test data.
  37509. 28:35:23So, this is um actually worse
  37510. 28:35:25performance just doing the train test
  37511. 28:35:27split.
  37512. 28:35:30This is a really higher.
  37513. 28:35:42Yeah, we can. We can. I'm going to I'm
  37514. 28:35:45going to show us how to how the scaling
  37515. 28:35:46will be done automatically. Yes, we can.
  37516. 28:35:50Um there's there's a really easy tool to
  37517. 28:35:53do that will scale it automatically.
  37518. 28:36:00It's going to be later in this notebook.
  37519. 28:36:01I'll show us it.
  37520. 28:36:10All right. So, just to recap this, this
  37521. 28:36:12is fitting the model.
  37522. 28:36:14This is fitting the model. This is
  37523. 28:36:16making the predictions, right?
  37524. 28:36:18Model.predict.
  37525. 28:36:20So, this is making the predictions. And
  37526. 28:36:21then this is calculating the error, the
  37527. 28:36:24mean squared error, which is looking at
  37528. 28:36:25our test labels versus our test
  37529. 28:36:27predictions, right? And this is
  37530. 28:36:29computing the distance, the average
  37531. 28:36:32distance away from these values to these
  37532. 28:36:34values,
  37533. 28:36:36right?
  37534. 28:36:38And then we can also compute the R squar
  37535. 28:36:40R R squar and we see that it's not a
  37536. 28:36:43very good R squar 65 uh is not a very
  37537. 28:36:46great model
  37538. 28:36:48um because it closer to one would be
  37539. 28:36:50better. So this is still this is not
  37540. 28:36:53very good.
  37541. 28:36:54We know that we knew that from the cross
  37542. 28:36:56file score but this is just doing um
  37543. 28:36:58this is just doing a hold out uh where
  37544. 28:37:01we do a train and test split. Right? So
  37545. 28:37:03it's a little bit simpler but it's not
  37546. 28:37:05quite as robust. Um,
  37547. 28:37:08it's not quite as robust as the cross
  37548. 28:37:10valve, but it works. Um, it's, you know,
  37549. 28:37:14we can do hold out. Um,
  37550. 28:37:17we can do hold out, uh, to to quickly
  37551. 28:37:19evaluate a model and see if we need to
  37552. 28:37:22make any adjustments.
  37553. 28:37:24It's a little bit quicker to run.
  37554. 28:37:30Okay. And any questions on it? Does it
  37555. 28:37:32make sense what we're doing here?
  37556. 28:37:33Model.fit fit to train it predict to get
  37557. 28:37:37our predictions. Um this is pretty
  37558. 28:37:40standard, right? To train is the
  37559. 28:37:41model.fit and then to use the model to
  37560. 28:37:44predict we predict on the test features.
  37561. 28:37:47Um so this is passing on on all of our
  37562. 28:37:49features into this model to generate
  37563. 28:37:53predictions for every row. That's
  37564. 28:37:55something I also want to point out that
  37565. 28:37:56may be a little bit confusing is this is
  37566. 28:37:58a data frame. So we're passing in a
  37567. 28:38:02bunch of rows of features with columns,
  37568. 28:38:04right? So um we're passing in a bunch of
  37569. 28:38:07data that looks like this. And what
  37570. 28:38:10we're doing is essentially making a
  37571. 28:38:12prediction for every row. So this will
  37572. 28:38:14generate a prediction. This row will
  37573. 28:38:17generate a prediction. This row will
  37574. 28:38:19generate a prediction and on and on and
  37575. 28:38:21on. So this this predict will predict
  37576. 28:38:24for every row. And so we end up with
  37577. 28:38:27this collection of predictions here for
  37578. 28:38:29each row. and we're comparing those to
  37579. 28:38:32the labels that we have for those rows
  37580. 28:38:35from our from our supervised learning,
  37581. 28:38:37right? From our data set. So that's
  37582. 28:38:40truly supervised learning, right? We
  37583. 28:38:42have the examples and we're comparing
  37584. 28:38:44those to what our model is predicting to
  37585. 28:38:47to get our performance.
  37586. 28:38:56All right.
  37587. 28:38:58So let's uh let's try the other just so
  37588. 28:39:01you can see it. The leave one out. Now
  37589. 28:39:03the leave one out cross validation is
  37590. 28:39:05going to actually work the same way
  37591. 28:39:07where we put in the leave one out um
  37592. 28:39:10strategy inside of the cross file score.
  37593. 28:39:13Now here we don't need to specify how
  37594. 28:39:14many folds there are because we know how
  37595. 28:39:17many they're going to be. It's going to
  37596. 28:39:18be the number of data points, right? So
  37597. 28:39:21which is actually going to be quite
  37598. 28:39:22large because there's 20,000 rows. So
  37599. 28:39:25this is going to be extremely
  37600. 28:39:28uh extremely um intensive because we are
  37601. 28:39:33doing um you know 20,000 examples and
  37602. 28:39:37leaving one example out to be our
  37603. 28:39:40validation and then um doing that across
  37604. 28:39:43every 20,000 uh examples.
  37605. 28:39:47So we could do it though just to see how
  37606. 28:39:48it works. Um we have this again leave
  37607. 28:39:51one out. We generate our crossfile score
  37608. 28:39:53from our model our data and then same
  37609. 28:39:56scoring that we had before and but this
  37610. 28:39:58time we change our cross file to be
  37611. 28:40:00instead of our kfold object we have our
  37612. 28:40:03leave one out object which is this
  37613. 28:40:06um and then we could run this. We can
  37614. 28:40:09compute our average uh across the all
  37615. 28:40:12the folds. Now this is going to be a lot
  37616. 28:40:14bigger of an array. It's going to be a
  37617. 28:40:1620,000 size array and we're going to
  37618. 28:40:19compute the average across it.
  37619. 28:40:23So, let's do that. It's going to take a
  37620. 28:40:25moment because there's lots. So, if you
  37621. 28:40:27notice it when you run, it's going to
  37622. 28:40:28take a little bit of time to run because
  37623. 28:40:30it's running across all 20,000 examples
  37624. 28:40:34and leaving one out. So, you have 20,000
  37625. 28:40:37and then one left out to uh test
  37626. 28:40:41against. So, it's quite intensive. You
  37627. 28:40:43can see it's taking a lot more time.
  37628. 28:40:52It's still running. It's taking a while.
  37629. 28:41:03Okay, just let that run. Still running.
  37630. 28:41:08So, if you guys try running this, it's
  37631. 28:41:10going to take a little bit of time.
  37632. 28:41:11Hopefully, that makes sense why it's
  37633. 28:41:13taking so long, right? It's because it's
  37634. 28:41:16instead of doing 10 folds, it's it's
  37635. 28:41:19putting every data point but one is the
  37636. 28:41:21training set and then iterating through
  37637. 28:41:23all 20,000 points.
  37638. 28:41:27This takes a while to do.
  37639. 28:41:44Let's see what our
  37640. 28:41:49RAM our memory is a little increased.
  37641. 28:42:00Okay,
  37642. 28:42:02still running. That's okay. I'll let it
  37643. 28:42:04run.
  37644. 28:42:07Come back when it's finished.
  37645. 28:42:15Yeah, exactly. This is a this is for
  37646. 28:42:17this is giving us a performance
  37647. 28:42:18evaluation. This is like the average
  37648. 28:42:21error across all of our uh different
  37649. 28:42:24folds. Um now this is the extreme case
  37650. 28:42:26where we have the number of folds equals
  37651. 28:42:28the number of points.
  37652. 28:42:31Right? So it's an extreme case but yes
  37653. 28:42:33it's just like kfold. It's giving us
  37654. 28:42:35that performance estimate.
  37655. 28:42:44Okay. It's about the same. Right. This
  37656. 28:42:47is still around 50,000.
  37657. 28:42:50Not much difference, right? Still right
  37658. 28:42:52around there. But look how much longer
  37659. 28:42:55it took. That took 2 minutes to run. The
  37660. 28:42:57other one was pretty instant, right? So
  37661. 28:42:59this this took about 2 minutes to run.
  37662. 28:43:02So um definitely uh
  37663. 28:43:09yeah, definitely don't want to run this
  37664. 28:43:11uh too often. I think that it's
  37665. 28:43:14generally preferred to do k-fold. If
  37666. 28:43:16you're going to do cross validation,
  37667. 28:43:17generally want to do k-fold or just the
  37668. 28:43:19regular hold out train test split. Uh
  37669. 28:43:22generally better than doing leave one
  37670. 28:43:24out. It's just going to take too long
  37671. 28:43:26and um it results in about the same kind
  37672. 28:43:29of score as the kfold.
  37673. 28:43:41Okay,
  37674. 28:43:44any questions about um the cross
  37675. 28:43:47validation that we just did.
  37676. 28:44:08Okay,
  37677. 28:44:10good. And as it says here that the
  37678. 28:44:12stratified kfold is usually used for
  37679. 28:44:14classification. Again, we're not doing
  37680. 28:44:15classification yet. That's in going to
  37681. 28:44:17be in lesson four. So, we don't need to
  37682. 28:44:19worry too much about that. Just for
  37683. 28:44:20regression, um regular k-fold is
  37684. 28:44:23preferred, right? Because we don't need
  37685. 28:44:25to um worry about distributing
  37686. 28:44:28categories amongst our folds uh in any
  37687. 28:44:31regression problems.
  37688. 28:44:34And as we see the error is kind of high.
  37689. 28:44:36Um there's going to be some things we
  37690. 28:44:37can do to improve that which will be uh
  37691. 28:44:40later on we'll learn about some more
  37692. 28:44:42advanced models. This signals that the
  37693. 28:44:45performance is bad. We probably need a
  37694. 28:44:47more complex model. Um one thing we
  37695. 28:44:50could try before we try a complex model
  37696. 28:44:52is to do scaling. We will try to do
  37697. 28:44:55scaling. I'm going to show us how we can
  37698. 28:44:57do that coming up um in a in a nice
  37699. 28:45:00streamlined fashion. Um, but uh outside
  37700. 28:45:04of that, if we still had bad
  37701. 28:45:06performance, we would likely need to use
  37702. 28:45:07a more advanced model. And we'll learn
  37703. 28:45:10about more advanced models uh in the
  37704. 28:45:13next lesson. And what's great is some of
  37705. 28:45:15those advanced models can actually be
  37706. 28:45:17used for regression. So they have
  37707. 28:45:19variations that can be used for both
  37708. 28:45:21classification and regression, which is
  37709. 28:45:23pretty cool. So I'll point those out
  37710. 28:45:25when we get to them. Um, okay.
  37711. 28:45:30So what I want to talk about now is a
  37712. 28:45:33way we can combat overfitting. So if we
  37713. 28:45:36have overfitting which remember that is
  37714. 28:45:38the case where the uh the we see good
  37715. 28:45:43performance on the training data but
  37716. 28:45:45then um it doesn't generalize over to
  37717. 28:45:47the test data. We get poor performance
  37718. 28:45:49on the test data. Um there's there's a
  37719. 28:45:52drop off there. Um that would signal
  37720. 28:45:55overfitting.
  37721. 28:45:58overfitting
  37722. 28:46:00and one way of um combating overfitting
  37723. 28:46:03is to do something called regularization
  37724. 28:46:06which we're going to talk about next. So
  37725. 28:46:09the key idea in regularization
  37726. 28:46:13is to
  37727. 28:46:15change our uh the change the way we
  37728. 28:46:19train. Essentially, what we're going to
  37729. 28:46:21do is modify our training
  37730. 28:46:27uh error function or sometimes called
  37731. 28:46:30the objective function or loss function.
  37732. 28:46:33We're going to change that to add a
  37733. 28:46:36penalty to penalize excessive complex
  37734. 28:46:40complexity. Essentially the the way that
  37735. 28:46:43we're going to penalize is by making
  37736. 28:46:45sure the size of the coefficients
  37737. 28:46:48doesn't grow too much which should
  37738. 28:46:51mitigate overfitting because remember in
  37739. 28:46:54linear regression what we are learning
  37740. 28:46:56are the coefficients right we're
  37741. 28:46:58learning the beta 0 the beta 1 the beta
  37742. 28:47:012 and on and on however many betas there
  37743. 28:47:04are beta n we're learning all of those
  37744. 28:47:06guys um through the regression error
  37745. 28:47:10function we're trying to minimize that
  37746. 28:47:11error function. That's how it trains. We
  37747. 28:47:13talked about that on Monday.
  37748. 28:47:16Um so what we're going to do is um
  37749. 28:47:21basically penalize the these guys
  37750. 28:47:24growing too big and making sure we kind
  37751. 28:47:27of keep them small so that no one
  37752. 28:47:31coefficient has a dominant uh effect on
  37753. 28:47:34the model. And this should help with
  37754. 28:47:36overfitting and complexity. It should
  37755. 28:47:38make the model simpler because all the
  37756. 28:47:40coefficients are going to be encouraged
  37757. 28:47:42to be smaller. They're not going to grow
  37758. 28:47:44too big. Um, and this this has the
  37759. 28:47:47effect of making the model so basically
  37760. 28:47:50make the model simpler.
  37761. 28:47:54Make the model simpler is what these
  37762. 28:47:57regularization techniques are
  37763. 28:47:58essentially trying to achieve is is
  37764. 28:48:00remove complexity, make them a little
  37765. 28:48:02bit simpler, make these coefficients
  37766. 28:48:04smaller so that you can generalize a bit
  37767. 28:48:07better and and prevent overfitting. So
  37768. 28:48:10we want to prevent
  37769. 28:48:13uh overfitting,
  37770. 28:48:15right, is what we want to do. Um so
  37771. 28:48:19there's going to be a penalty and I'll
  37772. 28:48:20show you where that penalty gets added
  37773. 28:48:22and kind of what it looks like.
  37774. 28:48:24Um but uh to control the level of that
  37775. 28:48:28penalty we are actually going to
  37776. 28:48:29introduce another parameter to our model
  37777. 28:48:33um called alpha.
  37778. 28:48:35Alpha is going to scale the penalty. So
  37779. 28:48:38if alpha is really high that imposes a
  37780. 28:48:42stronger penalty on the coefficients um
  37781. 28:48:45which will make the model a lot simpler.
  37782. 28:48:48So the higher the alpha the simpler the
  37783. 28:48:51model we will get and we the the risk
  37784. 28:48:54with that is we actually underfit. So if
  37785. 28:48:57alpha is too big we may underfit the
  37786. 28:49:00training data
  37787. 28:49:02um a bit too much because it will make
  37788. 28:49:04the model way too simple. Um and again
  37789. 28:49:08I'll show you what this means
  37790. 28:49:08mathematically in a moment. Um but on
  37791. 28:49:12the other hand if we have a lower alpha
  37792. 28:49:14this will have a lower penalty. it's a
  37793. 28:49:17weaker penalty term and that'll lead to
  37794. 28:49:20a model that is um a bit more complex.
  37795. 28:49:24Um which could um risk some level of
  37796. 28:49:27overfitting. Um so there's so there's
  37797. 28:49:31still the risk of overfitting if you
  37798. 28:49:33have a low alpha. And of course if alpha
  37799. 28:49:35goes all the way to zero there's no
  37800. 28:49:37penalty at all. So you're back to your
  37801. 28:49:39original linear regression um which
  37802. 28:49:42could risk a lot of overfitting.
  37803. 28:49:45Right? So you you generally want to pick
  37804. 28:49:47an alpha um effectively and actually
  37805. 28:49:50we're going to see h what's the best way
  37806. 28:49:52to pick alpha. Um we're actually going
  37807. 28:49:54to learn how to do that. I'm going to
  37808. 28:49:56show us how doing some tuning techniques
  37809. 28:49:58to pick what alpha should be. Um but um
  37810. 28:50:03a a pretty industry standard alpha that
  37811. 28:50:05most people default to is alpha equals
  37812. 28:50:08to one. So just just one which signals
  37813. 28:50:12that there should be some penalty. we
  37814. 28:50:14just have alpha equal to one is a
  37815. 28:50:16standard penalty. We don't want it to be
  37816. 28:50:18too high. We don't want it to be too
  37817. 28:50:19low. Like we don't want it to be a
  37818. 28:50:20fraction. Um but a penalty of one is
  37819. 28:50:23usually uh good enough.
  37820. 28:50:27Okay, I'm going to show you where that
  37821. 28:50:29comes into play in a moment.
  37822. 28:50:32Um but the whole purpose of doing this
  37823. 28:50:34is to mitigate overfitting, right? Um
  37824. 28:50:37that's what and and doing this penalty
  37825. 28:50:40is is called regularization. So adding
  37826. 28:50:43so going beyond just regular linear
  37827. 28:50:45regression adding this extra penalty to
  37828. 28:50:48to the training process um to penalize
  37829. 28:50:51large weights large coefficients
  37830. 28:50:54um is known as regularization.
  37831. 28:50:58Okay. Um and there's two common
  37832. 28:51:01penalties that are added. Um so there's
  37833. 28:51:04actually two different variations on the
  37834. 28:51:05penalty. Um we're going to study both of
  37835. 28:51:07them and um they're they're known as
  37836. 28:51:10lasso. So if you take linear regression
  37837. 28:51:12and add a particular type of penalty,
  37838. 28:51:14it's known as lasso. If you add another
  37839. 28:51:17type of penalty, it's known as ridge
  37840. 28:51:19regression. We're going to study both of
  37841. 28:51:21those and what their differences are.
  37842. 28:51:23But these are the primary two
  37843. 28:51:26uh regularization tech uh models that
  37844. 28:51:29are used um to take a regular both of
  37845. 28:51:32these take regular linear regression and
  37846. 28:51:34just modify the training process a
  37847. 28:51:37little bit in different ways. Two
  37848. 28:51:39different ways. um using that alpha
  37849. 28:51:43um to penalize the terms in slightly
  37850. 28:51:46different mathematical ways. So we're
  37851. 28:51:48going to learn about these two guys.
  37852. 28:51:50Lasso regression there. Both of these
  37853. 28:51:52are just offshoots of linear regression.
  37854. 28:51:54So underlying model is still linear
  37855. 28:51:56regression. It just adds different types
  37856. 28:51:59of penalties to the training process.
  37857. 28:52:02So both of these are still in the family
  37858. 28:52:05of linear regression. In fact, in um in
  37859. 28:52:09scikitlearn, they both come from they
  37860. 28:52:11both are still from the linear model
  37861. 28:52:13family in inside of the linear model
  37862. 28:52:16module, which is where linear regression
  37863. 28:52:18comes from. So there's still linear
  37864. 28:52:19regression. They just have different
  37865. 28:52:22styles of penalties added to them. Um
  37866. 28:52:25which we're going to see.
  37867. 28:52:28Okay, so just to recap that
  37868. 28:52:32regularization is the process of adding
  37869. 28:52:34a penalty to the training to discourage
  37870. 28:52:38complexity. In this case, we're going to
  37871. 28:52:40discourage large coefficients.
  37872. 28:52:44And um this should help prevent
  37873. 28:52:47overfitting.
  37874. 28:52:49And so uh these are going to lead us to
  37875. 28:52:52two different offshoots of linear
  37876. 28:52:53regression that have two different
  37877. 28:52:55penalties.
  37878. 28:52:56lasso and ridge regression, which we're
  37879. 28:52:59going to uh study next,
  37880. 28:53:01but they they function the same way as
  37881. 28:53:03linear regression. They will just have
  37882. 28:53:06different penalty terms added onto their
  37883. 28:53:08training process um to discourage
  37884. 28:53:12uh discourage um again those large
  37885. 28:53:15weights.
  37886. 28:53:20Okay, any questions about regularization
  37887. 28:53:22before we first look at our we're going
  37888. 28:53:24to look at our first uh variation on on
  37889. 28:53:27our first regularization technique which
  37890. 28:53:28is going to be called lasso regression.
  37891. 28:53:48Okay, let's look at lasso regression. So
  37892. 28:53:50what is lasso regression? It's actually
  37893. 28:53:53lasso is short for least absolute
  37894. 28:53:56shrinkage and selection operator
  37895. 28:53:58regression. Um and this will function by
  37896. 28:54:03adding a particular penalty to the
  37897. 28:54:07linear regression model. So again, it's
  37898. 28:54:09based on linear regression. That's the
  37899. 28:54:11underlying model. It's just that during
  37900. 28:54:13the training process, we are going to um
  37901. 28:54:16add a penalty which has the effect of
  37902. 28:54:21shrinkage of the weights. That's why
  37903. 28:54:23it's called shrinkage. It encourages
  37904. 28:54:25smaller weights through that penalty.
  37905. 28:54:28And it also will shrink some of them so
  37906. 28:54:30much that they'll become zero. And so it
  37907. 28:54:33has has an effect of kind of selection
  37908. 28:54:36which means that some of them get wiped
  37909. 28:54:38out to zero.
  37910. 28:54:40And this means that whatever is left
  37911. 28:54:43over is kind of what's selected as our
  37912. 28:54:45features because the other ones will
  37913. 28:54:48have zero weight applied to them. So
  37914. 28:54:50this penalty will really favor small
  37915. 28:54:54weights um and penalize really large
  37916. 28:54:58weights. In fact, it will favor small
  37917. 28:55:00weight so much that some of them will
  37918. 28:55:02actually um be shrunk to zero um during
  37919. 28:55:06the training process. And the ones that
  37920. 28:55:08are left over are the ones that um are
  37921. 28:55:12the ones that are what we call selected
  37922. 28:55:15because they are the ones that remain in
  37923. 28:55:17in the training um after the other ones
  37924. 28:55:20get uh coefficients of zero. Um now when
  37925. 28:55:24you make some of the coefficient zero
  37926. 28:55:27you are inherently making the model
  37927. 28:55:29simpler right there's less features
  37928. 28:55:31involved in the prediction that or less
  37929. 28:55:33features that have an effect on the
  37930. 28:55:35prediction. So this definitely makes the
  37931. 28:55:37model simpler. This lasso this shrinkage
  37932. 28:55:41and selection uh process makes makes the
  37933. 28:55:45model simpler for sure. Um
  37934. 28:55:48and this is supposed to reduce
  37935. 28:55:50overfitting. Right? If you make the
  37936. 28:55:52model simpler, it's not as complex. It
  37937. 28:55:54has less of a chance of memorizing
  37938. 28:55:56training data and not generalizing over
  37939. 28:55:59to test data. So our whole goal with uh
  37940. 28:56:03regularization is to make our model
  37941. 28:56:05better at generalization, right? Over to
  37942. 28:56:08test data from the original training
  37943. 28:56:10data.
  37944. 28:56:12Um so how does this happen? We have to
  37945. 28:56:15go back to the
  37946. 28:56:18uh training process. If you guys
  37947. 28:56:20remember, I I wrote out this equation a
  37948. 28:56:22little bit earlier, which is the
  37949. 28:56:24distance. This is the sum of squared
  37950. 28:56:27distance between our labels and our
  37951. 28:56:28prediction.
  37952. 28:56:30This is basically the mean squared error
  37953. 28:56:32uh calculation that we're trying to
  37954. 28:56:34reduce when we build our model using the
  37955. 28:56:36training data. Um so this is just in
  37956. 28:56:38standard linear regression. This is the
  37957. 28:56:41um uh sum of squares uh distance, right?
  37958. 28:56:45So this is this is what the model is
  37959. 28:56:47trying to minimize when it learns these
  37960. 28:56:50coefficients.
  37961. 28:56:51So when it learns these coefficients,
  37962. 28:56:53it's trying to minimize this guy
  37963. 28:56:58minimize. It's trying to find the betas
  37964. 28:57:01that minimize this quantity
  37965. 28:57:03mathematically. That's what it's doing.
  37966. 28:57:05Um and there's there's a algorithm that
  37967. 28:57:08will discover what the best betas are
  37968. 28:57:11that actually minimize uses that gives
  37969. 28:57:12us a line of best fit, right? That's
  37970. 28:57:14what we've been talking about for
  37971. 28:57:15regression.
  37972. 28:57:17Now, in regularization,
  37973. 28:57:21here's, by the way, here is that same
  37974. 28:57:22thing, but we've just inserted our model
  37975. 28:57:24for the predictions. This is our model.
  37976. 28:57:28Just a fancy way of writing down our
  37977. 28:57:29model, right? It's the beta 0 plus all
  37978. 28:57:32of these betas. So, beta 1 x1 plus beta
  37979. 28:57:362 x2
  37980. 28:57:38plus on and on and on, right? That's
  37981. 28:57:41that's what this uh means. If you're
  37982. 28:57:43unfamiliar with the sigma notation, it
  37983. 28:57:45just means sum. So it's the sum of all
  37984. 28:57:47these guys or this term. Um, so this is
  37985. 28:57:52this here is just a regular linear
  37986. 28:57:54regression
  37987. 28:57:58uh training regular linear regression
  37988. 28:58:01training. So we the training process
  37989. 28:58:05solves for these parameters, right? It
  37990. 28:58:07solves for these weights. We discover
  37991. 28:58:09what those are by minimizing this
  37992. 28:58:11quantity. That's the whole training
  37993. 28:58:13process. Um, but when we do lasso,
  37994. 28:58:18we add a penalty which is this.
  37995. 28:58:23Here is our penalty.
  37996. 28:58:27So basically um we take our linear
  37997. 28:58:31regression training which is this and we
  37998. 28:58:34add on a penalty which is this. And you
  37999. 28:58:37can see exactly what this penalty when
  38000. 28:58:40when you minimize this penalty. It's
  38001. 28:58:42when these weights are small. So this
  38002. 28:58:45encourages
  38003. 28:58:46So minimizing this quantity encourages
  38004. 28:58:50small weights
  38005. 28:58:53encourages small betas
  38006. 28:58:57beta I
  38007. 28:58:59right you or in this case beta j sorry
  38008. 28:59:04this encourages small beta js uh because
  38009. 28:59:07we want this thing to be minimized
  38010. 28:59:11minimized
  38011. 28:59:13so Um, what's going to make this minimal
  38012. 28:59:16is of course the line of best fit and
  38013. 28:59:18small weights, right? Are going to make
  38014. 28:59:20are going to bring this error down the
  38015. 28:59:24most.
  38016. 28:59:26So, um, and here's our alpha, right?
  38017. 28:59:28Here's our alpha. So, you can encourage
  38018. 28:59:30a higher penalty with a larger alpha or
  38019. 28:59:33a lower penalty. If alpha equals zero,
  38020. 28:59:36what happens to that term? It just goes
  38021. 28:59:39away. So if alpha equals zero, there's
  38022. 28:59:41no penalty and we're back to uh we're
  38023. 28:59:45back to regular
  38024. 28:59:47linear regression.
  38025. 28:59:50We just have regular linear regression
  38026. 28:59:51because we have no penalty at that point
  38027. 28:59:52when alpha equals zero. So the smaller
  38028. 28:59:56alpha is, the less penalty we're
  38029. 28:59:59enforcing and in the regularization.
  38030. 29:00:03Okay.
  38031. 29:00:04Now what happens is in reality when you
  38032. 29:00:07train with lasso. So this is lasso is
  38033. 29:00:10this particular penalty. This is called
  38034. 29:00:12the lasso penalty
  38035. 29:00:14or sometimes um people call this the L1
  38036. 29:00:17penalty.
  38037. 29:00:19Um L1 just comes from the fact that this
  38038. 29:00:23is the first power or absolute value. Um
  38039. 29:00:26so it's not a squared penalty, it's a
  38040. 29:00:28single uh single power penalty
  38041. 29:00:31um there. But when you add this lasso
  38042. 29:00:35penalty, what can happen is it it does
  38043. 29:00:38because the because you're minimizing
  38044. 29:00:40this, it does encourage some of these
  38045. 29:00:42weights to become zero.
  38046. 29:00:46So some if you're really trying to get
  38047. 29:00:48the lowest quantity of this,
  38048. 29:00:51the lower the better.
  38049. 29:00:54What makes this thing lower is of course
  38050. 29:00:57if some of these go away if some of
  38051. 29:00:58these go to zero then that of course
  38052. 29:01:01will lower this as much as we as much as
  38053. 29:01:03possible right so what happens during
  38054. 29:01:05the training is some of these
  38055. 29:01:08coefficients actually they're encouraged
  38056. 29:01:10to be small because of this penalty but
  38057. 29:01:12some of them will actually become will
  38058. 29:01:15actually become zero um in order to get
  38059. 29:01:18the best model the best fit some of
  38060. 29:01:20these will actually get so small that
  38061. 29:01:22they'll basically become zero
  38062. 29:01:24And that means that that that feature
  38063. 29:01:27basically has no effect anymore. It's
  38064. 29:01:30it's been the model has been simplified,
  38065. 29:01:33right? That feature no longer really has
  38066. 29:01:34an effect.
  38067. 29:01:40So just to call out the alpha again, um
  38068. 29:01:42if alpha zero some code, uh basically
  38069. 29:01:46you have your linear regression, you're
  38070. 29:01:48back to linear regression because alpha
  38071. 29:01:500 is just wiping this out and you're
  38072. 29:01:51back to linear regression.
  38073. 29:01:54um if alpha is infinity. Now if alpha is
  38074. 29:01:56infinity that's an extreme. So if alpha
  38075. 29:01:58is infinity the only way to make this
  38076. 29:02:00minimize is if all your coefficients are
  38077. 29:02:02zero. If every beta is zero then this
  38078. 29:02:05will lower the the error as as much as
  38079. 29:02:07possible. So you basically have no
  38080. 29:02:09model. So if all coefficients are zero
  38081. 29:02:12you have no model and that's useless. So
  38082. 29:02:14you don't want your penalty you don't
  38083. 29:02:17want your alpha to be huge is what this
  38084. 29:02:19is saying. You also don't want your
  38085. 29:02:20alpha to be small. you're basically back
  38086. 29:02:22to linear regression. So you want
  38087. 29:02:23something in between. Um and the typical
  38088. 29:02:27typical value is alpha equals 1.
  38089. 29:02:31Typical is alpha equals 1
  38090. 29:02:35to have some level of penalty there. So
  38091. 29:02:38just a regular kind of regular penalty
  38092. 29:02:40term.
  38093. 29:02:48But we are actually going to have a way
  38094. 29:02:50to test and evaluate which alphas are
  38095. 29:02:52the best.
  38096. 29:03:07Um,
  38097. 29:03:08basically you can yeah you can have a
  38098. 29:03:12you can have a penalty that's close to
  38099. 29:03:13zero. You can get rid of this if just a
  38100. 29:03:16regular linear regression performs
  38101. 29:03:18pretty well. You can basically have no
  38102. 29:03:20penalty in that case.
  38103. 29:03:23Yeah. So nearer zero or like it could be
  38104. 29:03:26that adding a little bit of penalty
  38105. 29:03:28actually helps the overfitting and it
  38106. 29:03:30could be really small. One thing that
  38107. 29:03:32we're basically going to do is have a
  38108. 29:03:34strategy to try out different alphas.
  38109. 29:03:40try different alphas
  38110. 29:03:43and evaluate performance
  38111. 29:03:47and then we can decide which so that's
  38112. 29:03:49what we're going to do is have a
  38113. 29:03:51strategy to just plug in different
  38114. 29:03:52alphas generate the like train the model
  38115. 29:03:56and then see what its performance is and
  38116. 29:03:58see if those alphas are good what what
  38117. 29:04:00which alpha is the best we can evaluate
  38118. 29:04:03that
  38119. 29:04:04because we can train the model and see
  38120. 29:04:06what it performance is
  38121. 29:04:11right.
  38122. 29:04:14Yeah. Yeah. So, we'll do that. We'll
  38123. 29:04:15practice that.
  38124. 29:04:24Okay. Great. Any other questions about
  38125. 29:04:26this lasso regression? So, remember this
  38126. 29:04:28is linear regression here. This is the
  38127. 29:04:30this is how you're training to find the
  38128. 29:04:33betas in linear regression. So this is
  38129. 29:04:35just linear regression uh um training
  38130. 29:04:40function there.
  38131. 29:04:42We're adding a penalty which is this is
  38132. 29:04:44the lasso penalty
  38133. 29:04:47lasso penalty there right we're adding
  38134. 29:04:49that this is known as regularization
  38135. 29:04:52and the goal of regularization is to
  38136. 29:04:56prevent overfitting. So you add a
  38137. 29:04:57penalty here this makes the model
  38138. 29:05:00simpler which prevents overfitting.
  38139. 29:05:03helps you generalize better when it's
  38140. 29:05:06simpler.
  38141. 29:05:16Any questions conceptually on this?
  38142. 29:05:18We're going to do a code example with it
  38143. 29:05:19coming up, but any questions on this?
  38144. 29:05:37Uh yeah, you you so that's the thing,
  38145. 29:05:39Ronald, is you may be willing to
  38146. 29:05:41sacrifice some accuracy in order to
  38147. 29:05:44generalize to unseen data because
  38148. 29:05:46remember that's what we're really trying
  38149. 29:05:48to get after is we may be willing to
  38150. 29:05:51sacrifice some accuracy on this training
  38151. 29:05:52data in order to have it perform better
  38152. 29:05:54on the test data, right? we may be
  38153. 29:05:57willing to do that. That's a willing
  38154. 29:05:59that's an okay sacrifice
  38155. 29:06:02as long like if if it generalizes
  38156. 29:06:05better. That's what we want. That's what
  38157. 29:06:07we're trying to do here is add a
  38158. 29:06:09penalty, make the model simpler, and
  38159. 29:06:13help it generalize better to new and
  38160. 29:06:16unseen data. Right?
  38161. 29:06:20That's that picture I've been using with
  38162. 29:06:21the with the um train and test split.
  38163. 29:06:25Where is the square?
  38164. 29:06:27So in the model there's no square. So
  38165. 29:06:30remember the model is the model is this
  38166. 29:06:36um equation uh that has no squares in
  38167. 29:06:38it, right? It's beta 0 plus beta 1 x1
  38168. 29:06:43plus beta 2 x2 plus beta n xn.
  38169. 29:06:50That's the that's the linear regression
  38170. 29:06:52model. This is the now this this is the
  38171. 29:06:55model but this is the equation that
  38172. 29:06:58helps us train and find the betas. This
  38173. 29:07:01is how this is what we find the betas
  38174. 29:07:03with. So we'll continue. Um we were
  38175. 29:07:07talking about the lasso regression which
  38176. 29:07:10uh adds it takes linear regression right
  38177. 29:07:13which is this optimization and adds in a
  38178. 29:07:16penalty um scaled by the alpha. Um, and
  38179. 29:07:20what that does in order to minimize this
  38180. 29:07:23whole thing, it encourages these to be
  38181. 29:07:26small uh as possible. Um, which makes
  38182. 29:07:30the model simpler, right? The weights
  38183. 29:07:32don't get overly big and complex. Um,
  38184. 29:07:35they they tend to stay small. In fact,
  38185. 29:07:37some of them can even go all the way to
  38186. 29:07:39zero. Um, which makes the model even
  38187. 29:07:41more simpler,
  38188. 29:07:43right? Um, so let's practice uh using it
  38189. 29:07:47in code. It's actually really easy to
  38190. 29:07:48use. It's going to be essentially the
  38191. 29:07:51same uh style and and code as linear
  38192. 29:07:55regression except we are um just going
  38193. 29:07:58to have to uh put in our alpha parameter
  38194. 29:08:01um when we use the lasso. So here we are
  38195. 29:08:06um from the linear model family right
  38196. 29:08:10which makes sense. It's a linear
  38197. 29:08:12regression offshoot that has this
  38198. 29:08:13penalty in it during the training. um we
  38199. 29:08:15are grabbing our lasso regression. Um it
  38200. 29:08:19also has a version of the lasso that
  38201. 29:08:22we're going to take a look at that is
  38202. 29:08:23used for cross validation which is
  38203. 29:08:26really um convenient as well. So it has
  38204. 29:08:29a cross validation lasso which is a
  38205. 29:08:31really convenient um combination of
  38206. 29:08:34basically cross val score and lasso um
  38207. 29:08:37all in one. So it actually is really
  38208. 29:08:39nice to use that way. Um so we'll take a
  38209. 29:08:42look at that example. Um, but we are
  38210. 29:08:45importing it. The main thing is going to
  38211. 29:08:47be the lasso model here. Um, we're going
  38212. 29:08:50to be using a different data set for
  38213. 29:08:51this one. So, not the ocean uh data, but
  38214. 29:08:54this hitters data, which is a baseball
  38215. 29:08:56data set. Um, so it has 322 rows um with
  38216. 29:09:0120 different columns and it looks like
  38217. 29:09:03this. So, you want to download that one.
  38218. 29:09:06Um, hopefully you guys have access to
  38219. 29:09:08that one.
  38220. 29:09:11Um,
  38221. 29:09:13so I will upload it into
  38222. 29:09:16this.
  38223. 29:09:19So give me a moment.
  38224. 29:09:25There's that. And then we can run this.
  38225. 29:09:29Okay. So we are displaying the data and
  38226. 29:09:32so it has um the the hitters names and
  38227. 29:09:36then it has a bunch of different
  38228. 29:09:37statistics. These are all baseball
  38229. 29:09:38statistics.
  38230. 29:09:40Um, if you're unfamiliar with with them,
  38231. 29:09:42that's okay. It's not a big deal. Um,
  38232. 29:09:44but just different baseball stats here.
  38233. 29:09:49Okay. Were you guys able to load that?
  38234. 29:09:51Um, if you're following along, were you
  38235. 29:09:52able to load that? You should have
  38236. 29:09:54access to this data. The hitters CSV.
  38237. 29:09:58This is the one we're going to use for
  38238. 29:09:59the lasso model
  38239. 29:10:02to build a lasso model.
  38240. 29:10:12Yeah.
  38241. 29:10:19Okay. Able to load that one. Perfect.
  38242. 29:10:22Okay. So, able to load that one. Um, and
  38243. 29:10:25we take a look at the the head. Um, so
  38244. 29:10:28we're actually going to uh drop this
  38245. 29:10:31unnamed column because we don't care
  38246. 29:10:33about their name. it's actually just the
  38247. 29:10:35batter's name which is not going to be
  38248. 29:10:37useful in modeling. Um so and remember
  38249. 29:10:40that's generally true like an ID, a user
  38250. 29:10:44ID, like a customer ID, a name, that's
  38251. 29:10:47usually not going to be useful in any
  38252. 29:10:48kind of modeling. So we're actually just
  38253. 29:10:50going to drop that uh column and we're
  38254. 29:10:52going to do it in place.
  38255. 29:10:54And access equals 1 means we're dropping
  38256. 29:10:56that column. Um, so we're going to drop
  38257. 29:10:59that and we should no longer have that
  38258. 29:11:02column and we have all of these guys
  38259. 29:11:03now. So you want to run that. This will
  38260. 29:11:05drop that. Um, this will drop drops the
  38261. 29:11:10column in place.
  38262. 29:11:14Um, and now we can see we have uh all we
  38263. 29:11:18have this data where um we have this
  38264. 29:11:21data where it's now removed. So, this
  38265. 29:11:24that column is now gone and now we have
  38266. 29:11:26these guys. Um, do you notice anything
  38267. 29:11:30about this
  38268. 29:11:32from the info?
  38269. 29:11:38Looks like we have a couple categorical
  38270. 29:11:39features, a few of them, league and
  38271. 29:11:42division
  38272. 29:11:44and new league. What do you notice about
  38273. 29:11:47this
  38274. 29:11:56nullles? Yep. So, there's definitely
  38275. 29:11:57some missing data there um that we're
  38276. 29:12:00going to have to deal with.
  38277. 29:12:05So, it looks like there are uh there are
  38278. 29:12:0959.
  38279. 29:12:11Um there are 59. Now we could we the
  38280. 29:12:15alternative to doing that is we could uh
  38281. 29:12:17we could just use our usual code where
  38282. 29:12:19we do dfis
  38283. 29:12:21uh isnull.
  38284. 29:12:24Um and then we do uh dot sum to total
  38285. 29:12:29those up across our different columns.
  38286. 29:12:31And we can see that uh we have 59 of
  38287. 29:12:34those in the salary column. That's this
  38288. 29:12:37is the standard way of doing that,
  38289. 29:12:38right?
  38290. 29:12:42standard way of doing that. And we have
  38291. 29:12:43so we have 59 of those.
  38292. 29:12:5359 of those. So we have to deal with it.
  38293. 29:12:55Any ideas on how to deal with it?
  38294. 29:13:02Any ideas on how to deal with it? This
  38295. 29:13:04is now this is 59 out of 300.
  38296. 29:13:08So,
  38297. 29:13:09what do you guys think about that? It's
  38298. 29:13:11a little bit different than 200 out of
  38299. 29:13:1220,000. A little bit different. We have
  38300. 29:13:16We have about 60 out of 300.
  38301. 29:13:21There's a decent amount.
  38302. 29:13:25Any ideas on how to handle this one?
  38303. 29:13:30Replace. Yep, we should replace. What do
  38304. 29:13:32you think we should replace with?
  38305. 29:13:35It's a float. It's a floating point uh
  38306. 29:13:38value.
  38307. 29:13:49By the way, something unique about this
  38308. 29:13:51that's a little different than usual,
  38309. 29:13:52too, is that the uh this is actually the
  38310. 29:13:56column we're going to use as our label.
  38311. 29:13:58So, we're actually going to predict the
  38312. 29:14:00salary based on the uh based on the um
  38313. 29:14:04rest of the features. So, we definitely
  38314. 29:14:07need to fill in these nles, right?
  38315. 29:14:09Because they're actually going to be the
  38316. 29:14:11labels.
  38317. 29:14:12We're missing some labels uh in our
  38318. 29:14:15data.
  38319. 29:14:16We definitely need to fill them in.
  38320. 29:14:17Yeah. So, we're going to replace them.
  38321. 29:14:29All right. So, we'll we will replace
  38322. 29:14:31them down below. That's going to be
  38323. 29:14:32coming up. Uh we'll come back and
  38324. 29:14:34replace them. um before we replace them,
  38325. 29:14:37we're actually going to get our uh one
  38326. 29:14:40hot encodings for those three different
  38327. 29:14:42um features we have. Um so we do uh get
  38328. 29:14:49dummies with this. Now um of course we
  38329. 29:14:52don't need to do this if we just so this
  38330. 29:14:55code we don't need to do if we just pass
  38331. 29:14:57in the dype here
  38332. 29:15:01um which is uh then we don't need to do
  38333. 29:15:04this. So we can comment this out.
  38334. 29:15:08Um so now what I want you guys to notice
  38335. 29:15:11is this is the alternative to what we
  38336. 29:15:13did before where we are purposely just
  38337. 29:15:15doing these columns not the whole data
  38338. 29:15:17frame but just doing these columns and
  38339. 29:15:20then we can um concatenate those these
  38340. 29:15:26one hot encodings. We're going to
  38341. 29:15:27concatenate back to the data frame.
  38342. 29:15:30Right? So if we do our dummies and then
  38343. 29:15:33do dummies.info info. Um, we can see
  38344. 29:15:36that we end up with six new columns. And
  38345. 29:15:39in fact, we can do dummies.head
  38346. 29:15:43and take a look at what those are.
  38347. 29:15:45Right? So, these are league A, league
  38348. 29:15:48uh, N, division E, W, division W, new
  38349. 29:15:51league A, new league N.
  38350. 29:15:54Okay.
  38351. 29:15:56So, um these are uh these are our one
  38352. 29:16:00hot encodings for these three different
  38353. 29:16:02features which are strings, right? So,
  38354. 29:16:04those those features were strings. If
  38355. 29:16:06you go back up, those were our only
  38356. 29:16:07string features we had. So, we've one
  38357. 29:16:10hot encoded those so we can use them in
  38358. 29:16:11our model. What we need to do is just
  38359. 29:16:14concatenate this back to our data frame.
  38360. 29:16:17Right? So, we just need to concatenate
  38361. 29:16:19it back into our data.
  38362. 29:16:28Okay. So, what we're going to do then is
  38363. 29:16:31we're going to grab um we're going to
  38364. 29:16:35grab Y as our salary. And of course,
  38365. 29:16:37we're going to fill nles on that Y
  38366. 29:16:39coming up shortly. But we're going to
  38367. 29:16:42grab Y as our salary and X new. Now
  38368. 29:16:45before building a full X, we're going to
  38369. 29:16:48take a look at X numerical as our data
  38370. 29:16:51frame minus these columns. The reason
  38371. 29:16:54we're doing minus those is because we
  38372. 29:16:57are going to concatenate our dummy
  38373. 29:17:00variables back into this that are going
  38374. 29:17:02to replace these guys. So we're going to
  38375. 29:17:05replace these anyways with our one hot
  38376. 29:17:07encodings. We don't want the strings. So
  38377. 29:17:09we're going to get rid of those. And
  38378. 29:17:12we're also going to get rid of the
  38379. 29:17:13salary because that's going to be part
  38380. 29:17:15of our that's just a label. So we don't
  38381. 29:17:17want that in the X, the eventual X.
  38382. 29:17:24Are you guys able to run this one?
  38383. 29:17:28Hope I'm not going too fast. You guys
  38384. 29:17:30able to run this? And does it make
  38385. 29:17:32sense? What we're doing is we're putting
  38386. 29:17:34our labels in Y, which is what we
  38387. 29:17:36usually do. So, we're going to predict
  38388. 29:17:38the salary
  38389. 29:17:40and we're getting ready to build the X.
  38390. 29:17:42But before we first want to get rid of
  38391. 29:17:44those one hot the strings. This is
  38392. 29:17:46getting rid of the strings
  38393. 29:17:49and this is getting rid of the label.
  38394. 29:17:51And that's going to be part of our
  38395. 29:17:52features. What we need to do is build
  38396. 29:17:54our final X by concatenating our dummies
  38397. 29:17:56with this. Do you guys see that? We're
  38398. 29:17:59going to concatenate our dummies with
  38399. 29:18:00this to build our final X.
  38400. 29:18:04But but prior to doing that, we need to
  38401. 29:18:06get rid of these string columns here. So
  38402. 29:18:09we're dropping those
  38403. 29:18:11dropping those from the uh data frame uh
  38404. 29:18:15and getting a numerical uh x numerical
  38405. 29:18:18here.
  38406. 29:18:20You can see the columns of that are just
  38407. 29:18:22these guys here. So the the results we
  38408. 29:18:25need to concatenate our we need to
  38409. 29:18:27concatenate this guy um into this and
  38410. 29:18:30then that'll be our full x all of our
  38411. 29:18:32features.
  38412. 29:18:37Okay. So you can see x is going to be
  38413. 29:18:40pd.con
  38414. 29:18:42of this with our dummies.
  38415. 29:18:46This with our dummies. And um
  38416. 29:18:51uh instead of doing this, I'm actually
  38417. 29:18:53going to do the full dummies. We don't
  38418. 29:18:55need to
  38419. 29:18:56um pick just a few columns. We're
  38420. 29:18:59actually going to do our full dummies
  38421. 29:19:01here and um do x equals 1. Now, the
  38422. 29:19:05reason that's the case is because um
  38423. 29:19:08this will get rid of one column per
  38424. 29:19:12feature and basically assume that if you
  38425. 29:19:15have a if you have a zero, the other one
  38426. 29:19:17should be a one. If you have a one, the
  38427. 29:19:19other one should be a zero. Um so it
  38428. 29:19:22basically makes that assumption because
  38429. 29:19:24we only have two of them. Um so whenever
  38430. 29:19:28there's a one, the other should be zero.
  38431. 29:19:30Um, so you can get away with just having
  38432. 29:19:33these three, but um I think it makes
  38433. 29:19:35more sense to just have to have the full
  38434. 29:19:38dummies,
  38435. 29:19:41but by process of elimination, you can
  38436. 29:19:43get away with just using two of them
  38437. 29:19:44because anytime you have a zero, the
  38438. 29:19:47other one should be the other feature
  38439. 29:19:48would have been would have been a one,
  38440. 29:19:51right? And vice versa, when there's a
  38441. 29:19:53one, the other feature would have been a
  38442. 29:19:54zero.
  38443. 29:20:00So we do that one.
  38444. 29:20:11And you can see all of our uh all of our
  38445. 29:20:14one hot encoding features end up back in
  38446. 29:20:16there.
  38447. 29:20:20So this is the code that I want you guys
  38448. 29:20:22to run. I think it makes more sense. It
  38449. 29:20:25follows along what we've been doing.
  38450. 29:20:27um which will concatenate our dummies
  38451. 29:20:30back to our features here to build out
  38452. 29:20:32our full X. So now X is all of our
  38453. 29:20:34features. Um remember X
  38454. 29:20:38X contains all of our features
  38455. 29:20:44now.
  38456. 29:20:46So X contains all of our features and so
  38457. 29:20:49we have all of this now.
  38458. 29:20:56Okay. Were you guys able to run this
  38459. 29:20:57one?
  38460. 29:21:00Damn. We have y, we have x. We still
  38461. 29:21:03need to deal with the nles in y. So that
  38462. 29:21:06something we still need to deal with.
  38463. 29:21:14But hopefully you have this. Now
  38464. 29:21:18all these are numerical.
  38465. 29:21:21So that should be good with the model.
  38466. 29:21:25That's one thing about X is you should
  38467. 29:21:27you our X should have all numerical
  38468. 29:21:31features, right? Because it's going to
  38469. 29:21:33go into a model to to learn those betas.
  38470. 29:21:36So it needs to have all numerical
  38471. 29:21:38features,
  38472. 29:21:40right? These are going to be all
  38473. 29:21:41numerical, which makes sense. We change
  38474. 29:21:43we did one hot encoding to change all
  38475. 29:21:45those guys to numerical.
  38476. 29:21:55Sorry, I'm scrolling down.
  38477. 29:22:00Okay, we do fill in the nil later. Okay.
  38478. 29:22:13Okay.
  38479. 29:22:15Any questions so far? So, we're just
  38480. 29:22:16getting our data ready. We haven't
  38481. 29:22:17applied the lasso yet, but we're just
  38482. 29:22:19doing some prep. Now, hopefully you guys
  38483. 29:22:22recognize th these are some standard
  38484. 29:22:25steps that we're taking when we do our
  38485. 29:22:28modeling. We have to do these data prep
  38486. 29:22:31steps. They're necessary. And so, if it
  38487. 29:22:34seems like it's a lot of work, that's
  38488. 29:22:35because it is. It is work that you do to
  38489. 29:22:39prepare your data to get ready for
  38490. 29:22:41modeling. You have to do that. Okay.
  38491. 29:22:47So, we're doing that here. Um, now we're
  38492. 29:22:50going to do our train test split because
  38493. 29:22:52we're just going to do uh we're going to
  38494. 29:22:54do hold out here. So, we're doing a
  38495. 29:22:56train test split with about with a test
  38496. 29:22:58size of about 0.25. So, again, anywhere
  38497. 29:23:00between 0.2 to.3 would be okay.
  38498. 29:23:04Um, so uh
  38499. 29:23:08it's our choice. We could do 0 2. We
  38500. 29:23:10could do 3. We could do anywhere in
  38501. 29:23:11between there. We're doing 0.25. That's
  38502. 29:23:14fine. Um, that's okay. So, we we build
  38503. 29:23:17our train test split right there. Um, so
  38504. 29:23:21pretty pretty simple and we've seen that
  38505. 29:23:23a bunch of times with our X and our Y
  38506. 29:23:27data frames. There we have our train
  38507. 29:23:30test split.
  38508. 29:23:33Okay.
  38509. 29:23:35Um, now what we're going to do is do our
  38510. 29:23:38our scaling. So, we're we didn't do this
  38511. 29:23:41last time, but we're going to do this
  38512. 29:23:43now as uh because we should get in the
  38513. 29:23:45habit of doing that. Um is um we're
  38514. 29:23:50going to um go ahead and scale our
  38515. 29:23:54features and we're going to use the
  38516. 29:23:56standard scaler here uh to do that
  38517. 29:23:59scaling. Okay. Now, we could use minmax
  38518. 29:24:02scaler. That's fine, too. We're just
  38519. 29:24:04going to use the standard scaler here.
  38520. 29:24:06Um and remember we are going to uh um
  38521. 29:24:11use the standard scaler from sklearn and
  38522. 29:24:14we're going to transform our features uh
  38523. 29:24:18uh according to our um according to our
  38524. 29:24:23training data. So we have our
  38525. 29:24:27pre-processing standard scaler here. So
  38526. 29:24:29we import that guy and then we um build
  38527. 29:24:33our standard scaler and fit it on the
  38528. 29:24:36training data only on the numerical
  38529. 29:24:39features. Um so that's which is going to
  38530. 29:24:44be uh all of these guys. So we're doing
  38531. 29:24:48the scaling on all of these guys. Now
  38532. 29:24:50something to note is that we are not
  38533. 29:24:53scaling all of these one hot encodings
  38534. 29:24:56mainly because it doesn't make sense to
  38535. 29:24:58scale those really. They're zero or one.
  38536. 29:25:01They don't need to be scaled, right?
  38537. 29:25:03They're already zero and one. So they're
  38538. 29:25:06they don't need to be even if we were
  38539. 29:25:07doing minmax scaling, it's going to put
  38540. 29:25:09them between zero and one. It wouldn't
  38541. 29:25:11affect it really, right? So these one
  38542. 29:25:14hot encoding features, we're not going
  38543. 29:25:15to scale because they're they're always
  38544. 29:25:17going to be zero or one.
  38545. 29:25:20There's no need to scale them really.
  38546. 29:25:22Um, but we're going to scale all the
  38547. 29:25:24other features here that are floats.
  38548. 29:25:27So that's these guys here. These
  38549. 29:25:30numerical features we're going to scale.
  38550. 29:25:33Okay.
  38551. 29:25:34Don't need to we don't really need to
  38552. 29:25:36scale the one hot encoding. Uh, it's
  38553. 29:25:38pretty much already scaled.
  38554. 29:25:46Oh, you should change that. Um, go back
  38555. 29:25:48and rerun go back and rerun this, but
  38556. 29:25:51make sure you have your data type as int
  38557. 29:25:53here.
  38558. 29:25:55Make sure you add that in there to
  38559. 29:25:57change that over to integer and rerun
  38560. 29:25:59that and then rerun the rerun the
  38561. 29:26:03concatenation.
  38562. 29:26:04So, make sure you run this
  38563. 29:26:07and then u make sure you rerun this and
  38564. 29:26:10rerun the concatenation part which is uh
  38565. 29:26:14this
  38566. 29:26:27Okay. So, we go ahead and fit the um
  38567. 29:26:32scaler to this data and then we're going
  38568. 29:26:34to transform our training features,
  38569. 29:26:36those numerical features. um we're and
  38570. 29:26:40then we're going to uh transform these
  38571. 29:26:42features uh uh the test features in the
  38572. 29:26:46same way. So we're going to perform the
  38573. 29:26:48same transformation from the scaler on
  38574. 29:26:50the test data. So that's something
  38575. 29:26:52really important I want to note here is
  38576. 29:26:53that we always scale both the training
  38577. 29:26:59and test data. We always scale both. Of
  38578. 29:27:02course, we're going to train the model
  38579. 29:27:03on the training data. Um, but we are
  38580. 29:27:07going to also test it on the testing
  38581. 29:27:10data and it also needs to be scaled
  38582. 29:27:12because our model that we build is going
  38583. 29:27:14to assume scaled features. The
  38584. 29:27:17coefficients that it learns are going to
  38585. 29:27:19be assuming scaled features.
  38586. 29:27:22So, we need to also scale our test data
  38587. 29:27:26in the same way. So, we're doing that as
  38588. 29:27:29well.
  38589. 29:27:33So, we scale that and now we have our uh
  38590. 29:27:37training and testing features have been
  38591. 29:27:39scaled.
  38592. 29:27:41No, we haven't replaced. We're going to
  38593. 29:27:42do that. We have not yet. We're going to
  38594. 29:27:45do that coming up in a minute. Yeah, we
  38595. 29:27:48haven't done that. Um, it is it is the
  38596. 29:27:51label. We definitely need to replace
  38597. 29:27:53NLES. We just haven't done it yet
  38598. 29:27:55because it's not in the features and
  38599. 29:27:56we're doing all of our uh uh
  38600. 29:27:58pre-processing to our pre-processing to
  38601. 29:28:00our features.
  38602. 29:28:10Yeah. So, we're definitely we need to
  38603. 29:28:12we're going to in a minute.
  38604. 29:28:16Okay. So, if you look at the data now,
  38605. 29:28:18it's all been scaled. So, these are all
  38606. 29:28:20um zcores. These are all on a much
  38607. 29:28:23better scale now. Um, and these are we
  38608. 29:28:28still have our one hot encoding features
  38609. 29:28:29which are zero or one. So this scaling
  38610. 29:28:34should lead to a better model than if we
  38611. 29:28:36didn't scale. So scaling is really
  38612. 29:28:39important. We can see that here.
  38613. 29:28:46Okay.
  38614. 29:28:48Now, um, let me ask you guys, were you
  38615. 29:28:50able to run the scaling? Are you caught
  38616. 29:28:53up to here? If you're following along,
  38617. 29:28:55were you able to run the scaling?
  38618. 29:29:08Okay, great. Great.
  38619. 29:29:16Awesome.
  38620. 29:29:19Okay. So, uh what we're going to do now
  38621. 29:29:22is replace NLES in the uh replace NLES
  38622. 29:29:28by calculating the median of the data.
  38623. 29:29:32So, what I want you to notice is that we
  38624. 29:29:35are taking the NLES now this is um this
  38625. 29:29:39is on purpose is we are purposely taking
  38626. 29:29:42the NLES um out of the median
  38627. 29:29:45calculation. So we're skipping the NLES
  38628. 29:29:47when we compute the median because we
  38629. 29:29:49don't want those NLES to affect the
  38630. 29:29:51median calculation.
  38631. 29:29:53Um so we compute a median salary here
  38632. 29:29:56and then we fill our NLES with the
  38633. 29:29:59median salary um from the training data.
  38634. 29:30:03So this is our choice. This is a choice
  38635. 29:30:06um to use the median and it's also a
  38636. 29:30:10choice to use the training set median
  38637. 29:30:14for both train and test. What we could
  38638. 29:30:18have done this is an alternative that we
  38639. 29:30:20could have done is use the entire column
  38640. 29:30:24and then um use the median of all of the
  38641. 29:30:28data to replace. That's really up to us.
  38642. 29:30:31Um this is one way of doing it. We could
  38643. 29:30:34have done before we did the split. We
  38644. 29:30:37could have um filled in with the median
  38645. 29:30:41earlier. We chose to do it here mainly
  38646. 29:30:44because it doesn't affect the features.
  38647. 29:30:46So we could have done this earlier and
  38648. 29:30:48did it before we did the split and
  38649. 29:30:50filled the NAS. Um really doesn't it's
  38650. 29:30:53doesn't matter that much which way we do
  38651. 29:30:54it. Um but we do need to fill in NLES.
  38652. 29:30:58We cannot have those be null when we
  38653. 29:31:00when we put it into our model. So some
  38654. 29:31:02way we need to fill in nulls. Um and so
  38655. 29:31:06in this strategy we're filling in our y
  38656. 29:31:08train um with the median salary from our
  38657. 29:31:12training data. And same with this we're
  38658. 29:31:14filling in with the median salary of the
  38659. 29:31:16training data as well. But that's a
  38660. 29:31:18choice. We could fill in with the mean
  38661. 29:31:21with the average. Um we could fill in
  38662. 29:31:25with the we could do it with all the
  38663. 29:31:27data together before we split it. we
  38664. 29:31:29could have filled in with all of the the
  38665. 29:31:31median across the whole data set. Um
  38666. 29:31:34either one works. You can do it either
  38667. 29:31:36way, but we we did it um later here to
  38668. 29:31:40show that it doesn't really affect the
  38669. 29:31:42features. So, we can choose when we do
  38670. 29:31:44it, right? It doesn't affect the
  38671. 29:31:46features at all. So, we can do all of
  38672. 29:31:48our pre-processing on the features and
  38673. 29:31:50then do our label uh filling in NLES um
  38674. 29:31:54if if we have them.
  38675. 29:32:00uh x numerical. Um make sure you're
  38676. 29:32:03running uh this
  38677. 29:32:07uh x numerical was defined here
  38678. 29:32:12when we split it apart um from
  38679. 29:32:17uh when we dropped these columns here.
  38680. 29:32:19So make sure you're running this. This
  38681. 29:32:21is x numerical.
  38682. 29:32:23It's defined there.
  38683. 29:32:25So go back up this uh this cell
  38684. 29:32:29where we split apart the y and we and we
  38685. 29:32:31have the x here x numerical.
  38686. 29:32:35Make sure you run this.
  38687. 29:32:44Make sure you run this. And then you can
  38688. 29:32:46run these. Then you run this to build x.
  38689. 29:32:59All right.
  38690. 29:33:01Are we up to here with this filling in
  38691. 29:33:04the labels?
  38692. 29:33:06Uh because then we can build our model
  38693. 29:33:09once we're up to here. We've scaled
  38694. 29:33:11everything. We filled in our NLES.
  38695. 29:33:14We've gotten one hot encoding.
  38696. 29:33:21Yeah, it is. That's why you know that's
  38697. 29:33:23why we spend a lot of uh time on model
  38698. 29:33:25on data preparation with pandas, right?
  38699. 29:33:27That's why we did all that pandas work
  38700. 29:33:29for sure. Yes, there is a lot of work
  38701. 29:33:31before we can build a model.
  38702. 29:33:34Yes, the mo do you guys notice that like
  38703. 29:33:37the modeling is relatively easy. It's
  38704. 29:33:38just a fit and predict. The modeling is
  38705. 29:33:41actually really easy. It's all the other
  38706. 29:33:43work that's that's more involved, right?
  38707. 29:33:47more code.
  38708. 29:33:50The modeling itself is really easy.
  38709. 29:33:53It's just it's just one line of like
  38710. 29:33:55ffit.
  38711. 29:33:58Yeah, pretty easy to do.
  38712. 29:34:03And then you do evaluation which is a
  38713. 29:34:06couple lines.
  38714. 29:34:27Yep. There's these are all the these are
  38715. 29:34:30the common steps. All these steps we're
  38716. 29:34:32doing are very very prototypical in
  38717. 29:34:34model building is you let's just go back
  38718. 29:34:37through this to see what we did, right?
  38719. 29:34:38We imported our data. Um we analy we
  38720. 29:34:43dropped this name column because it's
  38721. 29:34:45not useful to us. So we dropped that. Um
  38722. 29:34:48we filled in the NLES eventually. Um but
  38723. 29:34:52you know if there were any nulls in our
  38724. 29:34:54features we would have to deal with
  38725. 29:34:55those as well by replacing them or
  38726. 29:34:57dropping the rows like we did earlier.
  38727. 29:35:00Um
  38728. 29:35:01and then we do one hot encoding because
  38729. 29:35:04of course we can't have any string
  38730. 29:35:05columns going into our models. We got a
  38731. 29:35:07one hot encode.
  38732. 29:35:09Um we uh then build our X and Y by
  38733. 29:35:13concatenating the one hot encoded back
  38734. 29:35:16to the numerical features.
  38735. 29:35:19Then we train test split. Right? That's
  38736. 29:35:22pretty common. Or we could do cross
  38737. 29:35:23validation either way. Um a kfold cross
  38738. 29:35:27validation. Then we scale. So we didn't
  38739. 29:35:30do this last time, but this is something
  38740. 29:35:31we should get in the habit of is scaling
  38741. 29:35:33um our features. So we do that. And now
  38742. 29:35:37we're ready to model. So now we're ready
  38743. 29:35:39to model. Um so that's this part.
  38744. 29:35:44Okay. So let's do the model. Um the
  38745. 29:35:48model is actually uh pretty easy to do.
  38746. 29:35:50So we're going to use a lasso. So we
  38747. 29:35:52have a lasso model here. Notice what
  38748. 29:35:54we're setting our alpha to. So the big
  38749. 29:35:56parameter we really need, ignore this
  38750. 29:35:59iterations. We actually don't really
  38751. 29:36:00need the we don't really need that
  38752. 29:36:02parameter. Um so just ignore it for the
  38753. 29:36:04moment. But the big one that we're
  38754. 29:36:06setting here is the alpha. So when we
  38755. 29:36:09did linear regression, we didn't need
  38756. 29:36:11any parameters to go inside the linear
  38757. 29:36:13regression object. We didn't need any
  38758. 29:36:15parameters, right? Because there are
  38759. 29:36:17really no parameters of it. But for
  38760. 29:36:20lasso, the important one is the alpha.
  38761. 29:36:22And so we need to know what to set alpha
  38762. 29:36:24to. Um let's start with alpha equals 1.
  38763. 29:36:29That's a good starting place. So a
  38764. 29:36:31typical um starting point
  38765. 29:36:35for alpha
  38766. 29:36:38um
  38767. 29:36:39is uh is one. So that's a typical
  38768. 29:36:43starting point. And so we can set alpha
  38769. 29:36:45equals to one. This max iterations is
  38770. 29:36:48the the parameter that governs the
  38771. 29:36:51training process because it is
  38772. 29:36:53iterative. So if for some reason we we
  38773. 29:36:56can't converge to the right betas and
  38774. 29:36:59we've run it for 10,000 steps once we
  38775. 29:37:01pass 10,000 steps, uh it will stop and
  38776. 29:37:04just give us the betas at that point.
  38777. 29:37:07But it will likely never hit this
  38778. 29:37:09number. It'll converge before then. So
  38779. 29:37:12um we don't really need to um specify
  38780. 29:37:15it. So, I'm actually just going to get
  38781. 29:37:16rid of it. Um, it's not really a big
  38782. 29:37:18deal. It should converge before then.
  38783. 29:37:21Um, but if if we want to set like a
  38784. 29:37:24maximum step size in the optimization,
  38785. 29:37:26we definitely could there. Uh, but not
  38786. 29:37:30concerned about that too much. But
  38787. 29:37:32here's our lasso. And then we're just
  38788. 29:37:33going to do a fit on our data. So, look
  38789. 29:37:36how easy that is. Just like a linear
  38790. 29:37:38regression, lasso.fit,
  38791. 29:37:42right? So, we do fit. Um,
  38792. 29:37:50oh, I didn't run this. I'm sorry. I got
  38793. 29:37:52to run this. Okay. Actually, that's a
  38794. 29:37:56good example of what happens when you
  38795. 29:37:57don't when you have nles, right? So, it
  38796. 29:37:59says our our null contains nan. That's
  38797. 29:38:02because I didn't run this. But now that
  38798. 29:38:04should be filled in. Now, we should be
  38799. 29:38:06able to run this. Okay, perfect. So it
  38800. 29:38:09runs.
  38801. 29:38:16Okay. So you can see what the intercept
  38802. 29:38:18is. Um this is one of our coefficients,
  38803. 29:38:20right? The intercept is 457. And look
  38804. 29:38:23now what's really interesting about the
  38805. 29:38:24coefficients is look at what some of the
  38806. 29:38:27coefficients are. Some of them are
  38807. 29:38:32actually zero, which is really So some
  38808. 29:38:34of them ended up being at zero, which is
  38809. 29:38:36very very interesting. that means that
  38810. 29:38:39those features get cancelceled out and
  38811. 29:38:42they're basically not part of the model
  38812. 29:38:46which is really interesting. Um, so we
  38813. 29:38:48have all these coefficients and some of
  38814. 29:38:50them are zero.
  38815. 29:38:56Yeah, negative0 is just because of the
  38816. 29:38:59convergence like they started out
  38817. 29:39:01negative and worked their way up to
  38818. 29:39:03zero. it. Negative Z really just means
  38819. 29:39:07zero, but they just were coming from
  38820. 29:39:09they were like small negatives and ended
  38821. 29:39:12up at zero
  38822. 29:39:14during the training process. They were
  38823. 29:39:16negative at one point and it ended up
  38824. 29:39:18zero. Um
  38825. 29:39:20so yeah, negative 0 just obviously means
  38826. 29:39:23zero. Um it's still still zero there.
  38827. 29:39:34So what's interesting is some of these
  38828. 29:39:36features ended up uh being zero which
  38829. 29:39:39you don't usually see in a linear
  38830. 29:39:41regression. So if we were to train this
  38831. 29:39:43using a linear regression we typically
  38832. 29:39:45wouldn't see that but some of these
  38833. 29:39:47turned out to be zero because again
  38834. 29:39:49we're encouraging those betas to be
  38835. 29:39:53small. we're encouraging them to be uh
  38836. 29:39:55small and so um you know what happens is
  38837. 29:40:00some of them can be shrunk all the way
  38838. 29:40:02down to zero meaning those features
  38839. 29:40:04don't contribute that's a really simple
  38840. 29:40:06model at that point right so we've taken
  38841. 29:40:09something complex that includes all of
  38842. 29:40:12these features and actually reduced it
  38843. 29:40:14into something simple that only includes
  38844. 29:40:16these features
  38845. 29:40:18right
  38846. 29:40:21so that's what it does um now we need to
  38847. 29:40:24evaluate this to see how good of a model
  38848. 29:40:26it is. But that's what this is saying
  38849. 29:40:29here in this text is that um a positive
  38850. 29:40:32uh coefficient indicates that as the
  38851. 29:40:34independent variable increases the
  38852. 29:40:36dependent variable also increases.
  38853. 29:40:38Negative coefficient means as the
  38854. 29:40:40independent variable increases dependent
  38855. 29:40:43decreases because it's reducing the
  38856. 29:40:44value. Um and lasso is known for feature
  38857. 29:40:49selection by shrinking some of them to
  38858. 29:40:51zero effectively removing those
  38859. 29:40:53variables from the model from the
  38860. 29:40:55equation right
  38861. 29:40:57um
  38862. 29:40:59so that's what happens
  38863. 29:41:03some of them end up being zero
  38864. 29:41:07were you guys able to run this this
  38865. 29:41:10lasso uh fit which is the training of
  38866. 29:41:13the lasso Control.
  38867. 29:41:32No, it doesn't ensure there's no
  38868. 29:41:34overfit, but it helps with overfitting.
  38869. 29:41:36It's supposed to help by making the
  38870. 29:41:38model simpler. And this is definitely a
  38871. 29:41:40simpler model because it's removing some
  38872. 29:41:42of the features from the model
  38873. 29:41:43essentially, right? Because some of the
  38874. 29:41:46features aren't going to contribute.
  38875. 29:41:47It's a simpler model.
  38876. 29:41:50It doesn't it doesn't mean there's not
  38877. 29:41:51going to be any overfitting, but it
  38878. 29:41:53helps prevent it. That's what it's
  38879. 29:41:55designed to do to help prevent it.
  38880. 29:41:59Yeah. So higher coefficient. Yes. The
  38881. 29:42:02higher coefficient means it's a more
  38882. 29:42:04important feature towards the
  38883. 29:42:06prediction. Yes. That's what it means
  38884. 29:42:09for sure. The higher the magnitude, the
  38885. 29:42:12more of a contributor towards that
  38886. 29:42:14prediction. Uh it is. Yes.
  38887. 29:42:23And it's not just it's it could be
  38888. 29:42:25higher positive or negative there. Like
  38889. 29:42:27a higher negative is also a pretty big
  38890. 29:42:30factor,
  38891. 29:42:32right? So So you want to think about it
  38892. 29:42:33in terms of absolute value.
  38893. 29:42:41does not guarantee but helps. Yes, it
  38894. 29:42:43doesn't guarantee it but it's designed
  38895. 29:42:45to help overfitting, help prevent it.
  38896. 29:42:47Yes, absolutely.
  38897. 29:43:13Okay.
  38898. 29:43:15So let's do some evaluation. Um so let's
  38899. 29:43:19do in this case we are going to do our
  38900. 29:43:22predict
  38901. 29:43:42Oh, yeah. I'm not sure why that's the
  38902. 29:43:45case.
  38903. 29:43:47Interesting.
  38904. 29:44:07We could try increasing the um max
  38905. 29:44:11iterations.
  38906. 29:44:23Okay, that's why. Yeah. So then you get
  38907. 29:44:25that result with the with the higher max
  38908. 29:44:27iterations.
  38909. 29:44:29It doesn't get cut off there.
  38910. 29:44:33I think that's why you probably left
  38911. 29:44:35this in there,
  38912. 29:44:37which is fine. You get about the same
  38913. 29:44:38numbers.
  38914. 29:44:46Yeah.
  38915. 29:44:56All right. Let's evaluate this. So,
  38916. 29:44:57we're going to to to do evaluation. I
  38917. 29:44:59want you guys to see again. We should
  38918. 29:45:01get in the habit of doing evaluation,
  38919. 29:45:04which is taking our model and predicting
  38920. 29:45:07on the training and predicting on the
  38921. 29:45:09test sets, right? So we predict on the
  38922. 29:45:11train set and calculate our MSE
  38923. 29:45:15and we um calculate our R2 score um or R
  38924. 29:45:22squar score I should say. Uh but again
  38925. 29:45:25the MSE is the one we're really going to
  38926. 29:45:26use mostly. Um but we calculate so we do
  38927. 29:45:30our predictions and then we compare that
  38928. 29:45:32into our mean squared error with our
  38929. 29:45:34labels
  38930. 29:45:36and we uh go ahead and do the same thing
  38931. 29:45:40with the test. Right? So we do uh
  38932. 29:45:42lasso.predict
  38933. 29:45:44on our test features and we go ahead and
  38934. 29:45:47compare that with the test labels. And
  38935. 29:45:51so what we're doing there is generating
  38936. 29:45:52our MSE.
  38937. 29:45:55So, we we take a look at our MSE and we
  38938. 29:45:58get uh 84,000
  38939. 29:46:01MSE. Um, and so, of course, we could
  38940. 29:46:05take the um what we could do with that
  38941. 29:46:08is take a look at the um MSE on the uh
  38942. 29:46:14we could do um MP. Square root
  38943. 29:46:19and do the square root of the MSE test.
  38944. 29:46:25and we get um 340. So this would be in
  38945. 29:46:29the units of our label. So, we go back
  38946. 29:46:32and look at our label um for some of
  38947. 29:46:34those um
  38948. 29:46:48so uh we are in 300s and our data is
  38949. 29:46:52like right around the 500. So, of
  38950. 29:46:54course, if we describe this um we could
  38951. 29:46:56see what the statistics are of it. So,
  38952. 29:46:59we could do df.describe describe and
  38953. 29:47:01generate that. But that doesn't look
  38954. 29:47:02like a very good error, right? If these
  38955. 29:47:04are in the 400s, um that's that's not a
  38956. 29:47:07very good error.
  38957. 29:47:09So again, it's not a very great model.
  38958. 29:47:12But one thing I want you to see is that
  38959. 29:47:13it's it's not overfitting.
  38960. 29:47:16Um if anything, it's actually
  38961. 29:47:18underfitting, which is what this kind of
  38962. 29:47:21um MSE suggests, right? because our
  38963. 29:47:23error here is 84 uh excuse me 84,000.
  38964. 29:47:29Um
  38965. 29:47:34our our area here is 84,000
  38966. 29:47:38excuse me and on the test set it's
  38967. 29:47:41116,000.
  38968. 29:47:43Um so these two errors are both bad. So
  38969. 29:47:48it's not overfitting. This is actually
  38970. 29:47:51underfitting. So it's not overfitting,
  38971. 29:47:53it's actually underfitting. Um, and so
  38972. 29:47:56that's the risk with something like
  38973. 29:47:57lasso is that it's making the model a
  38974. 29:48:00bit too simple and we actually risk
  38975. 29:48:03underfitting, which is what happens. We
  38976. 29:48:06have too much error across both the
  38977. 29:48:09training and the test set. Overfitting
  38978. 29:48:12is when we do we have really good
  38979. 29:48:14performance on the training set, but bad
  38980. 29:48:16performance on the test set. We're not
  38981. 29:48:18overfitting.
  38982. 29:48:20um we are uh underfitting because our
  38983. 29:48:24performance is not good either way. Even
  38984. 29:48:26this R squar is pretty low. It's not
  38985. 29:48:28even at 50%.
  38986. 29:48:36Okay, so that's so we we do the
  38987. 29:48:39evaluation and again the evaluation just
  38988. 29:48:40comes down to making predictions and
  38989. 29:48:43computing our error amongst those
  38990. 29:48:45predictions to our labels. That's always
  38991. 29:48:47what the uh evaluation is going to be
  38992. 29:48:54for MSE.
  38993. 29:48:57What's the ideal MSE? What do you think
  38994. 29:49:00it should be? What is So, think about it
  38995. 29:49:03like this. The MSE represents the
  38996. 29:49:05average distance between our predictions
  38997. 29:49:10and the labels.
  38998. 29:49:13So, if we're getting it right all the
  38999. 29:49:16time, what's that distance going to be
  39000. 29:49:18if we're always right? What's our
  39001. 29:49:20distance from what's our distance from
  39002. 29:49:23our predictions to our labels going to
  39003. 29:49:26be if we're always getting it right?
  39004. 29:49:29Zero. Yeah, there's not going to be any
  39005. 29:49:31distance. It's going to be right. It's
  39006. 29:49:33going to be perfectly aligned, right?
  39007. 29:49:35There's going to be no distance there.
  39008. 29:49:37So, yeah, an ideal MSE is zero. That's
  39009. 29:49:41an ideal MSE.
  39010. 29:49:43So, anything close to like the smaller
  39011. 29:49:46the better for MSE. The smaller the
  39012. 29:49:49better. Um, for this R squared, uh, it's
  39013. 29:49:53it's a scale between 0 to one where one
  39014. 29:49:56is the best. So, one would be perfectly
  39015. 29:49:58aligned predictions. Um, so, and again,
  39016. 29:50:02this this is we actually multiply by 100
  39017. 29:50:05to get uh because it's it's a number
  39018. 29:50:07between 0 and one. So we get about 47%
  39019. 29:50:10which is not good.
  39020. 29:50:21Okay.
  39021. 29:50:36All right. Any questions on this
  39022. 29:50:39evaluation?
  39023. 29:50:54All right. I want to show you something
  39024. 29:50:56which is
  39025. 29:50:58Yeah, this that's true. The scale of it
  39026. 29:51:01matters on the data because we should be
  39027. 29:51:03you should always interpret your MSE in
  39028. 29:51:05the scale of
  39029. 29:51:07um your your labels because your labels
  39030. 29:51:12like in this case our labels um you know
  39031. 29:51:14we could take uh for example we could
  39032. 29:51:17easily let's actually do that let's take
  39033. 29:51:20the average
  39034. 29:51:22let's take the average of our labels on
  39035. 29:51:25the training data
  39036. 29:51:29and and we could see what those are. Um,
  39037. 29:51:32so the average is 500,
  39038. 29:51:35right? The average is 500. And look at
  39039. 29:51:38what our uh square root of our MSE is,
  39040. 29:51:41which is in the same units as our
  39041. 29:51:43original. Um, so we have uh quite a bit
  39042. 29:51:47of error. 340 when our units are right
  39043. 29:51:50around 500.
  39044. 29:51:53So that's quite a bit of error.
  39045. 29:52:03Yeah, MSE of zero means our our uh our
  39046. 29:52:06predictions are nearly identical to the
  39047. 29:52:10test labels. Yes, that's what MSE of
  39048. 29:52:13zero means. There's zero distance.
  39049. 29:52:18So closer to zero, the better.
  39050. 29:52:27But we talked about it as you you really
  39051. 29:52:29so the rule of thumb should be what is
  39052. 29:52:34your RMSSE as a percentage of your
  39053. 29:52:37typical value. So your typical value is
  39054. 29:52:40in the 500s. Our our RMSSE is 340.
  39055. 29:52:45That's just really high. That's over
  39056. 29:52:47like 60% of that value.
  39057. 29:52:51So that's just a lot. That's too much
  39058. 29:52:54error. What we would love this RMSSE to
  39059. 29:52:56be is under 20% of the typical value. So
  39060. 29:53:00that means on average we are 20% or less
  39061. 29:53:05off in our prediction. That would be
  39062. 29:53:08good. That would be pretty good. That
  39063. 29:53:10means we're like 80% accurate,
  39064. 29:53:14right? That'd be pretty ideal. So you
  39065. 29:53:16got to think about it in terms of this
  39066. 29:53:17RMSSE, which is in the same units as
  39067. 29:53:20your labels.
  39068. 29:53:23This is the
  39069. 29:53:27RMSSE
  39070. 29:53:30which is in the same units as the
  39071. 29:53:34labels.
  39072. 29:53:37So and then to interpret this we have
  39073. 29:53:40340
  39074. 29:53:42is compared to
  39075. 29:53:45typical
  39076. 29:53:47um salary unit of 500
  39077. 29:53:52right so this is uh quite a bit when the
  39078. 29:53:55typical value is 500 and we are off on
  39079. 29:53:58average by 340 units
  39080. 29:54:01that's so much relative to the typical
  39081. 29:54:04value
  39082. 29:54:06that's just too. That's a lot of error.
  39083. 29:54:08That's not a very good model, right?
  39084. 29:54:11It's underfitting. It's definitely
  39085. 29:54:13underfitting.
  39086. 29:54:25Yeah. So, that's a great question. What
  39087. 29:54:26should we do from here? So, um because
  39088. 29:54:29we're underfitting
  39089. 29:54:37um we should use a more complex model.
  39090. 29:54:42So uh we're going to learn about those
  39091. 29:54:45in lesson four, but we should use
  39092. 29:54:47something different. This linear
  39093. 29:54:48regression is still too basic. Even with
  39094. 29:54:50lasso, it's still too basic.
  39095. 29:54:56Yeah, we're underfitting because we But
  39096. 29:54:59it could also be we're underfitting with
  39097. 29:55:00a regular linear regression. We should
  39098. 29:55:02test that out. Um, and maybe it would be
  39099. 29:55:04an exercise for you guys um to test that
  39100. 29:55:08out yourself. It shouldn't be hard to
  39101. 29:55:10do. Um, you already have all the data
  39102. 29:55:12scaled. You So, do you see how you would
  39103. 29:55:15do that? You would just come in here and
  39104. 29:55:17build a linear regression rather than a
  39105. 29:55:18lasso and dofit and then you would
  39106. 29:55:21evaluate it the same way with a
  39107. 29:55:22dotpredict. It's really easy to do that.
  39108. 29:55:26And then we can compare that um to to
  39109. 29:55:29this. It shouldn't be that hard to do
  39110. 29:55:32that, right?
  39111. 29:55:34And something you guys could do for
  39112. 29:55:35sure. Um,
  39113. 29:55:38is build the linear regression and
  39114. 29:55:41actually compare it and see what kind of
  39115. 29:55:44difference it makes. I mean, we honestly
  39116. 29:55:46we could do it ourselves. We could do it
  39117. 29:55:48right now. Maybe it's worth trying that.
  39118. 29:55:52So, let's build a linear regression
  39119. 29:55:56for comparison.
  39120. 29:56:00So we have our linear regression
  39121. 29:56:04uh is linear regression and then we do
  39122. 29:56:08ffit linear regression.fit fit
  39123. 29:56:12right so so this will train it um and
  39124. 29:56:16then we can evaluate it so lin MSE is
  39125. 29:56:21mean squared error
  39126. 29:56:24and then we can do our um let's do our
  39127. 29:56:27training let's do the training and then
  39128. 29:56:32um let's predict
  39129. 29:56:34actually let me do that here
  39130. 29:56:37uh y prediction
  39131. 29:56:40train
  39132. 29:56:43linear
  39133. 29:56:45equals um linear regression.predict
  39134. 29:56:50and then we're going to predict on our
  39135. 29:56:52training features.
  39136. 29:56:57Okay, do you guys see what I'm doing?
  39137. 29:56:58I'm building a linear regression for
  39138. 29:57:00comparison.
  39139. 29:57:02I'm doing dofit here to train it and
  39140. 29:57:05then I'm making some predictions on the
  39141. 29:57:06training set and we're going to evaluate
  39142. 29:57:09those. I'm going to replace that here
  39143. 29:57:10with y prred
  39144. 29:57:14uh train
  39145. 29:57:17linear. So these predictions
  39146. 29:57:42Okay. So, if you guys want this code, I
  39147. 29:57:44can paste it in.
  39148. 29:57:55So, let's see what the RMSSE for just a
  39149. 29:57:58linear model is.
  39150. 29:58:01It's a little bit better. It's better
  39151. 29:58:03for sure.
  39152. 29:58:06So 289 is better than this 340. It's
  39153. 29:58:10better. It's getting closer to zero.
  39154. 29:58:13It's still underfitting though,
  39155. 29:58:17right? And that's just on the training
  39156. 29:58:18set. Let's look at the Let's do the same
  39157. 29:58:21thing, but on
  39158. 29:58:25Let's change this. Let's swap this out
  39159. 29:58:27for um test.
  39160. 29:58:31And then let's do test.
  39161. 29:58:35And then let's do test
  39162. 29:58:41test.
  39163. 29:58:44and then
  39164. 29:58:47test.
  39165. 29:58:58Okay, so this is producing test
  39166. 29:58:59predictions on the test set.
  39167. 29:59:02We are generating an MSE test
  39168. 29:59:07and then we're doing MSE test
  39169. 29:59:10which is using the test labels and our
  39170. 29:59:12test predictions
  39171. 29:59:14and then we take the square root of that
  39172. 29:59:15for RMSSE and then we're going to
  39173. 29:59:17generate that. So it's still under fit.
  39174. 29:59:20I mean this is still high. This is still
  39175. 29:59:23high um on the test set and versus on
  39176. 29:59:25the training set. So it's still pretty
  39177. 29:59:27high. Um, even the basic linear
  39178. 29:59:29regression is under is still
  39179. 29:59:31underfitting. Still underfitting, right?
  39180. 29:59:34Even without the lasso, which is lasso
  39181. 29:59:37is supposed to help with overfitting.
  39182. 29:59:39It's definitely not overfitting. Um,
  39183. 29:59:42it's definitely underfitting,
  39184. 29:59:44but this is a signal that it's kind of
  39185. 29:59:46overfitting because this is performing
  39186. 29:59:47better on the training data and then it
  39187. 29:59:49gets worse on the test data.
  39188. 29:59:53Definitely gets worse, right?
  39189. 30:00:03Did you guys follow?
  39190. 30:00:07I'm just running this above I'm running
  39191. 30:00:09this above this. It doesn't matter where
  39192. 30:00:12you put it. We could uh we could move it
  39193. 30:00:13down.
  39194. 30:00:20We could move it down to I just ran I
  39195. 30:00:23just picked a new cell right here.
  39196. 30:00:25and ran it. But we could move it.
  39197. 30:00:27Actually, let's do that. Let's move it
  39198. 30:00:29down
  39199. 30:00:33to
  39200. 30:00:36after the lasso evaluation.
  39201. 30:00:41Okay. So, I just moved it there.
  39202. 30:00:44And then let's move
  39203. 30:00:47this down.
  39204. 30:00:49So, I just put it here after the um
  39205. 30:00:52after this. So this is the um this is
  39206. 30:00:56basically the objective function right
  39207. 30:00:59of the training process. So during the
  39208. 30:01:02algorithm that runs when we call ffit in
  39209. 30:01:05scikitlearn it's going to find these
  39210. 30:01:08betas right it's actually going to learn
  39211. 30:01:10what these best betas are for our model.
  39212. 30:01:14Um this is our model here right it's the
  39213. 30:01:16combination of betas times our features
  39214. 30:01:19um plus an intercept beta. Uh so that's
  39215. 30:01:23our model but um we penalize those large
  39216. 30:01:26uh weights in absolute value by um
  39217. 30:01:30adding a penalty term like this um where
  39218. 30:01:33alpha is some level of penalty that we
  39219. 30:01:37want to provide. Usually alpha equals 1
  39220. 30:01:39is okay. But um actually what we're
  39221. 30:01:41going to learn uh to finish out this
  39222. 30:01:43section is there's going to be a
  39223. 30:01:44systematic way we can test out different
  39224. 30:01:46alphas um that represent the level of
  39225. 30:01:49penalty we want to uh apply to lasso or
  39226. 30:01:52even ridge
  39227. 30:01:54uh regression. So that was the lasso and
  39228. 30:01:57um if you guys remember using it was
  39229. 30:01:59super easy. Uh we worked through this
  39230. 30:02:02problem with this um baseball data um
  39231. 30:02:06and we had uh
  39232. 30:02:09let's see scrolling down we um split out
  39233. 30:02:12our numerical data and we did uh we one
  39234. 30:02:16hot encoded our our categorical data
  39235. 30:02:19combined it back together. Hopefully
  39236. 30:02:21that um rings a bell there. Um and we
  39237. 30:02:24actually scaled our data which is pretty
  39238. 30:02:26standard to do is we do some type of
  39239. 30:02:28scaling to our features especially our
  39240. 30:02:30numerical features right want to scale
  39241. 30:02:32those in some way whether it's minmax
  39242. 30:02:34scale or standard scaler um want to do
  39243. 30:02:37that and so we did that for this example
  39244. 30:02:39and then we um ran the lasso regression
  39245. 30:02:44which is pretty easy to use. You just
  39246. 30:02:45use the lasso object and you pick an
  39247. 30:02:47alpha here. Um, again, we are going to
  39248. 30:02:51have a way to test out different alphas
  39249. 30:02:54that could be candidates and we can see
  39250. 30:02:56which one's the best. Um, so I'm going
  39251. 30:02:59to show us that today coming up shortly.
  39252. 30:03:03But that was that was the lasso. If you
  39253. 30:03:05guys remember, we did that. Um, this it
  39254. 30:03:08we compared that to a basic linear
  39255. 30:03:10regression which is just this pretty
  39256. 30:03:12straightforward just a fit and then
  39257. 30:03:14predict and then we can generate mean
  39258. 30:03:16squared error. Um, still not a very good
  39259. 30:03:19mean squared error on this data, it's
  39260. 30:03:21still fairly large. Um, so it's still
  39261. 30:03:25not, no matter which model we use, it's
  39262. 30:03:27still not very good, but at least we can
  39263. 30:03:29practice doing that comparison. That's
  39264. 30:03:30what we did last time. We did this on
  39265. 30:03:33Wednesday.
  39266. 30:03:34Um
  39267. 30:03:36and then
  39268. 30:03:38we saw that the effect of different
  39269. 30:03:40alphas we had a lasso um
  39270. 30:03:43we had a lasso uh cross validation
  39271. 30:03:46example here. So beyond just using a
  39272. 30:03:48regular lasso model that um scikitlearn
  39273. 30:03:50has a lasso cv which allows you to try
  39274. 30:03:53out different alphas uh with cross
  39275. 30:03:56validation and um figure out what the
  39276. 30:03:58best alpha is. Um, now we're actually
  39277. 30:04:01going to have a different strategy
  39278. 30:04:03that'll instead of just picking random
  39279. 30:04:05ones, we can actually um supply multiple
  39280. 30:04:08parameters that we may want to test um
  39281. 30:04:11as many as the models may support. And
  39282. 30:04:13in some more complex models will have
  39283. 30:04:16more than one parameter like lasso only
  39284. 30:04:18has the alpha. Um, technically it also
  39285. 30:04:21has its max iterations, but really the
  39286. 30:04:23only one that matters is this alpha.
  39287. 30:04:26Other models have many more
  39288. 30:04:27hyperparameters that we can um uh change
  39289. 30:04:32and so we want a way to systematically
  39290. 30:04:34test out those different combinations
  39291. 30:04:37and to see which one leads to the best
  39292. 30:04:39uh version of that model. Let's say the
  39293. 30:04:41best results. So um we're going to
  39294. 30:04:44explore that coming up. So we had lasso.
  39295. 30:04:48Um now this is where we ended last time.
  39296. 30:04:50We had ridge regression. If you guys
  39297. 30:04:52remember, this one is just a slightly
  39298. 30:04:55different penalty. Um,
  39299. 30:04:58it takes the it I drew it out for us. It
  39300. 30:05:01takes the same penalty we had before.
  39301. 30:05:03So, it has that um residual sum of
  39302. 30:05:06squares error, which is the main one we
  39303. 30:05:09used for linear regression, but it has a
  39304. 30:05:11penalty with an alpha and then it has
  39305. 30:05:13the sum of the beta squares
  39306. 30:05:17beta squares. So it penalizes it has a
  39307. 30:05:22penalty but it penalizes slightly
  39308. 30:05:24differently where it uses the square not
  39309. 30:05:26the absolute value. That's the ridge
  39310. 30:05:28regression. And this has the similar
  39311. 30:05:30effect of you don't in order to minimize
  39312. 30:05:33this right because our goal in training
  39313. 30:05:34a model was to minimize this thing
  39314. 30:05:38minimize this um quantity and find the
  39315. 30:05:41best betas that minimize this. Um so
  39316. 30:05:44generally yes you want to encourage
  39317. 30:05:46lower values but the um once you get
  39318. 30:05:50values that are a fraction if you square
  39319. 30:05:52them they actually get smaller. Um so uh
  39320. 30:05:56it's it's not um it's not necessary to
  39321. 30:06:01shrink them all the way to zero. They
  39322. 30:06:03will get smaller as soon as they're kind
  39323. 30:06:05of below one. Um so they don't encourage
  39324. 30:06:08it to completely go away uh like the
  39325. 30:06:12absolute value does. It's just slightly
  39326. 30:06:13different minimization. Um so what we
  39327. 30:06:15see with the ridge is we don't see the
  39328. 30:06:18features kind of get wiped out
  39329. 30:06:19completely like we do with a lasso. In
  39330. 30:06:21lasso they get encouraged to be um to
  39331. 30:06:24become zero because that's kind of the
  39332. 30:06:26only way to minimize an absolute value.
  39333. 30:06:28But with squares they can keep getting
  39334. 30:06:30smaller and smaller and smaller um
  39335. 30:06:32fractions and they don't have to become
  39336. 30:06:35zero. It's not as harsh of a of a
  39337. 30:06:38penalty.
  39338. 30:06:39Um,
  39339. 30:06:41so, uh, the ridge was easy to use as
  39340. 30:06:45well. Um, and it also has an alpha that
  39341. 30:06:50we can set. So, it's literally the same
  39342. 30:06:52exact code, just a different model.
  39343. 30:06:55There's slightly different penalty and
  39344. 30:06:57it results in different coefficients.
  39345. 30:06:59You notice that none of them are exactly
  39346. 30:07:00zero. Like with the lasso, you can get
  39347. 30:07:02ones that are exactly zero. We don't see
  39348. 30:07:05that with the ridge. You remember that?
  39349. 30:07:08Um so we we s pointed out that last
  39350. 30:07:10time. Notice the coefficients aren't
  39351. 30:07:12zero. Um and then we can evaluate it. So
  39352. 30:07:15we did our MSE calculation which is a
  39353. 30:07:17pretty standard thing where we use our
  39354. 30:07:19model to predict on a training set,
  39355. 30:07:21predict on a test set, evaluate those um
  39356. 30:07:25by computing the metric like the mean
  39357. 30:07:27squed error and we can see if we're
  39358. 30:07:29overfitting, underfitting. This is
  39359. 30:07:31definitely the same kind of story we've
  39360. 30:07:33seen with all these models is
  39361. 30:07:34underfitting because the error is so big
  39362. 30:07:36across both sets
  39363. 30:07:38across training and tests. So it's it's
  39364. 30:07:40definitely underfitting.
  39365. 30:07:42Um
  39366. 30:07:44and same thing as lasso, it has a cross
  39367. 30:07:46validation uh variation on it that
  39368. 30:07:49allows you to try out different alphas
  39369. 30:07:52and um do different folds. So 10fold,
  39370. 30:07:56fivefolds, whatever and compute the um
  39371. 30:08:00try to find the best alpha that way.
  39372. 30:08:03Okay.
  39373. 30:08:07All right.
  39374. 30:08:09Any questions on this so far from last
  39375. 30:08:12time from reviewing that a little bit?
  39376. 30:08:15Hopefully that uh hopefully that is
  39377. 30:08:18jogging your memory a little bit on
  39378. 30:08:20ridge and lasso. Um, you know, where
  39379. 30:08:23we're going to pick it up today is to
  39380. 30:08:25finish out this lesson with one more
  39381. 30:08:27model,
  39382. 30:08:29which is going to be a combination of
  39383. 30:08:32ridge and lasso. So, you can actually
  39384. 30:08:34combine them together
  39385. 30:08:37um in a linear fashion, those penalties.
  39386. 30:08:40So, you can actually have both
  39387. 30:08:41penalties, the absolute value and the
  39388. 30:08:43square. And when you have both penalties
  39389. 30:08:47um that's a special model called the
  39390. 30:08:49elastic net uh regression or elastic net
  39391. 30:08:53model. Um so this is a combination of
  39392. 30:08:57lasso and ridge together. So you have
  39393. 30:08:59lasso, you have ridge and then you have
  39394. 30:09:01elastic net which combines both of those
  39395. 30:09:03penalties. Um let me show you the
  39396. 30:09:06equation.
  39397. 30:09:08So here is the uh so here is the the
  39398. 30:09:13model. This is the same that we've
  39399. 30:09:15always had. This is our usual u model
  39400. 30:09:20fitting for linear. This is a basic
  39401. 30:09:22linear regression um loss function or
  39402. 30:09:25objective function that we're trying to
  39403. 30:09:26minimize to find the betas. Notice how
  39404. 30:09:29we have both of our penalties though
  39405. 30:09:30this time. So instead of just having one
  39406. 30:09:32of the penalties, we actually have both.
  39407. 30:09:34So we have the lasso penalty
  39408. 30:09:38and then we have the ridge penalty here.
  39409. 30:09:40So we actually use both of them and um
  39410. 30:09:44try to find a balance of minimizing
  39411. 30:09:47those two uh those two penalties.
  39412. 30:09:51Okay. And notice how they instead of
  39413. 30:09:53just a single alpha, we kind of have a
  39414. 30:09:55balance on both of them.
  39415. 30:09:58So we can actually weight the lasso one
  39416. 30:10:01more. We can weight the ridge one more.
  39417. 30:10:03We can weight them the same. Uh we can
  39418. 30:10:07um change that around as much as we
  39419. 30:10:08want. So they have two different weights
  39420. 30:10:10there um that they could be.
  39421. 30:10:14Um now what happens in reality is uh
  39422. 30:10:19we're going to see this in the model is
  39423. 30:10:21that um usually what happens is these
  39424. 30:10:24get combined into a fraction. So there's
  39425. 30:10:28usually a ratio of lambda 1 to lambda 2
  39426. 30:10:32and this is known as the um this is
  39427. 30:10:35sometimes known as the L1 ratio
  39428. 30:10:38and this is a this is a a parameter
  39429. 30:10:40inside the model that we'll be able to
  39430. 30:10:42set um along with alpha. So we'll be
  39431. 30:10:45able to set an alpha and then this
  39432. 30:10:47ratio. Um the idea is is that um the
  39433. 30:10:52ratio will uh allow us to control which
  39434. 30:10:56one is more dominant. So if this number
  39435. 30:10:59is bigger the um this lasso penalty will
  39436. 30:11:03will be weighted more. If this ratio is
  39437. 30:11:06smaller if it's less than one for
  39438. 30:11:09example that means that the um ridge
  39439. 30:11:11regression is more uh dominant. Um but
  39440. 30:11:16the so we'll have this we'll have really
  39441. 30:11:18this and this at our disposal and alpha
  39442. 30:11:23is um
  39443. 30:11:26alpha is kind of like a a you can think
  39444. 30:11:28of it as a scale that is um so lambda 1
  39445. 30:11:33kind of like lambda 1 plus lambda 2 um
  39446. 30:11:37combined to equal alpha.
  39447. 30:11:40So it's like our total level of penalty
  39448. 30:11:43um our total level of penalty and we can
  39449. 30:11:46set that equal to one. We can set it
  39450. 30:11:48equal to whatever we want. Um and so
  39451. 30:11:51these will be in this ratio and there'll
  39452. 30:11:53be a total level of penalty that we can
  39453. 30:11:55apply. So the model will actually use
  39454. 30:11:58these two parameters when we when we do
  39455. 30:12:00it. But that's how they're that's how
  39456. 30:12:02they're all related.
  39457. 30:12:05Okay. So ridge uses both penalties.
  39458. 30:12:08That's the only difference between lasso
  39459. 30:12:10or sorry elastic net uses both
  39460. 30:12:12penalties. Um so one thing I want you to
  39461. 30:12:15notice is that uh if we um if we want we
  39462. 30:12:21could set this L1 ratio all the way to
  39463. 30:12:23zero.
  39464. 30:12:25Um which uh if we do that um the only
  39465. 30:12:30way this L1 ratio could be zero would be
  39466. 30:12:32if lambda 1 is zero. So it would just
  39467. 30:12:34revert back to ridge regression. So it
  39468. 30:12:36complet if if this is zero this will
  39469. 30:12:39wipe out this term and we'll be back to
  39470. 30:12:40ridge if the L1 ratio is zero.
  39471. 30:12:46Okay.
  39472. 30:12:50All right. So we have a elastic net
  39473. 30:12:53model. Um now it's used the exact same
  39474. 30:12:57way as we did the other models in the
  39475. 30:12:59code. So we have elastic net um uh from
  39476. 30:13:03the scikitlearn linear model family just
  39477. 30:13:06exactly where we had linear regression
  39478. 30:13:09lasso ridge all of those came from this
  39479. 30:13:12linear model um elastic net also comes
  39480. 30:13:15from there and then the cross validation
  39481. 30:13:16version also comes from there um
  39482. 30:13:21so let's see so when we build our model
  39483. 30:13:23it's going to be um very very simple
  39484. 30:13:26easy stuff because it's the same code
  39485. 30:13:28that we always have um we just use the
  39486. 30:13:31elastic net. We set an alpha alpha
  39487. 30:13:34equals 1 is pretty standard um just like
  39488. 30:13:37it is in in the last one ridge that's
  39489. 30:13:39industry standard is one and then an L
  39490. 30:13:42L1 ratio of.5
  39491. 30:13:44that's pretty standard as well. What the
  39492. 30:13:45L1 ratio of.5 is is kind of a um
  39493. 30:13:51uh kind of a that means that the lambda
  39494. 30:13:541 to lambda 2 ratio is 1/2. Um, so
  39495. 30:13:58that's that's a pretty standard uh ratio
  39496. 30:14:00as well. But again, we could set this
  39497. 30:14:03equal to one and they'd be kind of
  39498. 30:14:05equally weighted. Um, L1 ratio of a half
  39499. 30:14:08means that the uh ridge regard the the
  39500. 30:14:12ridge penalty is a little bit more
  39501. 30:14:14weighted uh in that in that situation.
  39502. 30:14:19Okay.
  39503. 30:14:21So uh once we have this model um we can
  39504. 30:14:24do ffit and we can run that on our
  39505. 30:14:26training data and we can um get we can
  39506. 30:14:30figure out what our parameters are like
  39507. 30:14:32our coefficients and our intercepts. Our
  39508. 30:14:33model will have that but more
  39509. 30:14:35importantly we can use our model to
  39510. 30:14:36predict right so we can predict on the
  39511. 30:14:38test set. Um let me go back and load our
  39512. 30:14:42data and actually run this.
  39513. 30:14:47So, we're going to be using the same
  39514. 30:14:48data that we did for uh lasso,
  39515. 30:14:54which is the I'm scrolling back up so I
  39516. 30:14:57can load it. It's the baseball data
  39517. 30:14:58here.
  39518. 30:15:02Um,
  39519. 30:15:05just run it from there.
  39520. 30:15:07It's this hitters.csv. So, hopefully you
  39521. 30:15:10have that one.
  39522. 30:15:22Let me load this.
  39523. 30:15:30Okay, so we loaded that and then that
  39524. 30:15:32should load.
  39525. 30:15:34Drop that unnamed column.
  39526. 30:15:41We will get our dummies
  39527. 30:15:50and then concatenate those split
  39528. 30:15:56scale. I'm just rerunning things. I'm
  39529. 30:15:58rerunning things so we can see our model
  39530. 30:16:00one more time.
  39531. 30:16:02So rerun that. Take a look at that. That
  39532. 30:16:04looks good. and then
  39533. 30:16:07fill in the nles on the on those.
  39534. 30:16:10Okay. So, we should be able to run our
  39535. 30:16:14uh elastic net now.
  39536. 30:16:23Okay. So, let's import that and then
  39537. 30:16:26let's build our model. So, there we go.
  39538. 30:16:28We build our model and the intercept is
  39539. 30:16:31that. Now, of course, we can look at our
  39540. 30:16:33coefficients. Let's look at that.
  39541. 30:16:39Look at our coefficients. So remember
  39542. 30:16:41the coefficients are the uh betas. These
  39543. 30:16:43are our betas that are in our model. Um
  39544. 30:16:46so we can take a look at those. Now um
  39545. 30:16:49they're it's somewhere in between. It's
  39546. 30:16:51not a full lasso where we're going to
  39547. 30:16:52see some of these be zero. It's not a
  39548. 30:16:54full ridge. Um so the coefficients we
  39549. 30:16:57get are different. They're somewhere in
  39550. 30:16:59between there those two models that
  39551. 30:17:01we've already built. So not quite the
  39552. 30:17:03same um somewhere in between there.
  39553. 30:17:10Um and then we can use our model to make
  39554. 30:17:12predictions and and compute the MSE
  39555. 30:17:15uh or the RMSSE I should say as well. So
  39556. 30:17:18we can take the mean squared error, pass
  39557. 30:17:20that into the square root and comput the
  39558. 30:17:21RMSSE. So still pretty bad. Um this is
  39559. 30:17:24right around that 300 range of what
  39560. 30:17:26we've gotten for our other RMSSE. So,
  39561. 30:17:28it's not like elastic net is any better
  39562. 30:17:31than those other like linear or lasso or
  39563. 30:17:34ridge. And that's not surprising because
  39564. 30:17:36it's just adding those extra penalties.
  39565. 30:17:38We don't expect it to magically get
  39566. 30:17:40better. It's actually a more complex
  39567. 30:17:42um when we add when we add those in,
  39568. 30:17:46we're actually reducing it and making it
  39569. 30:17:47simpler. And we need something more
  39570. 30:17:49complex, I should say. So, we're making
  39571. 30:17:51it simpler um by by making penalizing
  39572. 30:17:56our weights a little bit more. And so
  39573. 30:17:58it's still not a good fit. That's not
  39574. 30:18:01really surprising, right? It's still not
  39575. 30:18:03really a great fit.
  39576. 30:18:05And we can we can even double check
  39577. 30:18:07that. We know our RMSSE is pretty bad.
  39578. 30:18:10Um but we can double check it with this
  39579. 30:18:12R2 score. And it's, you know, still not
  39580. 30:18:15good. Remember, a one would be really
  39581. 30:18:16good. Um that'd be like a perfect linear
  39582. 30:18:18model. This is um still pretty bad.
  39583. 30:18:26Okay, so as we said, the alpha controls
  39584. 30:18:28the overall strength. Um, so the higher
  39585. 30:18:31the alpha, the more overall penalty
  39586. 30:18:34we're supplying, which makes the model
  39587. 30:18:37simpler. Um,
  39588. 30:18:40uh, but the L1 um ratio determines the
  39589. 30:18:43mix or that ratio of the lambdas, the
  39590. 30:18:47lasso to the ridge. Um, if you have it
  39591. 30:18:51be um exactly zero, you you revert all
  39592. 30:18:55the way back to um if you if you put it
  39593. 30:18:59at zero, you revert all the way back to
  39594. 30:19:00ridge one would be all the way to pure
  39595. 30:19:02lassos. Somewhere in between like one
  39596. 30:19:04half is is good.
  39597. 30:19:15Okay,
  39598. 30:19:17so this is another example of trying out
  39599. 30:19:21different values of alpha in the CV to
  39600. 30:19:23see which one works. Now again, I'm
  39601. 30:19:25going to show us in a minute a
  39602. 30:19:26systematic way to do this, but this is
  39603. 30:19:29just trying out um different alphas that
  39604. 30:19:31we set up in this uh in this um
  39605. 30:19:36uh range. So we have different uh values
  39606. 30:19:39between minus2 and two um
  39607. 30:19:42logarithmically.
  39608. 30:19:44Um so these are uh logarithm values that
  39609. 30:19:47are between this between minus2 and two
  39610. 30:19:49and we choose 100 different alphas and
  39611. 30:19:52then we choose 100 different um L1
  39612. 30:19:54ratios between 0.01 and one and we run
  39613. 30:19:58that we run this um cross validation
  39614. 30:20:00with 10 folds. So this is quite a bit.
  39615. 30:20:02So we're doing 10 folds and we're trying
  39616. 30:20:05out 100 different um options. Uh every
  39617. 30:20:09time we do an option we're trying out 10
  39618. 30:20:11folds to evaluate it. So, it's going to
  39619. 30:20:13take a minute to run.
  39620. 30:20:26It's still running here. But again, what
  39621. 30:20:29this is doing is trying out different
  39622. 30:20:30alphas and it's it's going to do a cross
  39623. 30:20:34validation. And you guys remember the
  39624. 30:20:36t-fold cross validation is where we take
  39625. 30:20:38our data and we divide it into 10 folds
  39626. 30:20:43and then we um train on nine of those
  39627. 30:20:46and then test on the remaining fold and
  39628. 30:20:48then we rotate all the folds 10 times.
  39629. 30:20:51and that we average those mean squared
  39630. 30:20:54error metrics together um against those
  39631. 30:20:5810 different uh fold options to generate
  39632. 30:21:02a basically like an average performance
  39633. 30:21:05for that value of alpha. And we're doing
  39634. 30:21:07that 100 times for all these different
  39635. 30:21:09100 alphas that there are and 100
  39636. 30:21:12different L1 ratios that we're trying
  39637. 30:21:13with them.
  39638. 30:21:19So that's quite a bit of processing but
  39639. 30:21:22uh it did finish.
  39640. 30:21:26So we can see what our best alpha is and
  39641. 30:21:28our best one ratio. So we get the best
  39642. 30:21:30alpha is this best one ratio is this. Um
  39643. 30:21:34and therefore we can uh build a model
  39644. 30:21:37with those with just these two guys as
  39645. 30:21:39the alpha and the L1 and um see how that
  39646. 30:21:44performs.
  39647. 30:21:46We build that model and then we predict
  39648. 30:21:48on the test set and we generate the
  39649. 30:21:50RMSSE. It's just a little bit better.
  39650. 30:21:52It's still not It's just a little bit
  39651. 30:21:54better, but it's still not good, right?
  39652. 30:21:56It's still 338. It is just way too big.
  39653. 30:22:00Remember, this is RMSSE, so it's in the
  39654. 30:22:03units of our uh target variable. So,
  39655. 30:22:07it's in the units of, if we go back to
  39656. 30:22:10our data, actually, I can just print it
  39657. 30:22:12out here.
  39658. 30:22:14um this RMSSE.
  39659. 30:22:18If I just do this, we could take a look
  39660. 30:22:20at um DF
  39661. 30:22:22or I could look at Y test
  39662. 30:22:27and you can see some of these values.
  39663. 30:22:28These are these salary values in the
  39664. 30:22:30hundreds, right? Some of them are in the
  39665. 30:22:31thousands. Um but an error of like 338
  39666. 30:22:36is just too big. That's a really big
  39667. 30:22:38error. That means we would be off by an
  39668. 30:22:39average of 300 when our our values if we
  39669. 30:22:42just do the mean
  39670. 30:22:46um
  39671. 30:22:48is only 550 as on average is 550 but we
  39672. 30:22:52have this amount of error on average um
  39673. 30:22:55so that's just a way too big of a
  39674. 30:22:57proportion of error right it's not a
  39675. 30:22:59very good model and again we can verify
  39676. 30:23:02that by looking at this R2 for.
  39677. 30:23:11So if we go down here,
  39678. 30:23:20still not very good.
  39679. 30:23:23Here's some of our coefficients. So
  39680. 30:23:25remember, you can always take your
  39681. 30:23:26coefficients and line them up to your
  39682. 30:23:28your data columns. Uh so that you can
  39683. 30:23:31get a sense of what coefficient belongs
  39684. 30:23:33with what feature. So that's all we're
  39685. 30:23:36doing here is just creating a series
  39686. 30:23:37where those coefficients instead of just
  39687. 30:23:39printing out the coefficients, we're
  39688. 30:23:40actually lining them up to the columns.
  39689. 30:23:42So this tells us um remember the larger
  39690. 30:23:45it is the more influence it kind of has
  39691. 30:23:47on the on the final result. Um either
  39692. 30:23:50way, so like this has a big negative
  39693. 30:23:52influence. Um, this has a large positive
  39694. 30:23:55influence.
  39695. 30:24:02Okay, let me pause there. Any questions
  39696. 30:24:05about the
  39697. 30:24:07elastic net model?
  39698. 30:24:11This is a really this model is a really
  39699. 30:24:13good one to use when you are building a
  39700. 30:24:16linear regression and it's performing
  39701. 30:24:18well, but it's overfitting. This is a
  39702. 30:24:20really good one to use because you can
  39703. 30:24:21balance
  39704. 30:24:23lasso and ridge you can get the best of
  39705. 30:24:25both worlds. So the the main strategy is
  39706. 30:24:28if you are using a linear regression and
  39707. 30:24:31you see overfitting
  39708. 30:24:33um meaning that it's performing decently
  39709. 30:24:36so on the training set
  39710. 30:24:39it's performing okay but then on the
  39711. 30:24:41test set like you know it's it's not
  39712. 30:24:44underfitting. it's performing pretty
  39713. 30:24:45well on the training set, but then on
  39714. 30:24:47the test set it's um performance is much
  39715. 30:24:51worse. That's overfitting. If you're
  39716. 30:24:54overfitting, then this is a great model
  39717. 30:24:55to use because we can try basically by
  39718. 30:24:58by rotating through different alphas and
  39719. 30:25:00different L1 ratios, we can try out
  39720. 30:25:03different strengths of penalty and
  39721. 30:25:06different variations on lasso and ridge
  39722. 30:25:08together. This is a really good model to
  39723. 30:25:11to use for those overfitting cases where
  39724. 30:25:13linear regression is doing decently
  39725. 30:25:17um but it's overfitting.
  39726. 30:25:20Right? So far we haven't ran that case
  39727. 30:25:22because so far no matter what model
  39728. 30:25:25we've used it's always underfit. So
  39729. 30:25:28anytime we have those underfitting cases
  39730. 30:25:31it signals that we should likely just
  39731. 30:25:33use a more complex model and we haven't
  39732. 30:25:36learned about those yet.
  39733. 30:25:38um we will coming up in lesson four, but
  39734. 30:25:42um that's for this data. That's ultim
  39735. 30:25:45ultimately what we'd want to do is
  39736. 30:25:46probably use a more advanced model
  39737. 30:25:48because it's underfitting um just using
  39738. 30:25:50a linear regression and and then using
  39739. 30:25:52the the overfitting variations of linear
  39740. 30:25:54regression like lasso ridge and elastic
  39741. 30:25:56net.
  39742. 30:26:03Okay.
  39743. 30:26:06Any questions on this on elastic then
  39744. 30:26:15the TV? Yeah, we Yeah, I think I have
  39745. 30:26:17it. I can share it with you.
  39746. 30:26:24I said that and now I can't find it. I
  39747. 30:26:26thought I had it.
  39748. 30:26:35I don't have it. I thought I had it, but
  39749. 30:26:37I don't.
  39750. 30:26:39If anyone does have that one.
  39751. 30:26:45Yeah, I'll look one more time. I thought
  39752. 30:26:47I had that one.
  39753. 30:26:51Um,
  39754. 30:26:57yeah, it's not in there. I had it. Let
  39755. 30:26:59me see.
  39756. 30:27:12Yeah, I don't have it either. I thought
  39757. 30:27:13I had it in here.
  39758. 30:27:22Yeah, I don't have that one.
  39759. 30:27:27I don't have that one. and I'll have to
  39760. 30:27:28find it. Uh I have this marketing data.
  39761. 30:27:30I don't think this is the same one.
  39762. 30:27:34I have this marketing data. I don't
  39763. 30:27:35think that's the right one, but you can
  39764. 30:27:36take a look at it.
  39765. 30:27:41No, we're using So, for this example,
  39766. 30:27:43we're using the same hitters data set
  39767. 30:27:45that we used earlier for lasso.
  39768. 30:27:48No, that's an earlier one.
  39769. 30:27:54That's from the uh very beginning of the
  39770. 30:27:58notebook. So that's the that's from this
  39771. 30:28:01one.
  39772. 30:28:05Oh, this Oh, this is where it is. Sorry.
  39773. 30:28:08This is where it is. You can find it
  39774. 30:28:09here.
  39775. 30:28:14That's right. It was from a URL.
  39776. 30:28:20It was used in the very beginning of the
  39777. 30:28:21notebook.
  39778. 30:28:23And we did we did this.
  39779. 30:28:27Okay.
  39780. 30:28:30That's right. That's why I didn't have
  39781. 30:28:31it downloaded.
  39782. 30:28:35Okay.
  39783. 30:28:39All right. Any other questions on the
  39784. 30:28:41elastic net before I move I'm going to
  39785. 30:28:43move on to uh finding those a systematic
  39786. 30:28:47way to find the best hyperparameters.
  39787. 30:28:50Um, I'm going to show you a couple
  39788. 30:28:51strategies to doing that. Um, so far
  39789. 30:28:54we've just ran CV with some random
  39790. 30:28:56choices. Um, I'm going to show you a
  39791. 30:28:59better, more systematic approach. That's
  39792. 30:29:00kind of the industry standard for doing
  39793. 30:29:02tuning. Um, so I'm going to I'm going to
  39794. 30:29:05show you that next. But any questions on
  39795. 30:29:07the elastic net?
  39796. 30:29:13Okay. And again like you know
  39797. 30:29:16scikitlearn makes it really easy for you
  39798. 30:29:17guys because it just behaves the same
  39799. 30:29:21way as any other model. You use the
  39800. 30:29:24object and then you do ffit and predict
  39801. 30:29:26right? So the ffit is going to train it
  39802. 30:29:29um and the predict is going to allow you
  39803. 30:29:32to use that model to predict. It's it's
  39804. 30:29:34super easy that way. Every scikitlearn
  39805. 30:29:36model is like that dofitit and predict.
  39806. 30:29:39So it provides a really simple way to
  39807. 30:29:41use basically every model.
  39808. 30:29:47Okay,
  39809. 30:29:49let's talk about let's finish up this
  39810. 30:29:51lesson with a couple things. Um, one of
  39811. 30:29:54those things is going to be
  39812. 30:29:55hyperparameter tuning. So what is this?
  39813. 30:29:59The hyperparameter tuning is a
  39814. 30:30:01systematic way to find the best
  39815. 30:30:05parameters in a machine learning model.
  39816. 30:30:08So a lot of machine learning models have
  39817. 30:30:10what are called hyperparameters.
  39818. 30:30:13These are not the betas that we learn
  39819. 30:30:15during the training that's learned from
  39820. 30:30:17the data. These are settings that we set
  39821. 30:30:20ahead of time like the alpha. That's a
  39822. 30:30:23perfect example like alpha L1 ratio in
  39823. 30:30:26in the elastic net. We set those up
  39824. 30:30:28ahead of time and depending on what we
  39825. 30:30:30pick for those we get different
  39826. 30:30:31performance, right? And so what we
  39827. 30:30:34really need is a systematic way to find
  39828. 30:30:37the best settings for those
  39829. 30:30:40hyperparameters as we are training our
  39830. 30:30:42models. Um the the the main like idea
  39831. 30:30:48behind this process though is going to
  39832. 30:30:50be to systematically try out different
  39833. 30:30:54combinations as many as we want to try.
  39834. 30:30:57And so we're we're basically going to
  39835. 30:30:59have a strategy for tuning that is going
  39836. 30:31:02to exhaust all the combinations of those
  39837. 30:31:06hyperparameters that we want to try
  39838. 30:31:08until we find the one that performs the
  39839. 30:31:11best. Um and and that strategy is known
  39840. 30:31:14as grid search. Um and essentially what
  39841. 30:31:19it does is it sets up a grid um where
  39842. 30:31:22which is basically like a matrix to say
  39843. 30:31:25okay which parameters do you want to
  39844. 30:31:27try? I want to try um alpha and I want
  39845. 30:31:30to try L1 ratio
  39846. 30:31:33um L1 ratio like let's say I want to try
  39847. 30:31:37these two. So we set these up in a grid
  39848. 30:31:39where we say, "Okay, I want to try this
  39849. 30:31:41value. I want to try this value. I want
  39850. 30:31:43to try this value. This one, this one,
  39851. 30:31:44this one, and on and as many as we want
  39852. 30:31:47to try." So we could set up set those up
  39853. 30:31:49systematically like a linear um a
  39854. 30:31:52linearly spaced like I want to try every
  39855. 30:31:55alpha between between 0 and 10 spaced by
  39856. 30:31:58one. Um whatever, you know, we can set
  39857. 30:32:01up different ranges of those, but that's
  39858. 30:32:03going to be in this grid. And then the
  39859. 30:32:05L1 ratio, same thing. We can try out
  39860. 30:32:07different values of these that we want
  39861. 30:32:08to try. Let's say there's many of those.
  39862. 30:32:12Um maybe every um tenth between 0 to one
  39863. 30:32:16I want to try out. Um so you set up your
  39864. 30:32:20parameters and you can set up as many as
  39865. 30:32:21you want in the grid. And then
  39866. 30:32:23essentially what you're going to do to
  39867. 30:32:25do grid search is you're going to work
  39868. 30:32:27your way through every combination of
  39869. 30:32:29those. You're going to try out this
  39870. 30:32:31combo. You're going to try out this
  39871. 30:32:33combo. You're going to try out this
  39872. 30:32:34combo.
  39873. 30:32:36this combo. So the first value of alpha
  39874. 30:32:40with every possible L1 ratio, then go to
  39875. 30:32:43the next, try out the next value of
  39876. 30:32:45alpha with every L1 ratio, and on and on
  39877. 30:32:48and on. So we're going to try
  39878. 30:32:51all combos
  39879. 30:32:54in the grid.
  39880. 30:32:56We're going to try all combos and we're
  39881. 30:32:58going to find the lowest MSE
  39882. 30:33:02combination. find lowest
  39883. 30:33:05MSE
  39884. 30:33:07combo.
  39885. 30:33:08So whatever leads to the best model is
  39886. 30:33:11going to be the um parameters that are
  39887. 30:33:14that are deemed to be the best. And the
  39888. 30:33:16idea is once we have found those we know
  39889. 30:33:20that we can use we can go ahead and
  39890. 30:33:22train a model with those best alpha and
  39891. 30:33:24len ratio and on and on and on.
  39892. 30:33:34Yeah, when you get an So this goes back
  39893. 30:33:36to the error. Remember that for a
  39894. 30:33:39regression,
  39895. 30:33:41the error is this measurement of how far
  39896. 30:33:44off we are, right? So if we have a bunch
  39897. 30:33:46of points and we draw we fit a line
  39898. 30:33:49through there, the the MSE is measuring
  39899. 30:33:52this distance, right? So what do you
  39900. 30:33:54think is a good distance? Like if our
  39901. 30:33:56model is perfect,
  39902. 30:33:59what's the best distance from our
  39903. 30:34:02predictions to the actual points? Zero.
  39904. 30:34:05Yes. So the lower the better. The lower
  39905. 30:34:09the better. Um so for an R RMSSE, the
  39906. 30:34:13lower the closer to zero the better.
  39907. 30:34:15However, the RMSSE can be it's its units
  39908. 30:34:19are interpreted in the units of our
  39909. 30:34:22target.
  39910. 30:34:23So what is deemed to be good is relative
  39911. 30:34:26to our target. Like let's say our target
  39912. 30:34:28is in the thousands. Like it averages in
  39913. 30:34:31the thousands. If we produce an MSE of
  39914. 30:34:3450 or sorry an RMSSE of 50, that's
  39915. 30:34:39pretty good, right? Because our units
  39916. 30:34:41are in the thousands
  39917. 30:34:43and we're only on average we are off by
  39918. 30:34:4750 units,
  39919. 30:34:49right? Our distance away is about 50
  39920. 30:34:51units. That's pretty good. So the RMSSE
  39921. 30:34:55is relative to your target variable.
  39922. 30:34:58Does that make sense? Yeah. It depends
  39923. 30:35:00on the target. It depends on what you're
  39924. 30:35:02trying to predict.
  39925. 30:35:05So that's why we got RMSSE that were in
  39926. 30:35:07the 300s for those hitters, but the
  39927. 30:35:09average was the average of the target
  39928. 30:35:11was in the 500s. So that's a really bad
  39929. 30:35:15proportion of error relative to the
  39930. 30:35:18average target value, right? If our
  39931. 30:35:21RMSSE was 300,
  39932. 30:35:24but the target was sitting in the 500s,
  39933. 30:35:27that's just too much error. Way too much
  39934. 30:35:30error, right? That's just too big of a
  39935. 30:35:32value. Um, our predictions are just off
  39936. 30:35:36way too much in terms of that distance.
  39937. 30:35:39So, this would be this was a bad model.
  39938. 30:35:42It was underfit.
  39939. 30:35:44We know that from the the RMSSE. So
  39940. 30:35:47yeah, the RMSSE closer to zero, no
  39941. 30:35:49matter what is good,
  39942. 30:35:52zero is being perfect. Um, but it to
  39943. 30:35:56know what's good, you need to know what
  39944. 30:35:58your target is on average and then think
  39945. 30:36:01of this as kind of a ratio to that
  39946. 30:36:03average target. I think that's the best
  39947. 30:36:06way to think about it.
  39948. 30:36:17Okay, so going back to this grid idea is
  39949. 30:36:22so the grid is just basically laying out
  39950. 30:36:25all possible parameter combinations and
  39951. 30:36:28trying them all out by fitting and
  39952. 30:36:30predicting until and generating an a
  39953. 30:36:33metric like an MSE
  39954. 30:36:36until we find the one with the lowest
  39955. 30:36:38MSE. So find the lowest MSE combination
  39956. 30:36:42and that will be the best
  39957. 30:36:44that will be the best combo. And then if
  39958. 30:36:47we once we know that best combo we can
  39959. 30:36:49use that we can use that alpha we can
  39960. 30:36:51use that L1 ratio and use that model
  39961. 30:36:54going forward. We can we can use those
  39962. 30:36:56parameters in our model. So this
  39963. 30:36:59strategy it has a name. It's known as
  39964. 30:37:02grid search. So it is a hyperparameter
  39965. 30:37:05tuning process that tries out all
  39966. 30:37:08combinations.
  39967. 30:37:11So what's the what's the uh benefit to
  39968. 30:37:15this is that we get to test out a lot of
  39969. 30:37:17different combo combos of those
  39970. 30:37:18parameters like the alpha and L1. So we
  39971. 30:37:21can be confident what the best model is,
  39972. 30:37:23right? So we can pick the alpha and L1
  39973. 30:37:26perfectly because we're trying out a
  39974. 30:37:27bunch of different combinations on the
  39975. 30:37:29data to see which one's the best. What's
  39976. 30:37:32the downside?
  39977. 30:37:34It's expensive, right? It's an
  39978. 30:37:37exhaustive search. So, if you have many
  39979. 30:37:40different parameters and you're trying
  39980. 30:37:43out many different combinations, it can
  39981. 30:37:46get exponentially
  39982. 30:37:48expensive
  39983. 30:37:49to perform this search. Okay, so grid
  39984. 30:37:53search is great except for the fact that
  39985. 30:37:56it can be expensive if you have many
  39986. 30:37:57parameters with with very wide ranges
  39987. 30:38:00that you're searching over because that
  39988. 30:38:02that's a lot of combinations you have to
  39989. 30:38:03test, right? And especially if you have
  39990. 30:38:06a lot of data, that's going to be
  39991. 30:38:08expensive
  39992. 30:38:10um to do.
  39993. 30:38:13So uh we're going to practice doing grid
  39994. 30:38:16search, but that is that's the pro and
  39995. 30:38:17con. The pro is that we get to try out
  39996. 30:38:19all these combinations and see which
  39997. 30:38:20one's the best. The downside is it can
  39998. 30:38:23be expensive to do that if you have a
  39999. 30:38:24lot of parameters um that you want to
  40000. 30:38:27tune for your model um and you have very
  40001. 30:38:32uh many different choices that you're
  40002. 30:38:34trying to evaluate for those and it just
  40003. 30:38:36creates a really big um collection of
  40004. 30:38:39combinations that you have to try out,
  40005. 30:38:42right? Um that's the only downside to
  40006. 30:38:45grid search.
  40007. 30:38:48Now on the opposite end of the spectrum
  40008. 30:38:50of that is a randomized search or random
  40009. 30:38:54search and this will basically just um
  40010. 30:38:58do a sampling of those parameters from
  40011. 30:39:03um kind of fixed uh specified
  40012. 30:39:05distribution. So essentially what you do
  40013. 30:39:08is similarly you define your range. So
  40014. 30:39:11you say I want to look at alphas um
  40015. 30:39:14between zero or sorry between let's say
  40016. 30:39:18yeah 0 to 10. I want to look at a bunch
  40017. 30:39:20of different alphas um and I want to
  40018. 30:39:22look at a bunch of different L1 ratios
  40019. 30:39:25that are between 0 to 1
  40020. 30:39:290 to one and um what we do is we say
  40021. 30:39:33okay I'm going to restrict only testing
  40022. 30:39:3720 30 40 times. I'm not going to do all
  40023. 30:39:40possible combinations. I'm just going to
  40024. 30:39:43randomly sample something in this range
  40025. 30:39:46and randomly sample something in this
  40026. 30:39:48range. And so, and I'm going to perform
  40027. 30:39:50that experiment a fixed number of times.
  40028. 30:39:53So, let's say I set the uh sampling
  40029. 30:39:56where I'm only going to do um 20
  40030. 30:40:00evaluations.
  40031. 30:40:02And so, 20 times we're going to pick a
  40032. 30:40:04combo randomly. So, I'm going to pick an
  40033. 30:40:08alpha and I'm going to pick an L1 ratio.
  40034. 30:40:14L1 ratio.
  40035. 30:40:16And um we are we are just going to uh
  40036. 30:40:20sample those randomly from this range.
  40037. 30:40:24Um and we're going to use those and test
  40038. 30:40:28those out. And then it's but otherwise
  40039. 30:40:29it's the same as grid search. Whatever
  40040. 30:40:31is the lowest MSE. Um, so whatever is
  40041. 30:40:34the lowest MSE is the best.
  40042. 30:40:37So we evaluate those, we sample, we
  40043. 30:40:40train the model, evaluate it. Whatever
  40044. 30:40:42is the lowest MSE
  40045. 30:40:45is the best is the best combo. Now
  40046. 30:40:49what's the benefit to this is it's a
  40047. 30:40:52much more controlled experiment in the
  40048. 30:40:56sense that we um aren't going to iterate
  40049. 30:40:59through every possible combination in
  40050. 30:41:00the grid. where we basically set up a
  40051. 30:41:03fixed number of times we're going to try
  40052. 30:41:05out stuff.
  40053. 30:41:06The risk to doing this is that you're
  40054. 30:41:08not
  40055. 30:41:10you're not exploring all combinations,
  40056. 30:41:12right? Because you're randomly sampling,
  40057. 30:41:14you may get unlucky and you may not
  40058. 30:41:17stumble into the best. You you can make
  40059. 30:41:21um samples and figure out what's the
  40060. 30:41:22best amongst your samples, but you may
  40061. 30:41:25not be covering all the combinations.
  40062. 30:41:27Does that make sense? The grid search is
  40063. 30:41:29going to try every combo. The random
  40064. 30:41:32search is going to randomly sample those
  40065. 30:41:35combos.
  40066. 30:41:36So, it's not going to try every single
  40067. 30:41:38one. It's going to try a limited number,
  40068. 30:41:40however many you set up. Now, if you set
  40069. 30:41:43that number really, really, really high,
  40070. 30:41:46now you're starting to approach a grid
  40071. 30:41:47search because now you're sampling so
  40072. 30:41:50many of those combos that you basically
  40073. 30:41:52are trying them all at that point,
  40074. 30:41:56right?
  40075. 30:41:58Um
  40076. 30:42:00so so that's the way the random search.
  40077. 30:42:02So by the way both of these use cross
  40078. 30:42:04validation in the sense that when you
  40079. 30:42:07evaluate accommodation you're actually
  40080. 30:42:09doing it with cross validation. So when
  40081. 30:42:11you do an evaluation, you're going to do
  40082. 30:42:14probably 10 or five folds where you
  40083. 30:42:16split your data and then you test it on
  40084. 30:42:19the rest of the folds and evaluate or
  40085. 30:42:21train it on the rest of the folds,
  40086. 30:42:22evaluate it on one of them and generate
  40087. 30:42:24an average MSE to get your evaluation.
  40088. 30:42:29So every evaluation is using cross
  40089. 30:42:31validation. That's why that's and
  40090. 30:42:34hopefully you can see why this would be
  40091. 30:42:35so expensive for a really big grid,
  40092. 30:42:38right? because you're trying out many
  40093. 30:42:40different combinations
  40094. 30:42:42and every combination is going to do a
  40095. 30:42:44cross validation procedure. So, it's
  40096. 30:42:47going to train 10 times and test against
  40097. 30:42:5010 different folds and average those
  40098. 30:42:52together. It's going to be a pretty
  40099. 30:42:53expensive operation
  40100. 30:42:55for a really big grid, right, of of
  40101. 30:42:58parameters.
  40102. 30:43:00Um, but these are the two kind of
  40103. 30:43:01systematic approaches we have at trying
  40104. 30:43:04out different hyperparameters. Remember
  40105. 30:43:07those those things are called
  40106. 30:43:08hyperparameters. These are those choices
  40107. 30:43:11that we have before we train our model.
  40108. 30:43:14Um those choices we have that affect the
  40109. 30:43:17performance of the model like the
  40110. 30:43:18alphas, the L1 ratios, those kind of
  40111. 30:43:20things. Um we have control over what
  40112. 30:43:23they're going to be. This is a
  40113. 30:43:24systematic approach to find out what the
  40114. 30:43:26best
  40115. 30:43:29uh value of those parameters is going to
  40116. 30:43:31be on our data,
  40117. 30:43:34right?
  40118. 30:43:37Okay. So, before we practice this, we're
  40119. 30:43:40going to practice a grid search first.
  40120. 30:43:42Um,
  40121. 30:43:44any questions?
  40122. 30:43:54Uh, I don't know if it has a built-in
  40123. 30:43:56That's a good question. by time limit. I
  40124. 30:43:57don't know if it has a built-in way of
  40125. 30:43:59doing it, but you could certainly set up
  40126. 30:44:00like a a a loop um to like to wrap
  40127. 30:44:05around, you know what I mean? Like you
  40128. 30:44:07could set up a loop where you check the
  40129. 30:44:08time. If it's if if the time elapsed as
  40130. 30:44:11you're doing the search, if the time
  40131. 30:44:12elapsed is greater than the the time
  40132. 30:44:15limit, then you can kind of break early.
  40133. 30:44:18Um so it's not hard to implement that,
  40134. 30:44:20but I don't know if it has that built
  40135. 30:44:22in. I don't think it does
  40136. 30:44:25because I don't think it really cares
  40137. 30:44:26how long every evaluation takes. It's
  40138. 30:44:29just going to exhaust all those
  40139. 30:44:30especially in a grid search.
  40140. 30:44:33But um yeah, I there's probably a way to
  40141. 30:44:36manually kind of set up a time time
  40142. 30:44:38loop.
  40143. 30:44:44So hyperparameters are um settings that
  40144. 30:44:49we have on the model itself and a really
  40145. 30:44:52good example of this is like the alpha
  40146. 30:44:53and L1 ratio in the in the elastic net.
  40147. 30:44:56So they're not things that we um learn
  40148. 30:45:00from the data directly like the betas in
  40149. 30:45:02the model like those get trained
  40150. 30:45:05directly by doing the um lease squares
  40151. 30:45:08process right um by doing that gradient
  40152. 30:45:11descent and all that optimization.
  40153. 30:45:14Um so these are not learned from that.
  40154. 30:45:16They're actually set ahead of time. And
  40155. 30:45:19so what we're saying is
  40156. 30:45:21the best way to understand the effects
  40157. 30:45:23of those is to try out different
  40158. 30:45:25combinations of those until we land on
  40159. 30:45:27the best one. Right? So hyperparameters
  40160. 30:45:30are those options we have in the model
  40161. 30:45:32like the alpha like the alpha and l1
  40162. 30:45:35ratio in the uh elastic net. Many models
  40163. 30:45:40have hyperparameters. Um we're actually
  40164. 30:45:42going to see that in in future models
  40165. 30:45:44that we study. they have options that
  40166. 30:45:46you can set that affect their
  40167. 30:45:48performance.
  40168. 30:45:49And so this this is just a strategy to
  40169. 30:45:52evaluate those different options to see
  40170. 30:45:53which one's the best.
  40171. 30:46:05Yeah. So again, hyperparameters, those
  40172. 30:46:08are settings on the model itself um that
  40173. 30:46:12affect the performance of it.
  40174. 30:46:16And basically we have the two two
  40175. 30:46:18strategies here. We can set up an
  40176. 30:46:20exhaustive grid and search through all
  40177. 30:46:22of those until we find the lowest MSE uh
  40178. 30:46:24option or we can randomly sample
  40179. 30:46:28potential options, try them out and see
  40180. 30:46:30which one's the lowest as well. And do
  40181. 30:46:32that a fixed number of times. Um
  40182. 30:46:37sort of like a fixed number of trials
  40183. 30:46:38almost. um which has a risk of not
  40184. 30:46:42trying out every option but but
  40185. 30:46:44hopefully you try out enough that you've
  40186. 30:46:46explored the space a bit and you get
  40187. 30:46:49some quality choices there but no
  40188. 30:46:52guarantees right no guarantees you try
  40189. 30:46:54everything which is what a grid search
  40190. 30:46:55will do it will try everything
  40191. 30:47:01okay now luckily per usual scikitlearn
  40192. 30:47:06has something to manage this process for
  40193. 30:47:08us in terms of grid search. Um so in
  40194. 30:47:13that way we will not need to manage this
  40195. 30:47:16process ourselves. We can just rely on
  40196. 30:47:18scikitlearn. And so if you're doing
  40197. 30:47:20hyperparameter tuning um this is going
  40198. 30:47:23to come from the model selection module
  40199. 30:47:25inside of sklearn. So we're going to
  40200. 30:47:28import from from skarn the model
  40201. 30:47:30selection module. We have our grid
  40202. 30:47:32search cross validation.
  40203. 30:47:35Okay. That's what the CV stands for.
  40204. 30:47:37grid search cross validation. So, this
  40205. 30:47:40is going to do that grid search
  40206. 30:47:41strategy. Um, we're going to set it up
  40207. 30:47:43with our dictionary essentially of
  40208. 30:47:46choices. So, we're going to say, hey,
  40209. 30:47:48here's the alphas I want to try. Here's
  40210. 30:47:49the L1 ratios I want to try. Um, and
  40211. 30:47:52here's my other settings like uh how
  40212. 30:47:55many folds I want to use, what my random
  40213. 30:47:57state is for the shuffling. So, we'll
  40214. 30:47:59set all that up. Um,
  40215. 30:48:02and then we'll just run the grid search.
  40216. 30:48:04And then what should come out of that is
  40217. 30:48:06the best options for our parameters from
  40218. 30:48:09the grid. And then we can use those
  40219. 30:48:11going forward in the we can build a
  40220. 30:48:13model with those best options, right? So
  40221. 30:48:16we're really doing some evaluation here
  40222. 30:48:19of what is going to be those best
  40223. 30:48:20alphas, those best1 ratios on our data
  40224. 30:48:23set, right? And the only way to really
  40225. 30:48:26know that is to evaluate them because
  40226. 30:48:28they're not things that are learned
  40227. 30:48:30during the training. Hopefully that
  40228. 30:48:32makes sense, right? They're not things
  40229. 30:48:33that we learn directly from training.
  40230. 30:48:36They're things that we have to set and
  40231. 30:48:38then kind of evaluate and see how they
  40232. 30:48:39affect things.
  40233. 30:48:43Okay. So, we have grid search CV. That's
  40234. 30:48:45going to be our primary um tool to do
  40235. 30:48:49the evaluations of the different
  40236. 30:48:50hyperparameter options.
  40237. 30:48:53Grid search CV. Um we're going to set up
  40238. 30:48:56our cross validation uh object here. Now
  40239. 30:48:59I want you to pay attention to this is
  40240. 30:49:01that um it's a slightly different
  40241. 30:49:04version than the kfold we had earlier.
  40242. 30:49:06So we've used kfold before with a
  40243. 30:49:08certain number of folds. This would be
  40244. 30:49:0910 folds and we can set a random state
  40245. 30:49:12for the shuffling um that happens in the
  40246. 30:49:14folds.
  40247. 30:49:16But this is actually a slight different
  40248. 30:49:17variation on it where it is a repeated
  40249. 30:49:19kfold where we do three repeated trials.
  40250. 30:49:23Now why would we do that? It's to be
  40251. 30:49:26extra extra extra careful with the
  40252. 30:49:29shuffling.
  40253. 30:49:31So this what this means is we do three
  40254. 30:49:33different shuffles. So we do k-fold, we
  40255. 30:49:36actually repeat it three times with
  40256. 30:49:38three different shufflings. That's all
  40257. 30:49:39that means. So the repeated k-fold is
  40258. 30:49:43actually a bit beyond the just basic
  40259. 30:49:46kfold. What basic kfold will do will
  40260. 30:49:49we'll shuffle and then do our splits
  40261. 30:49:52into 10 splits and then train on nine of
  40262. 30:49:54those. test on the other one and rotate
  40263. 30:49:57through all the splits.
  40264. 30:49:59We're actually going to do that process
  40265. 30:50:02three different times with three
  40266. 30:50:04different shuffles. So this and we're
  40267. 30:50:06going to average 30 results instead of
  40268. 30:50:09just 10. So repeated kfold is just going
  40269. 30:50:13ab above and beyond to do extra to
  40270. 30:50:16repeat the kfold three different times.
  40271. 30:50:18In this case only three. We could do
  40272. 30:50:19more.
  40273. 30:50:21But um now is that necessary to do? You
  40274. 30:50:24could argue not necessarily. Um but it
  40275. 30:50:27just provides extra robustness
  40276. 30:50:30uh beyond just our single shuffle and
  40277. 30:50:33then split and then rotation of those
  40278. 30:50:36folds. Right? We're doing it actually
  40279. 30:50:37three different shuffles. Um so we're
  40280. 30:50:40repeating our kfold three times uh for
  40281. 30:50:43every now is the thing is we're doing
  40282. 30:50:45that for every evaluation. So it is
  40283. 30:50:48going to be more expensive than just a
  40284. 30:50:49basic kfold.
  40285. 30:51:06So we have three different K-fold trials
  40286. 30:51:08that we're doing essentially.
  40287. 30:51:12Okay, hopefully that makes sense. This
  40288. 30:51:13is the repeated K-fold. We haven't
  40289. 30:51:15really seen that before. We've only
  40290. 30:51:16worked with the Kfold, which would get
  40291. 30:51:19rid of this repeats option and only have
  40292. 30:51:21uh 10 splits in a random state for the
  40293. 30:51:24for the single shuffle that we do. So we
  40294. 30:51:26can um recreate that same shuffle every
  40295. 30:51:29time. Um but now we're actually going to
  40296. 30:51:32do three random shuffles, uh three
  40297. 30:51:34different trials. So one shuffle creates
  40298. 30:51:37the and then create the 10 splits,
  40299. 30:51:38evaluate, then go back and do another
  40300. 30:51:40shuffle, another new 10 splits. So, one
  40301. 30:51:44thing that should be um clear is that we
  40302. 30:51:48get different splits every time because
  40303. 30:51:50we're going to shuffle once, right?
  40304. 30:51:52We're going to shuffle once and generate
  40305. 30:51:54our splits
  40306. 30:51:56and then we're going to shuffle again,
  40307. 30:51:58generate these splits which are going to
  40308. 30:52:00be different and then shuffle one more
  40309. 30:52:01time for for three different times,
  40310. 30:52:04right? And then get get these splits and
  40311. 30:52:06then we're going to get 10 metrics here,
  40312. 30:52:0810 metrics here, 10 metrics here, and
  40313. 30:52:11then average all of those together.
  40314. 30:52:14So, it's a bit more just going up extra
  40315. 30:52:17above and beyond for a kfold. Okay.
  40316. 30:52:22All right. So, here comes the fun of
  40317. 30:52:24when you do grid search. Now, the grid
  40318. 30:52:28is actually just a dictionary. It's a
  40319. 30:52:30Python dictionary where you declare what
  40320. 30:52:34your parameters are going to be inside
  40321. 30:52:36the dictionary and you set up a range of
  40322. 30:52:39values that you're go or a list. It can
  40323. 30:52:42be a list. It can be a range
  40324. 30:52:44but some declaration of what you are
  40325. 30:52:47going to test and evaluate inside of
  40326. 30:52:50your grid search. So the grid is
  40327. 30:52:52initialized as an empty dictionary.
  40328. 30:52:55And then what we do is we say okay in my
  40329. 30:52:58grid I want to check different alphas.
  40330. 30:53:00So we're going to add a collection of
  40331. 30:53:02alphas in here that we're going to test.
  40332. 30:53:06So let me make a comment there. We add a
  40333. 30:53:10add a range of alphas to test. And this
  40334. 30:53:16range is a this is just like the Python
  40335. 30:53:20range. Um
  40336. 30:53:23this is just like a Python range um uh
  40337. 30:53:26operator here where this is going to be
  40338. 30:53:29uh every so it's going to be um every
  40339. 30:53:34uh value between
  40340. 30:53:37zero and one um uh steps with a step
  40341. 30:53:43size
  40342. 30:53:45of 0.1. So it's going to try a bunch of
  40343. 30:53:49different alphas um between uh zero and
  40344. 30:53:530.1
  40345. 30:53:54sorry 0 and one stepping by 0.1. So it's
  40346. 30:53:57going to try zero.1
  40347. 30:53:592.3 point 4.5 6 right all the way up to
  40348. 30:54:02one.
  40349. 30:54:04So that's what this will do and it's a
  40350. 30:54:06numpy range. So it's just all those
  40351. 30:54:08decimals between 0 to one.
  40352. 30:54:11It you can use either that's valid.
  40353. 30:54:14Yeah, you can do you can do that to
  40354. 30:54:16create a dictionary or you can use the
  40355. 30:54:18keyword um dict. You can use either one.
  40356. 30:54:21Either one works.
  40357. 30:54:25Whatever whatever you want to use.
  40358. 30:54:26They're the same.
  40359. 30:54:29Yeah. The the reason people prefer
  40360. 30:54:33dictionary is because um sets are
  40361. 30:54:37created with the same braces.
  40362. 30:54:39So it it makes it clear what you're
  40363. 30:54:41creating as a dictionary. If you use if
  40364. 30:54:43you use this, that's the only advantage
  40365. 30:54:46is it's just plainly obvious what you're
  40366. 30:54:48making. Uh because technically you can
  40367. 30:54:50make a set with the curly braces as
  40368. 30:54:54well.
  40369. 30:54:59Yeah,
  40370. 30:55:02no worries. Um okay, so we have our
  40371. 30:55:05alphas here. So what I want you to
  40372. 30:55:07notice is that we are going to try out
  40373. 30:55:09different alphas and we are that's the
  40374. 30:55:12only parameter we are going to try in
  40375. 30:55:13our ridge regression. So we're going to
  40376. 30:55:16we're going to try ridge but just try
  40377. 30:55:18different alphas in the in this range um
  40378. 30:55:21in our grid search. So the grid search
  40379. 30:55:24CV takes in a model. It takes in our
  40380. 30:55:27grid dictionary which is really
  40381. 30:55:29critical. We need that dictionary to
  40382. 30:55:30declare what we're going to try.
  40383. 30:55:34um we need a scoring to say to find the
  40384. 30:55:36best. Now remember it uses the negative
  40385. 30:55:39to find the lowest which is going to be
  40386. 30:55:42the the least negative option.
  40387. 30:55:46Um otherwise it wouldn't um just based
  40388. 30:55:49on the optimization it would look for
  40389. 30:55:51the highest value. Um so the highest
  40390. 30:55:54would be closest to zero in this
  40391. 30:55:56situation. Um so we use negative and
  40392. 30:55:59again we could use squared error. It's
  40393. 30:56:02using absolute. We could use um squared
  40394. 30:56:05uh either either one works.
  40395. 30:56:09Um more typical would probably be
  40396. 30:56:11squared error, but um absolute is fine.
  40397. 30:56:15Here's where we have our repeated kfold.
  40398. 30:56:17So we pass in our um how we're doing CV.
  40399. 30:56:20That can be it can be a kfold object. It
  40400. 30:56:22can actually just be an integer, which
  40401. 30:56:24is say I just want to do 10 splits or
  40402. 30:56:26five splits um to to do every
  40403. 30:56:29evaluation. But these are the bare
  40404. 30:56:32minimum that you need just really the
  40405. 30:56:33model and the grid and your CV. Um what
  40406. 30:56:37metric you're using to evaluate what's
  40407. 30:56:39going to be the best. And then this end
  40408. 30:56:41jobs is to parallelize. If you have it
  40409. 30:56:43set to minus one, it's going to it's
  40410. 30:56:44going to try out all the grid options in
  40411. 30:56:46parallel. Um which is nice. It's going
  40412. 30:56:49to help speed up the overall search.
  40413. 30:56:52Okay. So let me mark that down as n
  40414. 30:56:57jobs equals minus one.
  40415. 30:57:01tries out the combos in parallel.
  40416. 30:57:06So in this situation, we actually don't
  40417. 30:57:08have more than one parameter. We only
  40418. 30:57:10have the alpha. So we're really just
  40419. 30:57:12going to be systematically working our
  40420. 30:57:14way through every alpha and evaluating
  40421. 30:57:16which one's the best right with this.
  40422. 30:57:19And notice that in order to use this
  40423. 30:57:22grid search, all we have to do is call
  40424. 30:57:23search.fit. So it works kind of like
  40425. 30:57:27every other model does, right? It's the
  40426. 30:57:29grid search.fit.
  40427. 30:57:32And we pass in our data.
  40428. 30:57:34And we um once we're once this prints
  40429. 30:57:38out the results, you get a results
  40430. 30:57:40object
  40431. 30:57:42um which has a best score and then a
  40432. 30:57:45dictionary with your best parameters.
  40433. 30:57:47So, whatever your best grid member was
  40434. 30:57:50or grid members, um it prints that out
  40435. 30:57:52and you can So, for from that, we can um
  40436. 30:57:56grab our best alpha, which it which
  40437. 30:57:58let's confirm what that ends up being.
  40438. 30:58:06Oops. We need to import repeated kfold.
  40439. 30:58:17So we'll import that.
  40440. 30:58:24Oh, I didn't. Let's do from
  40441. 30:58:27sklearn.linear
  40442. 30:58:32model import ridge.
  40443. 30:58:36Okay.
  40444. 30:58:44Okay. So, it completed the search and
  40445. 30:58:46what we found is this is the best score
  40446. 30:58:48is 238 for the mean absolute error and
  40447. 30:58:52the best alpha that we got was 0.9. So,
  40448. 30:58:56the best alpha that worked here, the one
  40449. 30:58:59that gave us the best score was actually
  40450. 30:59:010.9 as the alpha. So what it did is it
  40451. 30:59:04tried out everything between this range
  40452. 30:59:07and 0.9 was the best. So it did cross
  40453. 30:59:10validation tried out every single combo
  40454. 30:59:13in our grid.
  40455. 30:59:15So if we want we could actually print
  40456. 30:59:17out
  40457. 30:59:20print our grid so we can see
  40458. 30:59:23what our combinations were.
  40459. 30:59:29So, it tried out all of these guys and
  40460. 30:59:32the best one that we had was 0.9.
  40461. 30:59:40Okay, so pretty cool how that works. And
  40462. 30:59:44if we had other parameters, like if we
  40463. 30:59:46were doing a elastic net, we could add
  40464. 30:59:48those into our dictionary and it would
  40465. 30:59:50do all combinations of those. So if we
  40466. 30:59:53did um so for instance to add to our
  40467. 30:59:56grid we could do grid
  40468. 30:59:58um L1 ratio
  40469. 31:00:01this would be for like an elastic net
  40470. 31:00:02right now the ridge regression by itself
  40471. 31:00:04doesn't have an L1 ratio parameter but
  40472. 31:00:06just as an example um we could try out
  40473. 31:00:09different ranges
  40474. 31:00:11um similar range different one um maybe
  40475. 31:00:14an exact list whatever we want to do. So
  40476. 31:00:17this is going to try out different ones
  40477. 31:00:18between 0 to one as well. And so it's
  40478. 31:00:22going to try out every combination of
  40479. 31:00:24these from this grid.
  40480. 31:00:27Okay, if we did that. But again, this
  40481. 31:00:29the ridge regression doesn't have an L1
  40482. 31:00:32ratio. The elastic net does. So that the
  40483. 31:00:35ridge regression only has an alpha to to
  40484. 31:00:38as a hyperparameter. So we're only
  40485. 31:00:40testing out that one.
  40486. 31:00:45Okay, so that's grid search CV.
  40487. 31:00:49Pretty useful. This is pretty useful in
  40488. 31:00:51doing parameter tuning again when you
  40489. 31:00:53want to try out ranges of different
  40490. 31:00:55values and you can evaluate those to see
  40491. 31:00:58which one is your best and then we can
  40492. 31:01:01use that best going forward. So we can
  40493. 31:01:04for instance this is what this code does
  40494. 31:01:06below it is it fetches the best. Um you
  40495. 31:01:09can do it this way or you can do it um
  40496. 31:01:12the alternative is to do results.b best
  40497. 31:01:15params
  40498. 31:01:18and then you can just grab it like this
  40499. 31:01:21alpha.
  40500. 31:01:23Either way you can do get or like this
  40501. 31:01:26um and it this is just a dictionary,
  40502. 31:01:29right? And you can grab your alpha. So
  40503. 31:01:31that's the 0.9 um and we can pass that
  40504. 31:01:34alpha into the ridge regression and go
  40505. 31:01:36back and refit it to our data um and
  40506. 31:01:39then use that model going forward. So
  40507. 31:01:41the grid search really just evaluates
  40508. 31:01:44those different options, tells you
  40509. 31:01:46what's the best according to this score,
  40510. 31:01:50right?
  40511. 31:01:52And you should, by the way, you should
  40512. 31:01:54interpret this score in the positive
  40513. 31:01:56sense. It's only negative because we're
  40514. 31:01:59purposely making it negative to find out
  40515. 31:02:02what the lowest option is, right?
  40516. 31:02:04Because the lower is the better. So we
  40517. 31:02:06we purposely make it negative to make it
  40518. 31:02:08whatever is the least negative is the
  40519. 31:02:10winner. Um more negative is worse.
  40520. 31:02:15So it's the really positive version of
  40521. 31:02:18it is the is the true result for the
  40522. 31:02:21error. Um and they are a tool from
  40523. 31:02:24scikitlearn to put together your model
  40524. 31:02:26with your pre-processing steps. So they
  40525. 31:02:29kind of get automated together. Um and
  40526. 31:02:32they combine everything into kind of a
  40527. 31:02:34streamline process. You're going to see
  40528. 31:02:36what that looks like, but it's a really
  40529. 31:02:38nice um feature of scikitlearn. Um why
  40530. 31:02:42would we care about pipelines? They help
  40531. 31:02:45organize our code um so that we ensure
  40532. 31:02:48that we basically always run the
  40533. 31:02:50pre-processing steps before we train and
  40534. 31:02:52use a model to with the predictions. Um,
  40535. 31:02:56so it bundles those steps together,
  40536. 31:02:58minimizes the risk of forgetting a step
  40537. 31:03:00because one of the things that can
  40538. 31:03:02happen is when you do pre-processing, if
  40539. 31:03:04you're doing it on the training set, you
  40540. 31:03:05have to do it on new test data as well
  40541. 31:03:07when you put it through your model
  40542. 31:03:09because your model is training against
  40543. 31:03:11that pre-processed data.
  40544. 31:03:13So in order to make sure you never
  40545. 31:03:15forget that, you can bundle it all
  40546. 31:03:17together in a pipeline which is going to
  40547. 31:03:19make things really really easy to use
  40548. 31:03:22and and make sure that those steps
  40549. 31:03:25happen in a repeatable way. Um and it
  40550. 31:03:29makes things easier to uh deploy that
  40551. 31:03:32model as well because everything is
  40552. 31:03:34together in one pipeline. So in the in
  40553. 31:03:37the industry, I've seen this a lot. Um
  40554. 31:03:40you know, people will do their initial
  40555. 31:03:43exploration steps and initial model
  40556. 31:03:45building. They may not use pipelines
  40557. 31:03:47right away, but as they found their
  40558. 31:03:50model, um they'll generally move it into
  40559. 31:03:53a pipeline and all their steps into a
  40560. 31:03:54pipeline so that it's uh easier to work
  40561. 31:03:57with um when you're when you're
  40562. 31:03:58deploying it and actually using it uh in
  40563. 31:04:02in the real world. Um, so here's what a
  40564. 31:04:06pipeline generally looks like. It's from
  40565. 31:04:07scikitlearn. It's this pipeline object.
  40566. 31:04:10Um, and the pipeline is made up of steps
  40567. 31:04:13that we're going to see that that are
  40568. 31:04:15various um uh basically um kinds of
  40569. 31:04:20pre-processing we've seen before like a
  40570. 31:04:22scaler or um filling in missing values.
  40571. 31:04:26Those kind of things we can put here in
  40572. 31:04:28the steps which is basically a list. um
  40573. 31:04:31steps is just going to be a list of
  40574. 31:04:33scikitlearn functions that we can apply
  40575. 31:04:34to data. One of those being a model. Um
  40576. 31:04:38and then whenever we use the pipeline,
  40577. 31:04:40it's basically um you know it's going to
  40578. 31:04:43be something like pipeline.fit
  40579. 31:04:45or pipeline.predict.
  40580. 31:04:47So the pipeline kind of behaves like a
  40581. 31:04:50model. It's just going to contain many
  40582. 31:04:53more steps than that like the
  40583. 31:04:54pre-processing steps we've worked with
  40584. 31:04:56before. Um, and it also has some
  40585. 31:04:59capabilities for caching. So you can
  40586. 31:05:02like uh cache some of the data in
  40587. 31:05:04memory. Um, so that if you're reusing
  40588. 31:05:06the predictions, it kind of goes faster.
  40589. 31:05:09Um, so there's some options for that
  40590. 31:05:11too. I'm not too concerned about that at
  40591. 31:05:13this stage, but the main thing is going
  40592. 31:05:15to be filling out our steps and then
  40593. 31:05:17using the pipeline.
  40594. 31:05:20Okay.
  40595. 31:05:22Um, so some important bits of
  40596. 31:05:25information about the pipeline is that
  40597. 31:05:27it is going to be a sequence of data
  40598. 31:05:28transformations that will have at the
  40599. 31:05:31very end of the pipeline the model
  40600. 31:05:34because of course we're going to do
  40601. 31:05:35transformations and then train a model
  40602. 31:05:38or predict with a model. So every
  40603. 31:05:43Oh, can you guys hear me? Okay,
  40604. 31:05:46not able to hear me. Thanks for letting
  40605. 31:05:48me know. Can you guys were you able to
  40606. 31:05:49hear me so far?
  40607. 31:05:55Okay. Make sure. Yeah, it might be on
  40608. 31:05:56your internet or your your uh Yeah, it
  40609. 31:06:00seems like seems like it's good. So,
  40610. 31:06:04no, you can't hear me. Check your
  40611. 31:06:06volume. Check your headphones if you're
  40612. 31:06:08wearing headphones. Oh, no issues. Okay,
  40613. 31:06:11perfect.
  40614. 31:06:13Okay. Yeah, local internet issue. Yeah.
  40615. 31:06:17Okay.
  40616. 31:06:19Always let me know. always let me know
  40617. 31:06:20because it could be the case that it is
  40618. 31:06:22me. So, um always always make sure to
  40619. 31:06:26let me know. Um but sounds like yeah,
  40620. 31:06:29you may want to check on that. Um
  40621. 31:06:32so, okay, what I was saying is every
  40622. 31:06:35pipeline is going to have a uh a
  40623. 31:06:37sequence of steps that go first and then
  40624. 31:06:39the model at the end. Um so, the order
  40625. 31:06:42really matters. Um
  40626. 31:06:45uh so the order matters in the sense
  40627. 31:06:48that we want our transformations to go
  40628. 31:06:49first. Things like scaling, things like
  40629. 31:06:52filling in missing values, we want those
  40630. 31:06:53to be first and then we want our uh
  40631. 31:06:57model to be last because we want those
  40632. 31:06:59transformations to happen prior to
  40633. 31:07:01training or prior to prediction. So
  40634. 31:07:04usually what you'll see in these
  40635. 31:07:06pipelines is a model at the end, right?
  40636. 31:07:09a model that's going to be at the end of
  40637. 31:07:11the pipeline because we want basically
  40638. 31:07:14our processing steps then our training
  40639. 31:07:16or processing steps then our
  40640. 31:07:18predictions. Um so everything in the
  40641. 31:07:22pipeline though is going to be from
  40642. 31:07:23scikitlearn. Uh that's how it gets
  40643. 31:07:26automated in the sense that all of those
  40644. 31:07:28things are going to have fit and
  40645. 31:07:29transform functions built into them so
  40646. 31:07:31the pipeline can use them. Uh, and then
  40647. 31:07:34the last step is going to be a model
  40648. 31:07:36that has a fit and a predict. So it's
  40649. 31:07:39pretty standard that the last part of
  40650. 31:07:40the pipeline is just going to be a
  40651. 31:07:41model. Um,
  40652. 31:07:45uh, so we can um, as we do more
  40653. 31:07:50modeling, we're going to play around
  40654. 31:07:52with the pipelines quite a bit and see
  40655. 31:07:53how we can change up some of the
  40656. 31:07:55parameters. like if we want to change a
  40657. 31:07:57model's parameter um we can actually
  40658. 31:07:59adjust it to do things like uh grid
  40659. 31:08:02search or cross validation. So um we're
  40660. 31:08:05going to see some examples of some
  40661. 31:08:07pipelines but for right now mostly what
  40662. 31:08:09we're going to see is how to build one
  40663. 31:08:11and then how to use one. And then as we
  40664. 31:08:14get into lesson four, we'll get some
  40665. 31:08:16more practice with pipelines because
  40666. 31:08:18we're going to start using them quite a
  40667. 31:08:19bit uh to build our models rather than
  40668. 31:08:22do manual steps uh all the manual
  40669. 31:08:25pre-processing
  40670. 31:08:27um and then kind of building a model
  40671. 31:08:29from there. We'll just include all of it
  40672. 31:08:31together in a pipeline.
  40673. 31:08:35Okay, so the example we're going to do
  40674. 31:08:36is with this housing with ocean
  40675. 31:08:39proximity. So we've actually looked at
  40676. 31:08:40this data set before. Um so we have uh
  40677. 31:08:45this ocean proximity data set that has
  40678. 31:08:47the feature of like how close it is to
  40679. 31:08:49the ocean like the bay or the less than
  40680. 31:08:521 hour. Remember we had that and it had
  40681. 31:08:54the median house value for different
  40682. 31:08:56neighborhoods. Um so we're going to work
  40683. 31:08:58with that one again. Let me make sure I
  40684. 31:09:00have that one uploaded.
  40685. 31:09:04You guys should have this one. It should
  40686. 31:09:05be in your uh data sets.
  40687. 31:09:10Um, I'll I can upload it here in case
  40688. 31:09:11you don't have it though.
  40689. 31:09:22Does this use multi-threading? I think
  40690. 31:09:24it does. Yeah, I think in order to do it
  40691. 31:09:26can do uh um I think it can do
  40692. 31:09:29processing in parallel for some of the
  40693. 31:09:31pipeline steps. Um, now does it use that
  40694. 31:09:35all the time? Not necessarily because
  40695. 31:09:37some of it is sequential in nature where
  40696. 31:09:40you have to do one step and then you do
  40697. 31:09:41the next step and then you do the next
  40698. 31:09:43step. So it's not like you can do them
  40699. 31:09:44in parallel.
  40700. 31:09:46Um in terms of the like you need to know
  40701. 31:09:48the output of one step to compute the
  40702. 31:09:50the output of the next step. Um so it
  40703. 31:09:55can but it it doesn't always lend itself
  40704. 31:09:58well. The thing that will use
  40705. 31:10:00multi-threading is is like the training
  40706. 31:10:03process can be parallelized
  40707. 31:10:05like the fit um can be and for some
  40708. 31:10:08models it can be parallelized not every
  40709. 31:10:11model
  40710. 31:10:15it. So long answer is or the short
  40711. 31:10:18answer is that it depends
  40712. 31:10:20depends on what kind of transforms
  40713. 31:10:21you're doing and what kind of model
  40714. 31:10:22you're using. if you can really take
  40715. 31:10:24advantage of that.
  40716. 31:10:31Okay, so we load our data here and take
  40717. 31:10:33a look at that. Um, do you guys have
  40718. 31:10:36this data set? Are you able to load it
  40719. 31:10:38in? If you're following along, are you
  40720. 31:10:40able to load it?
  40721. 31:10:48Okay.
  40722. 31:10:50And and again, we've worked with this
  40723. 31:10:51data before, so hopefully it's somewhat
  40724. 31:10:53familiar. Remember, every row represents
  40725. 31:10:56a neighborhood, and it has a we're going
  40726. 31:10:58to end up trying to predict this median
  40727. 31:11:00house value as our target um variable,
  40728. 31:11:04our dependent variable. Um and we're
  40729. 31:11:06going to use the rest of these features.
  40730. 31:11:08Remember that um this feature is in
  40731. 31:11:11particular going to need to be one hot
  40732. 31:11:13encoded,
  40733. 31:11:15right? It's going to be one hot encoded
  40734. 31:11:17because it is currently a string. and we
  40735. 31:11:19need to turn that into a numerical
  40736. 31:11:22feature which is the one hot encoded
  40737. 31:11:24feature. So we're gonna have to do that
  40738. 31:11:26but we're going to do that as part of
  40739. 31:11:28our pipeline.
  40740. 31:11:30Okay. So we'll be able to include that
  40741. 31:11:32in our pipeline steps uh to to do one
  40742. 31:11:35hot encoding which is nice.
  40743. 31:11:39All right. So we're going to split apart
  40744. 31:11:41our data um as we normally do. So we're
  40745. 31:11:44going to uh create our feature uh data
  40746. 31:11:48frame which is everything but this
  40747. 31:11:50median house value. So we go ahead and
  40748. 31:11:52drop that column and then our target is
  40749. 31:11:54the median house value. So it is just
  40750. 31:11:56that column here. Pretty standard. Um
  40751. 31:12:00and then we're going to train test split
  40752. 31:12:03and um split it into 30%
  40753. 31:12:07uh test data. And again random state you
  40754. 31:12:10can choose whatever you want to be. that
  40755. 31:12:11just affects the shuffling. Um, so
  40756. 31:12:14whatever doesn't really matter what it
  40757. 31:12:16is. It's just so that when you rerun
  40758. 31:12:17this, you get the same result in the in
  40759. 31:12:19the shuffle.
  40760. 31:12:22Okay, so we have our train and our test.
  40761. 31:12:27So you want to make sure you run those.
  40762. 31:12:30All right, so what we're going to do is
  40763. 31:12:32take a look at our data
  40764. 31:12:35and see if we have any null values. Um,
  40765. 31:12:38if you guys remember this data actually
  40766. 31:12:41did have null values. You can see it
  40767. 31:12:42here in this this guy and exactly how
  40768. 31:12:46many there are is from this the sum. So
  40769. 31:12:48we have um 162 nles in in this data. Uh
  40770. 31:12:53and this is just a training data. So of
  40771. 31:12:55course you know the test data could have
  40772. 31:12:57that in there as well. Um so that's
  40773. 31:13:00something we're going to want to make
  40774. 31:13:01sure we fill in the blanks on any data
  40775. 31:13:03set we use whether we're using the
  40776. 31:13:05training or test set. Um, like if we're
  40777. 31:13:08doing training, we want to make sure
  40778. 31:13:09that gets filled in. If we're doing
  40779. 31:13:10predictions with the test set, want to
  40780. 31:13:12make sure that gets filled in. Um, so we
  40781. 31:13:16we should be doing that. Um, now
  40782. 31:13:21what we're going to do is use this data
  40783. 31:13:25to help uh train our pipeline or or use
  40784. 31:13:29with our pipeline. We need to construct
  40785. 31:13:31our pipeline. So far, we've just split
  40786. 31:13:33apart our data. We haven't done anything
  40787. 31:13:36with our processing steps in our model
  40788. 31:13:38yet. Um so roughly
  40789. 31:13:42it this should be the flow of our
  40790. 31:13:43pipeline. What should happen is we
  40791. 31:13:45should be doing some type of feature
  40792. 31:13:47scaling
  40793. 31:13:49um some type of uh feature um
  40794. 31:13:52manipulation. So that could be
  40795. 31:13:53engineering, that could be um that could
  40796. 31:13:57be uh doing the one hot encoding. Um so
  40797. 31:14:01extracting new features like one hot
  40798. 31:14:03encoding,
  40799. 31:14:06one hot encoding. Um we are going to be
  40800. 31:14:09doing that and and by the way this is
  40801. 31:14:11split up into this is when we use our
  40802. 31:14:13pipeline for training.
  40803. 31:14:16Um it's going to look like this where we
  40804. 31:14:17do our scaling, we do one hot encoding.
  40805. 31:14:20Um we have our model here. Um so that
  40806. 31:14:23could be a linear regression, that could
  40807. 31:14:24be a lasso, that could be a ridge, it
  40808. 31:14:26could be elastic net. Whatever model we
  40809. 31:14:28end up using is going to be last in the
  40810. 31:14:30pipeline. And we're going to run this
  40811. 31:14:33pipeline. Ultimately, we're going to run
  40812. 31:14:36pipeline.fit,
  40813. 31:14:40right? We're going to run a fit
  40814. 31:14:41function. and we get a fitted model as
  40815. 31:14:44the result of this pipeline.
  40816. 31:14:47Then when we use it when we use our
  40817. 31:14:50model for prediction,
  40818. 31:14:54we use our model for prediction in this
  40819. 31:14:56lower part, it's the same pipeline, same
  40820. 31:14:59exact pipeline, but it's this model has
  40821. 31:15:02now been trained.
  40822. 31:15:04So we now have a trained model here. So
  40823. 31:15:07the great thing about the pipeline is
  40824. 31:15:08it's the same this is the same pipeline
  40825. 31:15:11that we're using here. So it's just
  40826. 31:15:13going to it's going to repeat those same
  40827. 31:15:15transformations. It's going to do our
  40828. 31:15:17scaling. It's going to do our one hot
  40829. 31:15:19encoding. It's going to use our model
  40830. 31:15:21and it's going to generate predictions
  40831. 31:15:23and generate uh we can we can do
  40832. 31:15:26predictions. We can do evaluation like
  40833. 31:15:28an across validation. Um we can use it
  40834. 31:15:30however we want to use it. Uh but notice
  40835. 31:15:34that the pipeline makes it consistent
  40836. 31:15:37between training and test. We're using
  40837. 31:15:39the exact same transformations
  40838. 31:15:41and the model is last. It's it's either
  40839. 31:15:44being trained or it's being used for
  40840. 31:15:46prediction but it's last. Our
  40841. 31:15:47transformations are up front which are
  40842. 31:15:50things like our scaling, things like our
  40843. 31:15:52one hot encoding, right? Those happen
  40844. 31:15:54first. No matter what data we put
  40845. 31:15:57through there, we put our training data
  40846. 31:15:59through there, we put our test data
  40847. 31:16:00through there, they're going to go
  40848. 31:16:02through the same steps,
  40849. 31:16:05right?
  40850. 31:16:08So that's that's the design of the
  40851. 31:16:09pipeline. That's what it's supposed to
  40852. 31:16:11do. So our job is to create those steps.
  40853. 31:16:16So we need to create those relevant
  40854. 31:16:18steps and then put them together into
  40855. 31:16:20this pipeline. Okay, so that's going to
  40856. 31:16:23be the code we're going to see coming up
  40857. 31:16:24is we're going to build out these steps
  40858. 31:16:27and then put them together into the
  40859. 31:16:29pipeline.
  40860. 31:16:38Um, any questions on this diagram? Does
  40861. 31:16:41it make sense what we're trying to do
  40862. 31:16:42with this pipeline? We want to have
  40863. 31:16:44repeatable steps during the training,
  40864. 31:16:46during a prediction process.
  40865. 31:16:49Okay.
  40866. 31:16:55All right.
  40867. 31:17:01All right. So, um, a couple of things
  40868. 31:17:04we're going to need is, uh, to first of
  40869. 31:17:07all, let's jot down what steps we're
  40870. 31:17:09actually going to do. We're going to
  40871. 31:17:11need to deal with missing values. So,
  40872. 31:17:12we're going to fill in we're going to
  40873. 31:17:13need a pre-process pre-processing step
  40874. 31:17:16that fills in any nles. We always need
  40875. 31:17:19that, right? So, if there's nles, we're
  40876. 31:17:22going to fill them in somehow.
  40877. 31:17:24We're going to define how we do that in
  40878. 31:17:26our in our step. Um, and we also need to
  40879. 31:17:30one hot encode. And we need to scale,
  40880. 31:17:33right? Those are pretty standard steps
  40881. 31:17:36that we've dealt with whenever we're
  40882. 31:17:37building these models, right? So, pretty
  40883. 31:17:40standard things. fill in nles one hot
  40884. 31:17:42encode any categorical data whatever
  40885. 31:17:44however much we have and then go ahead
  40886. 31:17:47and um standardize which is the scaling.
  40887. 31:17:51So this this just is the same word for
  40888. 31:17:53scaling our numeric features. So we're
  40889. 31:17:56going to we're going to define those. Um
  40890. 31:17:59so that's why we're going to go ahead
  40891. 31:18:01and import from pre-processing. We're
  40892. 31:18:03going to import our scaler. Um again we
  40893. 31:18:06could use minmax scaler here. We're
  40894. 31:18:07going to use standard scaler. Um but we
  40895. 31:18:11could use minmax. Um we have our oneh
  40896. 31:18:13hot encoder here. Now usually when we do
  40897. 31:18:17oneh hot encoding we use pd.get dummies.
  40898. 31:18:22This does the same thing as that. But
  40899. 31:18:26because we're going to be building a
  40900. 31:18:27pipeline we actually want the
  40901. 31:18:29scikitlearn version of git dummies. So
  40902. 31:18:32this is the scikitlearn version of git
  40903. 31:18:34dummies here. And it and we have to use
  40904. 31:18:37that version in the pipeline because
  40905. 31:18:39everything in the pipeline needs to be
  40906. 31:18:41an sklearn object. It needs to be an
  40907. 31:18:43sklearn tool or object.
  40908. 31:18:46So um instead of using pandas get
  40909. 31:18:49dummies, we're using one hot encoder
  40910. 31:18:51which is does the same thing. Okay. In
  40911. 31:18:55fact, it just this basically just uses
  40912. 31:18:58pd.get dummies um under the hood.
  40913. 31:19:04Okay. So, it just uses that. Uh,
  40914. 31:19:06anyways, it's just code that builds on
  40915. 31:19:08builds on that.
  40916. 31:19:10Now, what's really nice here is we're
  40917. 31:19:12also going to use from sklearn.impute,
  40918. 31:19:15we're going to use a simple imper. Now,
  40919. 31:19:17what this is is an automated way to fill
  40920. 31:19:20in missing values. So, this is a fancy
  40921. 31:19:22way of basically doing the the fill na
  40922. 31:19:27on a data frame. So, simple imputer um
  40923. 31:19:30we are going to basically fill in the
  40924. 31:19:33blanks. What we're going to do when we
  40925. 31:19:35create this object is give it a strategy
  40926. 31:19:37of how to fill in blanks. Should you use
  40927. 31:19:39the average? Should you use the median?
  40928. 31:19:41Should you use the max? Should use the
  40929. 31:19:43men? Should you use a default value?
  40930. 31:19:45We're going to tell it what to do in
  40931. 31:19:47this object.
  40932. 31:19:50Okay. So, we're going to we're so we're
  40933. 31:19:52going to use this as our automated tool
  40934. 31:19:54for filling in missing values. So,
  40935. 31:19:56that's really nice. it has. So this is
  40936. 31:19:58going to be a critical part of our
  40937. 31:20:00pipeline an imputer that's going to fill
  40938. 31:20:03in missing values.
  40939. 31:20:06So we have that
  40940. 31:20:11yeah coding to reduce coding. Exactly.
  40941. 31:20:14Uh we have our pipeline now. So we have
  40942. 31:20:17our pipeline. So our pipeline is going
  40943. 31:20:18to hold everything. So we need the
  40944. 31:20:20pipeline object um to hold everything
  40945. 31:20:23and that comes from sklearn.pipeline.
  40946. 31:20:25Um, so everything's going to actually go
  40947. 31:20:27into a pipeline object. We're going to
  40948. 31:20:29see how that looks. Um, and finally,
  40949. 31:20:33we're going to from skarn.compose, we're
  40950. 31:20:36going to use a column transformer. The
  40951. 31:20:38reason we're going to do this is because
  40952. 31:20:40we are going to specify for some columns
  40953. 31:20:44like the numerical features, we should
  40954. 31:20:46be scaling.
  40955. 31:20:48For some columns like the categorical
  40956. 31:20:50features, we should be one hot encoding.
  40957. 31:20:53So the column transformer will allow us
  40958. 31:20:55to map different transformations to
  40959. 31:20:58different sections of columns which is
  40960. 31:21:00really useful. So this is actually going
  40961. 31:21:02to be a critical part of our pipeline to
  40962. 31:21:05apply to make sure we only apply this to
  40963. 31:21:07numerical features and only apply this
  40964. 31:21:10to categorical features. Right? So this
  40965. 31:21:14column transformer will help us um to to
  40966. 31:21:18apply pre-processing to particular
  40967. 31:21:20columns. Um like that ocean proximity is
  40968. 31:21:24the only one that really needs this but
  40969. 31:21:26every other column is going to need this
  40970. 31:21:28all the numerical features.
  40971. 31:21:31So we're going to use this column
  40972. 31:21:32transformer and again we're going to see
  40973. 31:21:34how this looks but just trying to give
  40974. 31:21:36you an idea of why we're importing all
  40975. 31:21:37these things.
  40976. 31:21:43Okay, so let's import those.
  40977. 31:21:46Uh, this mentions about the column
  40978. 31:21:48transformer. We just talked about it. It
  40979. 31:21:50allows us to have a particular column or
  40980. 31:21:53group of columns get the right
  40981. 31:21:54transformation. So again, uh, looking
  40982. 31:21:57ahead to our pipeline, the numerical
  40983. 31:22:00features are the ones that are going to
  40984. 31:22:02need scaling, but the categorical
  40985. 31:22:05features are the ones that are going to
  40986. 31:22:06need one hot encoding. However many
  40987. 31:22:07categoricals there are, in this case,
  40988. 31:22:09there's really only one, which is that
  40989. 31:22:10ocean proximity. to go back to our data.
  40990. 31:22:14Um, you can even see that in the info,
  40991. 31:22:16there's just that one. Um, and we see
  40992. 31:22:18that here, right? Just this one string
  40993. 31:22:20column that should be one hot encoded.
  40994. 31:22:22All these other guys should be scaled,
  40995. 31:22:25right? They should all be uh uh standard
  40996. 31:22:27scaled.
  40997. 31:22:29So, this will allow us to specify those
  40998. 31:22:32distinctions.
  40999. 31:22:36All right. So, let's get started
  41000. 31:22:40building our pipeline. So, this is going
  41001. 31:22:42to be really cool. We're going to build
  41002. 31:22:43out the pipeline. Um, let's extract our
  41003. 31:22:48numerical data and our categorical data.
  41004. 31:22:50Now, this is a really neat way of doing
  41005. 31:22:52that that I'm not sure we've seen
  41006. 31:22:53before. Um, so what this does is we'll
  41007. 31:22:58take our data frame, particularly our
  41008. 31:23:00training data frame, and select our
  41009. 31:23:04data.
  41010. 31:23:06That's what this select dtypes does is
  41011. 31:23:08select data from it. Um, which includes
  41012. 31:23:12only the object type columns. So only
  41013. 31:23:15the object types. Now what's that? The
  41014. 31:23:18object type is the string, right? So
  41015. 31:23:21this should select only this column
  41016. 31:23:24because it's in the include.
  41017. 31:23:27We go here include only object types in
  41018. 31:23:30the result. And so this should only have
  41019. 31:23:33our one categorical column which is the
  41020. 31:23:36ocean proximity. So housing cat is going
  41021. 31:23:39to have a reference to our uh it's going
  41022. 31:23:43to be a list that has a a basically just
  41023. 31:23:46our ocean proximity feature because this
  41024. 31:23:50select dtypes will make sure we only
  41025. 31:23:52pick object types and um
  41026. 31:23:56grab those columns. So this is a way to
  41027. 31:24:00neatly grab um our categorical features
  41028. 31:24:04here by including the object types. Now
  41029. 31:24:08on the flip side we can exclude object
  41030. 31:24:10types and get everything else. So this
  41031. 31:24:12is going to be all other columns which
  41032. 31:24:15is excluding the object. So this is
  41033. 31:24:18excluding this meaning we should get all
  41034. 31:24:21of our numerical features that way. So
  41035. 31:24:25this will be all of our numericals
  41036. 31:24:27by excluding the object type and this
  41037. 31:24:31will be our housing num which is short
  41038. 31:24:34for numerical. So this excludes
  41039. 31:24:38the uh object type
  41040. 31:24:42meaning all numerical
  41041. 31:24:46features
  41042. 31:24:49right all numerical features there.
  41043. 31:24:52Okay.
  41044. 31:24:55So, if we were to uh let's double check
  41045. 31:24:58this. Let's sanity check this. If we
  41046. 31:24:59were to print out the housing
  41047. 31:25:03cat, um this should be just the ocean
  41048. 31:25:07proximity feature, which it is. So, just
  41049. 31:25:10that one. If we were to print out the
  41050. 31:25:12housing num, this should be all the
  41051. 31:25:14numerical features, which are all these
  41052. 31:25:17guys. So it's just a reference to those
  41053. 31:25:19columns so that we can uh use those
  41054. 31:25:23later when we're mapping uh this
  41055. 31:25:25transform needs to go to this column
  41056. 31:25:27like the one hot encoding needs to go to
  41057. 31:25:30this column and the scaling needs to go
  41058. 31:25:32to these columns right so we have those
  41059. 31:25:36uh names of those columns already at our
  41060. 31:25:39disposal. So, we're just doing that.
  41061. 31:25:44And this is just a
  41062. 31:25:47simple check.
  41063. 31:25:51Uh, are you guys able to run this?
  41064. 31:25:57If you're following along, let me pause
  41065. 31:25:58there. Make sure I'm not going too fast.
  41066. 31:26:06Uh it so the the issue with a specific
  41067. 31:26:09data type like that is none of these are
  41068. 31:26:11ants. They're actually all floats. So we
  41069. 31:26:14did float. I think that should work. But
  41070. 31:26:16yes, that's the idea.
  41071. 31:26:21Great. I'm glad to hear that right there
  41072. 31:26:23with me. Great. Glad to hear that.
  41073. 31:26:34Okay. So, we have our columns picked out
  41074. 31:26:36here, which we're going to use later.
  41075. 31:26:39Okay.
  41076. 31:26:42All right. So, let's go ahead and build
  41077. 31:26:46out our steps for each of these types.
  41078. 31:26:50So, um for our numerical features, let's
  41079. 31:26:55build out our pipeline steps. So what
  41080. 31:26:57we're going to do is build out a
  41081. 31:26:58numerical pipeline. And it's going to be
  41082. 31:27:01a pipeline with a list
  41083. 31:27:05of tupils. And the reason these are
  41084. 31:27:08tupils is because every tupil has a
  41085. 31:27:11name. So here this is a name that we can
  41086. 31:27:13it can be whatever we want it to be. So
  41087. 31:27:16we're calling it imputer. We could call
  41088. 31:27:18it anything we want. We could call it
  41089. 31:27:19fill in the blanks. We could call it
  41090. 31:27:21null filling. Call it whatever you want.
  41091. 31:27:25We're calling it imputer because it's
  41092. 31:27:26that's a pretty um easy name for it. An
  41093. 31:27:30accurate name to what it's doing. Um but
  41094. 31:27:33the important thing is after the name
  41095. 31:27:36you give it, you put in the scikitlearn
  41096. 31:27:39object that you are going to use to
  41097. 31:27:41operate on your data. So in this case,
  41098. 31:27:44we're using a simple impery
  41099. 31:27:48of median. Now that's a choice. We could
  41100. 31:27:51use a strategy of mean, max.
  41101. 31:27:55Um, we could provide it a constant
  41102. 31:27:58default value. But what this means is we
  41103. 31:28:02are going to fill any blanks we find in
  41104. 31:28:04those columns with the median value of
  41105. 31:28:07that column. That's the strategy for the
  41106. 31:28:09imper. So that's pretty cool. This is
  41107. 31:28:11kind of an automated way to fill in the
  41108. 31:28:13blanks using for any column using its
  41109. 31:28:17median,
  41110. 31:28:19right? And so we could change that. We
  41111. 31:28:21could put mean here or max or min or
  41112. 31:28:23whatever. Um
  41113. 31:28:26but we are filling in the blank on any
  41114. 31:28:28column with its median. And the reason
  41115. 31:28:31this works is because we are going to
  41116. 31:28:33apply this pipeline only to these
  41117. 31:28:35numerical features. So that is fine.
  41118. 31:28:39We're we're not going to apply it to the
  41119. 31:28:41categorical features. We're going to
  41120. 31:28:42apply it to only those numerical. So it
  41121. 31:28:45should have a median value, right? So
  41122. 31:28:48that that's totally fine. So we're going
  41123. 31:28:51to now look at how we're constructing
  41124. 31:28:53the steps. We have a list of tupils.
  41125. 31:28:56Here's one tupil
  41126. 31:28:59which is the imper with a simple imper
  41127. 31:29:02of strategy median. And then we can have
  41128. 31:29:05as many tupils as we want which
  41129. 31:29:07represent processing steps. So every let
  41130. 31:29:10me write that down. Every tupil
  41131. 31:29:14represents
  41132. 31:29:16a pre-processing
  41133. 31:29:18step on our data.
  41134. 31:29:22Okay, so we have an imputer step named
  41135. 31:29:26imputer and the reason it has a name is
  41136. 31:29:29just so you can reference it in the
  41137. 31:29:31pipeline if you need to. So you so it
  41138. 31:29:33has like a a reference name um that you
  41139. 31:29:37give it. Um but this is the more
  41140. 31:29:39important part is the actual scikitlearn
  41141. 31:29:41object that's doing the processing. So
  41142. 31:29:44in this case a simple computer but
  41143. 31:29:46notice that we have a secondary step
  41144. 31:29:48which is our scaling. Now this makes
  41145. 31:29:50sense. This is something we should be
  41146. 31:29:51doing to our features is we should be
  41147. 31:29:54scaling them. So here we we say okay
  41148. 31:29:57let's fill in any blanks first.
  41149. 31:30:00By the way order
  41150. 31:30:03matters.
  41151. 31:30:06So, and what I mean by that is the
  41152. 31:30:10simple imper
  41153. 31:30:13is before the scaler. Now, that's
  41154. 31:30:17important because what that means is we
  41155. 31:30:20should be filling in any blanks before
  41156. 31:30:22we attempt scaling.
  41157. 31:30:25So, that order actually matters. We're
  41158. 31:30:27going to fill in blanks first in this
  41159. 31:30:30list. That's first. We're going to fill
  41160. 31:30:33in blanks. Then we are going to scale
  41161. 31:30:38right then we scale which makes sense
  41162. 31:30:41right so we we fill in blanks first then
  41163. 31:30:44we apply the scaler to scale our
  41164. 31:30:45features so those are our two steps
  41165. 31:30:50so so pretty simple um we are building
  41166. 31:30:54out our two steps now this is just one
  41167. 31:30:57piece of the puzzle we are going to put
  41168. 31:30:59this pipeline together with our one hot
  41169. 31:31:01encoding that's going to be coming up
  41170. 31:31:04next and build out our final pipeline.
  41171. 31:31:07But this is um a a pipeline that has two
  41172. 31:31:10steps that will actually be used with a
  41173. 31:31:12larger pipeline coming up where we we do
  41174. 31:31:15one hot encoding to our categoricals and
  41175. 31:31:17then we put a model in there at the end
  41176. 31:31:20to train and and use for prediction. So
  41177. 31:31:24um pipelines can actually be composed is
  41178. 31:31:28is uh something to realize there is that
  41179. 31:31:30we can have a pipeline that contains a
  41180. 31:31:33few steps. We can have another pipeline
  41181. 31:31:34over here that contains a few steps and
  41182. 31:31:36we can actually um kind of put them
  41183. 31:31:38together into a final pipeline that has
  41184. 31:31:40both pipelines uh kind of merged
  41185. 31:31:43together. Okay. So we're going to see
  41186. 31:31:45that coming up when we construct our
  41187. 31:31:47final one. Our final one, as you can
  41188. 31:31:49imagine, needs to handle this mapping of
  41189. 31:31:52basically saying, let's do one hot
  41190. 31:31:54encoding to these guys and then do this
  41191. 31:31:57pipeline here to these numerical
  41192. 31:32:00features. That's what our final pipeline
  41193. 31:32:03needs to handle, and it will. We're
  41194. 31:32:05going to build that out.
  41195. 31:32:08But let me pause here. Um, were you guys
  41196. 31:32:12able to run this? Are you with me on
  41197. 31:32:15this this pipeline here?
  41198. 31:32:18Does that make sense? Those two steps
  41199. 31:32:20one is filling in blanks with a median
  41200. 31:32:24whatever column. So where so this is
  41201. 31:32:27this is what's so amazing about this is
  41202. 31:32:30this is going to automatically search
  41203. 31:32:32for nles and if you come across a column
  41204. 31:32:36with a null, it's going to use the
  41205. 31:32:39median of that column
  41206. 31:32:42to fill in the blank, right? To fill in
  41207. 31:32:44those nles.
  41208. 31:32:58Okay,
  41209. 31:33:01great. Glad to hear. Glad to hear.
  41210. 31:33:05Okay.
  41211. 31:33:07All right. So we are going to now um put
  41212. 31:33:13this together with a column transformer
  41213. 31:33:17to basically say what steps are going to
  41214. 31:33:20be mapped to what columns.
  41215. 31:33:24Um so now you can see what we're doing
  41216. 31:33:27here is using the column transformer
  41217. 31:33:29which is going to be a list of tupils
  41218. 31:33:31again. So this is another um list of
  41219. 31:33:35tupils.
  41220. 31:33:37But the important thing is um
  41221. 31:33:41each tupil
  41222. 31:33:44has a name
  41223. 31:33:47followed by so it has a name uh which
  41224. 31:33:50again is is generic. You can say
  41225. 31:33:52whatever you want it to be. So here
  41226. 31:33:54we're kind of shortening this to
  41227. 31:33:55numerical. This is short for
  41228. 31:33:56categorical. But the important thing is
  41229. 31:33:59it's followed by a pipeline
  41230. 31:34:03slashstep
  41231. 31:34:06followed by a pipeline slashstep
  41232. 31:34:09um followed by a uh followed by a list
  41233. 31:34:15of columns that it applies to. So you
  41234. 31:34:19can see that pattern here. What we're
  41235. 31:34:22saying is we're going to apply that
  41236. 31:34:24numerical pipeline we just defined. So
  41237. 31:34:26this is saved in a numerical pipeline
  41238. 31:34:29object here. We're going to apply that
  41239. 31:34:32to those numerical features. So this is
  41240. 31:34:34that list
  41241. 31:34:36of numerical features here. So that's
  41242. 31:34:39how we do the mapping. We have a tupil
  41243. 31:34:41here that says okay apply these steps to
  41244. 31:34:44these columns.
  41245. 31:34:46Those go together in that tupole, right?
  41246. 31:34:49Apply these steps to this uh these
  41247. 31:34:52columns. And then apply this step. Now
  41248. 31:34:55what is the step? This is a one hot
  41249. 31:34:58encoder
  41250. 31:34:59which is going to uh uh encode um those
  41251. 31:35:04features and it's going to uh ignore um
  41252. 31:35:09basically nles for now. That's a choice
  41253. 31:35:11but it's going to ignore um uh basically
  41254. 31:35:16ignore nles and and uh skip over them
  41255. 31:35:19for now. We now we know there's no NLES
  41256. 31:35:23because we already did an is NA from
  41257. 31:35:25before and we know there's not any NLES
  41258. 31:35:28in that ocean proximity. So this isn't
  41259. 31:35:30going to be an issue. But that's what
  41260. 31:35:32that would do.
  41261. 31:35:34But we have a one hot encoder here which
  41262. 31:35:37we're going to apply to our categorical
  41263. 31:35:40features. Now of course that's just the
  41264. 31:35:43ocean proximity feature but that but
  41265. 31:35:46again you see the pattern in the tupil
  41266. 31:35:47is apply this transform which is a one
  41267. 31:35:50hot encoding to this column apply these
  41268. 31:35:53numerical transforms which is a whole
  41269. 31:35:55pipeline. So it's two steps in a
  41270. 31:35:59pipeline of um
  41271. 31:36:02uh an imputer and a scaler are going to
  41272. 31:36:05be applied to this
  41273. 31:36:08really nice. So those are going to be
  41274. 31:36:10all together in this column transformer.
  41275. 31:36:12And that is our way to signal that for
  41276. 31:36:14these numerical features, use these
  41277. 31:36:16steps. For our categorical features, use
  41278. 31:36:18this step. And and you know, if we had
  41279. 31:36:21more than one step, we were applying to
  41280. 31:36:23categorical. We could build a pipeline
  41281. 31:36:25for the categorical and it would and do
  41282. 31:36:27the same thing. We have more than one
  41283. 31:36:29step here. And so it's good practice
  41284. 31:36:32when you have more than one step to just
  41285. 31:36:33put that in a pipeline because we have
  41286. 31:36:35more than one step. We'll just put that
  41287. 31:36:37in this list inside of the pipeline and
  41288. 31:36:39we can map that pipeline to those
  41289. 31:36:42features. Here we only have one step. So
  41290. 31:36:45it's okay to just put that there um and
  41291. 31:36:48apply that to the categorical features.
  41292. 31:36:51But if we had more than one step um it
  41293. 31:36:54would be good practice to put that in a
  41294. 31:36:56pipeline
  41295. 31:36:57which is what we do here. Right? This
  41296. 31:36:58pipeline is being mapped to these
  41297. 31:37:00features. This step is being applied to
  41298. 31:37:03this feature.
  41299. 31:37:08Okay,
  41300. 31:37:10how about that? Are you guys able to run
  41301. 31:37:13that one? Does that make sense what we
  41302. 31:37:15have set up so far? So, we're almost
  41303. 31:37:17there. We almost have our final
  41304. 31:37:19pipeline. We have our pre-processing
  41305. 31:37:20basically done to say our numerical
  41306. 31:37:23features should be processed with that
  41307. 31:37:24other pipeline and our categorical
  41308. 31:37:27features should be one hot encoded.
  41309. 31:37:29We're getting close. The only thing
  41310. 31:37:31we're really missing here is a model.
  41311. 31:37:34The only thing we're really missing is
  41312. 31:37:36to have our final model training
  41313. 31:37:39pipeline is to actually include a model
  41314. 31:37:42which should come at the end.
  41315. 31:37:44Right? So it should we should be doing
  41316. 31:37:46these steps first
  41317. 31:37:49then doing modeling which we know right
  41318. 31:37:52we we've done that uh many times. We've
  41319. 31:37:54done our pre-processing and then we do
  41320. 31:37:55our modeling.
  41321. 31:38:00Any questions on that?
  41322. 31:38:17Okay.
  41323. 31:38:18Fantastic.
  41324. 31:38:22All right.
  41325. 31:38:25So, if we wanted to uh see if we wanted
  41326. 31:38:28to test this so far, um we could. So, we
  41327. 31:38:31could run the pre-processing and
  41328. 31:38:33actually run a fit transform on our data
  41329. 31:38:36and this will um basically apply that
  41330. 31:38:39pipeline to the data. Now, this would be
  41331. 31:38:41a sanity check. This is a good This is a
  41332. 31:38:44good kind of um This is a good sanity
  41333. 31:38:47check that our pre-processing
  41334. 31:38:52works. So, it's doing what we expected
  41335. 31:38:55to do. It's not our final pipeline
  41336. 31:38:57because we don't have our model in there
  41337. 31:38:59yet, but this is just to ensure that all
  41338. 31:39:01of the features are kind of behaving as
  41339. 31:39:03we expect. So, we can uh we can do that
  41340. 31:39:07and we can take a look at the um
  41341. 31:39:09results. This looks pretty good. this
  41342. 31:39:11all of our numerical features ended up
  41343. 31:39:13scaled
  41344. 31:39:15which is pretty good and we have one hot
  41345. 31:39:17encoded features for that ocean
  41346. 31:39:19proximity over here.
  41347. 31:39:22Okay, so this looks pretty this looks
  41348. 31:39:24reasonable of those steps being applied
  41349. 31:39:27to the right columns. But this is a good
  41350. 31:39:29kind of sanity check to just run our fit
  41351. 31:39:31transform on our data to ensure those
  41352. 31:39:35steps are actually happening and they
  41353. 31:39:37are. You can see here the result of the
  41354. 31:39:40scaling and the uh the one hot encoding.
  41355. 31:39:44So that that all looks pretty
  41356. 31:39:46reasonable,
  41357. 31:39:49right? And uh what we should also do is
  41358. 31:39:55make sure there are no nulls in this
  41359. 31:39:56which there shouldn't be because we did
  41360. 31:39:58the imputer. So we should be doing uh is
  41361. 31:40:01na dot
  41362. 31:40:05sum
  41363. 31:40:09And there is no NLES anymore. So that
  41364. 31:40:11looks pretty good, right? Those got
  41365. 31:40:13filled in uh by doing our steps. Our
  41366. 31:40:17pipeline steps executed really nicely on
  41367. 31:40:20our training data. Um and and we were
  41368. 31:40:24off and running. And there's nothing
  41369. 31:40:25unique about the training data. We could
  41370. 31:40:27do this to our test data as well
  41371. 31:40:30and verify that those steps are running
  41372. 31:40:32and they would, right? There's nothing
  41373. 31:40:34really that special about running it on
  41374. 31:40:35the training data. Um, it should also
  41375. 31:40:39work on the test features as well, and
  41376. 31:40:40it does. You can check that for
  41377. 31:40:43yourself.
  41378. 31:40:45Okay.
  41379. 31:40:50All right. So, that's pretty cool. We
  41380. 31:40:52can uh verify all that's working.
  41381. 31:41:01Any questions on that?
  41382. 31:41:04We're almost there with our full
  41383. 31:41:06pipeline. This this is this is not the
  41384. 31:41:08full pipeline, but this is something
  41385. 31:41:10that will run during our full pipeline.
  41386. 31:41:13Of course, our features are going to be
  41387. 31:41:14transformed according to those steps and
  41388. 31:41:16then it will be uh put into our model to
  41389. 31:41:19either predict or train with. Um
  41390. 31:41:22so let's do that. Let's actually build
  41391. 31:41:24out our final uh model here. So it's
  41392. 31:41:29actually going to be really easy to do.
  41393. 31:41:30All we need to do is um put in our
  41394. 31:41:34model. So here we're going to import the
  41395. 31:41:36ridge model here. Now we could use any
  41396. 31:41:39we could use linear regression, we could
  41397. 31:41:41use lasso, we could use elastic net. Um
  41398. 31:41:43we're just going to use ridge um uh um
  41399. 31:41:47just to test it out. And um we are going
  41400. 31:41:51to uh now put in a final pipeline. So
  41401. 31:41:55we're going to use our pipeline. And so
  41402. 31:41:58we're going to create a new one here.
  41403. 31:42:00and map our pre-processing to our
  41404. 31:42:03pre-processing that we've already built.
  41405. 31:42:05So, this is a column transformer that
  41406. 31:42:08already has all of our steps. And then
  41407. 31:42:10notice what comes after it is just the
  41408. 31:42:12model. Now, that's pretty pretty basic,
  41409. 31:42:14but it makes sense that it should come
  41410. 31:42:16after that model. Um, and of course,
  41411. 31:42:19this is a generic name. We could we can
  41412. 31:42:21name it whatever we want to. um model
  41413. 31:42:24ridge is pretty reasonable um to because
  41414. 31:42:28it is a ridge uh regression but uh of
  41415. 31:42:31course we could we could change that.
  41416. 31:42:35Okay, so that builds out our uh final um
  41417. 31:42:38pipeline. So now we have a pipeline and
  41418. 31:42:42what's great about that is this signals
  41419. 31:42:45that all of these steps should be
  41420. 31:42:47completed prior to doing anything with
  41421. 31:42:49this model. So all of those processing
  41422. 31:42:52steps are going to run and then we're
  41423. 31:42:54going to do ffit orpredict and that so
  41424. 31:42:57that's really great. It ensures that all
  41425. 31:42:59those steps are running together every
  41426. 31:43:01single time we call predict with this
  41427. 31:43:04with this model. So we're just going to
  41428. 31:43:06use the pipeline in place of the model
  41429. 31:43:10to ensure that all of those steps are
  41430. 31:43:12running together. And this is our this
  41431. 31:43:15is kind of our final pipeline that we
  41432. 31:43:17would use uh with like something like
  41433. 31:43:19ffit or predict.
  41434. 31:43:22So let me make that uh a note of that.
  41435. 31:43:24Now we can use this final pipeline just
  41436. 31:43:30like a regular model i.e. pipeline.fit
  41437. 31:43:36or pipeline.predict.
  41438. 31:43:40We could use it in ei in either fashion
  41439. 31:43:43uh to to train the pipeline would be
  41440. 31:43:47this guy and then use the pipeline to
  41441. 31:43:48predict would be this. And what we
  41442. 31:43:50should realize is under the hood these
  41443. 31:43:52steps are running first and then we
  41444. 31:43:54train it or these steps run first then
  41445. 31:43:57we use it for prediction.
  41446. 31:44:04Okay.
  41447. 31:44:06Questions on that? Does that make sense?
  41448. 31:44:09On this final pipeline here, it's just
  41449. 31:44:12now it it's really cool because we have
  41450. 31:44:14a pipeline
  41451. 31:44:16made up of a of a pipeline really,
  41452. 31:44:19right? A pipeline made up of a pipeline.
  41453. 31:44:21But that's scikitlearn allows you to do
  41454. 31:44:22that to compose pipelines in this way.
  41455. 31:44:26That's that's pretty uh pretty uh normal
  41456. 31:44:29there.
  41457. 31:44:39Okay,
  41458. 31:44:41what I want to show you is we can
  41459. 31:44:44actually use this pipeline in a grid
  41460. 31:44:46search. So that's pretty amazing. We can
  41461. 31:44:48use this pipeline in any way we can use
  41462. 31:44:51a mo like a regular model. It's just
  41463. 31:44:53that now our pre-processing steps have
  41464. 31:44:56kind of been packaged together with our
  41465. 31:44:58model to ensure that they always run
  41466. 31:45:01anytime we do any processing with this
  41467. 31:45:03model. Um so for instance we can do a
  41468. 31:45:07grid search just like we did with a
  41469. 31:45:09regular with with just a model right
  41470. 31:45:11with just this. Um we can do the same
  41471. 31:45:14thing with the whole pipeline. Um, so
  41472. 31:45:17the only catch is that you want to make
  41473. 31:45:20sure in your grid you name things in the
  41474. 31:45:24appropriate way inside of your your uh
  41475. 31:45:26keys in your dictionary. So uh for
  41476. 31:45:30instance um inside of the grid uh we're
  41477. 31:45:34going to set up the alpha that would be
  41478. 31:45:36used with this ridge regression by
  41479. 31:45:39referencing its name. So this is model
  41480. 31:45:41ridge is this is the name of the model
  41481. 31:45:44inside of the pipeline. So you want to
  41482. 31:45:46make sure that goes first.
  41483. 31:45:48And then what scikitlearn does is it
  41484. 31:45:51recognizes parameters that belong with
  41485. 31:45:53this model by using a double underscore.
  41486. 31:45:57So the so you have underscore underscore
  41487. 31:46:00alpha um here. So the double
  41488. 31:46:05uh underscore
  41489. 31:46:08signals a parameter
  41490. 31:46:11belonging to model ridge. in the in the
  41491. 31:46:17pipeline.
  41492. 31:46:19Okay, so we have a model ridge is just a
  41493. 31:46:22reference to the model in our pipeline.
  41494. 31:46:24That's the one we're going to test out
  41495. 31:46:26these parameters with. And
  41496. 31:46:28underscore_pha is just a way to say this
  41497. 31:46:31alpha belongs to this model. Okay, it
  41498. 31:46:36belongs so it's going to be used with
  41499. 31:46:38that model in our pipeline. Um otherwise
  41500. 31:46:42it's going to work exactly the same way.
  41501. 31:46:44It's just we need to line up this naming
  41502. 31:46:46convention of of scikitlearn.
  41503. 31:46:48You just have to reference this to
  41504. 31:46:51whatever name you provided here and then
  41505. 31:46:54underscore parameter. So L1 ratio alpha
  41506. 31:46:58whatever right would go there.
  41507. 31:47:01Okay. So there is a range from 0.1 to
  41508. 31:47:05two uh step size of 0.1
  41509. 31:47:09um and then we do our grid search CV. So
  41510. 31:47:12this is exactly the same setup as we had
  41511. 31:47:14before. It's just that our model is now
  41512. 31:47:18the pipeline. So our pipeline is going
  41513. 31:47:20in there. Um we have our grid going in
  41514. 31:47:23there. We have our scoring is the same,
  41515. 31:47:26you know, negative absolute error. Um
  41516. 31:47:28we're using five-fold cross validation
  41517. 31:47:31and we're parallelizing that search. Um,
  41518. 31:47:34so we're going to search through these
  41519. 31:47:35alphas and uh basically fit this to our
  41520. 31:47:40um data and find the best um find the
  41521. 31:47:45best alpha.
  41522. 31:47:47So it's going to try out all those
  41523. 31:47:49combinations and try to come up with the
  41524. 31:47:51best alpha.
  41525. 31:47:54So looks like the best alpha was 0.1 for
  41526. 31:47:57the ridge.
  41527. 31:48:00Okay, is the best. So then um if we
  41528. 31:48:04wanted to we could uh then predict using
  41529. 31:48:08the model um which would be doing
  41530. 31:48:11something like this. Um and we could
  41531. 31:48:14also go back and do something like so we
  41532. 31:48:18could
  41533. 31:48:22now use um this param. So we could do
  41534. 31:48:27model
  41535. 31:48:30um equals ridge
  41536. 31:48:35and then we could put in our alpha
  41537. 31:48:38um alpha is our results our best
  41538. 31:48:41parameters and then we get that model
  41539. 31:48:43ridge alpha and then we just rebuild our
  41540. 31:48:46our pipeline
  41541. 31:48:50equals um pipeline and then we uh put in
  41542. 31:48:54this new model here. So we could do
  41543. 31:48:57this. This would be going back and just
  41544. 31:49:00um putting in our best alpha here for
  41545. 31:49:04this model and then uh ensuring that's
  41546. 31:49:07part of our our pipeline. So we're just
  41547. 31:49:09overwriting that pipeline with the best
  41548. 31:49:10model there
  41549. 31:49:16to get the best model in our pipeline.
  41550. 31:49:23Okay,
  41551. 31:49:24so that's all this is doing is just
  41552. 31:49:26initializing a new um let me actually I
  41553. 31:49:30can put this code in here.
  41554. 31:49:36This is actually just getting this is
  41555. 31:49:38just getting a model with the best alpha
  41556. 31:49:40and then reinserting that into our our
  41557. 31:49:43uh we're just overwriting our final
  41558. 31:49:45pipeline there with the best model that
  41559. 31:49:47we have.
  41560. 31:49:50So pretty cool that pipeline can be used
  41561. 31:49:53basically exactly like a model, right?
  41562. 31:49:55It's it's going right here in the grid
  41563. 31:49:57search and being used uh entirely like a
  41564. 31:50:00basic model. So we do ffit
  41565. 31:50:04um and that allows us to use it. We
  41566. 31:50:06could dopredict. We could even do
  41567. 31:50:08pipeline.predict once we we could go
  41568. 31:50:10back and do final pipeline.fit
  41569. 31:50:13um with this and then final
  41570. 31:50:14pipeline.predict with this and evaluate
  41571. 31:50:21Okay,
  41572. 31:50:23so pretty cool that pipeline can be used
  41573. 31:50:25uh basically exactly like how a model
  41574. 31:50:27would be any way we' use a model.fit
  41575. 31:50:30model.predict, we can use a pipeline.
  41576. 31:50:34So grid search is for instance something
  41577. 31:50:37that can use a model in there. Um but
  41578. 31:50:39instead of just a model, we're ensuring
  41579. 31:50:41we have our pre-processing steps kind of
  41580. 31:50:43bundled with that model in this
  41581. 31:50:45pipeline.
  41582. 31:50:47Any
  41583. 31:50:50questions on
  41584. 31:50:52uh this example so far?
  41585. 31:50:58Were you guys able to run it up to here?
  41586. 31:51:00Were you able to run the grid search?
  41587. 31:51:16Okay, great.
  41588. 31:51:25Okay.
  41589. 31:51:27Okay. So, this is this is uh just
  41590. 31:51:30showing you what's actually happening
  41591. 31:51:31underneath the hood is uh you know,
  41592. 31:51:34we're doing some scaling. We're doing
  41593. 31:51:36some one hot encoding um
  41594. 31:51:40and we're doing some uh we're doing a
  41595. 31:51:43model here. And that's all part of our
  41596. 31:51:46pipeline. Um, and then we can use the
  41597. 31:51:50pipeline however we want. So for
  41598. 31:51:52example, I know it's not here, but for
  41599. 31:51:54an example, we could use um once we do
  41600. 31:51:57once we have this final pipeline um we
  41601. 31:52:00can can use the final um
  41602. 31:52:05pipeline to predict. So we can do um
  41603. 31:52:09predictions
  41604. 31:52:11equals final
  41605. 31:52:14pipeline.predict
  41606. 31:52:16and then we can pass in our test data.
  41607. 31:52:18Now what happens on this is once we have
  41608. 31:52:22ran our our pipeline.fit we have a
  41609. 31:52:25trained pipeline and then when we run
  41610. 31:52:27this final pipeline.predict uh this data
  41611. 31:52:30is going to be transformed.
  41612. 31:52:32It's going to go through those
  41613. 31:52:33transformation steps and then we would
  41614. 31:52:35apply our model to it at the end uh to
  41615. 31:52:39to make those predictions and then we
  41616. 31:52:40can evaluate those predictions which is
  41617. 31:52:42what we're doing kind of here.
  41618. 31:52:46Right?
  41619. 31:52:52Okay.
  41620. 31:52:56All right. So in conclusion uh we have
  41621. 31:52:59gone through a lot of stuff here. Um,
  41622. 31:53:02we've gone through regression, we've
  41623. 31:53:04done the regularization on regression.
  41624. 31:53:08So hopefully we have a good foundation
  41625. 31:53:09on regression. Um, what we're going to
  41626. 31:53:11do in a little bit is actually do some
  41627. 31:53:13additional practice with regression on a
  41628. 31:53:15new problem. We're going to do a
  41629. 31:53:16capstone problem and do some additional
  41630. 31:53:20regression work with that. Um, so we'll
  41631. 31:53:23do that next. Um but the other thing we
  41632. 31:53:27learned is how to evaluate the
  41633. 31:53:28regression using things like mean
  41634. 31:53:30squared error, RMSSE which is square
  41635. 31:53:32root of that. Um which is which is
  41636. 31:53:35really cool. So we have a sense of that
  41637. 31:53:38error which is our distance from our
  41638. 31:53:40prediction to the actual value. That's
  41639. 31:53:42always what these uh that's always what
  41640. 31:53:45these things are doing like this, right?
  41641. 31:53:48This mean absolute error metric from
  41642. 31:53:50scikitlearn is computing the average
  41643. 31:53:53distance from these predictions to these
  41644. 31:53:55test labels that we have, right? Those
  41645. 31:53:58actual values. Um, and that gives us a
  41646. 31:54:00sense of on average how far away are our
  41647. 31:54:03predictions
  41648. 31:54:04um to see how good of a model that we
  41649. 31:54:07have, right? And we should be evaluating
  41650. 31:54:10that error generally
  41651. 31:54:12um against the scale of our targets to
  41652. 31:54:16see, you know,
  41653. 31:54:19uh how far off we typically are.
  41654. 31:54:23Okay. Any questions at all on this
  41655. 31:54:25lesson on regression? Uh anything we
  41656. 31:54:27covered up to this point? We're going to
  41657. 31:54:30do some more practice with the next
  41658. 31:54:32we'll do the capstone. So we get so we
  41659. 31:54:35just do some more regression problems.
  41660. 31:54:47Yeah, it's a that's another bad score.
  41661. 31:54:49It's a little bit hard to interpret this
  41662. 31:54:50though because it's m ae. Um, so one
  41663. 31:54:54thing we could do is is compute mean
  41664. 31:54:57squared error and then take the square
  41665. 31:55:00root of it to get the RMSSE which is a
  41666. 31:55:02much better uh evaluation metric in
  41667. 31:55:05terms of our target. Um so we could
  41668. 31:55:08actually run that. Uh if we go back here
  41669. 31:55:11and um we could generate for instance we
  41670. 31:55:14could generate the MSE which is the mean
  41671. 31:55:17squared
  41672. 31:55:20error
  41673. 31:55:21and it's it's the same exact function uh
  41674. 31:55:25of using our predictions.
  41675. 31:55:28Um
  41676. 31:55:30and then we could just print that out.
  41677. 31:55:32Mean squared error.
  41678. 31:55:36So we have mean squared error and then
  41679. 31:55:38what we can do is let's take the um MP.
  41680. 31:55:43square root of that.
  41681. 31:55:46So that way we can generate the RMSSE.
  41682. 31:55:49So yeah that I mean that's pretty bad.
  41683. 31:55:50That's uh pretty bad. Uh now let's let's
  41684. 31:55:55go back and look at our
  41685. 31:55:58uh data though. So let's take a look at
  41686. 31:56:00the average for our y. Um remember one
  41687. 31:56:03thing we should be doing is taking a
  41688. 31:56:05look at um what our uh let's take a look
  41689. 31:56:08at y test mean
  41690. 31:56:11to get an average value. So the average
  41691. 31:56:14value is in the 200,000s. So
  41692. 31:56:18this isn't this isn't awful. This is
  41693. 31:56:2170,000. It's still a decent amount of
  41694. 31:56:23error. It's not as bad as the models we
  41695. 31:56:25have before though, right? This is an
  41696. 31:56:28average median price of the house is in
  41697. 31:56:32the 206,000 range and our error is off
  41698. 31:56:36by like 70,000,
  41699. 31:56:39right?
  41700. 31:56:41So, it's not good. Um, but it's not
  41701. 31:56:47hor like as bad as the it's not as
  41702. 31:56:50horrible as we've seen so far. Right.
  41703. 31:56:52This is a little bit better of a model.
  41704. 31:56:54A little bit better. closer to zero
  41705. 31:56:56would be better, right? Um but the
  41706. 31:56:59smaller the better. But uh remember this
  41707. 31:57:02is the um these even the mean absolute
  41708. 31:57:06error is is technically in similar units
  41709. 31:57:09as the as the uh um
  41710. 31:57:14as the target. So 50,000 60,000 here
  41711. 31:57:1770,000 it's still a decent amount of
  41712. 31:57:19error in terms of 200,000.
  41713. 31:57:24Uh so far we only come up with models
  41714. 31:57:25and test their accuracy with available
  41715. 31:57:27data. We haven't used a model to make
  41716. 31:57:28completely new predictions on No, we
  41717. 31:57:31haven't done that. Uh except we know how
  41718. 31:57:33to do that. Um it would so to make
  41719. 31:57:36predictions on new data would be exactly
  41720. 31:57:38how we're making them on our available
  41721. 31:57:40data because we actually do that all the
  41722. 31:57:43time. If we go back down to our model
  41723. 31:57:45building,
  41724. 31:57:47um it's it looks just like this, right?
  41725. 31:57:49where we take so for instance we do
  41726. 31:57:53predictions all the time on test data
  41727. 31:57:56that was never involved in the training.
  41728. 31:57:58So it's it's as if this data mimics new
  41729. 31:58:02data that we've never seen before. So if
  41730. 31:58:06we had new raw data it would just it
  41731. 31:58:08would be the same exact process. the new
  41732. 31:58:11now with our pipeline it makes it a
  41733. 31:58:14little bit easier because with the
  41734. 31:58:15pipeline
  41735. 31:58:17um the raw data will go through those
  41736. 31:58:19transformations which it should right
  41737. 31:58:21the raw data should because if it's
  41738. 31:58:22missing data it needs to be filled in if
  41739. 31:58:24it has categoricals it needs to be one
  41740. 31:58:26hot encoded so that's the purpose of the
  41741. 31:58:29pipeline actually is to make sure that
  41742. 31:58:33if we're dealing with raw data um those
  41743. 31:58:37steps can happen on the data before it
  41744. 31:58:40goes into the model. Right?
  41745. 31:58:43So we so
  41746. 31:58:45that's kind of the purpose of the
  41747. 31:58:47pipeline
  41748. 31:58:48is to ensure that we run those steps
  41749. 31:58:51ahead of using it using a model with it.
  41750. 31:58:56But but ultimately that's how it uh any
  41751. 31:58:58scikitlearn model is going to be doing
  41752. 31:59:00the predict even if it's a pipeline
  41753. 31:59:02right it's going to be uh we just go
  41754. 31:59:05back down here it's going to be um
  41755. 31:59:08predict it's always going to be that on
  41756. 31:59:10new data
  41757. 31:59:11>> hello everyone in this session we will
  41758. 31:59:14cover all the important supervised and
  41759. 31:59:16unsupervised learning algorithms that
  41760. 31:59:17are widely used with hands-on
  41761. 31:59:18demonstrations in Python our instructors
  41762. 31:59:21with rich experience in machine learning
  41763. 31:59:23will take us through this course but
  41764. 31:59:25before Before we begin, make sure to
  41765. 31:59:26subscribe to the SimplyLearn channel and
  41766. 31:59:28hit the bell icon to never miss an
  41767. 31:59:29update.
  41768. 31:59:31So, we will start by understanding the
  41769. 31:59:33basics of machine learning from a short
  41770. 31:59:35animated video followed by the
  41771. 31:59:37difference between supervised and
  41772. 31:59:38unsupervised learning. We will then jump
  41773. 31:59:40into learning the various algorithms
  41774. 31:59:43from scratch. So, we will understand
  41775. 31:59:45linear regression, logistic regression,
  41776. 31:59:47decision tree, and random forest. We'll
  41777. 31:59:50then look at support vector machines and
  41778. 31:59:52K nearest neighbor algorithm with a
  41779. 31:59:53hands-on demonstration in Python.
  41780. 31:59:56Finally, we'll get an idea about
  41781. 31:59:58unsupervised learning algorithms such as
  41782. 32:00:00K means clustering and principal
  41783. 32:00:01component analysis. We will conclude
  41784. 32:00:04this session with regularization in
  41785. 32:00:06machine learning. So let's get started.
  41786. 32:00:09>> We know humans learn from their past
  41787. 32:00:11experiences and machines follow
  41788. 32:00:13instructions given by humans.
  41789. 32:00:16But what if humans can train the
  41790. 32:00:18machines to learn from their past data
  41791. 32:00:20and do what humans can do and much
  41792. 32:00:22faster? Well, that's called machine
  41793. 32:00:24learning. But it's a lot more than just
  41794. 32:00:26learning. It's also about understanding
  41795. 32:00:28and reasoning. So today we will learn
  41796. 32:00:30about the basics of machine learning. So
  41797. 32:00:33that's Paul. He loves listening to new
  41798. 32:00:36songs.
  41799. 32:00:38He either likes them or dislikes them.
  41800. 32:00:40Paul decides this on the basis of the
  41801. 32:00:42song's tempo, genre, intensity, and the
  41802. 32:00:46gender of voice. For simplicity, let's
  41803. 32:00:49just use tempo and intensity for now.
  41804. 32:00:52So, here tempo is on the x-axis, ranging
  41805. 32:00:55from relaxed to fast, whereas intensity
  41806. 32:00:58is on the y-axis, ranging from light to
  41807. 32:01:01soaring. We see that Paul likes the song
  41808. 32:01:04with fast tempo and soaring intensity
  41809. 32:01:07while he dislikes the song with relaxed
  41810. 32:01:10tempo and light intensity. So now we
  41811. 32:01:13know Paul's choices. Let's say Paul
  41812. 32:01:15listens to a new song. Let's name it as
  41813. 32:01:17song A. Song A has fast tempo and a
  41814. 32:01:20soaring intensity. So it lies somewhere
  41815. 32:01:23here. Looking at the data, can you guess
  41816. 32:01:25whether Paul will like the song or not?
  41817. 32:01:27Correct. So Paul likes this song. By
  41818. 32:01:30looking at Paul's past choices, we were
  41819. 32:01:32able to classify the unknown song very
  41820. 32:01:35easily, right? Let's say now Paul
  41821. 32:01:37listens to a new song. Let's label it as
  41822. 32:01:40song B. So song B lies somewhere here
  41823. 32:01:44with medium tempo and medium intensity.
  41824. 32:01:47Neither relaxed nor fast, neither light
  41825. 32:01:50nor soaring. Now, can you guess whether
  41826. 32:01:52Paul likes it or not? Not able to guess
  41827. 32:01:54whether Paul will like it or dislike it.
  41828. 32:01:56Are the choices unclear? Correct. We
  41829. 32:01:59could easily classify song A. But when
  41830. 32:02:02the choice became complicated as in the
  41831. 32:02:04case of song B. Yes. And that's where
  41832. 32:02:06machine learning comes in. Let's see
  41833. 32:02:08how. In the same example for song B, if
  41834. 32:02:11we draw a circle around the song B, we
  41835. 32:02:13see that there are four votes for like
  41836. 32:02:15whereas one vote for dislike. If we go
  41837. 32:02:18for the majority votes, we can say that
  41838. 32:02:20Paul will definitely like the song.
  41839. 32:02:22That's all. This was a basic machine
  41840. 32:02:24learning algorithm also. It's called K
  41841. 32:02:26nearest neighbors. So this is just a
  41842. 32:02:28small example in one of the many machine
  41843. 32:02:30learning algorithms quite easy right
  41844. 32:02:33believe me it is but what happens when
  41845. 32:02:36the choices become complicated as in the
  41846. 32:02:39case of song B that's when machine
  41847. 32:02:40learning comes in it learns the data
  41848. 32:02:43builds the prediction model and when the
  41849. 32:02:45new data point comes in it can easily
  41850. 32:02:47predict for it more the data better the
  41851. 32:02:49model higher will be the accuracy there
  41852. 32:02:52are many ways in which the machine
  41853. 32:02:54learns it could be either supervised
  41854. 32:02:56learning unsupervised learning or
  41855. 32:02:58reinforcement learning. Let's first
  41856. 32:03:00quickly understand supervised learning.
  41857. 32:03:02Suppose your friend gives you 1 million
  41858. 32:03:05coins of three different currencies. Say
  41859. 32:03:071 rupee, 1 and 1 dirham. Each coin has
  41860. 32:03:10different weights. For example, a coin
  41861. 32:03:12of 1 rupee weighs 3 g. 1 euro weighs 7 g
  41862. 32:03:16and 1 dirham weighs 4 g. Your model will
  41863. 32:03:19predict the currency of the coin. Here
  41864. 32:03:21your weight becomes the feature of coins
  41865. 32:03:24while currency becomes their label. When
  41866. 32:03:26you feed this data to the machine
  41867. 32:03:28learning model, it learns which feature
  41868. 32:03:30is associated with which label. For
  41869. 32:03:33example, it will learn that if a coin is
  41870. 32:03:35of 3 g, it will be a 1 rupee coin. Let's
  41871. 32:03:38give a new coin to the machine. On the
  41872. 32:03:40basis of the weight of the new coin,
  41873. 32:03:42your model will predict the currency.
  41874. 32:03:44Hence, supervised learning uses labeled
  41875. 32:03:46data to train the model. Here, the
  41876. 32:03:48machine knew the features of the object
  41877. 32:03:50and also the labels associated with
  41878. 32:03:53those features. On this note, let's move
  41879. 32:03:55to unsupervised learning and see the
  41880. 32:03:57difference. Suppose you have cricket
  41881. 32:03:58data set of various players with their
  41882. 32:04:00respective scores and wickets taken.
  41883. 32:04:02When we feed this data set to the
  41884. 32:04:05machine, the machine identifies the
  41885. 32:04:07pattern of player performance. So, it
  41886. 32:04:09plots this data with the respective
  41887. 32:04:10wickets on the x-axis while runs on the
  41888. 32:04:13y-axis. While looking at the data,
  41889. 32:04:14you'll clearly see that there are two
  41890. 32:04:16clusters. The one cluster are the
  41891. 32:04:18players who scored high runs and took
  41892. 32:04:21less wickets while the other cluster is
  41893. 32:04:23of the players who scored less runs but
  41894. 32:04:26took many wickets. So here we interpret
  41895. 32:04:28these two clusters as batsmen and
  41896. 32:04:30bowlers. The important point to note
  41897. 32:04:32here is that there were no labels of
  41898. 32:04:34batsmen and bowlers. Hence the learning
  41899. 32:04:37with unlabeled data is unsupervised
  41900. 32:04:39learning. So we saw supervised learning
  41901. 32:04:41where the data was labeled and the
  41902. 32:04:43unsupervised learning where the data was
  41903. 32:04:45unlabeled. And then there is
  41904. 32:04:47reinforcement learning which is a
  41905. 32:04:48reward-based learning or we can say that
  41906. 32:04:50it works on the principle of feedback.
  41907. 32:04:52Here let's say you provide the system
  41908. 32:04:54with an image of a dog and ask it to
  41909. 32:04:56identify it. The system identifies it as
  41910. 32:04:59a cat. So you give a negative feedback
  41911. 32:05:01to the machine saying that it's a dog's
  41912. 32:05:03image. The machine will learn from the
  41913. 32:05:04feedback and finally if it comes across
  41914. 32:05:06any other image of a dog, it'll be able
  41915. 32:05:09to classify it correctly. That is
  41916. 32:05:11reinforcement learning. To generalize
  41917. 32:05:13machine learning model, let's see a
  41918. 32:05:14flowchart. Input is given to a machine
  41919. 32:05:16learning model which then gives the
  41920. 32:05:18output according to the algorithm
  41921. 32:05:20applied. If it's right, we take the
  41922. 32:05:22output as a final result. Else we
  41923. 32:05:24provide feedback to the training model
  41924. 32:05:26and ask it to predict until it learns. I
  41925. 32:05:29hope you've understood supervised and
  41926. 32:05:31unsupervised learning. So let's have a
  41927. 32:05:33quick quiz. You have to determine
  41928. 32:05:35whether the given scenarios uses
  41929. 32:05:37supervised or unsupervised learning.
  41930. 32:05:38Simple, right? Scenario one. Facebook
  41931. 32:05:41recognizes your friend in a picture from
  41932. 32:05:43an album of tagged photographs.
  41933. 32:05:46Scenario two, Netflix recommends new
  41934. 32:05:49movies based on someone's past movie
  41935. 32:05:51choices.
  41936. 32:05:53Scenario three, analyzing bank data for
  41937. 32:05:55suspicious transactions and flagging the
  41938. 32:05:58fraud transactions. Think wisely and
  41939. 32:06:00comment below your answers. Moving on,
  41940. 32:06:02don't you sometimes wonder how is
  41941. 32:06:04machine learning possible in today's
  41942. 32:06:06era? Well, that's because today we have
  41943. 32:06:08humongous data available. Everybody's
  41944. 32:06:11online either making a transaction or
  41945. 32:06:13just surfing the internet and that's
  41946. 32:06:15generating a huge amount of data every
  41947. 32:06:17minute. And that data my friend is the
  41948. 32:06:20key to analysis. Also, the memory
  41949. 32:06:22handling capabilities of computers have
  41950. 32:06:24largely increased which helps them to
  41951. 32:06:26process such huge amount of data at hand
  41952. 32:06:29without any delay. And yes, computers
  41953. 32:06:31now have great computational powers. So
  41954. 32:06:34there are a lot of applications of
  41955. 32:06:36machine learning out there. To name a
  41956. 32:06:37few, machine learning is used in
  41957. 32:06:39healthcare where diagnostics are
  41958. 32:06:41predicted for doctor's review. The
  41959. 32:06:43sentiment analysis that the tech giants
  41960. 32:06:45are doing on social media is another
  41961. 32:06:47interesting application of machine
  41962. 32:06:49learning. Fraud detection in the finance
  41963. 32:06:51sector and also to predict customer
  41964. 32:06:53churn in the e-commerce sector. While
  41965. 32:06:54booking a cab, you must have encountered
  41966. 32:06:57search pricing often where it says the
  41967. 32:06:59fair of your trip has been updated.
  41968. 32:07:01Continue booking. Yes, please. I'm
  41969. 32:07:03getting late for office. Well, that's an
  41970. 32:07:06interesting machine learning model which
  41971. 32:07:08is used by global taxi giant Uber and
  41972. 32:07:10others where they have differential
  41973. 32:07:12pricing in real time based on demand,
  41974. 32:07:14the number of cars available, bad
  41975. 32:07:16weather, rush hour, etc. So they use the
  41976. 32:07:19search pricing model to ensure that
  41977. 32:07:21those who need a cab can get one. Also,
  41978. 32:07:24it uses predictive modeling to predict
  41979. 32:07:27where the demand will be high with a
  41980. 32:07:29goal that drivers can take care of the
  41981. 32:07:31demand and search pricing can be
  41982. 32:07:33minimized. Great. Hey Siri, can you
  41983. 32:07:35remind me to book a cab at 6 p.m. today?
  41984. 32:07:37>> Okay, I'll remind you.
  41985. 32:07:39>> Thanks.
  41986. 32:07:40>> No problem.
  41987. 32:07:41>> Comment below some interesting everyday
  41988. 32:07:43examples around you where machines are
  41989. 32:07:45learning and doing amazing jobs. Hi
  41990. 32:07:48guys, this is Acha from SimplyLearn and
  41991. 32:07:51we're going to talk about the two types
  41992. 32:07:52of machine learning supervised and
  41993. 32:07:55unsupervised learning their types and
  41994. 32:07:57applications. But before we talk about
  41995. 32:07:59them, let's quickly understand what is
  41996. 32:08:01machine learning. These days
  41997. 32:08:02applications use artificial intelligence
  41998. 32:08:05in machine learning to optimize speech
  41999. 32:08:07recognition. I usually ask Siri things I
  42000. 32:08:10want to know like hey Siri how far is
  42001. 32:08:12the nearest subway? So whenever we ask
  42002. 32:08:14something to Siri, a powerful speech
  42003. 32:08:17recognition kicks off and converts the
  42004. 32:08:19audio into its corresponding textual
  42005. 32:08:21form which is then sent to the Apple
  42006. 32:08:23servers for further processing. Then
  42007. 32:08:26neural language processing algorithms
  42008. 32:08:28are run to understand the user's intent
  42009. 32:08:31and then finally Siri tells you the
  42010. 32:08:33answer. Well, this is what machine
  42011. 32:08:35learning is all about. making the
  42012. 32:08:37machines learn and act like humans by
  42013. 32:08:39feeding them with data and information
  42014. 32:08:42without being explicitly programmed. As
  42015. 32:08:44we saw in the previous example, when the
  42016. 32:08:46data comes in, machines immediately
  42017. 32:08:48starts analyzing the data and eventually
  42018. 32:08:51gets trained on it and learns it. Now
  42019. 32:08:53when a new data point comes in, machine
  42020. 32:08:56accurately makes prediction and
  42021. 32:08:58decisions based on the past data. Now
  42022. 32:09:00that you know what is machine learning,
  42023. 32:09:02let's talk about supervised and
  42024. 32:09:03unsupervised learning. Supervised
  42025. 32:09:05learning as the name suggests works
  42026. 32:09:07under supervision that is it's a
  42027. 32:09:09learning in which machine is trained
  42028. 32:09:11with data which is welllabeled and then
  42029. 32:09:13predicts with the help of the label data
  42030. 32:09:15set. But what is a label data set? Data
  42031. 32:09:18for which you already know the target
  42032. 32:09:19answer is called a label data. Like I
  42033. 32:09:22show you an image and tell you that it's
  42034. 32:09:25a dog then it's a label data. While if I
  42035. 32:09:28show you an image without telling you
  42036. 32:09:30what exactly it is, then it's an
  42037. 32:09:32unlabelled data. Now let's say we have
  42038. 32:09:35images which are labeled as spoon or
  42039. 32:09:37knife. We then feed it to the machine
  42040. 32:09:39which analyzes and learns the
  42041. 32:09:41association of these images with its
  42042. 32:09:43labels based on its features such as
  42043. 32:09:46shape, size, sharpness etc. Now when a
  42044. 32:09:50new image is fed to the machine without
  42045. 32:09:52any label with the help of the past data
  42046. 32:09:54the machine is able to predict
  42047. 32:09:56accurately and tell that it's a spoon.
  42048. 32:09:58Hence in supervised machine learning the
  42049. 32:10:00algorithm teaches the model to learn
  42050. 32:10:02from the labeled example that we
  42051. 32:10:04provide. So supervised learning can be
  42052. 32:10:07further divided into classification and
  42053. 32:10:09regression. It is a classification
  42054. 32:10:11problem when the output variable is
  42055. 32:10:13categoric such as red or blue, disease
  42056. 32:10:16or no disease, male or female. Whereas
  42057. 32:10:19it's a regression problem when the
  42058. 32:10:20output variable is a real or continuous
  42059. 32:10:23value. For example, salary based on work
  42060. 32:10:26experience, weight based on height. So
  42061. 32:10:29it creates a predictive model showing
  42062. 32:10:31trends in data. So now if I say will I
  42063. 32:10:34get a salary raise or not, that's
  42064. 32:10:36classification. But if I say how much
  42065. 32:10:38salary raise will I get, that is
  42066. 32:10:41regression. Now let's understand
  42067. 32:10:42classification with the help of an
  42068. 32:10:44example. In order to predict whether an
  42069. 32:10:46email is a spam or not, first we need to
  42070. 32:10:49teach our machine what a spam mail looks
  42071. 32:10:51like. This is done based on a lot of
  42072. 32:10:54spam filters like firstly reviewing the
  42073. 32:10:57content of the email then review the
  42074. 32:10:59email header and search if it contains
  42075. 32:11:01any falsified information. This is done
  42076. 32:11:03based on some keywords like free lottery
  42077. 32:11:08prize claim etc. Then general blacklist
  42078. 32:11:12filters to stop emails that come from
  42079. 32:11:14already blacklisted known spammers and
  42080. 32:11:18etc. So all these filters scores the
  42081. 32:11:20email which is known as spam score. The
  42082. 32:11:23lower the total spam score of the email,
  42083. 32:11:26it is more likely that the email will
  42084. 32:11:28land in subscribers inboxes. So based on
  42085. 32:11:31the content labels and spam score of the
  42086. 32:11:34new incoming mail, the algorithm decides
  42087. 32:11:37whether it should land in inbox or the
  42088. 32:11:40spam folder. Now let's quickly
  42089. 32:11:42understand regression. So let's say we
  42090. 32:11:44have two variables that is temperature
  42091. 32:11:46and humidity where temperature is the
  42092. 32:11:49independent variable and humidity is the
  42093. 32:11:51dependent variable such that as the
  42094. 32:11:54temperature increases humidity decreases
  42095. 32:11:56hence they're correlated. When we feed
  42096. 32:11:58this data to a regression model it will
  42097. 32:12:01understand the relationship between
  42098. 32:12:02these two variables and how one variable
  42099. 32:12:05depends on the other. After the machine
  42100. 32:12:07is trained it can easily predict the
  42101. 32:12:10humidity based on the given temperature.
  42102. 32:12:12Well, that was about regression. Now,
  42103. 32:12:15let's see some real life applications
  42104. 32:12:17where supervised learning is used. So,
  42105. 32:12:19supervised learning is used in risk
  42106. 32:12:21assessment to assess risk in financial
  42107. 32:12:24services or an insurance domain to
  42108. 32:12:26minimize the risk portfolio of the
  42109. 32:12:28companies. Image classification.
  42110. 32:12:30Facebook recognizes your friend in a
  42111. 32:12:32picture from an album of tact photos.
  42112. 32:12:34So, image classification is one of the
  42113. 32:12:36key use cases of demonstrating
  42114. 32:12:38supervised machine learning algorithms.
  42115. 32:12:40Obviously a lot more goes into all these
  42116. 32:12:43like convolutional neural networks etc.
  42117. 32:12:46Fraud detection whether the transactions
  42118. 32:12:48made by the user are authentic or not
  42119. 32:12:50and visual recognition the ability of a
  42120. 32:12:53machine learning model to identify
  42121. 32:12:55objects places people and actions in
  42122. 32:12:58images. Now let's quickly talk about
  42123. 32:13:00unsupervised learning. In unsupervised
  42124. 32:13:02learning, there is no supervision that
  42125. 32:13:04is no training will be given to the
  42126. 32:13:06machine allowing it to act on the data
  42127. 32:13:09which is not labeled. Hence, machine
  42128. 32:13:12tries to identify patterns and gives the
  42129. 32:13:14response. Let's take a similar example
  42130. 32:13:17as before. But this time we do not tell
  42131. 32:13:19the machine whether it's a spoon or a
  42132. 32:13:21knife. The machine identifies patterns
  42133. 32:13:24from the given set and groups them based
  42134. 32:13:27on their patterns, similarities, etc.
  42135. 32:13:30Again unsupervised learning can be
  42136. 32:13:32further grouped into clustering and
  42137. 32:13:34association. Clustering is basically
  42138. 32:13:36where the machine forms groups based on
  42139. 32:13:39the behavior of the data. Secondly,
  42140. 32:13:41association. It is a rule-based machine
  42141. 32:13:44learning to discover interesting
  42142. 32:13:45relation between variables in large data
  42143. 32:13:48sets. For example, which customer made
  42144. 32:13:51similar product purchases is clustering.
  42145. 32:13:54Whereas association is which products
  42146. 32:13:56were purchased together. Now let's
  42147. 32:13:59understand clustering with the help of
  42148. 32:14:00an example. To reduce their churn rate,
  42149. 32:14:03a telecom company studies the behavior
  42150. 32:14:05of the customers based on average call
  42151. 32:14:08duration and internet usage and observes
  42152. 32:14:11that while some customers call duration
  42153. 32:14:13is quite high, others have heavy
  42154. 32:14:15internet usage. The customers are
  42155. 32:14:17grouped based on their observed behavior
  42156. 32:14:19and a strategy is adopted to minimize
  42157. 32:14:22churn rate and maximize profit via
  42158. 32:14:24suitable promotions and campaigns. As
  42159. 32:14:27you can see in the chart on the right
  42160. 32:14:28hand side, customers in group A uses
  42161. 32:14:31more data and also have high
  42162. 32:14:33qualuration. Group B customers are heavy
  42163. 32:14:35internet users while group C customers
  42164. 32:14:38have high qualation. So group B will be
  42165. 32:14:41given more data benefit plans while
  42166. 32:14:43group C will be given cheaper call rates
  42167. 32:14:45to buy their loyalty. So this was the
  42168. 32:14:48example of clustering. Now let's
  42169. 32:14:50understand association with another
  42170. 32:14:52example. Let's say customer one goes to
  42171. 32:14:54a supermarket and buys these products.
  42172. 32:14:56say bread, milk, fruits, wheat. Then
  42173. 32:15:00customer two goes and buys bread, milk,
  42174. 32:15:03rice and butter. Now when customer three
  42175. 32:15:06goes and buys bread, it is highly likely
  42176. 32:15:09that he will also buy milk. Hence
  42177. 32:15:12relationship is established based on
  42178. 32:15:14customer behavior and recommendations
  42179. 32:15:16are made. Now let's look at some real
  42180. 32:15:19life applications of unsupervised
  42181. 32:15:20learning. Market basket analysis is a
  42182. 32:15:23machine learning model based on the
  42183. 32:15:25algorithm that if you buy a certain
  42184. 32:15:26group of items, you are less or more
  42185. 32:15:28likely to buy another group of items.
  42186. 32:15:30Semantic clustering. Semantically
  42187. 32:15:32similar words share similar context.
  42188. 32:15:35People post their queries on websites in
  42189. 32:15:37their own ways. Semantic clustering
  42190. 32:15:39groups all responses in a cluster with
  42191. 32:15:42same meaning to ensure that the customer
  42192. 32:15:44finds the information they want quickly
  42193. 32:15:47and easily. It plays an important role
  42194. 32:15:49in information retrieval. Good browsing
  42195. 32:15:51experience and comprehension. Delivery
  42196. 32:15:54store optimization. Machine learning
  42197. 32:15:56models are used to predict the demand
  42198. 32:15:58and keep up with the supply also to open
  42199. 32:16:00stores where demand is more and
  42200. 32:16:03optimizing routes for more efficient
  42201. 32:16:04deliveries according to past data and
  42202. 32:16:07behavior. We can also use unsupervised
  42203. 32:16:09machine learning models to identify
  42204. 32:16:11accidentprone areas based on the
  42205. 32:16:13intensity of those accidents and the
  42206. 32:16:15area in order to introduce safety
  42207. 32:16:17measures. By now I hope you've
  42208. 32:16:19understood supervised and unsupervised
  42209. 32:16:21learning. For a quick recap, let's see a
  42210. 32:16:24few differences between the two. The
  42211. 32:16:25most fundamental difference is that
  42212. 32:16:27supervised learning uses known and
  42213. 32:16:29labelled data and unsupervised learning
  42214. 32:16:31uses unlabelled data as their input.
  42215. 32:16:33Secondly, supervised learning follows a
  42216. 32:16:36feedback mechanism while unsupervised
  42217. 32:16:38learning does not. Also, the most
  42218. 32:16:40commonly used algorithms in supervised
  42219. 32:16:42learning are decision tree, logistic
  42220. 32:16:44regression, support vector machine, etc.
  42221. 32:16:47And in unsupervised learning there K
  42222. 32:16:50means clustering, hierarchical
  42223. 32:16:51clustering, a priori algorithm and many
  42224. 32:16:54more.
  42225. 32:16:55>> Welcome to linear regression. My name is
  42226. 32:16:57Richard Kersner. I'm with SimplyLearn.
  42227. 32:16:59Let's look at an example of a common use
  42228. 32:17:01for linear regression, profit estimation
  42229. 32:17:04of a company. If I was going to invest
  42230. 32:17:06in a company, I would like to know how
  42231. 32:17:08much money I could expect to make. So
  42232. 32:17:10we'll take a look at a venture
  42233. 32:17:11capitalist firm and try to understand
  42234. 32:17:14which companies they should invest in.
  42235. 32:17:16So we'll take the idea that we need to
  42236. 32:17:18decide the companies to invest in. We
  42237. 32:17:20need to predict the profit the company
  42238. 32:17:22makes and we're going to do it based on
  42239. 32:17:24the company's expenses and even just a
  42240. 32:17:27specific expense. In this case we have
  42241. 32:17:29our company, we have the different
  42242. 32:17:31expenses. So we have our R&D which is
  42243. 32:17:33your research and development. We have
  42244. 32:17:35our marketing. Uh we might have the
  42245. 32:17:37location. We might have what kind of
  42246. 32:17:39administration it's going through. Based
  42247. 32:17:41on all this different information, we
  42248. 32:17:43would like to calculate the profit. Now,
  42249. 32:17:45in actuality, there's usually about 23
  42250. 32:17:48to 27 different markers that they look
  42251. 32:17:50at if they're a heavy duty investor.
  42252. 32:17:52We're only going to take a look at one
  42253. 32:17:54basic one. We're going to come in and
  42254. 32:17:56for simplicity, let's consider a single
  42255. 32:17:58variable, R&D, and find out which
  42256. 32:18:00companies to invest in based on that.
  42257. 32:18:02So, we take our R&D and we're plotting
  42258. 32:18:04the profit based on the R&D expenditure,
  42259. 32:18:06how much money they put into the
  42260. 32:18:07research and development. And then we
  42261. 32:18:09look at the profit that goes with that.
  42262. 32:18:11We can predict a line to estimate the
  42263. 32:18:13profit. So we can draw a line right
  42264. 32:18:15through the data. And when you look at
  42265. 32:18:16that, you can see how much they invest
  42266. 32:18:18in the R&D is a good marker as to how
  42267. 32:18:20much profit they're going to have. We
  42268. 32:18:22can also note that companies spending
  42269. 32:18:23more on R&D make good profit. So let's
  42270. 32:18:26invest in the ones that spend a higher
  42271. 32:18:27rate in their R&D. What's in it for you?
  42272. 32:18:30First, we'll have an introduction to
  42273. 32:18:32machine learning followed by machine
  42274. 32:18:34learning algorithms. These will be
  42275. 32:18:36specific to linear regression and where
  42276. 32:18:38it fits into the larger model. Then
  42277. 32:18:40we'll take a look at applications of
  42278. 32:18:42linear regression, understanding linear
  42279. 32:18:44regression, and multiple linear
  42280. 32:18:46regression. Finally, we'll roll up our
  42281. 32:18:48sleeves and do a little programming in
  42282. 32:18:50use case profit estimation of companies.
  42283. 32:18:53Let's go ahead and jump in. Let's start
  42284. 32:18:55with our introduction to machine
  42285. 32:18:56learning along with some machine
  42286. 32:18:58learning algorithms and where that fits
  42287. 32:19:00in with linear regression. Let's look at
  42288. 32:19:02another example of machine learning.
  42289. 32:19:04Based on the amount of rainfall, how
  42290. 32:19:05much would be the crop yield? So we here
  42291. 32:19:08we have our crops, we have our rainfall,
  42292. 32:19:10and we want to know how much we're going
  42293. 32:19:12to get from our crops this year. So
  42294. 32:19:14we're going to introduce two variables,
  42295. 32:19:16independent and dependent. The
  42296. 32:19:18independent variable is a variable whose
  42297. 32:19:20value does not change by the effect of
  42298. 32:19:22other variables and is used to
  42299. 32:19:24manipulate the dependent variable. It is
  42300. 32:19:26often denoted as X. In our example,
  42301. 32:19:28rainfall is the independent variable.
  42302. 32:19:30This is a wonderful example because you
  42303. 32:19:32can easily see that we can't control the
  42304. 32:19:34rain, but the rain does control the
  42305. 32:19:36crop. So we talk about the independent
  42306. 32:19:38variable controlling the dependent
  42307. 32:19:40variable. Let's define dependent
  42308. 32:19:42variable as a variable whose value
  42309. 32:19:43change when there is any manipulation in
  42310. 32:19:46the values of the independent variables.
  42311. 32:19:48It is often denoted as y. And you can
  42312. 32:19:50see here our crop yield is dependent
  42313. 32:19:52variable and it is dependent on the
  42314. 32:19:54amount of rainfall received. Now that
  42315. 32:19:56we've taken a look at a real life
  42316. 32:19:57example, let's go a little bit into the
  42317. 32:19:59theory and some definitions on machine
  42318. 32:20:02learning and see how that fits together
  42319. 32:20:04with linear regression. numerical and
  42320. 32:20:06categorical values. Let's take our data
  42321. 32:20:09coming in and this is kind of random
  42322. 32:20:11data from any kind of project. We want
  42323. 32:20:14to divide it up into numerical and
  42324. 32:20:16categorical. So numerical is numbers,
  42325. 32:20:19age, salary, height, where categorical
  42326. 32:20:23would be a description, the color, a
  42327. 32:20:26dog's breed, gender. Categorical is
  42328. 32:20:29limited to very specific items where
  42329. 32:20:30numerical is a range of information. Now
  42330. 32:20:34that you've seen the difference between
  42331. 32:20:36numerical and categorical data, let's
  42332. 32:20:38take a look at some different machine
  42333. 32:20:40learning definitions. When we look at
  42334. 32:20:42our different machine learning
  42335. 32:20:43algorithms, we can divide them into
  42336. 32:20:46three areas. Supervised, unsupervised,
  42337. 32:20:50reinforcement. We're only going to look
  42338. 32:20:51at supervised today. Unsupervised means
  42339. 32:20:54we don't have the answers and we're just
  42340. 32:20:56grouping things. Reinforcement is where
  42341. 32:20:58we give positive and negative feedback
  42342. 32:21:01to our algorithm to program it. and it
  42343. 32:21:03doesn't have the information till after
  42344. 32:21:04the fact. But today we're just looking
  42345. 32:21:06at supervised because that's where
  42346. 32:21:07linear regression fits in. In supervised
  42347. 32:21:10data, we have our data already there and
  42348. 32:21:12our answers for a group. And then we use
  42349. 32:21:14that to program our model and come up
  42350. 32:21:17with an answer. The two most common uses
  42351. 32:21:19for that is through the regression and
  42352. 32:21:21classification. Now, we're doing linear
  42353. 32:21:23regression. So, we're just going to
  42354. 32:21:24focus on the regression side. And in the
  42355. 32:21:27regression we have simple linear
  42356. 32:21:29regression, we have multiple linear
  42357. 32:21:31regression and we have polomial linear
  42358. 32:21:34regression. Now on these three simple
  42359. 32:21:36linear regression is the examples we've
  42360. 32:21:38looked at so far where we have a lot of
  42361. 32:21:40data and we draw a straight line through
  42362. 32:21:42it. Multiple linear regression means we
  42363. 32:21:44have multiple variables. Remember where
  42364. 32:21:47we had the rainfall and the crops. We
  42365. 32:21:49might add additional variables in there
  42366. 32:21:51like how much food do we give our crops?
  42367. 32:21:53When do we harvest them? Those would be
  42368. 32:21:55additional information add into our
  42369. 32:21:57model and that's why it' be multiple
  42370. 32:21:59linear regression. And finally we have
  42371. 32:22:00polomial linear regression that is
  42372. 32:22:02instead of drawing a line we can draw a
  42373. 32:22:04curved line through it. Now that you see
  42374. 32:22:06where regression model fits into the
  42375. 32:22:08machine learning algorithms and we're
  42376. 32:22:11specifically looking at linear
  42377. 32:22:12regression. Let's go ahead and take a
  42378. 32:22:14look at applications for linear
  42379. 32:22:16regression. Let's look at a few
  42380. 32:22:18applications of linear regression.
  42381. 32:22:21Economic growth used to determine the
  42382. 32:22:23economic growth of a country or a state
  42383. 32:22:26in the coming quarter can also be used
  42384. 32:22:28to predict the GDP of a country. Product
  42385. 32:22:31price can be used to predict what would
  42386. 32:22:33be the price of a product in the future.
  42387. 32:22:34We can guess whether it's going to go up
  42388. 32:22:36or down or should I buy today. Housing
  42389. 32:22:38sales to estimate the number of houses a
  42390. 32:22:40builder would sell and what price in the
  42391. 32:22:42coming months. Score predictions.
  42392. 32:22:44Cricket fever to predict the number of
  42393. 32:22:46runs a player would score in the coming
  42394. 32:22:48matches based on the previous
  42395. 32:22:50performance. I'm sure you can figure out
  42396. 32:22:52other applications you could use linear
  42397. 32:22:54regression for. So let's jump in and
  42398. 32:22:57let's understand linear regression and
  42399. 32:22:59dig into the theory. Understanding
  42400. 32:23:01linear regression. Linear regression is
  42401. 32:23:04the statistical model used to predict
  42402. 32:23:07the relationship between independent and
  42403. 32:23:09dependent variables by examining two
  42404. 32:23:11factors. The first important one is
  42405. 32:23:14which variables in particular are
  42406. 32:23:16significant predictors of the outcome
  42407. 32:23:18variable. And the second one that we
  42408. 32:23:20need to look at closely is how
  42409. 32:23:22significant is the regression line to
  42410. 32:23:24make predictions with the highest
  42411. 32:23:25possible accuracy. If it's inaccurate,
  42412. 32:23:28we can't use it. So, it's very important
  42413. 32:23:30we find out the most accurate line we
  42414. 32:23:32can get. Since linear regression is
  42415. 32:23:35based on drawing a line through data,
  42416. 32:23:37we're going to jump back and take a look
  42417. 32:23:39at some uklitian geometry. The simplest
  42418. 32:23:41form of a simple linear regression
  42419. 32:23:43equation with one dependent and one
  42420. 32:23:45independent variable is represented by y
  42421. 32:23:49= m * x + c. And if you look at our
  42422. 32:23:52model here, we plotted two points on
  42423. 32:23:54here. Uh x1 and y1, x2 and y2. y being
  42424. 32:24:00the dependent variable, remember that
  42425. 32:24:02from before. And x being the independent
  42426. 32:24:05variable. So y depends on whatever x is.
  42427. 32:24:08M in this case is the slope of the line
  42428. 32:24:11where M equals the difference in the Y2
  42429. 32:24:15- Y1 and X2 - X1. And finally we have C
  42430. 32:24:19which is the coefficient of the line or
  42431. 32:24:22where happens to cross the zero axis.
  42432. 32:24:25Let's go back and look at an example we
  42433. 32:24:27used earlier of linear regression. We're
  42434. 32:24:29going to go back to plotting the amount
  42435. 32:24:31of crop yield based on the amount of
  42436. 32:24:33rainfall. And here we have our rainfall.
  42437. 32:24:36Remember, we cannot change rainfall. And
  42438. 32:24:39we have our crop yield, which is
  42439. 32:24:41dependent on the rainfall. So, we have
  42440. 32:24:43our independent and our dependent
  42441. 32:24:44variables. We're going to take this and
  42442. 32:24:46draw a line through it as best we can
  42443. 32:24:49through the middle of the data. And then
  42444. 32:24:50we look at that. We put the red point on
  42445. 32:24:52the y ais is the amount of crop yield
  42446. 32:24:55you can expect for the amount of
  42447. 32:24:56rainfall represented by the green dot.
  42448. 32:24:59So, if we have an idea what the rainfall
  42449. 32:25:00is for this year and what's going on,
  42450. 32:25:03then we can guess how good our crops are
  42451. 32:25:05going to be. and we've created a nice
  42452. 32:25:06line right through the middle to give us
  42453. 32:25:08a nice mathematical formula. Let's take
  42454. 32:25:10a look and see what the math looks like
  42455. 32:25:12behind this. Let's look at the intuition
  42456. 32:25:15behind the regression line. Now, before
  42457. 32:25:17we dive into the math and the formulas
  42458. 32:25:20that go behind this and what's going on
  42459. 32:25:22behind the scenes, I want you to note
  42460. 32:25:25that when we get into the case study and
  42461. 32:25:28we actually apply some Python script
  42462. 32:25:30that this math that you're going to see
  42463. 32:25:31here is already done automatically for
  42464. 32:25:33you. You don't have to have it
  42465. 32:25:34memorized. It is, however, good to have
  42466. 32:25:37an idea what's going on so if people
  42467. 32:25:40reference the different terms, you'll
  42468. 32:25:41know what they're talking about. Let's
  42469. 32:25:43consider a sample data set with five
  42470. 32:25:46rows and find out how to draw the
  42471. 32:25:48regression line. We're only going to do
  42472. 32:25:49five rows because if we did like the
  42473. 32:25:51rainfall with hundreds of points of
  42474. 32:25:53data, that would be very hard to see
  42475. 32:25:55what's going on with the mathematics.
  42476. 32:25:57So, we'll go ahead and create our own
  42477. 32:25:58two sets of data. And we have our
  42478. 32:26:01independent variable x and our dependent
  42479. 32:26:03variable y. And when x was 1, we got y =
  42480. 32:26:072. When x was uh 2, y was 4. And so on
  42481. 32:26:12and so on. If we go ahead and plot this
  42482. 32:26:14data on a graph, we can see how it forms
  42483. 32:26:17a nice line through the middle. You can
  42484. 32:26:19see where it's kind of grouped going
  42485. 32:26:21upwards to the right. The next thing we
  42486. 32:26:23want to know is what the means is of
  42487. 32:26:25each of the data coming in, the x and
  42488. 32:26:28the y. The means doesn't mean anything
  42489. 32:26:30other than the average. So, we add up
  42490. 32:26:32all the numbers and divide by the total.
  42491. 32:26:34So, 1 + 2 + 3 + 4 + 5 over 5 equals 3.
  42492. 32:26:39And the same for y, we get four. If we
  42493. 32:26:41go ahead and plot the means on the
  42494. 32:26:43graph, we'll see we get 3a 4, which
  42495. 32:26:46draws a nice line down the middle, a
  42496. 32:26:48good estimate. Here, we're going to dig
  42497. 32:26:50deeper into the math behind the
  42498. 32:26:52regression line. Now, remember before I
  42499. 32:26:54said you don't have to have all these
  42500. 32:26:56formulas memorized or fully understand
  42501. 32:26:59them, even though we're going to go into
  42502. 32:27:00a little more detail of how it works.
  42503. 32:27:02And if you're not a math wiz and you
  42504. 32:27:04don't know if you've never seen the
  42505. 32:27:05sigma character before, which looks a
  42506. 32:27:08little bit like an e that's opened up,
  42507. 32:27:10that just means summation. That's all
  42508. 32:27:12that is. So, when you see the sigma
  42509. 32:27:14character, it just means we're adding
  42510. 32:27:15everything in that row. And for
  42511. 32:27:17computers, this is great because as a
  42512. 32:27:19programmer, you can easily iterate
  42513. 32:27:21through each of the XY points and create
  42514. 32:27:24all the information you need. So in the
  42515. 32:27:26top half, you can see where we've broken
  42516. 32:27:27that down into pieces. And as it goes
  42517. 32:27:29through the first two points, it
  42518. 32:27:31computes the squared value of X, the
  42519. 32:27:33squared value of Y, and X * Y. And then
  42520. 32:27:36it takes all of X and adds them up. All
  42521. 32:27:38of Y adds them up. All of X squar adds
  42522. 32:27:40them up. And so on and so on. And you
  42523. 32:27:42can see we have the sum of equal to 15.
  42524. 32:27:45The sum is equal to 20. all the way up
  42525. 32:27:47to x * y where the sum equals 66. This
  42526. 32:27:50all comes from our formula for
  42527. 32:27:52calculating a straight line where y
  42528. 32:27:54equals the slope* x plus the coefficient
  42529. 32:27:57c. So we go down below and we're going
  42530. 32:28:00to compute more like the averages of
  42531. 32:28:02these and we're going to explain exactly
  42532. 32:28:04what that is in just a minute and where
  42533. 32:28:06that information comes from. It's called
  42534. 32:28:07the square means error, but we'll go
  42535. 32:28:09into that in detail in a few minutes.
  42536. 32:28:11All you need to do is look at the
  42537. 32:28:12formula and see how we've gone about
  42538. 32:28:14computing it line by line instead of
  42539. 32:28:17trying to have a huge set of numbers
  42540. 32:28:20pushed into it. And down here you'll see
  42541. 32:28:22where the slope m equals and then the
  42542. 32:28:25top part if you read through the
  42543. 32:28:27brackets you have the number of data
  42544. 32:28:29points times the sum of x * y which we
  42545. 32:28:34computed one line at a time there. And
  42546. 32:28:36that's just the 66. and take all that
  42547. 32:28:38and you subtract it from the sum of x
  42548. 32:28:41times the sum of y and those have both
  42549. 32:28:43been computed. So you have 15 * 20. And
  42550. 32:28:45on the bottom we have the number of
  42551. 32:28:48lines times the sum of x^2 easily
  42552. 32:28:51computed as 86 for the sum minus I'll
  42553. 32:28:54take all that and subtract the sum of
  42554. 32:28:57x^2. And we end up as we come across
  42555. 32:29:00with our formula. You can plug in all
  42556. 32:29:02those numbers which is very easy to do
  42557. 32:29:03on the computer. You don't have to do
  42558. 32:29:05the math on a piece of paper or
  42559. 32:29:07calculator. And you'll get a slope of 6
  42560. 32:29:09and you'll get your C coefficient. If
  42561. 32:29:11you continue to follow through that
  42562. 32:29:12formula, you'll see it comes out as
  42563. 32:29:14equal to 2.2. Continuing deeper into
  42564. 32:29:17what's going behind the scenes, let's
  42565. 32:29:19find out the predicted values of y for
  42566. 32:29:21corresponding values of x using the
  42567. 32:29:23linear equation where m=6 and c = 2.2.
  42568. 32:29:28We're going to take these values and
  42569. 32:29:30we're going to go ahead and plot them.
  42570. 32:29:32We're going to predict them. So y =6 * x
  42571. 32:29:35= 1 + 2.2 = 2.8 so on and so on. And
  42572. 32:29:39here the blue points represent the
  42573. 32:29:41actual y values and the brown points
  42574. 32:29:44represent the predicted yv values based
  42575. 32:29:46on the model we created. The distance
  42576. 32:29:48between the actual and predicted values
  42577. 32:29:50is known as residuals or errors. The
  42578. 32:29:54best fit line should have the least sum
  42579. 32:29:56of squares of these errors also known as
  42580. 32:29:59equare. If we put these into a nice
  42581. 32:30:01chart where you can see X and you can
  42582. 32:30:04see Y what the actual values were and
  42583. 32:30:06you can see Y predicted you can easily
  42584. 32:30:08see where we take Y minus Y predicted
  42585. 32:30:11and we get an answer. What is the
  42586. 32:30:12difference between those two and if we
  42587. 32:30:14square that Y - Y prediction squared we
  42588. 32:30:18can then sum those squared values.
  42589. 32:30:20That's where we get the 64 plus the.36 +
  42590. 32:30:231 all the way down until we have a
  42591. 32:30:25summation equals 2.4. So the sum of
  42592. 32:30:28squared errors for this regression line
  42593. 32:30:30is 2.4. We check this error for each
  42594. 32:30:32line and conclude the best fit line
  42595. 32:30:34having the least e value. In a nice
  42596. 32:30:37graphical representation, we can see
  42597. 32:30:39here where we keep moving this line
  42598. 32:30:41through the data points to make sure the
  42599. 32:30:43best fit line has the least squared
  42600. 32:30:45distance between the data points and the
  42601. 32:30:47regression line. Now we only looked at
  42602. 32:30:50the most commonly used formula for
  42603. 32:30:52minimizing the distance. There are lots
  42604. 32:30:55of ways to minimize the distance between
  42605. 32:30:57the line and the data points like sum of
  42606. 32:30:59squared errors, sum of absolute errors,
  42607. 32:31:02root mean square error, etc. What you
  42608. 32:31:04want to take away from this is whatever
  42609. 32:31:06formula is being used, you can easily
  42610. 32:31:09using a computer programming and
  42611. 32:31:11iterating through the data calculate the
  42612. 32:31:13different parts of it. That way, these
  42613. 32:31:16complicated formulas you see with the
  42614. 32:31:17different summations and absolute values
  42615. 32:31:20are easily computed one piece at a time.
  42616. 32:31:22Up until this point, we've only been
  42617. 32:31:24looking at two values, X and Y. Well, in
  42618. 32:31:28the real world, it's very rare that you
  42619. 32:31:29only have two values when you're
  42620. 32:31:31figuring out a solution. So, let's move
  42621. 32:31:33on to the next topic, multiple linear
  42622. 32:31:36regression. Let's take a brief look at
  42623. 32:31:38what happens when you have multiple
  42624. 32:31:39inputs. So, in multiple linear
  42625. 32:31:41regression, we have uh well, we'll start
  42626. 32:31:43with the simple linear regression where
  42627. 32:31:45we had y = m + x + c and we're trying to
  42628. 32:31:49find the value of y. Now with multiple
  42629. 32:31:51linear regression we have multiple
  42630. 32:31:53variables coming in. So instead of
  42631. 32:31:55having just x we have x1 x2 x3 and
  42632. 32:32:00instead of having just one slope each
  42633. 32:32:02variable has its own slope attached to
  42634. 32:32:04it. As you can see here we have m1 m2 m3
  42635. 32:32:08and we still just have the single
  42636. 32:32:09coefficient. So when you're dealing with
  42637. 32:32:11multiple linear regression you basically
  42638. 32:32:13take your single linear regression and
  42639. 32:32:15you spread it out. So you have y = m1 *
  42640. 32:32:18x1 + m2 * x2 so on all the way to m x to
  42641. 32:32:24the nth and then you add your
  42642. 32:32:25coefficient on there. Implementation of
  42643. 32:32:28linear regression. Now we get into my
  42644. 32:32:30favorite part. Let's understand how
  42645. 32:32:32multiple linear regression works by
  42646. 32:32:34implementing it in Python. If you
  42647. 32:32:36remember before we were looking at a
  42648. 32:32:38company and just based on its R&D trying
  42649. 32:32:41to figure out its profit. We're going to
  42650. 32:32:43start looking at the expenditure of the
  42651. 32:32:44company. We're going to go back to that.
  42652. 32:32:46We're going to predict his profit, but
  42653. 32:32:48instead of predicting it just on the
  42654. 32:32:50R&D, we're going to look at other
  42655. 32:32:51factors like administration costs,
  42656. 32:32:54marketing costs, and so on. And from
  42657. 32:32:56there, we're going to see if we can
  42658. 32:32:58figure out what the profit of that
  42659. 32:32:59company is going to be. To start our
  42660. 32:33:01coding, we're going to begin by
  42661. 32:33:03importing some basic libraries. And
  42662. 32:33:05we're going to be looking through the
  42663. 32:33:06data before we do any kind of linear
  42664. 32:33:08regression. We're going to take a look
  42665. 32:33:09at the data to see what we're playing
  42666. 32:33:10with. Then we'll go ahead and format the
  42667. 32:33:13data to the format we need to be able to
  42668. 32:33:15run it in the linear regression model.
  42669. 32:33:17And then from there we'll go ahead and
  42670. 32:33:18solve it and just see how valid our
  42671. 32:33:21solution is. So let's start with
  42672. 32:33:22importing the basic libraries. Now I'm
  42673. 32:33:24going to be doing this in Anaconda
  42674. 32:33:27Jupyter notebook, a very popular IDE. I
  42675. 32:33:30enjoy it because it's such a visual to
  42676. 32:33:32look at and so easy to use. Um just any
  42677. 32:33:34ID for Python will work just fine for
  42678. 32:33:36this. So break out your favorite Python
  42679. 32:33:38IDE. So, here we are in our Jupyter
  42680. 32:33:41notebook. Let me go ahead and paste our
  42681. 32:33:42first piece of code in there. And let's
  42682. 32:33:44walk through what libraries we're
  42683. 32:33:46importing. First, we're going to import
  42684. 32:33:48numpy as np. And then I want you to skip
  42685. 32:33:51one line and look at import pandas as
  42686. 32:33:53pd. These are very common tools that you
  42687. 32:33:55need with most of your linear
  42688. 32:33:56regression. The numpy, which stands for
  42689. 32:33:59number python, is usually denoted as np,
  42690. 32:34:01and you have to almost have that for
  42691. 32:34:03your sklearn toolbox. So, you always
  42692. 32:34:05import that right off the beginning.
  42693. 32:34:06pandas. Although you don't have to have
  42694. 32:34:08it for your sklearn libraries, it does
  42695. 32:34:10such a wonderful job of importing data,
  42696. 32:34:12setting it up into a data frame so we
  42697. 32:34:14can manipulate it rather easily and it
  42698. 32:34:16has a lot of tools also in addition to
  42699. 32:34:18that. So we usually like to use the
  42700. 32:34:20pandas when we can and I'll show you
  42701. 32:34:21what that looks like. The other three
  42702. 32:34:23lines are for us to get a visual of this
  42703. 32:34:26data and take a look at it. So we're
  42704. 32:34:28going to import mapplot library.pipplot
  42705. 32:34:30as plt and then seabour as sns. Seabor
  42706. 32:34:35works with the mattplot library. So you
  42707. 32:34:37have to always import mapplot library
  42708. 32:34:39and then seabour sits on top of it. And
  42709. 32:34:41we'll take a look at what that looks
  42710. 32:34:42like. You could use any of your own
  42711. 32:34:44plotting libraries you want. There's all
  42712. 32:34:46kinds of ways to look at the data. These
  42713. 32:34:48are just very common ones. And the
  42714. 32:34:49seabor is so easy to use. It just looks
  42715. 32:34:52beautiful. It's a nice representation
  42716. 32:34:53that you can actually take and show
  42717. 32:34:55somebody. And the final line is the
  42718. 32:34:57amberigned mattplot library inline. That
  42719. 32:35:01is only because I'm doing an inline IDE.
  42720. 32:35:03My interface in the Anaconda Jupiter
  42721. 32:35:05notebook requires I put that in there or
  42722. 32:35:08you're not going to see the graph when
  42723. 32:35:10it comes up. Let's go ahead and run
  42724. 32:35:11this. It's not going to be that
  42725. 32:35:12interesting because we're just setting
  42726. 32:35:13up variables. In fact, it's not going to
  42727. 32:35:15do anything that we can see, but it is
  42728. 32:35:17importing these different libraries and
  42729. 32:35:19setup. The next step is load the data
  42730. 32:35:22set and extract independent and
  42731. 32:35:24dependent variables. Now, here in the
  42732. 32:35:27slide, you'll see companies equals PD
  42733. 32:35:29read CSV. And it has a long line there
  42734. 32:35:32with the file at the end. 10,00
  42735. 32:35:34companies.csv. You're going to have to
  42736. 32:35:36change this to fit whatever setup you
  42737. 32:35:38have. And the file itself, you can
  42738. 32:35:41request. Just go down to the commentary
  42739. 32:35:43below this video and put a note in there
  42740. 32:35:45and SimplyLearn will try to get in
  42741. 32:35:47contact with you and supply you with
  42742. 32:35:48that file so you can try this coding
  42743. 32:35:50yourself. So, we're going to add this
  42744. 32:35:52code in here. And we're going to see
  42745. 32:35:53that I have companies equals
  42746. 32:35:55PD.reader_csv.
  42747. 32:35:57And I've changed this path to match my
  42748. 32:35:59computer. C/simplylearn/1000
  42749. 32:36:03companies.csv. And then below there,
  42750. 32:36:05we're going to set the x equals to
  42751. 32:36:07companies under the i location. And
  42752. 32:36:09because this is companies is a pd data
  42753. 32:36:12set, I can use this nice notation that
  42754. 32:36:14says take every row, that's what the
  42755. 32:36:16colon the first colon is, comma, except
  42756. 32:36:19for the last column. That's what the
  42757. 32:36:20second part is where we have a colon
  42758. 32:36:22minus one and we want the values set
  42759. 32:36:24into there. So x is no longer a data set
  42760. 32:36:27a pandas data set but we can easily
  42761. 32:36:29extract the data from our pandas data
  42762. 32:36:31set with this notation and then y we're
  42763. 32:36:33going to set equal to the last row. Well
  42764. 32:36:36the question is going to be what are we
  42765. 32:36:37actually looking at? So let's go ahead
  42766. 32:36:39and take a look at that and we're going
  42767. 32:36:41to look at the companies.
  42768. 32:36:43Which lists the first five rows of data
  42769. 32:36:45and I'll open up the file in just a
  42770. 32:36:47second so you can see where that's
  42771. 32:36:48coming from. But let's look at the data
  42772. 32:36:50in here as far as the way the pandas
  42773. 32:36:52sees it. When I hit run, you'll see it
  42774. 32:36:54breaks it out into a nice setup. This is
  42775. 32:36:56what pandas, one of the things pandas is
  42776. 32:36:58really good about is it looks just like
  42777. 32:36:59an Excel spreadsheet. You have your rows
  42778. 32:37:01and remember when we're programming, we
  42779. 32:37:03always start with zero. We don't start
  42780. 32:37:05with one. So it shows the first five
  42781. 32:37:08rows 0 1 2 3 4 and then it shows your
  42782. 32:37:11different columns. R&D spend,
  42783. 32:37:14administration, marketing spend, state,
  42784. 32:37:17profit. It even notes that the top are
  42785. 32:37:19column names. It was never told that,
  42786. 32:37:21but Pandas is able to recognize a lot of
  42787. 32:37:23things that they're not the same as the
  42788. 32:37:25data rows. Why don't we go ahead and
  42789. 32:37:26open this file up in a CSV so you can
  42790. 32:37:29actually see the raw data. So here I've
  42791. 32:37:31opened it up as a text editor. And you
  42792. 32:37:33can see at the top we have R&D spend,
  42793. 32:37:36administration, marketing spend, state,
  42794. 32:37:39profit, carriage return. I don't know
  42795. 32:37:41about you, but I'd go crazy trying to
  42796. 32:37:42read files like this. That's why we use
  42797. 32:37:45the pandas. You could also open this up
  42798. 32:37:47in an Excel and it would separate it
  42799. 32:37:48since it is a comma separated variable
  42800. 32:37:50file. But we don't want to look at this
  42801. 32:37:51one. We want to look at something we can
  42802. 32:37:53read rather easily. So let's flip back
  42803. 32:37:55and take a look at that top part, the
  42804. 32:37:57first five row. Now, as nice as this
  42805. 32:37:59format is where I can see the data, to
  42806. 32:38:01me it doesn't mean a whole lot. Maybe
  42807. 32:38:03you're an expert in business and
  42808. 32:38:05investments and you understand what
  42809. 32:38:08$165,34920
  42810. 32:38:11compared to the administration cost of
  42811. 32:38:14$136,897.80
  42812. 32:38:17so on so on helps to create the profit
  42813. 32:38:19of $192,26183.
  42814. 32:38:23That makes no sense to me whatsoever. No
  42815. 32:38:26pun intended. So let's flip back here
  42816. 32:38:27and take a look at our next set of code
  42817. 32:38:29where we're going to graph it so we can
  42818. 32:38:30get a better understanding of our data
  42819. 32:38:32and what it mean. So at this point we're
  42820. 32:38:35going to use a single line of code to
  42821. 32:38:37get a lot of information so we can see
  42822. 32:38:39where we're going with this. Let's go
  42823. 32:38:41ahead and paste that into our uh
  42824. 32:38:43notebook and see what we got going. And
  42825. 32:38:45so we have the visualization and again
  42826. 32:38:47we're using SNS which is pandas. As you
  42827. 32:38:50can see, we imported the mapplot
  42828. 32:38:51library.pipplot as plt, which then the
  42829. 32:38:55seabor uses and we imported the seabour
  42830. 32:38:57as sns. And then that final line of code
  42831. 32:39:00helps us show this in our um inline
  42832. 32:39:03coding. Without this, it wouldn't
  42833. 32:39:05display and you could display it to a
  42834. 32:39:06file and other means. And that's the map
  42835. 32:39:08plot library in line with the amber sign
  42836. 32:39:10at the beginning. So here we come down
  42837. 32:39:12to the single line of code. Seabor is
  42838. 32:39:14great because it actually recognizes the
  42839. 32:39:16panda data frame. So I can just take the
  42840. 32:39:18companies.core
  42841. 32:39:20for coordinates and I can put that right
  42842. 32:39:22into the seaborn. And when we run this,
  42843. 32:39:25we get this beautiful plot. And let's
  42844. 32:39:27just take a look at what this plot
  42845. 32:39:28means. If you look at this plot on mine,
  42846. 32:39:31the colors are probably a little bit
  42847. 32:39:32more purplish and blue than the original
  42848. 32:39:34one. Uh we have the columns and the
  42849. 32:39:36rows. We have R and D spending. We have
  42850. 32:39:38administration. We have marketing
  42851. 32:39:40spending and profit. And if you cross
  42852. 32:39:42index any two of these, since we're
  42853. 32:39:44interested in profit, if you cross-index
  42854. 32:39:46profit with profit, it's going to show
  42855. 32:39:48up, if you look at the scale on the
  42856. 32:39:49right, way up in the dark. Why? Because
  42857. 32:39:52those are the same data. They have an
  42858. 32:39:54exact correspondence. So R&D spending is
  42859. 32:39:57going to be the same as R&D spending.
  42860. 32:40:00And the same thing with administration
  42861. 32:40:01costs. But right down the middle, you
  42862. 32:40:03get this dark row or dark um diagonal
  42863. 32:40:06row that shows that this is the highest
  42864. 32:40:09corresponding data. That's exactly the
  42865. 32:40:11same. And as it becomes lighter, there's
  42866. 32:40:13less connections between the data. So we
  42867. 32:40:16can see with profit, obviously profit is
  42868. 32:40:18the same as profit. And next, it has a
  42869. 32:40:20very high correlation with R&D spending,
  42870. 32:40:23which we looked at earlier. And it has a
  42871. 32:40:25slightly less connection to marketing
  42872. 32:40:27spending and even less to how much money
  42873. 32:40:29we put into the administration. So now
  42874. 32:40:31that we have a nice look at the data,
  42875. 32:40:33let's go ahead and dig in and create
  42876. 32:40:35some actual useful linear regression
  42877. 32:40:37models so that we can predict values and
  42878. 32:40:39have a better profit. Now that we've
  42879. 32:40:42taken a look at the visualization of
  42880. 32:40:43this data, we're going to move on to the
  42881. 32:40:45next step. Instead of just having a
  42882. 32:40:47pretty picture, we need to generate some
  42883. 32:40:49hard data, some hard values. So let's
  42884. 32:40:52see what that looks like. We're going to
  42885. 32:40:54set up our linear regression model in
  42886. 32:40:56two steps. The first one is we need to
  42887. 32:40:58prepare some of our data so it fits
  42888. 32:41:01correctly. And let's go ahead and paste
  42889. 32:41:02this code into our Jupyter notebook. And
  42890. 32:41:04what we're bringing in is we're going to
  42891. 32:41:05bring in the sklearn pre-processing
  42892. 32:41:08where we're going to import the label
  42893. 32:41:10encoder and the one hot encoder. To use
  42894. 32:41:13the label encoder, we're going to create
  42895. 32:41:14a variable called label encoder and set
  42896. 32:41:16it equal to capital L label capital E
  42897. 32:41:19encoder. This creates a class that we
  42898. 32:41:21can reuse for transferring the labels
  42899. 32:41:24back and forth. Now about now you should
  42900. 32:41:25ask what labels are we talking about.
  42901. 32:41:27Let's go take a look at the data we
  42902. 32:41:29processed before and see what I'm
  42903. 32:41:31talking about here. If you remember when
  42904. 32:41:32we did the companies.head and we printed
  42905. 32:41:34the top five rows of data. We have our
  42906. 32:41:37columns going across. We have column
  42907. 32:41:39zero which is R&D spending, column one
  42908. 32:41:42which is administration, column two
  42909. 32:41:44which is marketing spending and column
  42910. 32:41:46three is state. And you'll see under
  42911. 32:41:49state we have New York, California,
  42912. 32:41:51Florida. Now to do a linear regression
  42913. 32:41:53model, it doesn't know how to process
  42914. 32:41:55New York. It knows how to process a
  42915. 32:41:57number. So the first thing we're going
  42916. 32:41:58to do is we're going to change that New
  42917. 32:42:00York, California, and Florida. And we're
  42918. 32:42:02going to change those to numbers. That's
  42919. 32:42:04what this line of code does here. X
  42920. 32:42:06equals and then it has the colon, 3 in
  42921. 32:42:09brackets. The first part, the colon,
  42922. 32:42:11comma, means that we're going to look at
  42923. 32:42:13all the different rows. So we're going
  42924. 32:42:15to keep them all together. But the only
  42925. 32:42:17row we're going to edit is the third
  42926. 32:42:18row. And in there, we're going to take
  42927. 32:42:20the label coder and we're going to fit
  42928. 32:42:22and transform the x also the third row.
  42929. 32:42:25So, we're going to take that third row,
  42930. 32:42:26we're going to set it equal to a
  42931. 32:42:28transformation. And that transformation
  42932. 32:42:29basically tells it that instead of
  42933. 32:42:31having a uh New York, it has a zero or a
  42934. 32:42:34one or a two. And then finally, we need
  42935. 32:42:36to do a one hot encoder, which equals
  42936. 32:42:39one hot encoder categorical features
  42937. 32:42:42equals three. And then we take the X and
  42938. 32:42:44we go ahead and do that equal to one hot
  42939. 32:42:46encoder fit transform X to array. This
  42940. 32:42:49final transformation preps our data for
  42941. 32:42:52us. So it's completely set the way we
  42942. 32:42:54need it as just a row of numbers. Even
  42943. 32:42:56though it's not in here, let's go ahead
  42944. 32:42:57and print X and just take a look what
  42945. 32:43:00this data is doing. You'll see I have an
  42946. 32:43:01array of arrays and then each array is a
  42947. 32:43:04row of numbers. And if I go ahead and
  42948. 32:43:06just do row zero, you'll see I have a
  42949. 32:43:08nice organized row of numbers that the
  42950. 32:43:10computer now understands. We'll go ahead
  42951. 32:43:12and take this out there because it
  42952. 32:43:14doesn't mean a whole lot to us. It's
  42953. 32:43:15just a row of numbers. Next on setting
  42954. 32:43:18up our data, we have avoiding dummy
  42955. 32:43:21variable trap. This is very important.
  42956. 32:43:24Why? Because the computer's
  42957. 32:43:26automatically transformed our header
  42958. 32:43:28into the setup and it's automatically
  42959. 32:43:30transformed all these different
  42960. 32:43:31variables. So when we did the encoder,
  42961. 32:43:34the encoder created two columns. And
  42962. 32:43:37what we need to do is just have the one
  42963. 32:43:39because it has both the variable and the
  42964. 32:43:41name. That's what this piece of code
  42965. 32:43:43does here. Let's go ahead and paste this
  42966. 32:43:45in here. And we have x= x colon, one
  42967. 32:43:49colon. All this is doing is removing
  42968. 32:43:51that one extra column we put in there
  42969. 32:43:54when we did our one hot encoder and our
  42970. 32:43:55label encoding. Let's go ahead and run
  42971. 32:43:57that. And now we get to create our
  42972. 32:43:59linear regression model. And let's see
  42973. 32:44:01what that looks like here. And we're
  42974. 32:44:03going to do that in two steps. The first
  42975. 32:44:06step is going to be in splitting the
  42976. 32:44:08data. Now, whenever we create a uh
  42977. 32:44:11predictive model of data, we always want
  42978. 32:44:14to split it up. So, we have a training
  42979. 32:44:15set and we have a testing set. That's
  42980. 32:44:17very important. Otherwise, we'd be very
  42981. 32:44:19unethical without testing it to see how
  42982. 32:44:21good our fit is. And then we'll go ahead
  42983. 32:44:23and create our multiple linear
  42984. 32:44:25regression model and train it and set it
  42985. 32:44:27up. Let's go ahead and paste this next
  42986. 32:44:29piece of code in here. And I'll go ahead
  42987. 32:44:30and shrink it down a size or two so it
  42988. 32:44:32all fits on one line. So from the
  42989. 32:44:34sklearn module selection, we're going to
  42990. 32:44:37import train test split. And you'll see
  42991. 32:44:40that we've created four completely
  42992. 32:44:42different variables. We have capital X
  42993. 32:44:44train capital X test smallercase Y train
  42994. 32:44:48smallerase Y test. That is the standard
  42995. 32:44:51way that they usually reference these
  42996. 32:44:53when we're doing different uh models.
  42997. 32:44:56usually see that a capital X and you see
  42998. 32:44:58the train and the test and the lowercase
  42999. 32:45:00Y. What this is is X is our data going
  43000. 32:45:02in. That's our R&D spin, our
  43001. 32:45:04administration, our marketing. And then
  43002. 32:45:07Y, which we're training, is the answer.
  43003. 32:45:09That's the profit because we want to
  43004. 32:45:10know the profit of an unknown entity. So
  43005. 32:45:12that's what we're going to shoot for in
  43006. 32:45:14this tutorial. The next part, train,
  43007. 32:45:16test, split. We take X and we take Y.
  43008. 32:45:20We've already created those. X has the
  43009. 32:45:22columns with the data in it and Y has a
  43010. 32:45:25column with profit in it. And then we're
  43011. 32:45:26going to set the test size equals 0.2.
  43012. 32:45:30That basically means 20%. So 20% of the
  43013. 32:45:33rows are going to be tested. We're going
  43014. 32:45:34to put them off to the side. So since
  43015. 32:45:36we're using a thousand lines of data,
  43016. 32:45:38that means that 200 of those lines we're
  43017. 32:45:41going to hold off to the side to test
  43018. 32:45:42for later. And then the random state
  43019. 32:45:44equals zero. We're going to randomize
  43020. 32:45:46which ones it picks to hold off to the
  43021. 32:45:48side. We'll go ahead and run this. It's
  43022. 32:45:50not overly exciting because it's setting
  43023. 32:45:51up our variables. But the next step is
  43024. 32:45:54the next step we actually create our
  43025. 32:45:55linear regression model. Now that we got
  43026. 32:45:57to the linear regression model, we get
  43027. 32:45:59that next piece of the puzzle. Let's go
  43028. 32:46:01ahead and put that code in there and
  43029. 32:46:02walk through it. So here we go. We're
  43030. 32:46:04going to paste it in there. And let's go
  43031. 32:46:05ahead and uh since this is a shorter
  43032. 32:46:07line of code, let's zoom up there so we
  43033. 32:46:09can get a good look. And we have from
  43034. 32:46:10the sklearn.linear_model,
  43035. 32:46:13we're going to import linear regression.
  43036. 32:46:15Now, I don't know if you recall from
  43037. 32:46:17earlier when we were doing all the math.
  43038. 32:46:19Let's go ahead and flip back there and
  43039. 32:46:20take a look at that. Do you remember
  43040. 32:46:22this where we had this long formula on
  43041. 32:46:24the bottom and we were doing all this
  43042. 32:46:25summization and then we also looked at
  43043. 32:46:28setting it up with the different lines
  43044. 32:46:30and then we also looked all the way down
  43045. 32:46:32to multiple linear regression where
  43046. 32:46:34we're adding all those formulas
  43047. 32:46:36together. All of that is wrapped up in
  43048. 32:46:38this one section. So what's going on
  43049. 32:46:40here is I'm going to create a variable
  43050. 32:46:42called regressor. And the regressor
  43051. 32:46:44equals the linear regression. That's a
  43052. 32:46:46linear regression model that has all
  43053. 32:46:47that math built in. So we don't have to
  43054. 32:46:50have it all memorized or have to compute
  43055. 32:46:51it individually. And then we do the
  43056. 32:46:53regressor.fit.
  43057. 32:46:54In this case, we do xrain and y train
  43058. 32:46:57because we're using the training data. X
  43059. 32:46:59being the data in and y being profit
  43060. 32:47:02what we're looking at. And this does all
  43061. 32:47:03that math for us. So within one click
  43062. 32:47:06and one line, we've created the whole
  43063. 32:47:07linear regression model and we fit the
  43064. 32:47:10data to the linear regression model. And
  43065. 32:47:12you can see that when I run the
  43066. 32:47:13regressor, it gives an output linear
  43067. 32:47:15regression. It says copy X equals true,
  43068. 32:47:18fit intercept equals true, in jobs equal
  43069. 32:47:201, normalize equals false. It's just
  43070. 32:47:22giving you some general information on
  43071. 32:47:24what's going on with that regressor
  43072. 32:47:26model. Now that we've created our linear
  43073. 32:47:28regression model, let's go ahead and use
  43074. 32:47:30it. And if you remember, we kept a bunch
  43075. 32:47:32of data aside. So, we're going to do a Y
  43076. 32:47:35predict variable and we're going to put
  43077. 32:47:37in the X test. And let's see what that
  43078. 32:47:39looks like. Scroll up a little bit.
  43079. 32:47:41Paste that in here. Predicting the test
  43080. 32:47:44set results. So, here we have Y predict
  43081. 32:47:47equals regressor.predict
  43082. 32:47:49X test going in. And this gives us Y
  43083. 32:47:52predict. Now, because I'm in Jupiter in
  43084. 32:47:54line, I can just put the variable up
  43085. 32:47:56there. And when I hit the run button,
  43086. 32:47:58it'll print that array out. I could have
  43087. 32:48:00just as easily done print y predict. So
  43088. 32:48:03if you're in a different IDE that's not
  43089. 32:48:05an inline setup like the Jupyter
  43090. 32:48:07notebook, you can do it this way. Print
  43091. 32:48:09y predict. And you'll see that for the
  43092. 32:48:11200 different test variables we kept off
  43093. 32:48:13to the side, it's going to produce 200
  43094. 32:48:16answers. This is what it says the profit
  43095. 32:48:18are for those 200 predictions. But let's
  43096. 32:48:20don't stop there. Let's keep going and
  43097. 32:48:22take a couple look. We're going to take
  43098. 32:48:24just a short detail here and calculating
  43099. 32:48:26the coefficients and the intercepts.
  43100. 32:48:28This gives us a quick flash at what's
  43101. 32:48:30going on behind the line. We're going to
  43102. 32:48:32take a short detour here and we're going
  43103. 32:48:34to be calculating the coefficient and
  43104. 32:48:36intercepts. So you can see what those
  43105. 32:48:38look like. What's really nice about our
  43106. 32:48:40regressor we created is it already has
  43107. 32:48:42the coefficients for us. And we can
  43108. 32:48:44simply just print regressor.coefficient
  43109. 32:48:47underscore. When I run this, you'll see
  43110. 32:48:49our coefficients here. And if we can do
  43111. 32:48:51the regressor coefficient, we can also
  43112. 32:48:54do the regressor intercept. And let's
  43113. 32:48:57run that and take a look at that. This
  43114. 32:48:58all came from the multiple regression
  43115. 32:49:00model. And we'll flip over so you can
  43116. 32:49:02remember where this is going into where
  43117. 32:49:04it's coming from. You can see the
  43118. 32:49:05formula down here where y = m1 * x1 + m2
  43119. 32:49:10* x2 and so on and so on plus c the
  43120. 32:49:12coefficient. So these variables fit
  43121. 32:49:14right into this formula. Y equ= slope 1
  43122. 32:49:18* column 1 variable plus slope 2 *
  43123. 32:49:21column 2 variable all the way to the m
  43124. 32:49:23into the n and x to the n + c the
  43125. 32:49:26coefficient or in this case you have -
  43126. 32:49:288.89 8 9 to the power of two etc etc
  43127. 32:49:32times the first column and the second
  43128. 32:49:34column and the third column and then our
  43129. 32:49:36intercept is the minus one3009
  43130. 32:49:39point. Boy, it gets kind of complicated
  43131. 32:49:41when you look at it. This is why we
  43132. 32:49:42don't do this by hand anymore. This is
  43133. 32:49:44why we have the computer to make these
  43134. 32:49:46calculations easy to understand and
  43135. 32:49:48calculate. Now, I told you that was a
  43136. 32:49:50short detour and we're coming towards
  43137. 32:49:52the end of our script. As you remember
  43138. 32:49:54from the beginning, I said if we're
  43139. 32:49:55going to divide this information, we
  43140. 32:49:57have to make sure it's a valid model,
  43141. 32:49:59that this model works and understand how
  43142. 32:50:01good it works. So calculating the R
  43143. 32:50:03squar value, that's what we're going to
  43144. 32:50:05use to predict how good our prediction
  43145. 32:50:08is. And let's take a look at what that
  43146. 32:50:09looks like in code. And so we're going
  43147. 32:50:11to use this from sklearn.metrics.
  43148. 32:50:14We're going to import R2 score. That's
  43149. 32:50:16the R squared value. We're looking at
  43150. 32:50:18the error. So in the R2 score, we take
  43151. 32:50:22our Y test versus our Y predict. Y test
  43152. 32:50:26is the actual values we're testing. That
  43153. 32:50:28was the one that was given to us. So we
  43154. 32:50:29know are true. The Y predict of those
  43155. 32:50:32200 values is what we think it was true.
  43156. 32:50:34And when we go ahead and run this, we
  43157. 32:50:36see we get a 9352.
  43158. 32:50:39That's the R2 score. Now, it's not
  43159. 32:50:41exactly a straight percentage. So it's
  43160. 32:50:43not saying it's 93% correct, but you do
  43161. 32:50:46want that in the upper 90s. O and higher
  43162. 32:50:49shows that this is a very valid
  43163. 32:50:51prediction based on the R2 score. And if
  43164. 32:50:53R squar value of N1 or 92 as we got on
  43165. 32:50:57our model remember it does have a random
  43166. 32:50:59generation involved. This proves the
  43167. 32:51:01model is a good model which means
  43168. 32:51:03success. Yay. We successfully trained
  43169. 32:51:06our model with certain predictors and
  43170. 32:51:08estimated the profit of the companies
  43171. 32:51:10using linear regression. What is
  43172. 32:51:12logistic regression? Let's say we have
  43173. 32:51:15to build a predictive model or a machine
  43174. 32:51:17learning model to predict whether the
  43175. 32:51:20passengers of the Titanic ship have
  43176. 32:51:23survived or not the shipwreck. So how do
  43177. 32:51:24we do that? So we use logistic
  43178. 32:51:26regression to build a model for this.
  43179. 32:51:29How do we use logistic regression? So we
  43180. 32:51:32have the information about the
  43181. 32:51:34passengers, their ID, whether they have
  43182. 32:51:36survived or not, their class and name
  43183. 32:51:39and so on and so forth. And we use this
  43184. 32:51:42information where we already know
  43185. 32:51:44whether the person has survived or not.
  43186. 32:51:46That is the labeled information and we
  43187. 32:51:48help the system to train based on this
  43188. 32:51:52information with based on this labeled
  43189. 32:51:54data. This is known as labeled data. And
  43190. 32:51:56during the process of building the
  43191. 32:51:58model, we probably will remove some of
  43192. 32:52:01the non-essential parameters or
  43193. 32:52:03attributes here. We only take those
  43194. 32:52:05attributes which are really required to
  43195. 32:52:07make these predictions. And once we
  43196. 32:52:09train the model, we run new data through
  43197. 32:52:12it whereby the model will predict
  43198. 32:52:14whether the passenger has survived or
  43199. 32:52:16not. All right. What is logistic
  43200. 32:52:18regression? As I mentioned earlier,
  43201. 32:52:20logistic regression is an algorithm for
  43202. 32:52:22performing binary classification. So
  43203. 32:52:25let's take an example and see how this
  43204. 32:52:27works. Let's say your car has not been
  43205. 32:52:30serviced for quite a few years and now
  43206. 32:52:32you want to find out if it is going to
  43207. 32:52:34break down in the near future. So this
  43208. 32:52:37is like a classification problem. Find
  43209. 32:52:39out whether your car will break down or
  43210. 32:52:41not. So how are we going to perform this
  43211. 32:52:44classification? So here's how it looks.
  43212. 32:52:47If we plot the information along the X
  43213. 32:52:50and Y axis, X is the number of years
  43214. 32:52:52since the last service was performed and
  43215. 32:52:55Y is the probability of your car
  43216. 32:52:58breaking down. And let's say this
  43217. 32:53:00information was this data rather was
  43218. 32:53:03collected from several car users. It's
  43219. 32:53:06not just your car but several car users.
  43220. 32:53:09So that is our labeled data. So the data
  43221. 32:53:11has been collected and um for for the
  43222. 32:53:15number of years and when the car broke
  43223. 32:53:18down and what was the probability and
  43224. 32:53:20that has been plotted along x and y
  43225. 32:53:22axis. So this provides an idea or from
  43226. 32:53:26this graph we can find out whether your
  43227. 32:53:29car will break down or not. We'll see
  43228. 32:53:31how. So first of all the probability can
  43229. 32:53:34go from 0 to one. As you all aware
  43230. 32:53:37probability can be between 0 and one.
  43231. 32:53:39And as we can imagine it is intuitive as
  43232. 32:53:43well. As the number of years are on the
  43233. 32:53:46lower side maybe 1 year, 2 years or 3
  43234. 32:53:49years till after the service the chances
  43235. 32:53:52of your car breaking down are very
  43236. 32:53:54limited. Right? So for example, chances
  43237. 32:53:57of your car breaking down or the
  43238. 32:53:58probability of your car breaking down
  43239. 32:54:00within 2 years of your last service are
  43240. 32:54:030.1 probability. Similarly 3 years is
  43241. 32:54:06maybe.3 and so on. But as the number of
  43242. 32:54:09years increases let's say if it was 6 or
  43243. 32:54:127 years there is almost a certainty that
  43244. 32:54:15your car is going to break down. That is
  43245. 32:54:17what this graph shows. So this is an
  43246. 32:54:19example of a application of the
  43247. 32:54:21classification algorithm and we will see
  43248. 32:54:24in little details how exactly logistic
  43249. 32:54:27regression is applied here. One more
  43250. 32:54:29thing needs to be added here is that the
  43251. 32:54:31dependent variables outcome is discrete.
  43252. 32:54:34So if we are talking about whether the
  43253. 32:54:37car is going to break down or not. So
  43254. 32:54:40that is a discrete value. The y that we
  43255. 32:54:42are talking about the dependent variable
  43256. 32:54:44that we are talking about what we are
  43257. 32:54:46looking at is whether the car is going
  43258. 32:54:48to break down or not yes or no that is
  43259. 32:54:51what we are talking about. So here the
  43260. 32:54:53outcome is discrete and not a continuous
  43261. 32:54:56value. So this is how the logistic
  43262. 32:54:58regression curve looks. Let me explain a
  43263. 32:55:00little bit what exactly and how exactly
  43264. 32:55:03we are going to uh determine the class
  43265. 32:55:07the outcome rather. So for a logistic
  43266. 32:55:10regression curve a threshold has to be
  43267. 32:55:12set saying that because this is a
  43268. 32:55:14probability calculation remember this is
  43269. 32:55:16a probability calculation and the
  43270. 32:55:19probability itself will not be zero or
  43271. 32:55:22one but based on the probability we need
  43272. 32:55:25to decide what the outcome should be. So
  43273. 32:55:28there has to be a threshold like for
  43274. 32:55:29example 0.5 can be the threshold let's
  43275. 32:55:32say in this case. So any value of the
  43276. 32:55:35probability below 0.5 is considered to
  43277. 32:55:38be zero and any value above.5 is
  43278. 32:55:41considered to be one. So an output of
  43279. 32:55:44let's say8
  43280. 32:55:46will mean that the car will break down.
  43281. 32:55:48So that is considered as an output of 1
  43282. 32:55:51and let's say an output of 29 is
  43283. 32:55:54considered as zero which means that the
  43284. 32:55:57car will not break down. So that's the
  43285. 32:55:59way logistic regression works. Now let's
  43286. 32:56:02do a quick comparison between logistic
  43287. 32:56:04regression and linear regression because
  43288. 32:56:06they both have the term regression in
  43289. 32:56:09them. So it can cause confusion. So
  43290. 32:56:11let's try to remove that confusion. So
  43291. 32:56:13what is linear regression? Linear
  43292. 32:56:15regression is a process is once again an
  43293. 32:56:18algorithm for supervised learning.
  43294. 32:56:21However, here you're going to find a
  43295. 32:56:24continuous value. You're going to
  43296. 32:56:25determine a continuous value. It could
  43297. 32:56:27be the price of a real estate property.
  43298. 32:56:30It could be your hike, how much hike
  43299. 32:56:32you're going to get or it could be a
  43300. 32:56:33stock price. These are all continuous
  43301. 32:56:36values. These are not discrete compared
  43302. 32:56:38to a yes or a no kind of a response that
  43303. 32:56:40we are looking for in logistic
  43304. 32:56:42regression. So this is one example of a
  43305. 32:56:44linear regression. Let's say the HR team
  43306. 32:56:47of a company tries to find out what
  43307. 32:56:49should be the salary hike of an
  43308. 32:56:51employee. So they collect all the
  43309. 32:56:53details of their existing employees,
  43310. 32:56:55their ratings and their salary hikes,
  43311. 32:56:57what has been given and that is the
  43312. 32:56:59labeled information that is available
  43313. 32:57:01and the system learns from this. It is
  43314. 32:57:04trained and it learns from this labeled
  43315. 32:57:06information so that when a new employees
  43316. 32:57:09information is fed based on the rating
  43317. 32:57:11it will determine what should be the
  43318. 32:57:13high. So this is a linear regression
  43319. 32:57:15problem and a linear regression example.
  43320. 32:57:18Now salary is a continuous value. You
  43321. 32:57:21can get 5,000, 5,500,
  43322. 32:57:255,600. It is not discrete like a cat or
  43323. 32:57:29a dog or an apple or a banana. These are
  43324. 32:57:32discrete or a yes or a no. These are
  43325. 32:57:34discrete values, right? So this where
  43326. 32:57:37you're trying to find continuous values
  43327. 32:57:40is where we use linear regression. So
  43328. 32:57:42let's say just to extend on this
  43329. 32:57:44scenario, we now want to find out
  43330. 32:57:47whether this employee is going to get a
  43331. 32:57:49promotion or not. So we want to find out
  43332. 32:57:53that is a discrete problem, right? A yes
  43333. 32:57:55or no kind of a problem. In this case,
  43334. 32:57:58we actually cannot use linear regression
  43335. 32:58:01even though we may have labeled data. So
  43336. 32:58:04this is the label data. So based on the
  43337. 32:58:06employee rating these are the ratings
  43338. 32:58:08and then some people got the promotion
  43339. 32:58:11and this is the ratings for which people
  43340. 32:58:13did not get promotion that is a no and
  43341. 32:58:16this is the rating for which people got
  43342. 32:58:18promotion we just plotted the data about
  43343. 32:58:21whether a person has got an employee has
  43344. 32:58:23got promotion or not yes no right so
  43345. 32:58:26there is nothing in between and what is
  43346. 32:58:28the employees rating okay and ratings
  43347. 32:58:30can be continuous that is not an issue
  43348. 32:58:32but the output is discrete in In this
  43349. 32:58:35case whether employee got promotion yes
  43350. 32:58:37no okay so if we try to plot that and we
  43351. 32:58:40try to find a straight line this is how
  43352. 32:58:43it would look and as you can see it
  43353. 32:58:45doesn't look very right because looks
  43354. 32:58:47like there will be lot of errors this
  43355. 32:58:49root mean square error if you remember
  43356. 32:58:51for linear regression would be very very
  43357. 32:58:54high and also the the values cannot go
  43358. 32:58:57beyond zero or beyond one. So the graph
  43359. 32:59:00should probably look somewhat like this
  43360. 32:59:02clipped at 0 and one. But still the
  43361. 32:59:06straight line doesn't look right.
  43362. 32:59:09Therefore instead of using a linear
  43363. 32:59:12equation we need to come up with
  43364. 32:59:14something different and therefore the
  43365. 32:59:16logistic regression model looks somewhat
  43366. 32:59:19like this. So we calculate the
  43367. 32:59:21probability and if we plot that
  43368. 32:59:23probability not in the form of a
  43369. 32:59:25straight line but we need to use some
  43370. 32:59:27other equation. And we will see very
  43371. 32:59:28soon what that equation is. Then it is a
  43372. 32:59:31gradual process. Right? So you see here
  43373. 32:59:34people with some of these ratings are
  43374. 32:59:35not getting any promotions and then
  43375. 32:59:38slowly uh at certain rating they get
  43376. 32:59:41promotion. So that is a gradual process
  43377. 32:59:44and uh this is how the math behind
  43378. 32:59:47logistic regression looks. So we are
  43379. 32:59:49trying to find the odds for a particular
  43380. 32:59:53event happening and this is the formula
  43381. 32:59:55for finding the odds. So the probability
  43382. 32:59:57of an event happening divided by the
  43383. 33:00:00probability of the event not happening.
  43384. 33:00:03So P if it is the probability of the
  43385. 33:00:05event happening probability of the
  43386. 33:00:07person getting a promotion and divided
  43387. 33:00:09by the probability of the person not
  43388. 33:00:11getting a promotion that is 1 minus P.
  43389. 33:00:15So this is how you measure the odds. Now
  43390. 33:00:17the values of the odds range from 0 to
  43391. 33:00:21infinity. So when this probability is
  43392. 33:00:24zero then the odds will the value of the
  43393. 33:00:27odds is equal to zero and when the
  43394. 33:00:29probability becomes 1 then the value of
  43395. 33:00:32the odds is 1 by 0 that will be infinity
  43396. 33:00:35but the probability itself remains
  43397. 33:00:37between 0 and 1. Now this is how an
  43398. 33:00:40equation of a straight line looks. So y
  43399. 33:00:42is equal to beta 0 plus beta 1x where
  43400. 33:00:45beta 0 is the y intercept and beta 1 is
  43401. 33:00:48the slope of the line. If we take the
  43402. 33:00:51odds equation and take a log of both
  43403. 33:00:54sides, then this would look somewhat
  43404. 33:00:56like this. And the term logistic is
  43405. 33:00:59actually derived from the fact that we
  43406. 33:01:01are doing this. We take a log of px by 1
  43407. 33:01:04minus px. This is an extension of the
  43408. 33:01:07calculation of odds that we have seen,
  43409. 33:01:09right? And that is equal to beta 0 plus
  43410. 33:01:12beta 1x which is the equation of the
  43411. 33:01:13straight line. And now from here if you
  43412. 33:01:16want to find out the value of px we will
  43413. 33:01:20see we can take the exponential on both
  43414. 33:01:22sides and then if we solve that equation
  43415. 33:01:25we will get the equation of px like this
  43416. 33:01:28px is equal to 1 by 1 + e ^ of minus
  43417. 33:01:33beta 0 + beta 1x and recall this is
  43418. 33:01:36nothing but the equation of the line
  43419. 33:01:37which is equal to y is equal to beta 0 +
  43420. 33:01:40beta 1x. So that this is the equation
  43421. 33:01:44also known as the sigmoid function and
  43422. 33:01:46this is the equation of the logistic
  43423. 33:01:48regression alg. All right and if this is
  43424. 33:01:50plotted this is how the sigmoid curve is
  43425. 33:01:53obtained. So let's compare linear and
  43426. 33:01:57logistic regression how they are
  43427. 33:01:59different from each other. Let's go
  43428. 33:02:00back. So linear regression is solved or
  43429. 33:02:03used to solve regression problems and
  43430. 33:02:07logistic regression is used to solve
  43431. 33:02:10classification problems. So both are
  43432. 33:02:12called regression. But linear regression
  43433. 33:02:14is used for solving regression problems
  43434. 33:02:17where we predict continuous values.
  43435. 33:02:19Whereas logistic regression is used for
  43436. 33:02:22solving classification problems where we
  43437. 33:02:24have had to predict discrete values. The
  43438. 33:02:27response variables in case of linear
  43439. 33:02:29regression are continuous in nature.
  43440. 33:02:31Whereas here they are categorical or
  43441. 33:02:33discrete in nature. And the linear
  43442. 33:02:36regression helps to estimate the
  43443. 33:02:38dependent variable when there is a
  43444. 33:02:40change in the independent variable.
  43445. 33:02:42Whereas here in case of logistic
  43446. 33:02:44regression it helps to calculate the
  43447. 33:02:46probability or the possibility of a
  43448. 33:02:49particular event happening. And linear
  43449. 33:02:51regression as the name suggests is a
  43450. 33:02:53straight line. That's why it's called
  43451. 33:02:54linear regression. Whereas logistic
  43452. 33:02:57regression is a sigmoid function and the
  43453. 33:03:00curve is the shape of the curve is S.
  43454. 33:03:03It's an S-shaped curve. This is another
  43455. 33:03:05example of application of logistic
  43456. 33:03:07regression in weather prediction.
  43457. 33:03:09Whether it's going to rain or not rain.
  43458. 33:03:12Now keep in mind both are used in
  43459. 33:03:14weather prediction. If we want to find
  43460. 33:03:16the discrete values like whether it's
  43461. 33:03:18going to rain or not rain that is a
  43462. 33:03:21classification problem. We use logistic
  43463. 33:03:23regression. But if we want to determine
  43464. 33:03:25what is going to be the temperature
  43465. 33:03:26tomorrow, then we use linear regression.
  43466. 33:03:29So just keep in mind that in weather
  43467. 33:03:31prediction, we actually use both. But
  43468. 33:03:33these are some examples of logistic
  43469. 33:03:34regression. So we want to find out
  43470. 33:03:36whether it's going to be rain or not,
  43471. 33:03:38it's going to be sunny or not, whether
  43472. 33:03:40it's going to snow or not. These are all
  43473. 33:03:42logistic regression examples. A few more
  43474. 33:03:44examples. Classification of objects.
  43475. 33:03:47This is a again another example of
  43476. 33:03:49logistic regression. Now here of course
  43477. 33:03:52one distinction is that these are
  43478. 33:03:54multiclass classification. So logistic
  43479. 33:03:57regression is not used in its original
  43480. 33:04:01form but it is used in a slightly
  43481. 33:04:03different form. So we say whether it is
  43482. 33:04:05a dog or not a dog. I hope you
  43483. 33:04:08understand. So instead of saying is it a
  43484. 33:04:09dog or a cat or elephant we convert this
  43485. 33:04:12into saying so because we need to keep
  43486. 33:04:15it to binary classification. So we say
  43487. 33:04:18is it a dog or not a dog? Is it a cat or
  43488. 33:04:22not a cat? So that's the way logistic
  43489. 33:04:24regression can be used for classifying
  43490. 33:04:26objects. Otherwise there are other
  43491. 33:04:28techniques which can be used for
  43492. 33:04:30performing multiclass classification. In
  43493. 33:04:32healthcare logistic regression is used
  43494. 33:04:34to find the survival rate of a patient.
  43495. 33:04:37So they take multiple parameters like
  43496. 33:04:40trauma score and age and so on and so
  43497. 33:04:42forth and they try to predict the rate
  43498. 33:04:45of survival. All right. Now finally
  43499. 33:04:48let's take an example and see how we can
  43500. 33:04:50apply logistic regression to predict the
  43501. 33:04:53number that is shown in the image. So
  43502. 33:04:56this is actually a live demo. I will
  43503. 33:04:58take you into Jupyter notebook and u
  43504. 33:05:01show the code. But before that let me
  43505. 33:05:03take you through a couple of slides to
  43506. 33:05:05explain what we're trying to do. So
  43507. 33:05:07let's say you have an 8x8 image and the
  43508. 33:05:10the image has a number 1 2 3 4 and you
  43509. 33:05:13need to train your model to predict what
  43510. 33:05:16this number is. So how do we do this? So
  43511. 33:05:19the first thing is obviously in any
  43512. 33:05:21machine learning process you train your
  43513. 33:05:23model. So in this case we are using
  43514. 33:05:25logistic regression. So and then we
  43515. 33:05:28provide a training set to train the
  43516. 33:05:30model and then we test how accurate our
  43517. 33:05:33model is with the test data which means
  43518. 33:05:36that like any machine learning process.
  43519. 33:05:38We split our initial data into two parts
  43520. 33:05:41training set and test set. With the
  43521. 33:05:42training set we train our model and then
  43522. 33:05:44with the test set we we test the model
  43523. 33:05:46till we get good accuracy and then we
  43524. 33:05:49use it for for inference. Right? So that
  43525. 33:05:52is typical methodology of uh uh
  43526. 33:05:56training, testing and then deploying of
  43527. 33:05:58machine learning models. So let's uh
  43528. 33:06:00take a look at the code and uh see what
  43529. 33:06:02we are doing. So I'll not go line by
  43530. 33:06:04line but just take you through some of
  43531. 33:06:06the blocks. So first thing we do is
  43532. 33:06:08import all the libraries and then we
  43533. 33:06:11basically take a look at the images and
  43534. 33:06:14see what is the total number of images.
  43535. 33:06:16We can display using mattplot lip some
  43536. 33:06:19of the images or a sample of these
  43537. 33:06:21images and um then we split the data
  43538. 33:06:24into training and test as I mentioned
  43539. 33:06:25earlier and we can do some exploratory
  43540. 33:06:28analysis and uh then we build our model.
  43541. 33:06:32We train our model with the training set
  43542. 33:06:34and then we test it with our test set
  43543. 33:06:37and find out how accurate our model is
  43544. 33:06:40using the confusion matrix the heat map
  43545. 33:06:43and use heat map for visualizing this
  43546. 33:06:45and uh I will show you in the code what
  43547. 33:06:48exactly is the confusion matrix and how
  43548. 33:06:50it can be used for finding the accuracy
  43549. 33:06:53in our example we got we get an accuracy
  43550. 33:06:56of about 94 which is pretty good or 94%
  43551. 33:06:59which is pretty good all right so what
  43552. 33:07:01is the confusion matrix. This is an
  43553. 33:07:03example of a confusion matrix and uh
  43554. 33:07:06this is used for identifying the
  43555. 33:07:09accuracy of a classification model or
  43556. 33:07:12like a logistic regression model. So the
  43557. 33:07:15most important part in a confusion
  43558. 33:07:17matrix is that first of all this as you
  43559. 33:07:19can see this is a matrix and the size of
  43560. 33:07:22the matrix depends on how many outputs
  43561. 33:07:24uh we are expecting right. So the the
  43562. 33:07:27most important part here is that the
  43563. 33:07:29model will be most accurate when we have
  43564. 33:07:32the maximum numbers in its diagonal like
  43565. 33:07:36in this case that's why it has almost 93
  43566. 33:07:3994% because the diagonals should have
  43567. 33:07:42the maximum numbers and the others other
  43568. 33:07:45than diagonals the cells other than the
  43569. 33:07:47diagonals should have very few numbers.
  43570. 33:07:50So here that's what is happening. So
  43571. 33:07:51there is a two here. There are there's a
  43572. 33:07:53one here. But most of them are along the
  43573. 33:07:57diagonal. This what does this mean? This
  43574. 33:07:59means that the number that has been fed
  43575. 33:08:03is zero and the number that has been
  43576. 33:08:06detected is also zero. So the predicted
  43577. 33:08:09value and the actual value are the same.
  43578. 33:08:11So along the diagonals that is true.
  43579. 33:08:14Which means that let's let's take this
  43580. 33:08:16diagonal right. If if the maximum number
  43581. 33:08:18is here that means that uh like here in
  43582. 33:08:21this case it is 34 which means that 34
  43583. 33:08:24of the images that have been fed or
  43584. 33:08:26rather actually there are two
  43585. 33:08:27mclassifications in there. So 36 images
  43586. 33:08:31have been fed which have number four and
  43587. 33:08:33out of which 34 have been predicted
  43588. 33:08:36correctly as number four and one has
  43589. 33:08:39been predicted as number eight and
  43590. 33:08:41another one has been predicted as number
  43591. 33:08:43nine. So these are two mclassifications.
  43592. 33:08:46Okay. So that is the meaning of saying
  43593. 33:08:48that the maximum number should be in the
  43594. 33:08:50diagonal. So if you have all of them so
  43595. 33:08:52for an ideal model which has let's say
  43596. 33:08:55100% accuracy everything will be only in
  43597. 33:08:58the diagonal. There will be no numbers
  43598. 33:09:00other than zero in all other cells. So
  43599. 33:09:02that is like a 100% accurate model.
  43600. 33:09:05Okay. So that's the gist of how to use
  43601. 33:09:08this matrix. How to use this uh
  43602. 33:09:10confusion matrix. So I know the name uh
  43603. 33:09:13is a little funny sounding confusion
  43604. 33:09:15matrix but actually it is not very
  43605. 33:09:17confusing. It's very straightforward. So
  43606. 33:09:19you are just plotting what has been
  43607. 33:09:21predicted and what is the labeled
  43608. 33:09:24information or what is the actual data
  43609. 33:09:26that's also known as the ground truth
  43610. 33:09:28sometimes. Okay, these are some fancy
  43611. 33:09:29terms that are used. So predicted label
  43612. 33:09:31and the actual label that's all it is.
  43613. 33:09:34Okay. Yeah. So we are showing a little
  43614. 33:09:35bit more information here. So 38 have
  43615. 33:09:38been predicted and here you will see
  43616. 33:09:40that all of them have been predicted
  43617. 33:09:41correctly. There have been 38 zeros and
  43618. 33:09:44the predicted value and the actual value
  43619. 33:09:47is is exactly the same. Whereas in this
  43620. 33:09:49case right it has uh there are I think
  43621. 33:09:5237 + 5 yeah 42 have been fed the images
  43622. 33:09:5742 images are of digit three and uh the
  43623. 33:10:01accuracy is only 37 of them have been
  43624. 33:10:04accurately predicted. Three of them have
  43625. 33:10:07been predicted as number seven and two
  43626. 33:10:09of them have been predicted as number
  43627. 33:10:11eight and so on and so forth. Okay. All
  43628. 33:10:13right. So with that let's go into
  43629. 33:10:15Jupyter notebook and see how the code
  43630. 33:10:17looks. So this is the code in in Jupyter
  43631. 33:10:22notebook for logistic regression. In
  43632. 33:10:25this particular demo, what we are going
  43633. 33:10:27to do is train our model to recognize
  43634. 33:10:31digits which are the images which have
  43635. 33:10:34digits from let's say 0 to 5 or 0 to 9
  43636. 33:10:38and um and then we will see how well it
  43637. 33:10:41is trained and whether it is able to
  43638. 33:10:43predict these numbers correctly or not.
  43639. 33:10:46So let's get started. So the first part
  43640. 33:10:48is as usual we are importing some
  43641. 33:10:51libraries that are required and uh then
  43642. 33:10:55the last line in this block is to load
  43643. 33:10:58the digits. So let's go ahead and run
  43644. 33:11:02this code. Then here we will visualize
  43645. 33:11:06the shape of these uh digits. So we can
  43646. 33:11:08see here if we take a look this is how
  43647. 33:11:11the shape is 1797 by 64. These are like
  43648. 33:11:158 by8 images. So that's that's what is
  43649. 33:11:17reflected in this uh shape. Now from
  43650. 33:11:20here onwards we are basically once again
  43651. 33:11:22importing some of the libraries that are
  43652. 33:11:24required like numpy and map plot and we
  43653. 33:11:27will take a look at uh some of the
  43654. 33:11:29sample images that we have loaded. So th
  43655. 33:11:33this one for example creates a figure uh
  43656. 33:11:36and then we go ahead and take a few
  43657. 33:11:38sample images to see how they look. So
  43658. 33:11:41let me run this code and so that it
  43659. 33:11:43becomes easy to understand. So these are
  43660. 33:11:46about five images sample images that we
  43661. 33:11:48are looking at 0 1 2 3 4. So this is how
  43662. 33:11:52the images this is how the data is.
  43663. 33:11:54Okay. And uh based on this we will
  43664. 33:11:57actually train our logistic regression
  43665. 33:12:00model and then we will test it and see
  43666. 33:12:03how well it is able to recognize. So the
  43667. 33:12:06way it works is the pixel information.
  43668. 33:12:09So as you can see here this is an 8x 8
  43669. 33:12:12pixel kind of a image and uh the each
  43670. 33:12:16pixel whether it is activated or not
  43671. 33:12:19activated that is the information
  43672. 33:12:20available for each pixel. Now based on
  43673. 33:12:23the pattern of this activation and
  43674. 33:12:26non-activation of the various pixels
  43675. 33:12:28this will be identified as a zero for
  43676. 33:12:31example right similarly as you can see
  43677. 33:12:34so overall each of these numbers
  43678. 33:12:37actually has a different pattern of the
  43679. 33:12:40pixel activation and that's pretty much
  43680. 33:12:42that our model needs to learn for which
  43681. 33:12:45number what is the pattern of the
  43682. 33:12:47activation of the pixels right so that
  43683. 33:12:50is what we are going to train our model.
  43684. 33:12:52Okay. So the first thing we need to do
  43685. 33:12:55is to split our data into training and
  43686. 33:12:59test data set. Right? So whenever we
  43687. 33:13:01perform any training, we split the data
  43688. 33:13:04into training and test. So that the
  43689. 33:13:06training data set is used to train the
  43690. 33:13:08system. So we pass this probably
  43691. 33:13:11multiple times. Uh and then we test it
  43692. 33:13:14with the test data set. And the split is
  43693. 33:13:16usually in the form of there and there
  43694. 33:13:18are various ways in which you can split
  43695. 33:13:20this data. It is up to the individual
  43696. 33:13:22preferences. In our case here we are
  43697. 33:13:25splitting in the form of 23 and 77. So
  43698. 33:13:29when we say test size as 2023
  43699. 33:13:33that means 23% of the entire data is
  43700. 33:13:38used for testing and the remaining 77%
  43701. 33:13:41is used for training. So there is a
  43702. 33:13:43readily available function which is uh
  43703. 33:13:46called train test split. So we don't
  43704. 33:13:49have to write any special code for the
  43705. 33:13:51splitting. It will automatically split
  43706. 33:13:53the data based on the proportion that we
  43707. 33:13:56give here which is test size. So we just
  43708. 33:13:58give the test size automatically
  43709. 33:14:00training size will be determined and uh
  43710. 33:14:02we pass the data that we want to split
  43711. 33:14:05and the the results will be stored in x
  43712. 33:14:09train and y train for the training data
  43713. 33:14:13set. And what is x train? This are these
  43714. 33:14:16are the features right which is like the
  43715. 33:14:18independent variable and y train is the
  43716. 33:14:22label right so in this case what happens
  43717. 33:14:25is we have the input value which is or
  43718. 33:14:28the features value which is in x train
  43719. 33:14:30and since this is a labeled data for
  43720. 33:14:33each of them each of the observations we
  43721. 33:14:36already have the label information
  43722. 33:14:38saying whether this digit is a zero or a
  43723. 33:14:41one or a two so that this this is what
  43724. 33:14:43will be used for comparison to find out
  43725. 33:14:46whether the the system is able to
  43726. 33:14:48recognize it correctly or there is an
  43727. 33:14:50error for each observation it will
  43728. 33:14:52compare with this right so this is the
  43729. 33:14:54label so the same way x train y train is
  43730. 33:14:58for the training data set x test y test
  43731. 33:15:02is for the test data set okay so let me
  43732. 33:15:05go ahead and execute this code as well
  43733. 33:15:06and then we can go and check quickly
  43734. 33:15:09what is the how many entries are there
  43735. 33:15:11and in each of this so x train the shape
  43736. 33:15:14is 1383x
  43737. 33:15:1764 and y train has 1383 because there is
  43738. 33:15:22uh nothing like the second part is not
  43739. 33:15:24required here and then x test shape we
  43740. 33:15:28see is 414 so actually there are 414
  43741. 33:15:32observations in test and 1383
  43742. 33:15:35observations in train so that's
  43743. 33:15:37basically what these four lines of code
  43744. 33:15:39are are saying okay then we import the
  43745. 33:15:42uh logistic regression
  43746. 33:15:44library and uh which is a part of
  43747. 33:15:47scikitlearn. So we we don't have to
  43748. 33:15:50implement the logistic regression
  43749. 33:15:51process itself. We just call these uh
  43750. 33:15:53the function and uh let me go ahead and
  43751. 33:15:56execute that so that uh we have the
  43752. 33:15:59logistic regression library imported.
  43753. 33:16:01Now we create an instance of logistic
  43754. 33:16:04regression. Right? So logistic regr is a
  43755. 33:16:07is an instance of logistic regression
  43756. 33:16:10and then we use that for training our
  43757. 33:16:12model. So let me first execute this
  43758. 33:16:15code. So these two lines. So the first
  43759. 33:16:17line basically creates an instance of
  43760. 33:16:19logistic regression model and then the
  43761. 33:16:22second line is where we are passing our
  43762. 33:16:25data the training data set. Right? This
  43763. 33:16:27is our the the predictors and uh this is
  43764. 33:16:31our target. We are passing this data set
  43765. 33:16:33to train our model. All right. So once
  43766. 33:16:36we do this in this case the data is not
  43767. 33:16:39large but by and large uh the training
  43768. 33:16:42is what takes usually a lot of time. So
  43769. 33:16:44we spend in machine learning activities
  43770. 33:16:47in machine learning projects we spend a
  43771. 33:16:50lot of time for the training part of it.
  43772. 33:16:52Okay. So here the data set is relatively
  43773. 33:16:54small so it was pretty quick. So all
  43774. 33:16:56right so now our model has been trained
  43775. 33:16:59using the training data set and uh we
  43776. 33:17:02want to see how accurate this is. So
  43777. 33:17:04what we'll do is we will test it out in
  43778. 33:17:07probably faces. So let me first try out
  43779. 33:17:10how well this is working for one image.
  43780. 33:17:14Okay, I will just try it out with one
  43781. 33:17:16image my the first entry in my test data
  43782. 33:17:19set and see whether it is uh correctly
  43783. 33:17:21predicting or not. So and in order to
  43784. 33:17:24test it so for training purpose we use
  43785. 33:17:26the fit method. There is a method called
  43786. 33:17:29fit which is for training the model and
  43787. 33:17:32once the training is done if you want to
  43788. 33:17:34test for uh a particular value new input
  43789. 33:17:37you use the predict method. Okay. So
  43790. 33:17:39let's run the predict method and we pass
  43791. 33:17:42this particular image and uh we see that
  43792. 33:17:46the shape is or the prediction is four.
  43793. 33:17:50So let's try a few more. Let me see for
  43794. 33:17:53the next 10 uh seems to be fine. So let
  43795. 33:17:56me just go ahead and test the entire
  43796. 33:17:58data set. Okay, that's basically what we
  43797. 33:18:00will do. So now we want to find out how
  43798. 33:18:02accurately this has u performed. So we
  43799. 33:18:06use the score method to find what is the
  43800. 33:18:09percentages of accuracy and we see here
  43801. 33:18:11that it has performed up to 94%
  43802. 33:18:14accurate. Okay. So that's uh on this
  43803. 33:18:17part. Now what we can also do is we can
  43804. 33:18:20um also see this accuracy using what is
  43805. 33:18:23known as confusion matrix. So let us go
  43806. 33:18:26ahead and uh try that as well. Uh so
  43807. 33:18:29that we can also visualize how well uh
  43808. 33:18:32this model has uh done. So let me
  43809. 33:18:35execute this piece of code which will
  43810. 33:18:37basically import some of the libraries
  43811. 33:18:39that are required and um we we basically
  43812. 33:18:42create a confusion matrix an instance of
  43813. 33:18:46confusion matrix by running confusion
  43814. 33:18:48matrix and passing these uh values. So
  43815. 33:18:52we have so this confusion_matrix
  43816. 33:18:55method takes two parameters one is the y
  43817. 33:18:59test and the other is uh the prediction.
  43818. 33:19:02So what is a y test? These are the
  43819. 33:19:04labeled values which we already know for
  43820. 33:19:06the test data set and predictions are
  43821. 33:19:10what the system has predicted for the
  43822. 33:19:13test data set. Okay. So this is known to
  43823. 33:19:16us and this is what the system has uh
  43824. 33:19:19the model has generated. So we kind of
  43825. 33:19:22create the confusion matrix and we will
  43826. 33:19:24print it. And uh this is how the
  43827. 33:19:26confusion matrix looks. As the name
  43828. 33:19:28suggests it is a matrix and um the key
  43829. 33:19:32point out here is that the accuracy of
  43830. 33:19:34the model is determined by how many
  43831. 33:19:38numbers are there in the diagonal. The
  43832. 33:19:40more the numbers in the diagonal, the
  43833. 33:19:43better the accuracy is. Okay. And first
  43834. 33:19:46of all, the total sum of all the numbers
  43835. 33:19:48in this whole matrix is equal to the
  43836. 33:19:50number of observations in the test data
  43837. 33:19:53set. That is the first thing, right? So
  43838. 33:19:55if you add up all these numbers, that
  43839. 33:19:57will be equal to the number of
  43840. 33:19:59observations in the test data set. And
  43841. 33:20:01then out of that, the maximum number of
  43842. 33:20:04them should be in the diagonal. That
  43843. 33:20:06means the accuracy is pretty good. If
  43844. 33:20:08the the numbers in the diagonal are less
  43845. 33:20:10and in all other places there are a lot
  43846. 33:20:12of numbers uh which means the accuracy
  43847. 33:20:15is very low. The diagonal indicates a
  43848. 33:20:17correct prediction that this means that
  43849. 33:20:19the actual value is same as the
  43850. 33:20:22predicted value. Here again actual value
  43851. 33:20:24is same as the predicted value and so
  43852. 33:20:25on. Right? So the moment you see a
  43853. 33:20:27number here that means the actual value
  43854. 33:20:29is something and the predicted value is
  43855. 33:20:32something else. Right? Similarly here
  43856. 33:20:34the actual value is something and the
  43857. 33:20:36predicted value is something else. So
  43858. 33:20:38that is basically how we read the
  43859. 33:20:41confusion matrix. Now how do we find the
  43860. 33:20:45accuracy? You can actually add up the
  43861. 33:20:47total values in the diagonal. So it it's
  43862. 33:20:50like 38 + 44 + 43 and so on and divide
  43863. 33:20:54that by the total number of test
  43864. 33:20:56observations that will give you the
  43865. 33:20:58percentage accuracy using a confusion
  43866. 33:21:01matrix. Now let us visualize this
  43867. 33:21:03confusion matrix in a slightly more
  43868. 33:21:06sophisticated way uh using a heat map.
  43869. 33:21:09So we will create a heat map with some
  43870. 33:21:11we'll add some colors as well. It's uh
  43871. 33:21:14it's like a more visually visually more
  43872. 33:21:17appealing. So that's the whole idea. So
  43873. 33:21:19if we let me run this piece of code and
  43874. 33:21:21this is how the heat map looks. Uh and
  43875. 33:21:25as you can see here the diagonals again
  43876. 33:21:28are all the values are here most of the
  43877. 33:21:30values. So which means reasonably this
  43878. 33:21:32seems to be reasonably accurate and yeah
  43879. 33:21:35basically the accuracy score is 94%.
  43880. 33:21:38This is calculated as I mentioned by
  43881. 33:21:40adding all these numbers divided by the
  43882. 33:21:43total test values or the total number of
  43883. 33:21:45observations in test data set. Okay. So
  43884. 33:21:49this is the confusion matrix for
  43885. 33:21:51logistic regression.
  43886. 33:21:54All right. So now that we have seen the
  43887. 33:21:57confusion matrix, let's take a quick
  43888. 33:22:00sample and see how well uh the system
  43889. 33:22:02has classified and we will take a a few
  43890. 33:22:05examples of the data. So if we see here
  43891. 33:22:09we we picked up randomly a few of them.
  43892. 33:22:11So this is uh number four which is the
  43893. 33:22:14actual value and also the predicted
  43894. 33:22:16value both are four. This is an image of
  43895. 33:22:19zero. So the predicted value is also
  43896. 33:22:22zero. Actual value is of course zero.
  43897. 33:22:24Then this is the image of nine. So this
  43898. 33:22:27has also been predicted correctly 9 and
  43899. 33:22:30actual value is 9. And this is the image
  43900. 33:22:32of one. And again this has been
  43901. 33:22:34predicted correctly as like the actual
  43902. 33:22:37value. Okay. So this was a quick demo of
  43903. 33:22:40logistic regression. How to use logistic
  43904. 33:22:43regression to identify images.
  43905. 33:22:46>> What is a decision tree? Let's go
  43906. 33:22:48through a very simple example before we
  43907. 33:22:50dig in deep. Decision tree is a
  43908. 33:22:52treeshaped diagram used to determine a
  43909. 33:22:54course of action. Each branch of the
  43910. 33:22:56tree represents a possible decision or
  43911. 33:22:58occurrence or reaction. Let's start with
  43912. 33:23:00a simple question. How to identify a
  43913. 33:23:02random vegetable from a shopping bag?
  43914. 33:23:04So, we have this group of vegetables in
  43915. 33:23:06here. And we can start off by asking a
  43916. 33:23:07simple question. Is it red? And if it's
  43917. 33:23:09not, then it's going to be the purple
  43918. 33:23:12fruit to the left, probably an eggplant.
  43919. 33:23:14If it's true, it's going to be one of
  43920. 33:23:15the red fruits. Is the diameter greater
  43921. 33:23:17than two? If false, it's going to be a
  43922. 33:23:20what looks to be a red chili. And if
  43923. 33:23:21it's true, it's going to be a bell
  43924. 33:23:24pepper from the capsicum family. So,
  43925. 33:23:26it's a capsicum.
  43926. 33:23:28Problems that decision tree can solve.
  43927. 33:23:30So, let's look at the two different
  43928. 33:23:31categories the decision tree can be used
  43929. 33:23:33on. It can be used on the
  43930. 33:23:35classification, the true false, yes, no,
  43931. 33:23:37and it can be used on regression where
  43932. 33:23:39we figure out what the next value is in
  43933. 33:23:41a series of numbers or a group of data.
  43934. 33:23:43In classification, the classification
  43935. 33:23:45tree will determine a set of logical if
  43936. 33:23:49then conditions to classify problems.
  43937. 33:23:51For example, discriminating between
  43938. 33:23:53three types of flowers based on certain
  43939. 33:23:55features. In regression, a regression
  43940. 33:23:58tree is used when the target variable is
  43941. 33:24:00numerical or continuous in nature. We
  43942. 33:24:02fit the regression model to the target
  43943. 33:24:04variable using each of the independent
  43944. 33:24:06variables. Each split is made based on
  43945. 33:24:08the sum of squared error. Before we dig
  43946. 33:24:12deeper into the mechanics of the
  43947. 33:24:14decision tree, let's take a look at the
  43948. 33:24:16advantages of using a decision tree and
  43949. 33:24:18we'll also take a glimpse at the
  43950. 33:24:20disadvantages. The first thing you'll
  43951. 33:24:22notice is that it's simple to
  43952. 33:24:24understand, interpret, and visualize. It
  43953. 33:24:27really shines here because you can see
  43954. 33:24:28exactly what's going on in a decision
  43955. 33:24:30tree. Little effort is required for data
  43956. 33:24:32preparation. So, you don't have to do
  43957. 33:24:34special scaling. There's a lot of things
  43958. 33:24:36you don't have to worry about when using
  43959. 33:24:37a decision tree. It can handle both
  43960. 33:24:39numerical and categorical data as we
  43961. 33:24:42discovered earlier and nonlinear
  43962. 33:24:44parameters don't affect its performance.
  43963. 33:24:46So even if the data doesn't fit an easy
  43964. 33:24:49curved graph, you can still use it to
  43965. 33:24:51create an effective decision or
  43966. 33:24:54prediction. If we're going to look at
  43967. 33:24:56the advantages of a decision tree, we
  43968. 33:24:59also need to understand the
  43969. 33:25:00disadvantages of a decision tree. The
  43970. 33:25:02first disadvantage is overfitting.
  43971. 33:25:04Overfitting occurs when the algorithm
  43972. 33:25:06captures noise in the data. That means
  43973. 33:25:08you're solving for one specific instance
  43974. 33:25:10instead of a general solution for all
  43975. 33:25:13the data. High variance. The model can
  43976. 33:25:15get unstable due to small variation in
  43977. 33:25:17data. Low bias tree. A highly
  43978. 33:25:20complicated decision tree tends to have
  43979. 33:25:22a low bias which makes it difficult for
  43980. 33:25:24the model to work with new data.
  43981. 33:25:26Decision tree important terms. Before we
  43982. 33:25:30dive in further, we need to look at some
  43983. 33:25:32basic terms. We need to have some
  43984. 33:25:35definitions to go with our decision tree
  43985. 33:25:37in the different parts we're going to be
  43986. 33:25:38using. We'll start with entropy. Entropy
  43987. 33:25:40is a measure of randomness or
  43988. 33:25:42unpredictability in the data set. For
  43989. 33:25:45example, we have a group of animals in
  43990. 33:25:47this picture. There's four different
  43991. 33:25:48kinds of animals. And this data set is
  43992. 33:25:50considered to have a high entropy. You
  43993. 33:25:52really can't pick out what kind of
  43994. 33:25:54animal it is based on looking at just
  43995. 33:25:55the four animals as a big clump of of uh
  43996. 33:25:58entities. So as we start splitting it
  43997. 33:26:02into subgroups, we come up with our
  43998. 33:26:04second definition which is information
  43999. 33:26:07gain. Information gain it is a measure
  44000. 33:26:09of decrease in entropy after the data
  44001. 33:26:12set is split. So in this case based on
  44002. 33:26:14the color yellow, we've split one group
  44003. 33:26:16of animals on one side as true and those
  44004. 33:26:19who aren't yellow as false. As we
  44005. 33:26:21continue down the yellow side, we split
  44006. 33:26:23based on the height. True or false
  44007. 33:26:24equals 10. And on the other side, height
  44008. 33:26:26is less than 10. True or false? And as
  44009. 33:26:29you see as we split it, the entropy
  44010. 33:26:31continues to be less and less and less.
  44011. 33:26:33And so our information gain is simply
  44012. 33:26:35the entropy E1 from the top and how it's
  44013. 33:26:38changed to E2 in the bottom. And we'll
  44014. 33:26:40look at the uh deeper math, although you
  44015. 33:26:42really don't need to know a huge amount
  44016. 33:26:44of math when you actually do the
  44017. 33:26:45programming in Python because it'll do
  44018. 33:26:47it for you. But we'll look on the actual
  44019. 33:26:48math of how they compute entropy.
  44020. 33:26:50Finally, we want to know the different
  44021. 33:26:51parts of our tree and they call the leaf
  44022. 33:26:54node. Leaf node carries the
  44023. 33:26:55classification or the decision. So it's
  44024. 33:26:57the final end at the bottom. The
  44025. 33:26:59decision node has two or more branches.
  44026. 33:27:02This is where we're breaking the group
  44027. 33:27:03up into different parts. And finally,
  44028. 33:27:06you have the root node. The topmost
  44029. 33:27:08decision node is known as the root node.
  44030. 33:27:11How does a decision tree work? Wonder
  44031. 33:27:13what kind of animals I'll get in the
  44032. 33:27:15jungle today? Maybe you're the hunter
  44033. 33:27:17with the gun. Or if you're more into
  44034. 33:27:18photography, you're a photographer with
  44035. 33:27:20a camera. So let's look at this group of
  44036. 33:27:22animals and let's try to classify
  44037. 33:27:24different types of animals based on
  44038. 33:27:26their features using a decision tree. So
  44039. 33:27:28the problem statement is to classify the
  44040. 33:27:31different types of animals based on
  44041. 33:27:32their features using a decision tree.
  44042. 33:27:34The data set is looking quite messy and
  44043. 33:27:36the entropy is high in this case. So
  44044. 33:27:38let's look at a training set or a
  44045. 33:27:40training data set and we're looking at
  44046. 33:27:42color. We're looking at height and then
  44047. 33:27:44we have our different animals. We have
  44048. 33:27:46our elephants, our giraffes, our
  44049. 33:27:48monkeys, and our tigers. And they're of
  44050. 33:27:50different colors and shapes. Let's see
  44051. 33:27:52what that looks like. And how do we
  44052. 33:27:53split the data? We have to frame the
  44053. 33:27:55conditions that split the data in such a
  44054. 33:27:58way that the information gain is the
  44055. 33:28:00highest. Note, gain is the measure of
  44056. 33:28:02decrease in entropy after splitting. So
  44057. 33:28:05the formula for entropy is the sum
  44058. 33:28:08that's what this symbol looks like. That
  44059. 33:28:09looks like kind of like a uh e funky e
  44060. 33:28:12of k where i equals 1 to k. K would
  44061. 33:28:15represent the number of animal the
  44062. 33:28:17different animals in there where value
  44063. 33:28:19or P value of I would be the percentage
  44064. 33:28:23of that animal times the log base 2 of
  44065. 33:28:26the same the percentage of that animal.
  44066. 33:28:28Let's try to calculate the entropy for
  44067. 33:28:29the current data set and take a look at
  44068. 33:28:32what that looks like. And don't be
  44069. 33:28:33afraid of the math. You don't really
  44070. 33:28:35have to memorize this math. Just be
  44071. 33:28:37aware that it's there and this is what's
  44072. 33:28:38going on in the background. And so we
  44073. 33:28:40have three giraffes, two tigers, one
  44074. 33:28:43monkey, two elephants, a total of eight
  44075. 33:28:45animals gathered. And if we plug that
  44076. 33:28:46into the formula, we get an entropy that
  44077. 33:28:49equals 3 over8. So we have three
  44078. 33:28:51giraffes, a total of 8 times the log.
  44079. 33:28:54Usually they use base 2 on the log. So
  44080. 33:28:56log base 2 of 3 over8 plus in this case,
  44081. 33:29:00let's say it's the elephants, 2 over 8.
  44082. 33:29:02Two elephants over total of 8 time log
  44083. 33:29:04base 2 2 over 8 plus one monkey over
  44084. 33:29:07total of 8. log base 2 1 over 8 and plus
  44085. 33:29:112 over 8 of the tigers log base 2 over 8
  44086. 33:29:15and if we plug that into our computer or
  44087. 33:29:17calculator I obviously can't do logs in
  44088. 33:29:19my head we get an entropy equal to.571
  44089. 33:29:23the program will actually calculate the
  44090. 33:29:25entropy of the data set similarly after
  44091. 33:29:27every split to calculate the gain now
  44092. 33:29:30we're not going to go through each set
  44093. 33:29:31one at a time to see what those numbers
  44094. 33:29:34are just want you to be aware that this
  44095. 33:29:36is a formula or the mathematics behind
  44096. 33:29:38It gain can be calculated by finding the
  44097. 33:29:39difference of the subsequent entropy
  44098. 33:29:41values after a split. Now we will try to
  44099. 33:29:43choose a condition that gives us the
  44100. 33:29:45highest gain. We will do that by
  44101. 33:29:47splitting the data using each condition
  44102. 33:29:49and checking that the gain we get out of
  44103. 33:29:51them. The condition that gives us the
  44104. 33:29:52highest gain will be used to make the
  44105. 33:29:54first split. Can you guess what that
  44106. 33:29:56first split will be just by looking at
  44107. 33:29:58this image? As a human, it's probably
  44108. 33:30:00pretty easy to split it. Let's see if
  44109. 33:30:02you're right. If you guessed the color
  44110. 33:30:04yellow, you're correct. Let's say the
  44111. 33:30:06condition that gives us the maximum gain
  44112. 33:30:08is yellow. So we will split the data
  44113. 33:30:10based on the color yellow. If it's true,
  44114. 33:30:12that group of animals goes to the left.
  44115. 33:30:14If it's false, it goes to the right. The
  44116. 33:30:16entropy after the splitting has
  44117. 33:30:18decreased considerably. However, we
  44118. 33:30:21still need some splitting at both the
  44119. 33:30:23branches to attain an entropy value
  44120. 33:30:25equal to zero. So we decide to split
  44121. 33:30:27both the nodes using height as a
  44122. 33:30:29condition. Since every branch now
  44123. 33:30:30contains single label type, we can say
  44124. 33:30:33that entropy in this case has reached
  44125. 33:30:35the least value. And here you see we
  44126. 33:30:37have the giraffes, the tigers, the
  44127. 33:30:39monkey and the elephants all separated
  44128. 33:30:40into their own groups. This tree can now
  44129. 33:30:42predict all the classes of animals
  44130. 33:30:44present in the data set with 100%
  44131. 33:30:46accuracy. That was easy. Use case loan
  44132. 33:30:50repayment prediction. Let's get into my
  44133. 33:30:52favorite part and open up some Python
  44134. 33:30:54and see what the programming code and
  44135. 33:30:56the scripting looks like. In here, we're
  44136. 33:30:58going to want to do a prediction. And we
  44137. 33:31:00start with this individual here who's
  44138. 33:31:01requesting to find out how good his
  44139. 33:31:03customers are going to be, whether
  44140. 33:31:04they're going to repay their loan or not
  44141. 33:31:06for this bank. And from that, we want to
  44142. 33:31:08generate a problem statement to predict
  44143. 33:31:11if a customer will repay loan amount or
  44144. 33:31:13not. And then we're going to be using
  44145. 33:31:14the decision tree algorithm in Python.
  44146. 33:31:16Let's see what that looks like. And
  44147. 33:31:18let's dive into the code. In our first
  44148. 33:31:20few steps of implementation, we're going
  44149. 33:31:22to start by importing the necessary
  44150. 33:31:24packages that we need from Python. and
  44151. 33:31:26we're going to load up our data and take
  44152. 33:31:28a look at what the data looks like. So,
  44153. 33:31:29the first thing I need is I need
  44154. 33:31:31something to edit my Python and run it
  44155. 33:31:33in. So, let's flip on over. And here I'm
  44156. 33:31:35using the Anaconda Jupiter notebook.
  44157. 33:31:39Now, you can use any Python IDE you like
  44158. 33:31:41to run it in, but I find the Jupyter
  44159. 33:31:43Notebook's really nice for doing things
  44160. 33:31:44on the fly. And let's go ahead and just
  44161. 33:31:46paste that code in the beginning. And
  44162. 33:31:48before we start, let's talk a little bit
  44163. 33:31:50about what we're bringing in. And then
  44164. 33:31:52we're going to do a couple things in
  44165. 33:31:53here. where I have to make a couple
  44166. 33:31:54changes as we go through this first part
  44167. 33:31:56of the import. The first thing we bring
  44168. 33:31:57in is numpy as np. That's very standard
  44169. 33:32:01when we're dealing with mathematics,
  44170. 33:32:03especially with uh very complicated
  44171. 33:32:05machine learning tools. You'll almost
  44172. 33:32:06always see the numpy come in for your
  44173. 33:32:08num your numbers. It's called number
  44174. 33:32:10python. It has your mathematics in
  44175. 33:32:12there. In this case, we actually could
  44176. 33:32:14take it out, but generally you'll need
  44177. 33:32:15it for most of your different things you
  44178. 33:32:17work with. And then we're going to use
  44179. 33:32:18pandas as pd. That's also a standard.
  44180. 33:32:21The pandas is a dataf frame setup and
  44181. 33:32:24you can liken this to uh taking your
  44182. 33:32:26basic data and storing it in a way that
  44183. 33:32:28looks like an Excel spreadsheet. So as
  44184. 33:32:31we come back to this when you see np or
  44185. 33:32:33pd those are very standard uses you'll
  44186. 33:32:35know that that's the pandas and I'll
  44187. 33:32:37show you a little bit more when we
  44188. 33:32:38explore the data in just a minute. Then
  44189. 33:32:40we're going to need to split the data.
  44190. 33:32:41So I'm going to bring in our train test
  44191. 33:32:43and split and this is coming from the
  44192. 33:32:45sklearn package cross validation. In
  44193. 33:32:48just a minute, we're going to change
  44194. 33:32:50that and we'll go over that, too. And
  44195. 33:32:52then there's also the sktree import
  44196. 33:32:54decision tree classifier. That's the
  44197. 33:32:56actual tool we're using. Remember, I
  44198. 33:32:57told you don't be afraid of the
  44199. 33:32:58mathematics. It's going to be done for
  44200. 33:33:00you. Well, the decision tree classifier
  44201. 33:33:02has all that mathematics in there for
  44202. 33:33:03you, so you don't have to figure it back
  44203. 33:33:05out again. And then we have
  44204. 33:33:06sklearn.metrics
  44205. 33:33:08for accuracy score. We need to score our
  44206. 33:33:10our setup. That's the whole reason we're
  44207. 33:33:12splitting it between the training and
  44208. 33:33:13testing data. And finally, we still need
  44209. 33:33:15the sklearn import tree. And that's just
  44210. 33:33:18the basic tree function that's needed
  44211. 33:33:20for the decision tree classifier. And
  44212. 33:33:22finally, we're going to load our data
  44213. 33:33:23down here. And I'm going to run this and
  44214. 33:33:25we're going to get two things on here.
  44215. 33:33:27One, we're going to get an error. And
  44216. 33:33:29two, we're going to get a warning. Let's
  44217. 33:33:30see what that looks like. So the first
  44218. 33:33:32thing we had is we have an error. Why is
  44219. 33:33:35this error here? Well, it's looking at
  44220. 33:33:36this. It says I need to read a file. And
  44221. 33:33:38when this was written, the person who
  44222. 33:33:41wrote it, this is their path where they
  44223. 33:33:43stored the file. So let's go ahead and
  44224. 33:33:45fix that.
  44225. 33:33:48And I'm going to put in here my file
  44226. 33:33:50path. I'm just going to call it full
  44227. 33:33:52file name. And you'll see it's on my C
  44228. 33:33:54drive. And there's this very lengthy
  44229. 33:33:56setup on here where I stored the data
  44230. 33:33:582.csv file.
  44231. 33:34:01Don't worry too much about the full path
  44232. 33:34:02because on your computer it'll be
  44233. 33:34:04different. The data.2 CSV file was
  44234. 33:34:08generated by SimplyLearn. If you want a
  44235. 33:34:11copy of that, you can comment down below
  44236. 33:34:13and request it here in the YouTube.
  44237. 33:34:16And then if I'm going to give it a name,
  44238. 33:34:18full file name, I'm going to go ahead
  44239. 33:34:21and change it here to full
  44240. 33:34:25file name. So let's go ahead and run it
  44241. 33:34:27now and see what happens.
  44242. 33:34:32And we get a warning
  44243. 33:34:37when you're coding. Understanding these
  44244. 33:34:39different warnings and these different
  44245. 33:34:41errors that come up is probably the
  44246. 33:34:43hardest lesson to learn. So let's just
  44247. 33:34:45go ahead and take a look at this and use
  44248. 33:34:47this as a uh opportunity to understand
  44249. 33:34:49what's going on here. If you read the
  44250. 33:34:52warning, it says the cross validation is
  44251. 33:34:55depreciated. So it's a warning on it's
  44252. 33:34:57being removed and it's going to be moved
  44253. 33:34:59in favor of the model selection. So if
  44254. 33:35:02we go up here, we have
  44255. 33:35:03sklearn.crossvalidation.
  44256. 33:35:06And if you research this and go to
  44257. 33:35:07sklearn site, you'll find out that you
  44258. 33:35:10can actually just swap it right in there
  44259. 33:35:11with model selection.
  44260. 33:35:15And so when I come in here and I run it
  44261. 33:35:17again, that removes a warning. What
  44262. 33:35:20they've done is they've had two
  44263. 33:35:22different developers develop it in two
  44264. 33:35:24different branches and then they decided
  44265. 33:35:26to keep one of those and eventually get
  44266. 33:35:28rid of the other one. That's all that is
  44267. 33:35:31and very easy and quick to fix.
  44268. 33:35:34Before we go any further, I went ahead
  44269. 33:35:36and opened up the data from this file.
  44270. 33:35:40Remember the the data file we just
  44271. 33:35:41loaded on here, the data_2.c
  44272. 33:35:44CSV. Let's talk a little bit more about
  44273. 33:35:46that and see what that looks like both
  44274. 33:35:47as a text file because it's a
  44275. 33:35:49commaepparated variable file and in a
  44276. 33:35:52spreadsheet. This is what it looks like
  44277. 33:35:54as a basic text file. You can see at the
  44278. 33:35:56top they've created a header and it's
  44279. 33:35:58got 1 2 3 4 five columns and each column
  44280. 33:36:01has data in it. And let me flip this
  44281. 33:36:03over cuz we're also going to look at
  44282. 33:36:04this uh in an actual spreadsheet so you
  44283. 33:36:06can see what that looks like. And here
  44284. 33:36:08I've opened it up in the open office
  44285. 33:36:10calc, which is pretty much the same as
  44286. 33:36:11um Excel and zoomed in. And you can see
  44287. 33:36:14we've got our columns and our rows of
  44288. 33:36:16data. A little easier to read in here.
  44289. 33:36:18We have a result, yes, yes, no. We have
  44290. 33:36:20initial payment, last payment, credit
  44291. 33:36:22score, house number. If we scroll way
  44292. 33:36:26down,
  44293. 33:36:28we'll see that this occupies a 101 lines
  44294. 33:36:31of code or lines of data with uh the
  44295. 33:36:34first one being a column and then 1,000
  44296. 33:36:36lines of data.
  44297. 33:36:41Now, as a programmer,
  44298. 33:36:43if you're looking at a small amount of
  44299. 33:36:44data, I usually start by pulling it up
  44300. 33:36:46in different sources so I can see what
  44301. 33:36:48I'm working with.
  44302. 33:36:50But in larger data, you won't have that
  44303. 33:36:52option. it would just be um too too
  44304. 33:36:54large. So you need to either bring in a
  44305. 33:36:56small amount that you can look at it
  44306. 33:36:57like we're doing right now or we can
  44307. 33:36:59start looking at it through the Python
  44308. 33:37:01code. So let's go ahead and move on and
  44309. 33:37:03take the next couple steps to explore
  44310. 33:37:05the data using Python. Let's go ahead
  44311. 33:37:07and see what it looks like in Python to
  44312. 33:37:10print the length and the shape of the
  44313. 33:37:12data. So let's start by printing the
  44314. 33:37:14length of the database. We can use a
  44315. 33:37:16simple lin function from Python. And
  44316. 33:37:19when I run this, you'll see that it's a
  44317. 33:37:22thousand long. And that's what we
  44318. 33:37:23expected. There's a thousand lines of
  44319. 33:37:25data in there. If you subtract the
  44320. 33:37:27column head, and this is one of the nice
  44321. 33:37:29things when we did the uh balance data
  44322. 33:37:31from the panda read CSV, you'll see that
  44323. 33:37:35the header is row zero. So, it
  44324. 33:37:37automatically removes a row and then
  44325. 33:37:40shows the data separate. It does a good
  44326. 33:37:42job sorting that data out for us. And
  44327. 33:37:45then we can use a different function.
  44328. 33:37:47And let's take a look at that. And
  44329. 33:37:49again, we're going to utilize the tools
  44330. 33:37:51in Panda.
  44331. 33:37:53And since the balance data was loaded as
  44332. 33:37:55a Panda data frame,
  44333. 33:37:58we can do a shape on it. And let's go
  44334. 33:38:00ahead and run the shape and see what
  44335. 33:38:02that looks like.
  44336. 33:38:04What's nice about the shape is not only
  44337. 33:38:06does it give me the length of the data,
  44338. 33:38:07we have a th00and lines, it also tells
  44339. 33:38:09me there's five columns. So when we were
  44340. 33:38:11looking at the data, we had five columns
  44341. 33:38:13of data. And then let's take one more
  44342. 33:38:15step to explore the data using Python.
  44343. 33:38:18And now that we've taken a look at the
  44344. 33:38:19length and the shape, let's go ahead and
  44345. 33:38:22use the uh pandas module for head.
  44346. 33:38:25Another beautiful thing in the data set
  44347. 33:38:27that we can utilize. So let's put that
  44348. 33:38:29on our sheet here. And we have print
  44349. 33:38:31data set and balance data.head.
  44350. 33:38:35And this is a pandas print statement of
  44351. 33:38:37its own. So it has its own print feature
  44352. 33:38:39in there. And then we went ahead and
  44353. 33:38:41gave a label for our print job here of
  44354. 33:38:43data set. Just a simple print statement.
  44355. 33:38:45And we run that. And let's just take a
  44356. 33:38:48closer look at that. Let me zoom in
  44357. 33:38:49here.
  44358. 33:38:51There we go.
  44359. 33:38:53Pandas does such a wonderful job of
  44360. 33:38:55making this a very clean readable data
  44361. 33:38:59set. So you can look at the data, you
  44362. 33:39:01can look at the column headers, you can
  44363. 33:39:02have it uh when you put it as the head,
  44364. 33:39:05it prints the first five lines of the
  44365. 33:39:07data. And we always start with zero. So
  44366. 33:39:09we have five lines. We have 0 1 2 3 4
  44367. 33:39:12instead of 1 2 3 4 5. That's a standard
  44368. 33:39:15scripting and programming set is you
  44369. 33:39:17want to start with the zero position.
  44370. 33:39:19And that is what the data head does. It
  44371. 33:39:20pulls the first five rows of data. Puts
  44372. 33:39:22it in a nice format that you can look at
  44373. 33:39:24and view. Very powerful tool to view the
  44374. 33:39:27data. So instead of having to flip and
  44375. 33:39:29open up an Excel spreadsheet or open
  44376. 33:39:31Office Cal or trying to look at a word
  44377. 33:39:34doc where it's all scrunched together
  44378. 33:39:36and hard to read, you can now get a nice
  44379. 33:39:38open view of what you're working with.
  44380. 33:39:40We're working with a shape of a thousand
  44381. 33:39:42long, five wide. So we have five columns
  44382. 33:39:45and we do the full data head. You can
  44383. 33:39:47actually see what this data looks like.
  44384. 33:39:48The initial payment, last payment,
  44385. 33:39:50credit scores, house number. So let's
  44386. 33:39:52take this now that we've explored the
  44387. 33:39:54data and let's start digging into the
  44388. 33:39:56decision tree. So in our next step,
  44389. 33:39:59we're going to train and build our data
  44390. 33:40:02tree. And to do that, we need to first
  44391. 33:40:05separate the data out. We're going to
  44392. 33:40:06separate into two groups so that we have
  44393. 33:40:08something to actually train the data
  44394. 33:40:10with. And then we have some data on the
  44395. 33:40:12side to test it to see how good our
  44396. 33:40:14model is. Remember with any of the
  44397. 33:40:15machine learning, you always want to
  44398. 33:40:17have some kind of test set to to weigh
  44399. 33:40:19it against so you know how good your
  44400. 33:40:20model is when you distribute it. Let's
  44401. 33:40:22go ahead and break this code down and
  44402. 33:40:24look at it in pieces. So first we have
  44403. 33:40:27our X and Y.
  44404. 33:40:30Where do X and Y come from? Well, X is
  44405. 33:40:32going to be our data and Y is going to
  44406. 33:40:35be the answer or the target. You can
  44407. 33:40:37look at it source and target. In this
  44408. 33:40:39case, we're using X and Y to denote the
  44409. 33:40:41data in and the data that we're actually
  44410. 33:40:43trying to guess what the answer is going
  44411. 33:40:45to be. And so to separate it, we can
  44412. 33:40:46simply put in X equals the balance of
  44413. 33:40:49the data values. The first brackets
  44414. 33:40:53means that we're going to select all the
  44415. 33:40:55lines in the database. So, it's all the
  44416. 33:40:57data. And the second one says we're only
  44417. 33:41:00going to look at columns 1 through five.
  44418. 33:41:02Remember, always start with zero. Zero
  44419. 33:41:04is a yes or no. And that's whether the
  44420. 33:41:06loan went default or not. So, we want to
  44421. 33:41:08start with one. If we go back up here,
  44422. 33:41:10that's the initial payment and it goes
  44423. 33:41:12all the way through the house number.
  44424. 33:41:15Well, if we want to look at uh 1 through
  44425. 33:41:17five, we can do the same thing for y,
  44426. 33:41:20which is the answers. And we're going to
  44427. 33:41:22set that just equal to the zero row. So,
  44428. 33:41:25it's just the zero row and then it's all
  44429. 33:41:27rows going in there. So, now we've
  44430. 33:41:28divided this into two different data
  44431. 33:41:31sets. One of them with the
  44432. 33:41:34data going in and one with the answers.
  44433. 33:41:40Next, we need to split the data.
  44434. 33:41:43And here you'll see that we have it
  44435. 33:41:45split into four different parts. The
  44436. 33:41:48first one is your X training, your X
  44437. 33:41:51test, your Y train, your Y test.
  44438. 33:41:56Simply put, we have X going in where
  44439. 33:41:58we're going to train it and we have to
  44440. 33:42:00know the answer to train it with. And
  44441. 33:42:02then we have X test where we're going to
  44442. 33:42:04test that data and we have to know in
  44443. 33:42:07the end what the Y was supposed to be.
  44444. 33:42:10And that's where this train test split
  44445. 33:42:12comes in that we loaded earlier in the
  44446. 33:42:14modules. This does it all for us. And
  44447. 33:42:16you can see they set the test size equal
  44448. 33:42:18to.3. So that's roughly 30% will be used
  44449. 33:42:21in the test. And then we use a random
  44450. 33:42:22state. So it's completely random which
  44451. 33:42:24rows it takes out of there. And then
  44452. 33:42:26finally we get to actually build our
  44453. 33:42:28decision tree. And they've called it
  44454. 33:42:29here CLF entropy. That's the actual
  44455. 33:42:33decision tree or decision tree
  44456. 33:42:34classifier. And in here, they've added a
  44457. 33:42:37couple variables which we'll explore in
  44458. 33:42:39just a minute. And then finally, we need
  44459. 33:42:41to fit the data to that. So, we take our
  44460. 33:42:44CLF entropy that we created and we fit
  44461. 33:42:46the X train. And since we know the
  44462. 33:42:48answers for X-ray or the Y train, we go
  44463. 33:42:50ahead and put those in. And let's go
  44464. 33:42:52ahead and run this. And what most of
  44465. 33:42:54these sklearn modules do is when you set
  44466. 33:42:57up the variable, in this case, when we
  44467. 33:42:59set the CLF entropy equal decision tree
  44468. 33:43:01classifier, it automatically prints out
  44469. 33:43:03what's in that decision tree. There's a
  44470. 33:43:05lot of variables you can play with in
  44471. 33:43:06here. And it's quite beyond the scope of
  44472. 33:43:08this tutorial to go through all of these
  44473. 33:43:11and how they work. But we're working on
  44474. 33:43:13entropy. That's one of the options.
  44475. 33:43:14We've added that it's completely a
  44476. 33:43:16random state of 100, so 100%. And we
  44477. 33:43:19have a max depth of three. Now, the max
  44478. 33:43:22depth, if you remember above when we
  44479. 33:43:23were doing the different graphs of
  44480. 33:43:25animals, means it's only going to go
  44481. 33:43:27down three layers before it stops. And
  44482. 33:43:30then we have minimal samples of leaves
  44483. 33:43:31is five. So, it's going to have at least
  44484. 33:43:33five leaves at the end. So, I'll have at
  44485. 33:43:35least three splits or have no more than
  44486. 33:43:37three layers and at least five end
  44487. 33:43:40leaves with the final result at the
  44488. 33:43:42bottom. Now that we've created our
  44489. 33:43:45decision tree classifier, not only
  44490. 33:43:47created it, but trained it, let's go
  44491. 33:43:49ahead and apply it and see what that
  44492. 33:43:51looks like. So, let's go ahead and make
  44493. 33:43:53a prediction and see what that looks
  44494. 33:43:55like. We're going to paste our predict
  44495. 33:43:57code in here. And before we run it,
  44496. 33:43:59let's just take a quick look at what's
  44497. 33:44:01doing here. We have a variable y predict
  44498. 33:44:04that we're going to do. And we're going
  44499. 33:44:06to use our variable CLF entropy that we
  44500. 33:44:10created.
  44501. 33:44:12And then you'll see predict. And it's
  44502. 33:44:14very common in the sklearn modules that
  44503. 33:44:17their different tools have the predict
  44504. 33:44:19when you're actually running a
  44505. 33:44:20prediction. In this case, we're going to
  44506. 33:44:22put our X test data in here. Now, if you
  44507. 33:44:26delivered this for use, an actual
  44508. 33:44:28commercial use, and distributed it, this
  44509. 33:44:31would be the new loans you're putting in
  44510. 33:44:33here to guess whether the person's going
  44511. 33:44:36to be uh pay them back or not. In this
  44512. 33:44:38case though, we need to test out the
  44513. 33:44:40data and just see how good our sample
  44514. 33:44:42is, how good of our tree does at
  44515. 33:44:45predicting the loan payments. And
  44516. 33:44:46finally, since Anaconda Jupyter notebook
  44517. 33:44:49is works as a command line for Python,
  44518. 33:44:51we can simply put the y predict en to
  44519. 33:44:54print it. I could just as easily have
  44520. 33:44:56put the print
  44521. 33:44:58and put brackets around y predict en to
  44522. 33:45:01print it out. We'll go ahead and do
  44523. 33:45:02that. It doesn't matter which way you do
  44524. 33:45:03it. And you'll see right here that it
  44525. 33:45:06runs a prediction. This is roughly 300
  44526. 33:45:09in here. Remember, it's 30% of a
  44527. 33:45:11thousand. So, you should have about 300
  44528. 33:45:13answers in here. And this tells you
  44529. 33:45:16which each one of those lines of ourh
  44530. 33:45:18test went in there. And this is what our
  44531. 33:45:20y predict came out. So, let's move on to
  44532. 33:45:23the next step where we're going to take
  44533. 33:45:25this data and try to figure out just how
  44534. 33:45:27good a model we have. So, here we go.
  44535. 33:45:29Since sklearn does all the heavy lifting
  44536. 33:45:31for you and all the math, we have a
  44537. 33:45:33simple line of code to let us know what
  44538. 33:45:35the accuracy is. And let's go ahead and
  44539. 33:45:37go through that and see what that means
  44540. 33:45:39and what that looks like. Let's go ahead
  44541. 33:45:40and paste this in. And let me zoom in a
  44542. 33:45:42little bit. There we go.
  44543. 33:45:45So you have a nice full picture. And
  44544. 33:45:47we'll see here. We're just going to do a
  44545. 33:45:48print accuracy is.
  44546. 33:45:51And then we do the accuracy score. And
  44547. 33:45:54this was something we imported um
  44548. 33:45:56earlier. If you remember at the very
  44549. 33:45:58beginning, let me just scroll up there
  44550. 33:45:59real quick so you can see where that's
  44551. 33:46:01coming from. That's coming from here
  44552. 33:46:03down here from sklearn.metrics metrics
  44553. 33:46:06import accuracy score. And you could
  44554. 33:46:09probably run a script, make your own
  44555. 33:46:10script to do this very easily. How
  44556. 33:46:12accurate is it? How many out of 300 do
  44557. 33:46:15we get right? And so we put in our y
  44558. 33:46:17test. That's the one we ran the predict
  44559. 33:46:19on. And then we put in our y predict en
  44560. 33:46:22that's the answers we got. And we're
  44561. 33:46:24just going to multiply that by 100
  44562. 33:46:26because this is just going to give us an
  44563. 33:46:27answer as a decimal and we want to see
  44564. 33:46:29it as a percentage. And let's run that
  44565. 33:46:31and see what it looks like. And if you
  44566. 33:46:34see here, we got an accuracy of
  44567. 33:46:3593.666667.
  44568. 33:46:38So when we look at the number of loans
  44569. 33:46:40and we look at how good our model fit,
  44570. 33:46:42we can tell people it has about a 93.6
  44571. 33:46:46fitting to it. So just a quick recap on
  44572. 33:46:49that. We now have accuracy set up on
  44573. 33:46:52here. And so we have created a model
  44574. 33:46:54that uses the decision tree algorithm to
  44575. 33:46:56predict whether a customer will repay
  44576. 33:46:57the loan or not. The accuracy of the
  44577. 33:46:59model is about 94.6%.
  44578. 33:47:02The bank can now use this model to
  44579. 33:47:03decide whether it should approve the
  44580. 33:47:05loan request from a particular customer
  44581. 33:47:07or not. And so this information is
  44582. 33:47:09really powerful. We may not be able to
  44583. 33:47:11as individuals understand all these
  44584. 33:47:12numbers because they have thousands of
  44585. 33:47:14numbers that come in, but you can see
  44586. 33:47:16that this is a smart decision for the
  44587. 33:47:17bank to use a tool like this to help
  44588. 33:47:20them to predict how good their uh
  44589. 33:47:22profit's going to be off of the loan
  44590. 33:47:24balances and how many are going to
  44591. 33:47:25default or not. We're going to be
  44592. 33:47:27looking at random forest, one of the
  44593. 33:47:28many powerful tools in the machine
  44594. 33:47:30learning library. Before we dive into
  44595. 33:47:33the topic, let's start by looking at a
  44596. 33:47:35few of the uses for random forest.
  44597. 33:47:38Currently today, it's used in remote
  44598. 33:47:40sensing. Uh for example, they're used in
  44599. 33:47:43the ETM devices. If you're a space buff,
  44600. 33:47:46that's the enhanced thermatic mapper
  44601. 33:47:48they use on satellites which see uh far
  44602. 33:47:51outside the human spectrum for looking
  44603. 33:47:53at land masses. and they acquire images
  44604. 33:47:55of the earth's surface. The accuracy is
  44605. 33:47:57higher and training time is less than
  44606. 33:48:00many other machine learning tools out
  44607. 33:48:01there. Also, object detection,
  44608. 33:48:03multiclass object detection is done
  44609. 33:48:05using random forest algorithms. A good
  44610. 33:48:07example is a traffic where you're trying
  44611. 33:48:09to sort out the different cars, buses,
  44612. 33:48:11and things. And it provides a better
  44613. 33:48:12detection in complicated environments.
  44614. 33:48:14They're very complicated up there. And
  44615. 33:48:16then we have uh another example connect.
  44616. 33:48:19And let's take a little closer look at
  44617. 33:48:20connect. Connect. They use a random
  44618. 33:48:23forest as part of the game console and
  44619. 33:48:26what it does is it tracks a body
  44620. 33:48:27movements and it recreates it in the
  44621. 33:48:29game and let's see what that looks like.
  44622. 33:48:32Uh we have a user who performs a step.
  44623. 33:48:34In this case it looks like Elvis Presley
  44624. 33:48:36going there that is then recorded so
  44625. 33:48:39that connect registers the movement and
  44626. 33:48:41then it marks the user based on
  44627. 33:48:43accuracy. And it looks like we have uh
  44628. 33:48:46Prince going on this one from Elvis
  44629. 33:48:48Presley to Prince. It's great. Uh so it
  44630. 33:48:51marks user base on the accuracy. If we
  44631. 33:48:53look at that a little closer, we have a
  44632. 33:48:55training set to identify body parts.
  44633. 33:48:57Where are the hands? Where are the feet?
  44634. 33:48:59Uh what's going on with the body? That
  44635. 33:49:02then goes into a random forest
  44636. 33:49:04classifier that learns from it. Once
  44637. 33:49:06we've trained the classifier, it then
  44638. 33:49:09identifies the body parts while the
  44639. 33:49:11person's dancing. It's able to represent
  44640. 33:49:13that in a computer format. And then
  44641. 33:49:16based on that, it scores the game and
  44642. 33:49:18how accurate you are as being Elvis
  44643. 33:49:20Presley or Prince in your dancing. So
  44644. 33:49:22why random forest? It's always important
  44645. 33:49:26to understand why we use this tool over
  44646. 33:49:29the other ones. What are the benefits
  44647. 33:49:31here? And so with the random forest, the
  44648. 33:49:34first one is there's no overfitting. If
  44649. 33:49:36you use of multiple trees, reduce the
  44650. 33:49:39risk of overfitting. Training time is
  44651. 33:49:41less. Overfitting means that we have fit
  44652. 33:49:44the data so close to what we have as our
  44653. 33:49:46sample that we pick up on all the weird
  44654. 33:49:49parts and instead of predicting the
  44655. 33:49:51overall data, you're predicting the
  44656. 33:49:52weird stuff which you don't want. High
  44657. 33:49:55accuracy runs efficiently on large
  44658. 33:49:57database. For large data, it produces
  44659. 33:49:59highly accurate predictions. In today's
  44660. 33:50:02world of uh big data, this is really
  44661. 33:50:04important. And this is probably where it
  44662. 33:50:07really shines. This is where Y random
  44663. 33:50:09forest really comes in. It estimates
  44664. 33:50:11missing data. Data in today's world is
  44665. 33:50:14very messy. So when you have a random
  44666. 33:50:15forest, it can maintain the accuracy
  44667. 33:50:17when a large proportion of the data is
  44668. 33:50:19missing. What that means is if you have
  44669. 33:50:21data that comes in from uh five or six
  44670. 33:50:23different areas and maybe they took one
  44671. 33:50:26set of statistics in one area and they
  44672. 33:50:28took a slightly different set of
  44673. 33:50:30statistics in the other. So they have
  44674. 33:50:31some of the sh same shared data, but one
  44675. 33:50:34is missing like the uh number of
  44676. 33:50:36children in the house if you're doing
  44677. 33:50:37something over demographics. and the
  44678. 33:50:40other one is missing the size of the
  44679. 33:50:42house. It will look at both of those
  44680. 33:50:44separately and build two different trees
  44681. 33:50:46and then it can do a very good job of
  44682. 33:50:48guessing which one fits better even
  44683. 33:50:50though it's missing that data. Let us
  44684. 33:50:52dig deep into the theory of exactly how
  44685. 33:50:55it works. And let's look at what is
  44686. 33:50:58random forest. Random forest or random
  44687. 33:51:01decision forest is a method that
  44688. 33:51:04operates by constructing multiple
  44689. 33:51:05decision trees. The decision of the
  44690. 33:51:08majority of the trees is chosen by the
  44691. 33:51:10random forest as the final decision. And
  44692. 33:51:12let's uh we have some nice graphics
  44693. 33:51:14here. We have a decision tree and they
  44694. 33:51:15actually use a real tree to denote the
  44695. 33:51:17decision tree which I love. And given a
  44696. 33:51:20random some kind of picture of a fruit.
  44697. 33:51:22This decision tree decides that the
  44698. 33:51:24output is it's an apple. And we have a
  44699. 33:51:26decision tree too where we have that
  44700. 33:51:28picture of the fruit goes in and this
  44701. 33:51:30one decides that it's a lemon. And the
  44702. 33:51:31decision three tree gets another image
  44703. 33:51:33and it decides it's an apple. And then
  44704. 33:51:35this all comes together in what they
  44705. 33:51:37call the random forest. And this random
  44706. 33:51:39forest then looks at it and says, "Okay,
  44707. 33:51:41I got two votes for apple, one vote for
  44708. 33:51:43lemon. The majority is apples. So the
  44709. 33:51:46final decision is apples." To understand
  44710. 33:51:49how the random forest works, we first
  44711. 33:51:52need to dig a little deeper and take a
  44712. 33:51:54look at the random forest and the actual
  44713. 33:51:56decision tree and how it builds that
  44714. 33:51:58decision tree. In looking closer at how
  44715. 33:52:00the individual decision trees work,
  44716. 33:52:02we'll go ahead and continue to use the
  44717. 33:52:04fruit example since we're talking about
  44718. 33:52:06trees and forests. A decision tree is a
  44719. 33:52:08treerehaped diagram used to determine a
  44720. 33:52:11course of action. Each branch of the
  44721. 33:52:13tree represents a possible decision,
  44722. 33:52:15occurrence, or reaction. So in here we
  44723. 33:52:17have a bowl of fruit and if you look at
  44724. 33:52:18that it looks like um they switch from
  44725. 33:52:20lemons to oranges. So we have oranges,
  44726. 33:52:23cherries, and apples. And the first
  44727. 33:52:25decision of the decision tree might be
  44728. 33:52:27is a diameter greater than or equal to
  44729. 33:52:29three. And if it says false, it knows
  44730. 33:52:31that they're cherries because everything
  44731. 33:52:32else is bigger than that. So all the
  44732. 33:52:34cherries fall into that decision. So we
  44733. 33:52:36have all that data we're training. We
  44734. 33:52:38can look at that. We know that that's
  44735. 33:52:39what's going to come up. Is the color
  44736. 33:52:41orange? Well, goes, hm, orange or red?
  44737. 33:52:44Well, if it's true, then it comes out as
  44738. 33:52:47the orange. And if it's false, that
  44739. 33:52:49leaves apples. So in this example, it
  44740. 33:52:52sorts out the fruit in the bowl or the
  44741. 33:52:54images of the fruit. A decision tree.
  44742. 33:52:57These are very important terms to know
  44743. 33:52:59because these are very central to
  44744. 33:53:00understanding the decision tree and when
  44745. 33:53:01working with them. The first is entropy.
  44746. 33:53:04Everything on the decision tree and how
  44747. 33:53:06it makes a decision is based on entropy.
  44748. 33:53:09Entropy is a measure of randomness or
  44749. 33:53:11unpredictability in the data set. uh
  44750. 33:53:14then they also have information gain,
  44751. 33:53:17the leaf node, the decision node and the
  44752. 33:53:20root node. We'll cover these other four
  44753. 33:53:22terms as we go down the tree, but let's
  44754. 33:53:25start with entropy. So starting with
  44755. 33:53:28entropy, we have here a high amount of
  44756. 33:53:30randomness. What that means is that
  44757. 33:53:34whatever is coming out of this decision,
  44758. 33:53:35if it was going to guess based on this
  44759. 33:53:38data, it wouldn't be able to tell you
  44760. 33:53:40whether it's a lemon or an apple. it
  44761. 33:53:42would just say it's a fruit. Uh so the
  44762. 33:53:45first thing we want to do is we want to
  44763. 33:53:47split this apart and we take the initial
  44764. 33:53:49data set. We're going to set create a
  44765. 33:53:50data set one and a data set two. We just
  44766. 33:53:53split it in two. And if you look at
  44767. 33:53:54these new data sets after splitting
  44768. 33:53:57them, the entropy of each of those sets
  44769. 33:53:59is much less. So for the first one,
  44770. 33:54:02whatever comes in there, it's going to
  44771. 33:54:04sort that data and it's going to say,
  44772. 33:54:05okay, if this data goes this direction,
  44773. 33:54:07it's probably an apple. And if it goes
  44774. 33:54:09into the other direction, it's probably
  44775. 33:54:11a lemon. So that brings us up to
  44776. 33:54:13information gain. It is the measure of
  44777. 33:54:15decrease in the entropy after the data
  44778. 33:54:17set is split. What that means in here is
  44779. 33:54:20that we've gone from one set which has a
  44780. 33:54:23very high entropy to two lower sets of
  44781. 33:54:26entropy and we've added in the values of
  44782. 33:54:28E1 for the first one and E2 for the
  44783. 33:54:31second two which are much lower. And so
  44784. 33:54:33that information gain is increased
  44785. 33:54:36greatly in this example. And so you can
  44786. 33:54:38find that the information grain simply
  44787. 33:54:40equals uh decision E1 minus E2. As we're
  44788. 33:54:45going down our list of uh definitions,
  44789. 33:54:48we'll look at the leaf node. And the
  44790. 33:54:49leaf node carries the classification or
  44791. 33:54:52the decision. So we look down here to
  44792. 33:54:55the leaf node. We finally get to our set
  44793. 33:54:58one or our set two. When it comes down
  44794. 33:55:00there and it says, "Okay, this object's
  44795. 33:55:02gone into set one." If it's gone into
  44796. 33:55:05set one, it's going to be split by some
  44797. 33:55:08means and we'll either end up with
  44798. 33:55:09apples on the leaf node or a lemon on
  44799. 33:55:12the leaf node. And on the right, it'll
  44800. 33:55:13either be an apple or lemons. Those leaf
  44801. 33:55:16nodes are those final decisions or
  44802. 33:55:18classifications. Uh that's the
  44803. 33:55:20definition of leaf node in here. If
  44804. 33:55:22we're going to have a final leaf where
  44805. 33:55:24we make the decision, we should have a
  44806. 33:55:27name for the nodes above it. And they
  44807. 33:55:29call those decision nodes. A decision
  44808. 33:55:32node. decision node has two or more
  44809. 33:55:34branches and you can see here where we
  44810. 33:55:36have the uh five apples and one lemon
  44811. 33:55:39and in the other case the five lemons
  44812. 33:55:41and one apple. They have to make a
  44813. 33:55:43choice of which tree it goes down based
  44814. 33:55:45on some kind of measurement or
  44815. 33:55:48information given to the tree. And that
  44816. 33:55:51brings us to our last definition. The
  44817. 33:55:53root node, the topmost decision node is
  44818. 33:55:56known as the root node. And this is
  44819. 33:55:58where you have all of your data and you
  44820. 33:56:01have your first decision. it has to make
  44821. 33:56:02or the first split in information. So
  44822. 33:56:05far, we've looked at a very general
  44823. 33:56:07image um with the fruit being split.
  44824. 33:56:10Let's look and see exactly what that
  44825. 33:56:12means to split the data and how do we
  44826. 33:56:14make those decisions on there. Uh let's
  44827. 33:56:17go in there and find out how does a
  44828. 33:56:19decision tree work. So let's try to
  44829. 33:56:22understand this and let's use a simple
  44830. 33:56:24example and we'll stay with the fruit.
  44831. 33:56:27We have a bowl of fruit and so let's
  44832. 33:56:29create a problem statement and the
  44833. 33:56:31problem is we want to classify the
  44834. 33:56:34different types of fruits in the bowl
  44835. 33:56:35based on different features. The data
  44836. 33:56:38set in the bowl is looking quite messy
  44837. 33:56:40and the entropy is high in this case. So
  44838. 33:56:42if this bowl was our decision maker, it
  44839. 33:56:45wouldn't know what choice to make. It
  44840. 33:56:47has so many choices. Which one do you
  44841. 33:56:49pick? Apple, grapes, or lemons. And so
  44842. 33:56:52we look in here. We're going to start
  44843. 33:56:53with a d a training set. So this is our
  44844. 33:56:56data that we're training our data with
  44845. 33:56:58and we have a number of options here. We
  44846. 33:57:00have the color and under the color we
  44847. 33:57:01have red yellow purple uh we have a
  44848. 33:57:04diameter uh 331 331 and we have a label
  44849. 33:57:08apple lemon grapes apple lemon grapes
  44850. 33:57:11and how do we split the data? We have to
  44851. 33:57:13frame the conditions to split the data
  44852. 33:57:15in such a way that the information gain
  44853. 33:57:17is the highest. It's very key to note
  44854. 33:57:20that we're looking for the best gain. We
  44855. 33:57:22don't want to just start sorting out the
  44856. 33:57:23smallest piece in there. We want to
  44857. 33:57:25split it the biggest way we can. And so
  44858. 33:57:27we measure this decrease in entropy.
  44859. 33:57:30That's what they call it, entropy.
  44860. 33:57:31There's our entropy after splitting. And
  44861. 33:57:33now we'll try to choose a condition that
  44862. 33:57:35gives us the highest gain. We will do
  44863. 33:57:37that by splitting the data using each
  44864. 33:57:39condition and checking the gain that we
  44865. 33:57:41get out of them. The conditions that
  44866. 33:57:42give us the highest gain will be used to
  44867. 33:57:44make the first split. So let's take a
  44868. 33:57:46look at these different conditions. We
  44869. 33:57:48have color, we have diameter, and if we
  44870. 33:57:50look underneath that, we have a couple
  44871. 33:57:51different values. is we have diameter
  44872. 33:57:52equals 3, color equals yellow, red,
  44873. 33:57:55diameter equals 1. And when we look at
  44874. 33:57:57that, you'll see over here we have 1 2 3
  44875. 33:58:014 threes. That's a pretty hardy
  44876. 33:58:04selection. So let's say the condition
  44877. 33:58:06gives us a maximum gain of three. So we
  44878. 33:58:09have the most pieces fall into that
  44879. 33:58:12range. So our first split from our
  44880. 33:58:15decision node is we split the data based
  44881. 33:58:18on the diameter. Is it greater than or
  44882. 33:58:20equal to three? If it's not, that's
  44883. 33:58:23false. It goes into the grape bowl. And
  44884. 33:58:25if it's true, it goes into a bowl fold
  44885. 33:58:27of lemon and apples. The entropy after
  44886. 33:58:30splitting has decreased considerably. So
  44887. 33:58:32now we can make two decisions. If you
  44888. 33:58:34look at they're very uh much less chaos
  44889. 33:58:37going on there. This node has already
  44890. 33:58:39attain an entropy value of zero. As you
  44891. 33:58:41can see, there's only one kind of label
  44892. 33:58:43left for this branch. So no further
  44893. 33:58:45splitting is required for this node.
  44894. 33:58:48However, this node on the right is still
  44895. 33:58:51requires a split to decrease the entropy
  44896. 33:58:53further. So, we split the right node
  44897. 33:58:55further based on color. If you look at
  44898. 33:58:57this, if I split it on color, that
  44899. 33:58:59pretty much cuts it right down the
  44900. 33:59:01middle. And it's the only thing we have
  44901. 33:59:02left in our choices of color and
  44902. 33:59:03diameter, too. And if the color is
  44903. 33:59:06yellow, it's going to go to the right
  44904. 33:59:08bowl. And if it's false, it's going to
  44905. 33:59:09go to the left bowl. So, the entropy in
  44906. 33:59:11this case is now zero. So, now we have
  44907. 33:59:14three bowls with zero entropy. There's
  44908. 33:59:16only one type of data in each one of
  44909. 33:59:18those bowls. So, we can predict a lemon
  44910. 33:59:20with 100% accuracy. And we can predict
  44911. 33:59:23the apple also with 100% accuracy along
  44912. 33:59:26with our grapes up there. So, we've
  44913. 33:59:28looked at kind of a basic tree in our
  44914. 33:59:31forest. But what we really want to know
  44915. 33:59:33is how does a random forest work as a
  44916. 33:59:36whole. So to begin our um random forest
  44917. 33:59:40classifier, let's say we already have
  44918. 33:59:42built three trees. And we're going to
  44919. 33:59:44start with the first tree that looks
  44920. 33:59:46like this. Just like we did in the
  44921. 33:59:48example, this tree looks at the
  44922. 33:59:49diameter. If it's greater than or equal
  44923. 33:59:51to three, it's true. Otherwise, it's
  44924. 33:59:53false. So one side goes to the smaller
  44925. 33:59:56diameter, one side goes to larger
  44926. 33:59:58diameter. And if the color is orange,
  44927. 34:00:01it's going to go to the right. True.
  44928. 34:00:02We're using oranges now instead of
  44929. 34:00:04lemons. And if it's red, it's going to
  44930. 34:00:06go to the left. False. We build a second
  44931. 34:00:08tree very similar, but it's split
  44932. 34:00:10differently. Instead of the first one
  44933. 34:00:12being split by a diameter, uh this one
  44934. 34:00:15when they created it, if you look at
  44935. 34:00:16that first bowl, it has a lot of red
  44936. 34:00:18objects. So it says, is the color red?
  44937. 34:00:21Because that's going to bring our
  44938. 34:00:22entropy down the fastest. And so, of
  44939. 34:00:25course, if it's true, it goes to the
  44940. 34:00:26left. If it's false, it goes to the
  44941. 34:00:28right. And then it looks at the shape,
  44942. 34:00:30false or true, and so on and so on. And
  44943. 34:00:33tree three is the diameter equal to one.
  44944. 34:00:36And it came up with this because there's
  44945. 34:00:38a lot of cherries in this bowl. So that
  44946. 34:00:39would be the biggest split on there is
  44947. 34:00:41is the diameter equal to one. That's
  44948. 34:00:43going to drop the entropy the quickest.
  44949. 34:00:45And as you can see, it splits it into
  44950. 34:00:46true. If it goes false, and they've
  44951. 34:00:48added another category, does it grow in
  44952. 34:00:50the summer? And if it's false, it goes
  44953. 34:00:53off to the left. If it's true, it goes
  44954. 34:00:54off to the right. Let's go ahead and
  44955. 34:00:56bring these three trees so you can see
  44956. 34:00:58them all in one image. So this would be
  44957. 34:01:00three completely different trees
  44958. 34:01:02categorizing a fruit. And let's take a
  44959. 34:01:04fruit. Now let's try this. And this
  44960. 34:01:06fruit, if you look at it, we've
  44961. 34:01:08blackened it out. You can't see the
  44962. 34:01:10color on it. So it's missing data.
  44963. 34:01:12Remember one of the things we talked
  44964. 34:01:13about earlier is that a random forest
  44965. 34:01:16works really good if you're missing
  44966. 34:01:18data, if you're missing pieces. So this
  44967. 34:01:20fruit has an image, but maybe it's a
  44968. 34:01:22person had a black and white camera when
  44969. 34:01:24they took the picture. And we're going
  44970. 34:01:25to take a look at this. And it's going
  44971. 34:01:27to have um they put the color in there,
  44972. 34:01:29so ignore the color down there. But the
  44973. 34:01:31diameter equals three. We find out it
  44974. 34:01:33grows in the summer equals yes. And the
  44975. 34:01:35shape is a circle. And if you go to the
  44976. 34:01:37right, you can look at what one of the
  44977. 34:01:39decision trees did. This is the third
  44978. 34:01:41one. Is the diameter greater than equal
  44979. 34:01:43to three? Is a color orange? Well, it
  44980. 34:01:46doesn't really know on this one, but it
  44981. 34:01:48if you look at the value, it' say true,
  44982. 34:01:49and it go to the right. Tree two
  44983. 34:01:51classifies it as cherries. Is a color
  44984. 34:01:54equal red? Is the shape a circle? True.
  44985. 34:01:57It is a circle. So, this would look at
  44986. 34:01:59it and say, "Oh, that's a cherry." And
  44987. 34:02:01then we go to the other classifier and
  44988. 34:02:02it says, "Is the diameter equal one?"
  44989. 34:02:05Well, that's false. Does it grow in the
  44990. 34:02:07summer? True. So, it goes down and looks
  44991. 34:02:09at as oranges. So, how does this random
  44992. 34:02:12forest work? The first one says it's an
  44993. 34:02:14orange. The second one said it was a
  44994. 34:02:16cherry. And the third one says, hm, it's
  44995. 34:02:19an orange. And you can guess that if you
  44996. 34:02:21have two oranges and one says it's a
  44997. 34:02:23cherry, uh, when you add that all
  44998. 34:02:24together, the majority of the vote says
  44999. 34:02:27orange. So, the answer is it's
  45000. 34:02:29classified as an orange, even though we
  45001. 34:02:31didn't know the color and we're missing
  45002. 34:02:32data on it. I don't know about you, but
  45003. 34:02:34I'm getting tired of fruit. So, let's
  45004. 34:02:37switch. And I did promise you we'd start
  45005. 34:02:39looking at a case example and get into
  45006. 34:02:41some Python coding. Today, we're going
  45007. 34:02:43to use the case the iris flower
  45008. 34:02:45analysis.
  45009. 34:02:47This is the exciting part as we roll up
  45010. 34:02:49our sleeves and actually look at some
  45011. 34:02:51Python coding. Before we start the
  45012. 34:02:53Python coding, we need to go ahead and
  45013. 34:02:55create a problem statement. Wonder what
  45014. 34:02:57species of iris do these flowers belong
  45015. 34:02:59to? Let's try to predict the species of
  45016. 34:03:01the flowers using machine learning in
  45017. 34:03:03Python. Let's see how it can be done. So
  45018. 34:03:06here we begin to go ahead and implement
  45019. 34:03:08our Python code. And you'll find that
  45020. 34:03:11the first half of our implementation is
  45021. 34:03:13all about organizing and exploring the
  45022. 34:03:16data coming in. Let's go ahead and take
  45023. 34:03:18this first step, which is loading the
  45024. 34:03:20different modules into Python. And let's
  45025. 34:03:22go ahead and put that in our favorite
  45026. 34:03:24editor, whatever your favorite editor
  45027. 34:03:26is. In this case, I'm going to be using
  45028. 34:03:28the Anaconda Jupiter Notebook, which is
  45029. 34:03:31one of my favorites. Certainly, there's
  45030. 34:03:33Notepad++ and Eclipse and dozens of
  45031. 34:03:36others, or just even using the Python
  45032. 34:03:38terminal window. any of those will work
  45033. 34:03:40just fine to go ahead and explore this
  45034. 34:03:43Python coding. So, here we go. Let's go
  45035. 34:03:45ahead and flip over to our Jupyter
  45036. 34:03:46notebook. And I've already opened up a
  45037. 34:03:49new page for Python 3 code. And I'm just
  45038. 34:03:52going to paste this right in there. And
  45039. 34:03:53let's take a look and see what we're
  45040. 34:03:55bringing into our Python. The first
  45041. 34:03:57thing we're going to do is from the
  45042. 34:03:58sklearn.data sets import load iris. Now,
  45043. 34:04:02this isn't the actual data. So this is
  45044. 34:04:04just the module that allows us to bring
  45045. 34:04:06in the data, the load iris. And the iris
  45046. 34:04:09is so popular. It's been around since
  45047. 34:04:111936 when Ronald Fiser published a paper
  45048. 34:04:14on it. And they're measuring the
  45049. 34:04:16different parts of the flower. And based
  45050. 34:04:18on those measurements, predicting what
  45051. 34:04:19kind of flower it is. And then if we're
  45052. 34:04:21going to do a random forest classifier,
  45053. 34:04:24we need to go ahead and import a random
  45054. 34:04:25forest classifier from the sklearn
  45055. 34:04:27module. So sklearn.semble
  45056. 34:04:30import random forest classifier. And
  45057. 34:04:32then we want to bring in two more
  45058. 34:04:33modules. Um, and these are probably the
  45059. 34:04:36most commonly used modules in Python and
  45060. 34:04:38data science with any of the um, other
  45061. 34:04:42modules that we bring in. And one is
  45062. 34:04:43going to be pandas. We're going to
  45063. 34:04:45import pandas as pd. PD is the common
  45064. 34:04:47term used for pandas. And pandas is
  45065. 34:04:50basically creates a data format for us
  45066. 34:04:53where when you create a pandas data
  45067. 34:04:56frame, it looks like an Excel
  45068. 34:04:58spreadsheet. And you'll see that in a
  45069. 34:05:00minute when we start digging deeper into
  45070. 34:05:01the code. Panda is just wonderful
  45071. 34:05:03because it plays nice with all the other
  45072. 34:05:05modules in there. And then we have
  45073. 34:05:06Numpy, which is our numbers Python. And
  45074. 34:05:09the numbers Python allows us to do
  45075. 34:05:12different mathematical sets on here.
  45076. 34:05:14We'll see right off the bat, we're going
  45077. 34:05:16to take our NP and we're going to go
  45078. 34:05:18ahead and seed the randomness with it
  45079. 34:05:19with zero. So NP.random seed is seeding
  45080. 34:05:22that as zero. This code doesn't actually
  45081. 34:05:24show anything. We're going to go ahead
  45082. 34:05:26and run it because I need to make sure I
  45083. 34:05:28have all those loaded. And then let's
  45084. 34:05:29take a look at the next module on here.
  45085. 34:05:31The next six slides, including this one,
  45086. 34:05:34are all about exploring the data.
  45087. 34:05:36Remember, I told you half of this is
  45088. 34:05:38about looking at the data and getting it
  45089. 34:05:40all set. So, let's go ahead and take
  45090. 34:05:42this code right here, the script, and
  45091. 34:05:44let's get that over into our Jupyter
  45092. 34:05:46notebook. And here we go. We've gone
  45093. 34:05:47ahead and uh run the imports. Now I'm
  45094. 34:05:51going to paste the code down here
  45095. 34:05:54and let's take a look and see what's
  45096. 34:05:55going on. The first thing we're doing is
  45097. 34:05:57we're actually loading the iris data.
  45098. 34:06:00And if you remember up here, we loaded
  45099. 34:06:02the module that tells it how to get the
  45100. 34:06:04iris data. Now we're actually assigning
  45101. 34:06:06that data to the variable iris. And then
  45102. 34:06:08we're going to go ahead and use the df
  45103. 34:06:11to define dataf frame. And that's going
  45104. 34:06:13to equal pd. And if you remember that's
  45105. 34:06:15pandas as pd. So that's our pandas and
  45106. 34:06:19panda dataf frame. And then we're
  45107. 34:06:20looking at iris data and columns equals
  45108. 34:06:24iris feature names. And we're going to
  45109. 34:06:26do the DF head. And let's run this so
  45110. 34:06:29you can understand what's going on here.
  45111. 34:06:32The first thing you want to notice is
  45112. 34:06:34that our DF has created uh what looks
  45113. 34:06:36like an Excel spreadsheet. And in this
  45114. 34:06:39Excel spreadsheet, we have set the
  45115. 34:06:40columns. So up on the top, you can see
  45116. 34:06:42the four different columns. And then we
  45117. 34:06:45have the data iris.data down below. It's
  45118. 34:06:47a little confusing without knowing where
  45119. 34:06:49this data is coming from. So let's look
  45120. 34:06:51at the bigger picture and I'm going to
  45121. 34:06:53go print. I'm just going to change this
  45122. 34:06:55for a moment and we're going to print
  45123. 34:06:57all of Iris and see what that looks
  45124. 34:06:59like. So when I print all of Iris I get
  45125. 34:07:02this long list of information. And you
  45126. 34:07:05can scroll through here and see all the
  45127. 34:07:07different titles on there. What's
  45128. 34:07:10important to notice is that first off
  45129. 34:07:11there's a brackets at the beginning. So
  45130. 34:07:13this is a Python dictionary
  45131. 34:07:16and in a Python dictionary you'll have a
  45132. 34:07:19key or a label and this label pulls up
  45133. 34:07:23whatever information comes after it. So
  45134. 34:07:25feature names which we actually used
  45135. 34:07:27over here under columns is equal to an
  45136. 34:07:30array of sele length sele width pedal
  45137. 34:07:32length pedal width. These are the
  45138. 34:07:34different names they have for the four
  45139. 34:07:36different columns. And if you scroll
  45140. 34:07:38down far enough you'll also see data
  45141. 34:07:40down here. Oh goodness, it came up right
  45142. 34:07:42towards the top. And uh data is equal to
  45143. 34:07:44the different data we're looking at.
  45144. 34:07:47Now, there's a lot of other things in
  45145. 34:07:48here like target. We're going to be
  45146. 34:07:50pulling that up in a minute. And there's
  45147. 34:07:51also the names uh the target names which
  45148. 34:07:54is further down. And we'll show you that
  45149. 34:07:55also in a minute. Let's go ahead and set
  45150. 34:07:57that back to the head. And this is one
  45151. 34:08:01of the neat features of pandas and panda
  45152. 34:08:04dataf frames is when you do df.ad or the
  45153. 34:08:08panda dataf frame. head. It'll print the
  45154. 34:08:11first five lines of the data set in
  45155. 34:08:14there along with the headers if you have
  45156. 34:08:16them. In this case, we have the column
  45157. 34:08:18headers set to iris features. And in
  45158. 34:08:20here, you'll see that we have 0 1 2 3 4.
  45159. 34:08:24In Python, most arrays always start at
  45160. 34:08:26zero. So, when you look at the first
  45161. 34:08:28five, it's going to be 0 1 2 3 4, not 1
  45162. 34:08:312 3 4 5. So, now we've got our iris data
  45163. 34:08:34imported into a data frame. Let's take a
  45164. 34:08:36look at the next piece of code in here.
  45165. 34:08:38And so in this section here of the code,
  45166. 34:08:41we're going to take a look at the
  45167. 34:08:43target. And let's go ahead and get this
  45168. 34:08:45into our notebook, this piece of code,
  45169. 34:08:47so we can discuss it a little bit more
  45170. 34:08:48in detail. So here we are in our Jupyter
  45171. 34:08:51notebook. I'm going to put the code in
  45172. 34:08:52here. And before I run it, I want to
  45173. 34:08:55look at a couple things going on. So we
  45174. 34:08:57have uh DF species. And this is
  45175. 34:09:00interesting because right here you'll
  45176. 34:09:02see where I have DF species in brackets
  45177. 34:09:05which is uh the key code for creating
  45178. 34:09:07another column. And here we have
  45179. 34:09:09iris.target.
  45180. 34:09:11Now these are both in the pandas setup
  45181. 34:09:13on here. So in pandas we can do either
  45182. 34:09:16one. I could have just as easily done
  45183. 34:09:18iris and then in brackets target
  45184. 34:09:21depending on what I'm working on. Both
  45185. 34:09:23are um acceptable. Let's go ahead and
  45186. 34:09:26run this code and see how this changes.
  45187. 34:09:28And what we've done is we've added the
  45188. 34:09:30target from the iris data set as another
  45189. 34:09:33column on the end.
  45190. 34:09:36Now what species is this is what we're
  45191. 34:09:38trying to predict. So we have our data
  45192. 34:09:40which tells us the answer for all these
  45193. 34:09:42different pieces. And then we've added a
  45194. 34:09:44column with the answer. So that way when
  45195. 34:09:46we do our final setup, we'll have the
  45196. 34:09:48ability to program our our neural
  45197. 34:09:50network to look for these this different
  45198. 34:09:52data and know what a satossa is or a
  45199. 34:09:55veraricolor which we'll see in just a
  45200. 34:09:57minute or virginica. Those are the three
  45201. 34:09:59that are in there. And now we're going
  45202. 34:10:00to add one more column. I know we're
  45203. 34:10:03organizing all this data over and over
  45204. 34:10:05again. It's kind of fun. There's a lot
  45205. 34:10:07of ways to organize it. What's nice
  45206. 34:10:09about putting everything onto one data
  45207. 34:10:12frame is I can then do a print out and
  45208. 34:10:15it shows me exactly what I'm looking at.
  45209. 34:10:16And I'll show you where you where that's
  45210. 34:10:18different where you can alter that and
  45211. 34:10:20do it slightly differently. But let's go
  45212. 34:10:22ahead and put this into our script up to
  45213. 34:10:24now. And here we go. We're going to put
  45214. 34:10:26that down here and we're going to run
  45215. 34:10:29that. And let's talk a little bit about
  45216. 34:10:32what we're doing. Now we're exploring
  45217. 34:10:34data. And one of the challenges is
  45218. 34:10:37knowing how good your model is. Did your
  45219. 34:10:40model work? And to do this, we need to
  45220. 34:10:42split the data. And we split it into two
  45221. 34:10:44different parts. They usually call it
  45222. 34:10:46the training and the testing. And so in
  45223. 34:10:48here, we're going to go ahead and put
  45224. 34:10:50that in our database so you can see it
  45225. 34:10:52clearly. And we've set it df. And
  45226. 34:10:55remember, you can put brackets. This is
  45227. 34:10:56creating another column. Is train. So
  45228. 34:10:58we're going to use part of it for
  45229. 34:10:59training. And this equals np. Remember
  45230. 34:11:01that stands for numpy.random.uniform.
  45231. 34:11:04So we're generating a random number
  45232. 34:11:07between zero and one. And we're going to
  45233. 34:11:09do it for each of the rows. That's where
  45234. 34:11:12the length df comes from. So each row
  45235. 34:11:14gets a generated number. And if it's
  45236. 34:11:16less than 75, it's true. And if it's
  45237. 34:11:19greater than 75, it's false. This means
  45238. 34:11:22we're going to take 75% of the data
  45239. 34:11:26roughly because there's a randomness
  45240. 34:11:27involved. And we're going to use that to
  45241. 34:11:29train it. And then the other 25% we're
  45242. 34:11:32going to hold off to the side and use
  45243. 34:11:33that to test it later on. So let's flip
  45244. 34:11:36back on over and see what the next step
  45245. 34:11:37is. So now that we've labeled our
  45246. 34:11:39database for which is training and which
  45247. 34:11:41is testing, let's go ahead and sort that
  45248. 34:11:44into two different variables, train and
  45249. 34:11:46test. And let's take this code and let's
  45250. 34:11:48bring it into our project. And here we
  45251. 34:11:50go. Let's paste it on down here. And
  45252. 34:11:53before I run this, let's just take a
  45253. 34:11:55quick look at what's going on here. is
  45254. 34:11:58we have up above we created remember
  45255. 34:12:00there's our def head which prints the
  45256. 34:12:02first five rows and we've added a column
  45257. 34:12:04is train at the end and so we're going
  45258. 34:12:06to take that we're going to create two
  45259. 34:12:07variables we're going to create two new
  45260. 34:12:09data frames one's called train one's
  45261. 34:12:12called test 75% in train 25% in test and
  45262. 34:12:18then to sort that out we're going to do
  45263. 34:12:21that by doing df our main original data
  45264. 34:12:24frame with the iris data in it and if df
  45265. 34:12:27F is train equals true, that's going to
  45266. 34:12:30go in the train. And if DF is train
  45267. 34:12:32equals false, it goes in the test. And
  45268. 34:12:35so when I run this, we're going to print
  45269. 34:12:37out the number in each one. Let's see
  45270. 34:12:39what that looks like. And you'll see
  45271. 34:12:41that it puts 118 in the training module
  45272. 34:12:43and it puts 32 in the testing module,
  45273. 34:12:46which lets us know that there was 150
  45274. 34:12:48lines of data in here. So if you went
  45275. 34:12:49and looked at the original data, you
  45276. 34:12:51could see that there's 150 lines and
  45277. 34:12:53that's roughly 75% in one and 25% for us
  45278. 34:12:56to test our model on afterward. So let's
  45279. 34:12:59jump back to our code and see where this
  45280. 34:13:01goes. In the next two steps, we want to
  45281. 34:13:04do one more thing with our data, and
  45282. 34:13:06that's make it readable to humans. Um, I
  45283. 34:13:09don't know about you, but I hate looking
  45284. 34:13:10at zeros and ones. So, let's start with
  45285. 34:13:14the features and let's go ahead and take
  45286. 34:13:17those and make those readable to humans
  45287. 34:13:19and let's put that in our code.
  45288. 34:13:23Let's see. Here we go. Paste it in. And
  45289. 34:13:25you'll see here we've done a couple very
  45290. 34:13:28basic things. We know that the columns
  45291. 34:13:31in our data frame, again, this is a
  45292. 34:13:33panda thing, the DF columns, and we know
  45293. 34:13:37the first four of them, 01, 2, 3, that'd
  45294. 34:13:40be the first four are going to be the
  45295. 34:13:42features or the titles of those columns.
  45296. 34:13:44And so when I run this, you'll see down
  45297. 34:13:47here that it creates an index, sea
  45298. 34:13:49length, sea width, pedal length, and
  45299. 34:13:51pedal width. And this should be familiar
  45300. 34:13:53because if you look up here, here's our
  45301. 34:13:55column titles going across. And here's
  45302. 34:13:57the first four. One thing I want you to
  45303. 34:13:59notice here is that when you're in a
  45304. 34:14:02command line, whether it's Jupyter
  45305. 34:14:03notebook or you're running command line
  45306. 34:14:05in the uh terminal window, if you just
  45307. 34:14:08put the name of it, it'll print it out.
  45308. 34:14:10This is the same as doing print
  45309. 34:14:13features.
  45310. 34:14:15And the shortand is you just put
  45311. 34:14:17features in here. If you're actually
  45312. 34:14:19writing a code and saving the script and
  45313. 34:14:22running it by remote, you really need to
  45314. 34:14:24put the print in there. But for this,
  45315. 34:14:26when I run it, you'll see it gives me
  45316. 34:14:27the same thing.
  45317. 34:14:30But for this, we want to go ahead and
  45318. 34:14:31we'll just leave it as features because
  45319. 34:14:33it doesn't really matter. And this is
  45320. 34:14:34one of the fun thing about Jupyter
  45321. 34:14:36Notebooks is I'm just building the code
  45322. 34:14:37as we go. And then we need to go ahead
  45323. 34:14:39and create the labels for the other
  45324. 34:14:41part. So, let's take a look and see what
  45325. 34:14:42that for. Our final step in prepping our
  45326. 34:14:45data before we actually start running
  45327. 34:14:47the training and the testing is we're
  45328. 34:14:49going to go ahead and convert the
  45329. 34:14:51species on here into something the
  45330. 34:14:53computer understands. So, let's put this
  45331. 34:14:56code into our script and see where that
  45332. 34:14:58takes us.
  45333. 34:15:00All right, here we go. We've set y equal
  45334. 34:15:02to pd.factorize
  45335. 34:15:05train species of zero. So, let's break
  45336. 34:15:09this down just a little bit. We have our
  45337. 34:15:11pandas right here. PD factoriize. What
  45338. 34:15:14is factorized doing? I'm going to come
  45339. 34:15:16back to that in just a second. Let's
  45340. 34:15:18look at what train species is and why
  45341. 34:15:21we're looking at the group zero on
  45342. 34:15:23there. And let's go up here. And here is
  45343. 34:15:26our species.
  45344. 34:15:29Remember this on that? We created this
  45345. 34:15:30whole column here for species. And then
  45346. 34:15:33it has satossa, satossa, satossa,
  45347. 34:15:35satossa. And if you scroll down enough,
  45348. 34:15:37you'd also see virginica and
  45349. 34:15:39veraricolor.
  45350. 34:15:41We need to convert that into something
  45351. 34:15:42the computer understands. Zeros and
  45352. 34:15:44ones. So the train species of zero
  45353. 34:15:48because this is in the format of a of an
  45354. 34:15:51array of arrays. So you have to have the
  45355. 34:15:53zero on the end. And then species is
  45356. 34:15:55just that column. Factoriize goes in
  45357. 34:15:58there and looks at the fact that there's
  45358. 34:16:00only three of them. So when I run this,
  45359. 34:16:02you'll see that Y generates an array
  45360. 34:16:05that's equal to, in this case, it's the
  45361. 34:16:07training set, and it's zeros, ones, and
  45362. 34:16:10twos representing the three different
  45363. 34:16:12kinds of flowers we have. So now we have
  45364. 34:16:14something the computer understands, and
  45365. 34:16:16we have a nice table that we can read
  45366. 34:16:18and understand. And now finally we get
  45367. 34:16:21to actually start doing the predicting.
  45368. 34:16:23So here we go. Uh we have two lines of
  45369. 34:16:27code. Oh my goodness, that was a lot of
  45370. 34:16:29work to get to two lines of code. But
  45371. 34:16:31there is a lot in these two lines of
  45372. 34:16:33code. So let's take a look and see
  45373. 34:16:34what's going on here and put this into
  45374. 34:16:36our full script that we're running. And
  45375. 34:16:39let's paste this in here. And let's take
  45376. 34:16:41a look and see what this is. We have
  45377. 34:16:44we're creating a variable CLF. And we're
  45378. 34:16:46going to set this equal to the random
  45379. 34:16:48forest classifier. And we're passing two
  45380. 34:16:51variables in here. And there's a lot of
  45381. 34:16:52variables you can play with. As far as
  45382. 34:16:55these two are concerned, they're very
  45383. 34:16:56standard. In jobs, all that does is to
  45384. 34:16:59prioritize it. Not something to really
  45385. 34:17:01worry about. Usually when you're doing
  45386. 34:17:03this on your own computer, you do end
  45387. 34:17:05jobs equals 2. If you're working in a
  45388. 34:17:07larger or big data and you need to
  45389. 34:17:09prioritize it differently, this is what
  45390. 34:17:11that number does is it changes your
  45391. 34:17:12priorities and how it's going to run
  45392. 34:17:14across the system and things like that.
  45393. 34:17:16And then the random state is just how it
  45394. 34:17:18starts. Zero is fine for here.
  45395. 34:17:21But uh let's go ahead and run this.
  45396. 34:17:24We also have clf.fit train features, y.
  45397. 34:17:29And before we run it, let's talk about
  45398. 34:17:31this a little bit more. CLF.fit.
  45399. 34:17:35So, we're fitting, we're training it. We
  45400. 34:17:37are actually creating our random forest
  45401. 34:17:40classifier right here. This is the code
  45402. 34:17:43that does everything. And we're going to
  45403. 34:17:44take our training set. Remember, we kept
  45404. 34:17:46our test off to the side. And we're
  45405. 34:17:48going to take our training set with the
  45406. 34:17:50features. And then we're going to go
  45407. 34:17:51ahead and put that in. And here's our
  45408. 34:17:53target, the Y. So, the Y is 0, 1, and
  45409. 34:17:57two that we just created. And the
  45410. 34:17:59features is the actual data going in
  45411. 34:18:02that we put into the training set. And
  45412. 34:18:04let's go ahead and run that.
  45413. 34:18:07And this is kind of an interesting thing
  45414. 34:18:08because it printed out the random force
  45415. 34:18:11classifier
  45416. 34:18:13and everything around it. And so when
  45417. 34:18:16you're running this in your terminal
  45418. 34:18:18window or in a script like this, this
  45419. 34:18:20automatically treats this like just like
  45420. 34:18:22when we were up here and I typed in y
  45421. 34:18:24and it printed out y instead of print y.
  45422. 34:18:27This does the same thing. It treats this
  45423. 34:18:29as a variable and prints it out. But if
  45424. 34:18:32you were actually running your code,
  45425. 34:18:33that wouldn't be the case. And what is
  45426. 34:18:35printed out is it shows us all the
  45427. 34:18:37different variables we can change. And
  45428. 34:18:40if we go down here, you can actually see
  45429. 34:18:41in jobs equals 2. You can see the random
  45430. 34:18:44state equals zero. Those are the two
  45431. 34:18:46that we sent in there. You would really
  45432. 34:18:47have to dig deep to find out all these
  45433. 34:18:50different meanings of all these
  45434. 34:18:51different settings on here. Some of them
  45435. 34:18:53are self-explanatory if you kind of
  45436. 34:18:55think about it a little bit. Like max
  45437. 34:18:56features is auto. So all the features
  45438. 34:18:58that we're putting in there, it's just
  45439. 34:19:00going to automatically take all four of
  45440. 34:19:02them. Whatever we send it, it'll take.
  45441. 34:19:03Some of them might have so many features
  45442. 34:19:05because you're processing words. There
  45443. 34:19:07might be like 1.4 million features in
  45444. 34:19:10there because you're doing legal
  45445. 34:19:11documents and that's how many different
  45446. 34:19:12words are in there. At that point, you
  45447. 34:19:14probably want to limit the maximum
  45448. 34:19:16features that you're going to process.
  45449. 34:19:17And leaf nodes, that's the end nodes.
  45450. 34:19:19Remember, we had the fruit and we're
  45451. 34:19:20talking about the leaf nodes. Like I
  45452. 34:19:22said, there's a lot in this. We're
  45453. 34:19:24looking at a lot of stuff here. So you
  45454. 34:19:25might have uh in this case there's
  45455. 34:19:27probably only think three leaf nodes,
  45456. 34:19:29maybe four. You might have thousands of
  45457. 34:19:31leaf nodes at which point you do need to
  45458. 34:19:32put a cap on that and say, "Okay, you
  45459. 34:19:34can only go so far and then we're going
  45460. 34:19:35to use all of our resources on
  45461. 34:19:37processing this." And that really is
  45462. 34:19:39what most of these are about is limiting
  45463. 34:19:42the process and making sure we don't uh
  45464. 34:19:45overwhelm a system. And there's some
  45465. 34:19:47other settings in here. Again, we're not
  45466. 34:19:48going to go over all of them. Warm start
  45467. 34:19:50equals false. or start as if you're
  45468. 34:19:52programming it one piece at a time
  45469. 34:19:54externally since we're not we're not
  45470. 34:19:56going to have like we're not going to
  45471. 34:19:58continually to train this particular
  45472. 34:19:59learning tree and again like I said
  45473. 34:20:01there's a lot of things in here that
  45474. 34:20:02you'll want to look up more detail from
  45475. 34:20:04the sklearn and if you're digging in
  45476. 34:20:07deep and running a major project on here
  45477. 34:20:09for today though all we need to do is
  45478. 34:20:11fit or train our features and our target
  45479. 34:20:13Y. So now we have our training model.
  45480. 34:20:16What's next? If we're going to create a
  45481. 34:20:18model,
  45482. 34:20:20we now need to test it. Remember, we set
  45483. 34:20:23aside the test features, test group, 25%
  45484. 34:20:27of the data. So let's go ahead and take
  45485. 34:20:28this code and let's put it into our uh
  45486. 34:20:31script and see what that looks like.
  45487. 34:20:33Okay, here we go. And we're going to run
  45488. 34:20:35this.
  45489. 34:20:38And it's going to come out with a bunch
  45490. 34:20:39of zeros, ones, and twos, which
  45491. 34:20:42represents the three type of flowers,
  45492. 34:20:44the satossa, the virginica, and the
  45493. 34:20:45versa color. And what we're putting into
  45494. 34:20:47our predict is the test features. And I
  45495. 34:20:51always kind of like to know what it is I
  45496. 34:20:53am looking at. So, real quick, we're
  45497. 34:20:56going to do test
  45498. 34:20:58features. And remember, features is an
  45499. 34:21:01array
  45500. 34:21:03of sele
  45501. 34:21:05width, pedal length, pedal width. So
  45502. 34:21:07when we put it in this way, it actually
  45503. 34:21:09loads all these different columns that
  45504. 34:21:11we loaded into features. So if we did
  45505. 34:21:13just features, let me just do features
  45506. 34:21:15in here so you can see what features
  45507. 34:21:16looks like. This is just playing with
  45508. 34:21:18the with Panda's data frames. You'll see
  45509. 34:21:21that it's an index. So when you put an
  45510. 34:21:23index in like this
  45511. 34:21:27into test features into test, it then
  45512. 34:21:31takes those columns and creates a Panda
  45513. 34:21:34data frames from those columns. And in
  45514. 34:21:36this case, we're going to go ahead and
  45515. 34:21:39put those into our predict. So, we're
  45516. 34:21:42going to put each one of these lines of
  45517. 34:21:43data, the 5.0, 3.4, 1.5, point2, and
  45518. 34:21:48we're going to put those in, and we're
  45519. 34:21:49going to predict what our new um forest
  45520. 34:21:53classifier is going to come up with. And
  45521. 34:21:55this is what it predicts. It predicts uh
  45522. 34:21:570000121122.
  45523. 34:22:00and and uh again this is the flower type
  45524. 34:22:04satossa vica and versa color. So now
  45525. 34:22:07that we've taken our test features let's
  45526. 34:22:10explore that. Let's see exactly what
  45527. 34:22:12that data means to us. So the first
  45528. 34:22:14thing we can do with our predicts is we
  45529. 34:22:17can actually generate a different
  45530. 34:22:19prediction model. When I say different,
  45531. 34:22:21we're going to view it differently. It's
  45532. 34:22:23not that the data itself is different.
  45533. 34:22:25So let's take this next piece of code
  45534. 34:22:26and put it into our script.
  45535. 34:22:29So we're pasting it in here and you'll
  45536. 34:22:31see that we're doing uh predict and
  45537. 34:22:33we've added underscore proba for
  45538. 34:22:36probability. So there's our clff.predict
  45539. 34:22:39probability. So we're we're running it
  45540. 34:22:41just like we ran it up here, but this
  45541. 34:22:43time with this we're going to get a
  45542. 34:22:45slightly different result and we're only
  45543. 34:22:47going to look at the first 10. So you'll
  45544. 34:22:50see down here instead of looking at all
  45545. 34:22:51of them uh which was uh what 27 you'll
  45546. 34:22:54see right down here that this generates
  45547. 34:22:56a much larger field on the probability
  45548. 34:22:59and let's take a look and see what that
  45549. 34:23:00looks like and what that means. So when
  45550. 34:23:04we do the predict underscore probaba for
  45551. 34:23:07probability it generates three numbers.
  45552. 34:23:10So we had three leaf nodes at the end
  45553. 34:23:12and if you remember from all the theory
  45554. 34:23:14we did this is the predictors. The first
  45555. 34:23:17one is predicting a one for satossa. It
  45556. 34:23:21predicts a zero for virginica. And it
  45557. 34:23:24predicts a zero for versol. And so on
  45558. 34:23:27and so on and so on. And let's um you
  45559. 34:23:29know what? I'm going to change this just
  45560. 34:23:30a little bit. Let's look at 10
  45561. 34:23:33to 20 just because we can.
  45562. 34:23:37And we start to get in a little
  45563. 34:23:38different of data. And you'll see right
  45564. 34:23:40down here it gets to this one. This line
  45565. 34:23:42right here. And this line has zero 0.5
  45566. 34:23:460.5.
  45567. 34:23:48And so if we're going to vote and we
  45568. 34:23:50have two equal votes, it's going to go
  45569. 34:23:51with the first one. So it says uh
  45570. 34:23:53Satossa gets zero votes, virginica
  45571. 34:23:56gets.5 votes, VersaColor gets.5 votes,
  45572. 34:23:59but let's just go with the virginica
  45573. 34:24:02since these two are equal and so on and
  45574. 34:24:04so on down the list. You can see how
  45575. 34:24:05they vary on here. So now we've looked
  45576. 34:24:07at both how to do a basic predict of the
  45577. 34:24:09features and we've looked at the predict
  45578. 34:24:12probability. Let's see what's next on
  45579. 34:24:14here. So now we want to go ahead and
  45580. 34:24:17start mapping names for the plants. We
  45581. 34:24:19want to attach names so that it makes a
  45582. 34:24:21little more sense for us. And that's
  45583. 34:24:23what we're going to do in these next two
  45584. 34:24:24steps. We're going to start by setting
  45585. 34:24:27up our predictions and mapping them to
  45586. 34:24:30the name. So let's see what that looks
  45587. 34:24:32like. And let's go ahead and paste that
  45588. 34:24:35code in here and run it. And this goes
  45589. 34:24:37along with the next piece of code. So
  45590. 34:24:39we'll skip through this quickly and then
  45591. 34:24:40come back to it a little bit. So, here's
  45592. 34:24:42iris.target
  45593. 34:24:44names.
  45594. 34:24:47And uh if you remember correctly, this
  45595. 34:24:49was the the names that we've been
  45596. 34:24:51talking about this whole time, the
  45597. 34:24:52Satossa, Vica, VersaColor. And then
  45598. 34:24:55we're going to go ahead and do the
  45599. 34:24:56prediction again. We've run it. We could
  45600. 34:24:58have just set a variable equal to this
  45601. 34:25:00instead of rerunning it each time, but
  45602. 34:25:01we're going ahead and run it again.
  45603. 34:25:02CLF.predict test features. Remember that
  45604. 34:25:06returns the zeros, the ones, and the
  45605. 34:25:07twos. And then we're going to set that
  45606. 34:25:09equal to predictions. So this time we're
  45607. 34:25:12actually putting it in a variable. And
  45608. 34:25:14when I run this,
  45609. 34:25:16it distributes and it comes out as an
  45610. 34:25:18array. And the array is satossa,
  45611. 34:25:20satossa, satossa, satossa, satossa.
  45612. 34:25:22We're only looking at the first five. We
  45613. 34:25:24could actually do let's do the first 25
  45614. 34:25:27just so we can see a little bit more on
  45615. 34:25:28there. And you'll see that it starts
  45616. 34:25:30mapping it to all the different flower
  45617. 34:25:32types, the versa color and the virginica
  45618. 34:25:34in there. And let's see how this goes
  45619. 34:25:36with the next one. So, let's take a look
  45620. 34:25:38at the top part of our species in here.
  45621. 34:25:41And we'll take this code and put it in
  45622. 34:25:43our script.
  45623. 34:25:45And let's put that down here and paste
  45624. 34:25:47it. There we go. And we'll go ahead and
  45625. 34:25:49run it. And let's talk about both these
  45626. 34:25:51sections of code here and how they go
  45627. 34:25:54together. The first one is our
  45628. 34:25:57predictions. And I went ahead and did uh
  45629. 34:25:59predictions through 25. Let's just do
  45630. 34:26:01five.
  45631. 34:26:03And so we have stosis, satossis, stosis,
  45632. 34:26:05satossis. That's what we're predicting
  45633. 34:26:06from our test model. And then we come
  45634. 34:26:09down here and we look at test species.
  45635. 34:26:12And remember, I could have just done
  45636. 34:26:13test.species.head.
  45637. 34:26:15And you'll see it says Satossa, Satossa,
  45638. 34:26:17Satossa, Satossa. And they match. So the
  45639. 34:26:20first one is what our forest is doing
  45640. 34:26:24and the second one is what the actual
  45641. 34:26:27data is. Now is we need to combine these
  45642. 34:26:30so that we can understand what that
  45643. 34:26:31means. We need to know how good our
  45644. 34:26:33forest is, how good it is at predicting
  45645. 34:26:34the features. So that's where we come up
  45646. 34:26:36to the next step, which is lots of fun.
  45647. 34:26:39We're going to use a single line of code
  45648. 34:26:41to combine our predictions and our
  45649. 34:26:43actuals so we have a nice chart to look
  45650. 34:26:46at. And let's go ahead and put that in
  45651. 34:26:47our script in our Jupyter notebook here.
  45652. 34:26:50Let's see. Let's go ahead and paste that
  45653. 34:26:51in. And then I'm going to because I'm on
  45654. 34:26:54the Jupyter notebook, I can do a control
  45655. 34:26:56minus so we can see the whole line
  45656. 34:26:57there.
  45657. 34:26:59There we go. resize it and let's take a
  45658. 34:27:02look and see what's going on here. We're
  45659. 34:27:04going to create in pandas. Remember PD
  45660. 34:27:06stands for pandas and we're doing a
  45661. 34:27:08cross tab. This function takes two sets
  45662. 34:27:11of data and creates a chart out of them.
  45663. 34:27:13So when I run it, you'll get a nice
  45664. 34:27:14chart down here. And we have the
  45665. 34:27:17predicted species.
  45666. 34:27:19So across the top you'll see the satossa
  45667. 34:27:21versus color virginica and the actual
  45668. 34:27:24species satossa versus color virginica.
  45669. 34:27:27And so the way to read this chart and
  45670. 34:27:29let's go ahead and take a look on how to
  45671. 34:27:31read this chart here. When you read this
  45672. 34:27:33chart, you have satossa where they meet,
  45673. 34:27:35you have versolar where they meet, and
  45674. 34:27:37you have virginica where they meet. And
  45675. 34:27:39they're meeting where the actual and the
  45676. 34:27:41predicted agree. So this is the number
  45677. 34:27:44of accurate predictions. So in this
  45678. 34:27:46case, it equals 30. If you add 13 + 5 +
  45679. 34:27:4912, you get 30. And then we notice here
  45680. 34:27:52where it says virginica, but it was
  45681. 34:27:54supposed to be versol. This is
  45682. 34:27:55inaccurate. So now we have two two
  45683. 34:27:58inaccurate predictions and 30 accurate
  45684. 34:28:01predictions. So we'll say that the model
  45685. 34:28:03accuracy is 93. That's just 30 divided
  45686. 34:28:07by 32. And if we multiply it by 100, we
  45687. 34:28:11can say that it is 93% accurate. So we
  45688. 34:28:14have a 93% accuracy with our model. I
  45689. 34:28:18did want to add one more quick thing in
  45690. 34:28:20here on our scripting before we wrap it
  45691. 34:28:22up. So let's flip back on over to my
  45692. 34:28:24script. in here. We're going to take
  45693. 34:28:26this uh line of code from up above. I
  45694. 34:28:29don't know if you remember it, but
  45695. 34:28:30predicts equals the iris.target_names.
  45696. 34:28:34So, we're going to map it to the names
  45697. 34:28:36and we're going to run the prediction.
  45698. 34:28:38And we read it on test features. But,
  45699. 34:28:40you know, we're not just testing it. We
  45700. 34:28:42want to actually deploy it. So, at this
  45701. 34:28:43point, I would go ahead and change this.
  45702. 34:28:46And this is an array of arrays. This is
  45703. 34:28:48really important when you're running
  45704. 34:28:50these to know that. So, you need the
  45705. 34:28:52double brackets. And I could actually
  45706. 34:28:54create data. Maybe let's let's just do
  45707. 34:28:56two flowers. So maybe I'm processing
  45708. 34:28:58more data coming in. And we'll put two
  45709. 34:29:00flowers in here. And then uh I actually
  45710. 34:29:03want to see what the answer is. So let's
  45711. 34:29:06go ahead and type in PRS and print that
  45712. 34:29:08out. And when I run this, you'll see
  45713. 34:29:11that I've now predicted two flowers that
  45714. 34:29:13maybe I measured in my front yard as
  45715. 34:29:15VersaColor and VersaColor.
  45716. 34:29:18Not surprising since I put the same data
  45717. 34:29:20in for each one. This would be the
  45718. 34:29:22actual uh end product going out to be
  45719. 34:29:25used on data that you don't know the
  45720. 34:29:27answer for.
  45721. 34:29:30So that's going to conclude our
  45722. 34:29:31scripting part of this. Introducing
  45723. 34:29:33naive base classifier. Have you ever
  45724. 34:29:36wondered how your mail provider
  45725. 34:29:38implements spam filtering or how online
  45726. 34:29:40news channels perform news text
  45727. 34:29:42classification or how companies perform
  45728. 34:29:44sentimental analysis of their audience
  45729. 34:29:46on social media? All of this and more is
  45730. 34:29:48done through a machine learning
  45731. 34:29:50algorithm called naive bay classifier.
  45732. 34:29:53Welcome to Naive Bay tutorial. My name
  45733. 34:29:56is Richard Kersner. I'm with the
  45734. 34:29:58SimplyLearn team. That's
  45735. 34:29:59www.simplearn.com.
  45736. 34:30:02Get certified get ahead. What's in it
  45737. 34:30:04for you? We'll start with what is naive
  45738. 34:30:07bays? A basic overview of how it works.
  45739. 34:30:10We'll get into naive bays and machine
  45740. 34:30:12learning where it fits in with our other
  45741. 34:30:14machine learning tools. Why do we need
  45742. 34:30:16naive bays and understanding naive bays
  45743. 34:30:18classifier a much more in-depth of how
  45744. 34:30:20the math works in the background?
  45745. 34:30:22Finally, we'll get into the advantages
  45746. 34:30:24of the naive bay classifier in the
  45747. 34:30:26machine learning setup. And then we'll
  45748. 34:30:28roll up our sleeves and do my favorite
  45749. 34:30:30part. We'll actually do some Python
  45750. 34:30:31coding and do some text classification
  45751. 34:30:34using the naive bays. What is naive
  45752. 34:30:36bays? Let's start with a basic
  45753. 34:30:38introduction to the bay theorem named
  45754. 34:30:40after Thomas Bae from the 1700s who
  45755. 34:30:43first coined this in the western
  45756. 34:30:44literature. Naive bay classifier works
  45757. 34:30:47on the principle of conditional
  45758. 34:30:48probability as given by the bay theorem.
  45759. 34:30:50Before we move ahead, let us go through
  45760. 34:30:52some of the simple concepts in the
  45761. 34:30:54probability that we will be using. Let
  45762. 34:30:56us consider the following example of
  45763. 34:30:57tossing two coins. Here we have two
  45764. 34:31:00quarters and if we look at all the
  45765. 34:31:02different possibilities of what they can
  45766. 34:31:03come up as, we get that they could come
  45767. 34:31:05up as head heads. come up as head, tail,
  45768. 34:31:07tail, head and tell tail. When doing the
  45769. 34:31:09math on probability, we usually denote
  45770. 34:31:12probability as a P, a capital P. So the
  45771. 34:31:15probability of getting two heads equals
  45772. 34:31:161/4. You can see in our data set, we
  45773. 34:31:19have two heads and this occurs once out
  45774. 34:31:21of the four possibilities. And then the
  45775. 34:31:23probability of at least one tail occurs
  45776. 34:31:25three/arters of the time. You'll see on
  45777. 34:31:27three of the coin tosses, we have tails
  45778. 34:31:29in them. And out of four, that's
  45779. 34:31:30three/4s. And then the probability of
  45780. 34:31:32the second coin being a head given the
  45781. 34:31:35first coin is tail is 1/2. And the
  45782. 34:31:38probability of getting two heads given
  45783. 34:31:40the first coin is a head is 1/2. We'll
  45784. 34:31:42demonstrate that in just a minute and
  45785. 34:31:44show you how that math works. Now when
  45786. 34:31:45we're doing it with two coins, it's easy
  45787. 34:31:47to see. But when you have something more
  45788. 34:31:49complex, you can see where these pro
  45789. 34:31:50these formulas really come in and work.
  45790. 34:31:53So the base theorem gives us the
  45791. 34:31:55conditional probability of an event A
  45792. 34:31:57given another event B has occurred. In
  45793. 34:32:00this case, the first coin toss will be B
  45794. 34:32:03and the second coin toss A. This could
  45795. 34:32:05be confusing because we've actually
  45796. 34:32:06reversed the order of them and go from B
  45797. 34:32:09to A instead of A to B. You'll see this
  45798. 34:32:11a lot when you work in probabilities.
  45799. 34:32:13The reason is we're looking for event A,
  45800. 34:32:15we want to know what that is. So, we're
  45801. 34:32:17going to label that A since that's our
  45802. 34:32:18focus. And then given another event B
  45803. 34:32:21has occurred. In the Baze theorem, as
  45804. 34:32:23you can see on the left, the probability
  45805. 34:32:25of A occurring given B has occurred
  45806. 34:32:28equals the probability of B occurring
  45807. 34:32:30given A has occurred times the
  45808. 34:32:32probability of A over the probability of
  45809. 34:32:34B. This simple formula can be moved
  45810. 34:32:36around just like any algebra formula.
  45811. 34:32:38And we could do the probability of A
  45812. 34:32:40after given B times probability of B
  45813. 34:32:43equals the probability of B given A
  45814. 34:32:46times probability of A. You can easily
  45815. 34:32:47move that around and multiply it and
  45816. 34:32:49divide it out. Let us apply B theorem to
  45817. 34:32:52our example. Here we have our two
  45818. 34:32:53quarters and we'll notice that the first
  45819. 34:32:55two probabilities of getting two heads
  45820. 34:32:58and at least one tail we compute
  45821. 34:33:00directly off the data. So you can easily
  45822. 34:33:02see that we have one example hh out of
  45823. 34:33:06four 1/4 and we have three with tails in
  45824. 34:33:09them giving us three quarters or 3/4
  45825. 34:33:1175%. The second condition the second uh
  45826. 34:33:15set three and four we're going to
  45827. 34:33:16explore a little bit more in detail.
  45828. 34:33:18Now, we stick to a simple example with
  45829. 34:33:20two coins because you can easily
  45830. 34:33:21understand the math. The probability of
  45831. 34:33:23throwing a tail doesn't matter what
  45832. 34:33:25comes before it. And the same with the
  45833. 34:33:26heads. So, it's still going to be 50% or
  45834. 34:33:291/2. But when that come when that
  45835. 34:33:31probability gets more complicated, let's
  45836. 34:33:32say you have a d6 dice or some other
  45837. 34:33:35instance, then this formula really comes
  45838. 34:33:37in handy. But let's stick to the simple
  45839. 34:33:38example for now. In this sample space,
  45840. 34:33:41let A be the event that the second coin
  45841. 34:33:43is head and b be the event that the
  45842. 34:33:45first coin is tails. Again, we reversed
  45843. 34:33:47it because we want to know what the
  45844. 34:33:48second event's going to be. So, we're
  45845. 34:33:49going to be focusing on A. And we write
  45846. 34:33:51that out as the probability of A given
  45847. 34:33:54B. And we know this from our formula
  45848. 34:33:56that that equals the probability of B
  45849. 34:33:58given A times the probability of A over
  45850. 34:34:00the probability of B. And when we plug
  45851. 34:34:02that in, we plug in the probability of
  45852. 34:34:04the first coin being tails given the
  45853. 34:34:06second coin is heads and the probability
  45854. 34:34:08of the second coin being heads given the
  45855. 34:34:10first coin being over the probability of
  45856. 34:34:12the first coin being tails. When we plug
  45857. 34:34:14that data in and we have the probability
  45858. 34:34:16of the first coin being tails given the
  45859. 34:34:18second coin is heads times the
  45860. 34:34:20probability of the second coin being
  45861. 34:34:22heads over the probability of the first
  45862. 34:34:24coin being tails. You can see it's a
  45863. 34:34:25simple formula to calculate. We have 1/2
  45864. 34:34:28* 1/2 over 1/2 or 1/2 =.5 or 1/4. So the
  45865. 34:34:34B theorem basically calculates the
  45866. 34:34:36conditional probability of the
  45867. 34:34:38occurrence of an event based on prior
  45868. 34:34:40knowledge of conditions that might be
  45869. 34:34:42related to the event. We will explore
  45870. 34:34:44this in detail when we take up an
  45871. 34:34:45example of online shopping further in
  45872. 34:34:47this tutorial. Understanding naive bays
  45873. 34:34:49and machine learning. Like with any of
  45874. 34:34:51our other machine learning tools, it's
  45875. 34:34:53important to understand where the naive
  45876. 34:34:55bays fits in the hierarchy. So under the
  45877. 34:34:57machine learning, we have supervised
  45878. 34:34:59learning and there is other things like
  45879. 34:35:00unsupervised learning. There's also
  45880. 34:35:02reward system. This falls under the
  45881. 34:35:04supervised learning. And then under the
  45882. 34:35:06supervised learning, there's
  45883. 34:35:07classification. There's also regression.
  45884. 34:35:09But we're going to be in the
  45885. 34:35:10classification side. And then under
  45886. 34:35:12classification is your naive bays. Let's
  45887. 34:35:15go ahead and glance into where is naive
  45888. 34:35:18bays used. Let's look at some of the use
  45889. 34:35:20scenarios for it. As a classifier, we
  45890. 34:35:22use it in face recognition. Is this
  45891. 34:35:24Cindy or is it not Cindy or whoever? Or
  45892. 34:35:27it might be used to identify parts of
  45893. 34:35:29the face that they then feed into
  45894. 34:35:30another part of the face recognition
  45895. 34:35:32program. This is the eye. This is the
  45896. 34:35:34nose. This is the mouth. Weather
  45897. 34:35:36prediction. Is it going to be rainy or
  45898. 34:35:37sunny? Medical recognition. News
  45899. 34:35:39prediction. It's also used in medical
  45900. 34:35:41diagnosis. We might diagnose somebody as
  45901. 34:35:44either as high risk or not as high risk
  45902. 34:35:46for cancer or heart disease or other
  45903. 34:35:49ailments. And news classification you
  45904. 34:35:51look at the Google news and it says well
  45905. 34:35:53is this political or is this world news
  45906. 34:35:56or a lot of that's all done with the
  45907. 34:35:58naive bays. Understanding naive bay
  45908. 34:36:01classifier. Now we already went through
  45909. 34:36:03a basic understanding with the coins and
  45910. 34:36:06the two heads and two tails and head
  45911. 34:36:08tail tail heads etc. We're going to do
  45912. 34:36:10just a quick review on that and remind
  45913. 34:36:12you that the naive bay classifier is
  45914. 34:36:14based on the bay theorem which gives a
  45915. 34:36:16conditional probability of event A given
  45916. 34:36:20event B. And that's where the
  45917. 34:36:21probability of A given B equals the
  45918. 34:36:24probability of B given A times
  45919. 34:36:26probability of A over probability of B.
  45920. 34:36:28Remember this is an algebraic function
  45921. 34:36:30so we can move these different entities
  45922. 34:36:32around. We could multiply by the
  45923. 34:36:34probability of B. So it goes to the left
  45924. 34:36:36hand side and then we could divide by
  45925. 34:36:38the probability of A given B and just as
  45926. 34:36:40easily come up with a new formula for
  45927. 34:36:42the probability of B. To me staring at
  45928. 34:36:44these algebraic functions kind of gives
  45929. 34:36:46me a slight headache. It's a lot better
  45930. 34:36:49to see if we can actually understand how
  45931. 34:36:50this data fits together in a table. And
  45932. 34:36:52let's go ahead and start applying it to
  45933. 34:36:54some actual data so you can see what
  45934. 34:36:55that looks like. So, we're going to
  45935. 34:36:57start with the shopping demo problem
  45936. 34:36:59statement. And remember, we're going to
  45937. 34:37:00solve this first in a table form so you
  45938. 34:37:02can see what the math looks like. And
  45939. 34:37:04then we're going to solve it in Python.
  45940. 34:37:06And in here, we want to predict whether
  45941. 34:37:07the person will purchase a product. Are
  45942. 34:37:09they going to buy or don't buy? Very
  45943. 34:37:11important. If you're running a business,
  45944. 34:37:12you want to know how to maximize your
  45945. 34:37:14profits or at least maximize the
  45946. 34:37:16purchase of the people coming into your
  45947. 34:37:17store. And we're going to look at a
  45948. 34:37:19specific combination of different
  45949. 34:37:21variables. In this case, we're going to
  45950. 34:37:23look at the day, the discount, and the
  45951. 34:37:25free delivery. And you can see here
  45952. 34:37:26under the day we want to know whether
  45953. 34:37:28it's uh on the weekday, you know,
  45954. 34:37:29somebody's working, they come in after
  45955. 34:37:31work or maybe they don't work. Weekend,
  45956. 34:37:33you can see the bright colors coming
  45957. 34:37:34down there celebrating not being in work
  45958. 34:37:36or holiday. And did we offer a discount
  45959. 34:37:39that day? Yes or no. Did we offer free
  45960. 34:37:41delivery that day? Yes or no. And from
  45961. 34:37:43this, we want to know whether the
  45962. 34:37:44person's going to buy based on these
  45963. 34:37:45traits so we can maximize them and find
  45964. 34:37:48out the best system for getting somebody
  45965. 34:37:49to come in and purchase our goods and
  45966. 34:37:51products from our store. Now, having a
  45967. 34:37:53nice visual is great, but we do need to
  45968. 34:37:55dig into the data. So, let's go ahead
  45969. 34:37:57and take a look at the data set. We have
  45970. 34:37:59a small sample data set of 30 rows.
  45971. 34:38:01We're showing you the first 15 of those
  45972. 34:38:03rows for this demo. Now, the actual data
  45973. 34:38:05file you can request. Just type in below
  45974. 34:38:08under the comments on the YouTube video
  45975. 34:38:10and we'll send you some more information
  45976. 34:38:11and send you that file. As you can see
  45977. 34:38:13here, the file is very simple columns
  45978. 34:38:15and rows. We have the day, the discount,
  45979. 34:38:18the free delivery, and did the person
  45980. 34:38:20purchase or not. And then we have under
  45981. 34:38:22the day whether it was a weekday, a
  45982. 34:38:24holiday, was it the weekend? This is a
  45983. 34:38:26pretty simple set of data. And long
  45984. 34:38:29before computers, people used to look at
  45985. 34:38:30this data and calculate this all by
  45986. 34:38:32hand. So let's go ahead and walk through
  45987. 34:38:34this and see what that looks like when
  45988. 34:38:36we put that into tables. Also note in
  45989. 34:38:38today's world, we're not usually looking
  45990. 34:38:40at three different variables and 30
  45991. 34:38:42rows. Nowadays, because we're able to
  45992. 34:38:44collect data so much, we're usually
  45993. 34:38:45looking at 27, 30 variables across
  45994. 34:38:48hundreds of rows. The first thing we
  45995. 34:38:51want to do is we're going to take this
  45996. 34:38:52data and uh based on the data set
  45997. 34:38:55containing our three inputs day,
  45998. 34:38:57discount, and free delivery, we're going
  45999. 34:38:58to go ahead and populate that to
  46000. 34:39:00frequency tables for each attribute. So,
  46001. 34:39:02we want to know if they had a discount,
  46002. 34:39:04how many people buy and did not buy. Uh
  46003. 34:39:07did they have a discount? Yes or no. Do
  46004. 34:39:09we have a free delivery? Yes or no. On
  46005. 34:39:11those days, how many people made a
  46006. 34:39:13purchase and how many people didn't? And
  46007. 34:39:14the same with the three days of the
  46008. 34:39:16week. Was it a weekday, a weekend, a
  46009. 34:39:17holiday? And did they buy? Yes or no? As
  46010. 34:39:20we dig in deeper to this table for our
  46011. 34:39:22bay theorem, let the event buy be a. Now
  46012. 34:39:25remember when we looked at the coins, I
  46013. 34:39:26said we really want to know what the
  46014. 34:39:27outcome is. Did the person buy or not?
  46015. 34:39:30And that's usually event A is what
  46016. 34:39:32you're looking for. And the independent
  46017. 34:39:33variables, discount, free delivery, and
  46018. 34:39:35day be B. So we'll call that probability
  46019. 34:39:38of B. Now let us calculate the
  46020. 34:39:40likelihood table for one of the
  46021. 34:39:41variables. Let's start with day, which
  46022. 34:39:44includes weekday, weekend, and holiday.
  46023. 34:39:46And let us start by summing all of our
  46024. 34:39:48rows. So, we have the uh weekday row,
  46025. 34:39:51and out of the weekdays, there's 9 plus
  46026. 34:39:532, so there's 11 weekdays. There's eight
  46027. 34:39:55weekend days and 11 holidays. Wow,
  46028. 34:39:58that's a lot of holidays. And then we
  46029. 34:39:59want to sum up the total number of days.
  46030. 34:40:01So, we're looking at a total of 30 days.
  46031. 34:40:04Let's start pulling some information
  46032. 34:40:06from our chart and see where that takes
  46033. 34:40:08us. And when we fill in the chart on the
  46034. 34:40:10right, you can see that nine out of 24
  46035. 34:40:13purchases are made on the weekday, 7 out
  46036. 34:40:15of 24 purchases on the weekend, and
  46037. 34:40:18eight out of 24 purchases on a holiday.
  46038. 34:40:20And out of all the people who come in,
  46039. 34:40:2224 out of 30 purchase. You can also see
  46040. 34:40:24how many people do not purchase. On the
  46041. 34:40:26weekday, it's two out of six didn't
  46042. 34:40:28purchase and so on and so on. We can
  46043. 34:40:30also look at the totals and you'll see
  46044. 34:40:31on the right, we put together some of
  46045. 34:40:33the formulas. The probability of making
  46046. 34:40:35a purchase on the weekend comes out 11
  46047. 34:40:38out of 30. So out of the 30 people who
  46048. 34:40:40came into the store throughout the
  46049. 34:40:41weekend, weekday and holiday, 11 of
  46050. 34:40:44those purchases were made on the
  46051. 34:40:45weekday. And then you can also see the
  46052. 34:40:47probability of them not making a
  46053. 34:40:49purchase. And this is done for doesn't
  46054. 34:40:52matter which day of the week. So we call
  46055. 34:40:53that probability of no buy would be 6
  46056. 34:40:56over 30 or 0.2. So there's a 20% chance
  46057. 34:40:59that they're not going to make a
  46058. 34:41:00purchase no matter what day of the week
  46059. 34:41:01it is. And finally, we look at the
  46060. 34:41:03probability of B if A. In this case,
  46061. 34:41:06we're going to look at the probability
  46062. 34:41:07of the weekday and not buying. Two of
  46063. 34:41:10the no buys were done out of the weekend
  46064. 34:41:11out of the six people who did not make
  46065. 34:41:13purchases. So when we look at that,
  46066. 34:41:15probability of the week day without a
  46067. 34:41:17purchase is going to be.33 or 33%. Let's
  46068. 34:41:21take a look at this at different
  46069. 34:41:22probabilities. And uh based on this
  46070. 34:41:25likelihood table, let's go ahead and
  46071. 34:41:27calculate conditional probabilities as
  46072. 34:41:28below. The first three we just did. The
  46073. 34:41:31probability of making a purchase on the
  46074. 34:41:32weekday is 11 out of 30 or roughly 36 or
  46075. 34:41:3637%
  46076. 34:41:37367. The probability of not making a
  46077. 34:41:40purchase at all doesn't matter what day
  46078. 34:41:41of the week is roughly.2 or 20%. And the
  46079. 34:41:45probability of a weekday no purchase is
  46080. 34:41:49roughly two out of six. So two out of
  46081. 34:41:51six of our no purchases were made on the
  46082. 34:41:53weekday. And then finally we take our P
  46083. 34:41:56of A. If you looked we've kept the
  46084. 34:41:58symbols up there. So we got P of
  46085. 34:42:00probability of B, probability of A,
  46086. 34:42:01probability of B if A. We should
  46087. 34:42:04remember that the probability of A if B
  46088. 34:42:07is equal to the first one times the
  46089. 34:42:10probability of no per buys over the
  46090. 34:42:13probability of the weekday. So we could
  46091. 34:42:15calculate it both off the uh table we
  46092. 34:42:17created. We can also calculate this by
  46093. 34:42:19the formula and we get the.367
  46094. 34:42:22which equals or.33
  46095. 34:42:24*2 over.367 which equals.179
  46096. 34:42:28or roughly uh 17 to 18%. And that'd be
  46097. 34:42:32the probability of no purchase done on
  46098. 34:42:34the weekday. And this is important
  46099. 34:42:36because we can look at this and say as
  46100. 34:42:38the probability of buying on the weekday
  46101. 34:42:41is more than the probability of not
  46102. 34:42:43buying on the weekday, we can conclude
  46103. 34:42:45that customers will most likely buy the
  46104. 34:42:47product on a weekday. Now, we've kept
  46105. 34:42:49our chart simple and we're only looking
  46106. 34:42:51at one aspect. So, you should be able to
  46107. 34:42:53look at the table and come up with the
  46108. 34:42:54same information or the same conclusion.
  46109. 34:42:56That should be kind of intuitive at this
  46110. 34:42:58point. Next, we can take the same setup.
  46111. 34:43:01We have the frequency tables of all
  46112. 34:43:03three independent variables. Now we can
  46113. 34:43:05construct the likelihood tables for all
  46114. 34:43:07three of the variables we're working
  46115. 34:43:09with. We can take our day like we did
  46116. 34:43:12before. We have weekday, weekend, and
  46117. 34:43:13holiday. And we filled in this table.
  46118. 34:43:15And then we can come in and also do that
  46119. 34:43:16for the discount. Yes or no. Did they
  46120. 34:43:19buy? Yes or no. And we fill in that full
  46121. 34:43:21table. So now we have our probabilities
  46122. 34:43:24for a discount and whether the discount
  46123. 34:43:27leads to a purchase or not. And the
  46124. 34:43:29probability for free delivery. Does that
  46125. 34:43:31lead to a purchase or not? And this is
  46126. 34:43:33where it starts getting really exciting.
  46127. 34:43:34Let us use these three likelihood tables
  46128. 34:43:37to calculate whether a customer will
  46129. 34:43:38purchase a product on a specific
  46130. 34:43:40combination of day, discount, and free
  46131. 34:43:43delivery or not purchase. Here, let us
  46132. 34:43:46take a combination of these factors. Day
  46133. 34:43:48equals holiday, discount equals yes,
  46134. 34:43:50free delivery equals yes. Let's dig
  46135. 34:43:52deeper into the math and actually see
  46136. 34:43:54what this looks like. And we're going to
  46137. 34:43:56start with looking for the probability
  46138. 34:43:58of them not purchasing on the following
  46139. 34:44:01combinations of days. We are actually
  46140. 34:44:03looking for the probability of A equal
  46141. 34:44:05no buy. No purchase. And our probability
  46142. 34:44:08of B we're going to set equal to is it a
  46143. 34:44:10holiday? Did they get a discount? Yes.
  46144. 34:44:12And was it a free delivery? Yes. Before
  46145. 34:44:14we go further, let's look at the
  46146. 34:44:16original equation. the probability of a
  46147. 34:44:19if b equals the probability of b given
  46148. 34:44:22the condition a and the probability
  46149. 34:44:24times probability of a over the
  46150. 34:44:26probability of b occurring. Now this is
  46151. 34:44:28basic algebra so we can multiply this
  46152. 34:44:31information together. So when you see
  46153. 34:44:33the probability of a given b in this
  46154. 34:44:36case the condition is b c and d or the
  46155. 34:44:39three different variables we're looking
  46156. 34:44:41at. And when you see the probability of
  46157. 34:44:43B, that would be the conditions. We're
  46158. 34:44:45actually going to multiply those three
  46159. 34:44:47separate conditions out. Probability of
  46160. 34:44:50you'll see that in just a second in the
  46161. 34:44:51formula times the full probability of A
  46162. 34:44:54over the full probability of B. So here
  46163. 34:44:57we are back to this and we're going to
  46164. 34:44:59have let A equal no purchase. And we're
  46165. 34:45:01looking for the probability of B on the
  46166. 34:45:03condition A where A sets for three
  46167. 34:45:06different things. Remember that equals
  46168. 34:45:08the probability of A given the condition
  46169. 34:45:10B. And in this case, we just multiply
  46170. 34:45:13those three different variables
  46171. 34:45:15together. So we have the probability of
  46172. 34:45:17the discount times the probability of
  46173. 34:45:21free delivery times the probability is
  46174. 34:45:23the day equal a holiday. Those are our
  46175. 34:45:26three variables of the probability of A
  46176. 34:45:28if B. And then that is going to be
  46177. 34:45:30multiplied by the probability of them
  46178. 34:45:32not making a purchase. And then we want
  46179. 34:45:34to divide that by the total
  46180. 34:45:36probabilities and they're multiplied
  46181. 34:45:38together. So we have the probability of
  46182. 34:45:39a discount, the probability of a free
  46183. 34:45:42delivery, and the probability of it
  46184. 34:45:43being on a holiday. When we plug those
  46185. 34:45:45numbers in, we see that one out of six
  46186. 34:45:48were no purchase on a discounted day,
  46187. 34:45:50two out of six were a no purchase on a
  46188. 34:45:53free delivery day, and three out of six
  46189. 34:45:55were a no purchase on a holiday. Those
  46190. 34:45:58are our three probabilities of A of B
  46191. 34:46:01multiplied out. And then that has to be
  46192. 34:46:03multiplied by the probability of a no
  46193. 34:46:04purchase. And remember the prob
  46194. 34:46:06probability of a noby is across all the
  46195. 34:46:08data. So that's where we get the 6 out
  46196. 34:46:10of 30. We divide that out by the
  46197. 34:46:13probability of each category over the
  46198. 34:46:16total number. So we get the 20 out of 30
  46199. 34:46:19had a discount, 23 out of 30 had a yes
  46200. 34:46:22for free delivery, and 11 out of 30 were
  46201. 34:46:25on a holiday. We plug all those numbers
  46202. 34:46:27in, we get.178.
  46203. 34:46:30So in our probability math, we have
  46204. 34:46:32a.178
  46205. 34:46:34if it's a no-by for a holiday, a
  46206. 34:46:36discount, and a free delivery. Let's
  46207. 34:46:38turn that around and see what that looks
  46208. 34:46:40like if we have a purchase. I promise
  46209. 34:46:42this is the last page of math before we
  46210. 34:46:44dig into the Python script. So here
  46211. 34:46:46we're calculating the probability of the
  46212. 34:46:47purchase using the same math we did to
  46213. 34:46:49find out if they didn't buy. Now we want
  46214. 34:46:51to know if they did buy. And again,
  46215. 34:46:52we're going to go by the day equals a
  46216. 34:46:54holiday, discount equals yes, free
  46217. 34:46:56delivery equals yes, and let a equal
  46218. 34:46:58buy. Now, right about now, you might be
  46219. 34:47:00asking, why are we doing both
  46220. 34:47:02calculations? Why why would we want to
  46221. 34:47:04know the no buys and buys for the same
  46222. 34:47:07data going in? Well, we're going to show
  46223. 34:47:09you that in just a moment, but we have
  46224. 34:47:11to have both of those pieces of
  46225. 34:47:12information so that we can figure it out
  46226. 34:47:14as a percentage as opposed to a
  46227. 34:47:16probability equation. And we'll get to
  46228. 34:47:18that normalization here in just a
  46229. 34:47:20moment. Let's go ahead and walk through
  46230. 34:47:21this calculation. And as you can see
  46231. 34:47:23here, the probability of A on the
  46232. 34:47:25condition of B, B being all three
  46233. 34:47:27categories, did we have a discount with
  46234. 34:47:30a purchase, did we have a free delivery
  46235. 34:47:32with a purchase, and did we is a day
  46236. 34:47:34equal to holiday. And when we plug this
  46237. 34:47:36all into that formula and multiply it
  46238. 34:47:38all out, we get our probability of a
  46239. 34:47:40discount, probability of a free
  46240. 34:47:42delivery, probability of the day being a
  46241. 34:47:44holiday times the overall probability of
  46242. 34:47:47it being a purchase divided by again
  46243. 34:47:50multiplying the three variables out. The
  46244. 34:47:52full probability of there being a
  46245. 34:47:53discount, the full probability of being
  46246. 34:47:55a free delivery, and the full
  46247. 34:47:57probability of there being a day equal
  46248. 34:47:58holiday. And that's where we get this 19
  46249. 34:48:00over 24 * 21 over 24 * 8 over 24 * the p
  46250. 34:48:05of a 24 over 30 divided by the
  46251. 34:48:09probability of the discount the free
  46252. 34:48:11delivery times the day or 20 over 30 23
  46253. 34:48:15over 30 * 11 over 30 and that gives us
  46254. 34:48:17our 986.
  46255. 34:48:20So what are we going to do with these
  46256. 34:48:22two pieces of data we just generated?
  46257. 34:48:24Well, let's go ahead and go over them.
  46258. 34:48:25We have a probability of purchase
  46259. 34:48:27equals.986.
  46260. 34:48:29We have a probability of no purchase
  46261. 34:48:31equals.178.
  46262. 34:48:34So finally we have a conditional
  46263. 34:48:36probabilities of purchase on this day.
  46264. 34:48:37Let us take that we're going to
  46265. 34:48:38normalize it and we're going to take
  46266. 34:48:40these probabilities and turn them into
  46267. 34:48:42percentages. This is simply done by
  46268. 34:48:44taking the sum of probabilities which
  46269. 34:48:46equals 98686 plus.178
  46270. 34:48:50and that equals the 1.164.
  46271. 34:48:53If we divide each probability by the
  46272. 34:48:56sum, we get the percentage. And so the
  46273. 34:48:58likelihood of a purchase is 84.71%.
  46274. 34:49:02And the likelihood of no purchase is
  46275. 34:49:0415.29%
  46276. 34:49:06given these three different variables.
  46277. 34:49:09So it's if it's on a holiday, if it's a
  46278. 34:49:11with a discount and has free delivery,
  46279. 34:49:13then there's an 84.71%
  46280. 34:49:15chance that the customer is going to
  46281. 34:49:16come in and make a purchase. Hooray,
  46282. 34:49:18they purchased our stuff. We're making
  46283. 34:49:20money. If you were owning a shop, that's
  46284. 34:49:21like is the bottom line is you want to
  46285. 34:49:23make some money so you can keep your
  46286. 34:49:24shop open and have a living. Now, I
  46287. 34:49:26promised you that we were going to be
  46288. 34:49:27finishing up the math here with a few
  46289. 34:49:29pages. So, we're going to move on and
  46290. 34:49:31we're going to do two steps. The first
  46291. 34:49:33step is I want you to understand why you
  46292. 34:49:36want to why you want to use the naive
  46293. 34:49:38bays. What are the advantages of naive
  46294. 34:49:40bays? And then once we understand those
  46295. 34:49:42advantages, we just look at that
  46296. 34:49:44briefly. Then we're going to dive in and
  46297. 34:49:46do some Python coding. Advantages of
  46298. 34:49:48naive bay classifier. So let's take a
  46299. 34:49:51look at the six advantages of the naive
  46300. 34:49:53bay classifier. And we're going to walk
  46301. 34:49:55around this lovely wheel. Looks like an
  46302. 34:49:57origami folded paper. The first one is
  46303. 34:49:59very simple and easy to implement.
  46304. 34:50:01Certainly you could walk through the
  46305. 34:50:03tables and do this by hand. You got to
  46306. 34:50:05be a little careful because the
  46307. 34:50:06notations can get confusing. You have
  46308. 34:50:08all these different probabilities and I
  46309. 34:50:10certainly mess those up as I put them
  46310. 34:50:11on, you know, is it on the top or the
  46311. 34:50:13bottom? We got to really pay close
  46312. 34:50:14attention to that. When you put it into
  46313. 34:50:16Python, it's really nice because you
  46314. 34:50:18don't have to worry about any of that.
  46315. 34:50:19You let the Python handle that, the
  46316. 34:50:21Python module. But understanding it, you
  46317. 34:50:23can put it on a table and you can easily
  46318. 34:50:24see how it works. And it's a simple
  46319. 34:50:26algebraic function. It needs less
  46320. 34:50:28training data. So if you have smaller
  46321. 34:50:29amounts of data, this is great powerful
  46322. 34:50:31tool for that. Handles both continuous
  46323. 34:50:34and discrete data. It's highly scalable
  46324. 34:50:36with number of predictors and data
  46325. 34:50:38points. So, as you can see, you can just
  46326. 34:50:40keep multiplying different probabilities
  46327. 34:50:42in there and you can cover not just
  46328. 34:50:43three different variables or sets. You
  46329. 34:50:46can now expand this to even more
  46330. 34:50:47categories. Number five, it's fast. It
  46331. 34:50:50can be used in real time predictions.
  46332. 34:50:52This is so important. This is why it's
  46333. 34:50:54used in a lot of our predictions on
  46334. 34:50:56online shopping carts, uh, referrals,
  46335. 34:50:59spam filters, is because there's no time
  46336. 34:51:01delay as it has to go through and figure
  46337. 34:51:03out a neural network or one of the other
  46338. 34:51:06mini setups where you're doing
  46339. 34:51:07classification. And certainly there's a
  46340. 34:51:09lot of other tools out there in the
  46341. 34:51:10machine learning that can handle these,
  46342. 34:51:13but most of them are not as fast as the
  46343. 34:51:15naive bays. And then finally, it's not
  46344. 34:51:17sensitive to irrelevant features. So it
  46345. 34:51:20picks up on your different
  46346. 34:51:21probabilities. And if you're short on
  46347. 34:51:23data on one probability, you can kind of
  46348. 34:51:25it automatically adjusts for that. Those
  46349. 34:51:27formulas are very automatic. And so you
  46350. 34:51:29can still get a very solid
  46351. 34:51:31predictability even if you're missing
  46352. 34:51:33data or you have overlapping data for
  46353. 34:51:35two completely different areas. We see
  46354. 34:51:36that a lot in doing census and studying
  46355. 34:51:39of people and habits where they might
  46356. 34:51:42have one study that covers one aspect
  46357. 34:51:44and another one that overlaps and
  46358. 34:51:45because the two overlap they can then
  46359. 34:51:47predict the unknowns for the group that
  46360. 34:51:49they haven't done the second study on or
  46361. 34:51:51vice versa. So it's very powerful in
  46362. 34:51:53that it is not sensitive to the
  46363. 34:51:55irrelevant features and in fact you can
  46364. 34:51:57use it to help predict features that
  46365. 34:51:59aren't even in there. So now we're down
  46366. 34:52:01to my favorite part. We're going to roll
  46367. 34:52:03up our sleeves and do some actual
  46368. 34:52:05programming. We're going to do the use
  46369. 34:52:06case text classification. Now, I would
  46370. 34:52:10challenge you to go back and send us a
  46371. 34:52:12note on the notes below underneath the
  46372. 34:52:14video and request the data for the
  46373. 34:52:16shopping cart. So, you can plug that
  46374. 34:52:18into Python code and do that on your own
  46375. 34:52:20time. So, you can walk through it since
  46376. 34:52:22we walk through all the information on
  46377. 34:52:24it. But, we're going to do a Python code
  46378. 34:52:26doing text classification. Very popular
  46379. 34:52:28for doing the naive bays. So, we're
  46380. 34:52:31going to use our new tool to perform a
  46381. 34:52:33text classification of news headlines
  46382. 34:52:35and classify news into different topics
  46383. 34:52:37for a news website. As you can see here,
  46384. 34:52:39we have a nice image of the Google News
  46385. 34:52:42and then related on the right subgroups.
  46386. 34:52:45I'm not sure where they actually pulled
  46387. 34:52:46the actual data we're going to use from.
  46388. 34:52:48It's one of the standard sets, but
  46389. 34:52:50certainly this can be used on any of our
  46390. 34:52:52news headlines in classification. So,
  46391. 34:52:54let's see how it can be done using the
  46392. 34:52:56naive base classifier. Now, we're at my
  46393. 34:52:58favorite part. We're actually going to
  46394. 34:53:00write some Python script, roll up our
  46395. 34:53:02sleeves, and we're going to start by
  46396. 34:53:04doing our imports. These are very basic
  46397. 34:53:06imports, including our news group. And
  46398. 34:53:08we'll take a quick glance at the target
  46399. 34:53:10names. Then we're going to go ahead and
  46400. 34:53:11start training our data set and putting
  46401. 34:53:13it together. We'll put together a nice
  46402. 34:53:15graph because it's always good to have a
  46403. 34:53:17graph to show what's going on. And once
  46404. 34:53:19we've trained it and we've shown you a
  46405. 34:53:20graph of what's going on, then we're
  46406. 34:53:22going to explore how to use it and see
  46407. 34:53:24what that looks like. Now I'm going to
  46408. 34:53:26open up my favorite editor or inline
  46409. 34:53:28editor for Python. You don't have to use
  46410. 34:53:30this. You can use whatever your editor
  46411. 34:53:31that you like, whatever uh interface IDE
  46412. 34:53:34you want. This just happens to be the
  46413. 34:53:36Anaconda Jupiter notebook. And I'm going
  46414. 34:53:39to paste that first piece of code in
  46415. 34:53:40here so we can walk through it. Let's
  46416. 34:53:42make it a little bigger on the screen so
  46417. 34:53:44you have a nice view of what's going on.
  46418. 34:53:45Uh and we're using Python 3, in this
  46419. 34:53:47case 3.5. So this would work in any of
  46420. 34:53:50your 3X if you have it set up correctly.
  46421. 34:53:53should also work in a lot of the 2x. You
  46422. 34:53:55just have to make sure all of the the
  46423. 34:53:56versions of the modules match your
  46424. 34:53:58Python version. And in here, you'll
  46425. 34:54:00notice the first line is your percentage
  46426. 34:54:02mattplot library in line. Now, three of
  46427. 34:54:05these lines of code are all about
  46428. 34:54:07plotting the graph. This one lets the
  46429. 34:54:11notebook know and is inline setup that
  46430. 34:54:14we want the graphs to show up on this
  46431. 34:54:16page. Without it, in a notebook like
  46432. 34:54:18this, which is an explorer interface, it
  46433. 34:54:20won't show up. Now, a lot of IDEs don't
  46434. 34:54:23require that. A lot of them, like on if
  46435. 34:54:25I'm working on one of my other setups,
  46436. 34:54:27it just has a popup and the graph pops
  46437. 34:54:29up on there. So, you have a that setup
  46438. 34:54:31also. But for this, we want the mattplot
  46439. 34:54:33library in line. And then we're going to
  46440. 34:54:35import numpy as np. That's number
  46441. 34:54:39python, which has a lot of different
  46442. 34:54:40formulas in it that we use for both of
  46443. 34:54:43our sklearn module. And we also use it
  46444. 34:54:46for any of the upper math functions in
  46445. 34:54:48python. And it's very common to see that
  46446. 34:54:49as NP numpy as NP. The next two lines
  46447. 34:54:53are all about our graphing. Remember I
  46448. 34:54:55said three of these were about graphing.
  46449. 34:54:56Well, we need our mattplot
  46450. 34:54:58library.pipplot
  46451. 34:54:59as plt. And you'll see that plt is a
  46452. 34:55:02very common setup as is the sns and just
  46453. 34:55:04like the np. And we're going to import
  46454. 34:55:06seabor as sns and we're going to do the
  46455. 34:55:09sns set. Now seabor sits on top of
  46456. 34:55:13pipplot and it just makes a really nice
  46457. 34:55:15heat map. It's really good for heat
  46458. 34:55:17maps. And if you're not familiar with
  46459. 34:55:18heat maps, that just means we give it a
  46460. 34:55:19color scale. The term comes from the
  46461. 34:55:22brighter red it is, the hotter it is in
  46462. 34:55:25some form of data. And you can set it to
  46463. 34:55:26whatever you want. And we'll see that
  46464. 34:55:28later on. So those you'll see that those
  46465. 34:55:30three lines of code here are just
  46466. 34:55:32importing the graph function so we can
  46467. 34:55:33graph it. And as a data scientist, you
  46468. 34:55:36always want to graph your data and have
  46469. 34:55:38some kind of visual. It's really hard
  46470. 34:55:39just to shove numbers in front of people
  46471. 34:55:41and they look at it and it doesn't mean
  46472. 34:55:42anything. And then from the sklearn data
  46473. 34:55:45sets, we're going to import the fetch 20
  46474. 34:55:47news groups. Very common one for
  46475. 34:55:50analyzing tokenizing words and setting
  46476. 34:55:52them up and exploring how the words work
  46477. 34:55:54and how do you categorize different
  46478. 34:55:55things when you're dealing with
  46479. 34:55:56documents. And then we set our data
  46480. 34:55:58equal to fetch 20 news groups. So our
  46481. 34:56:01data variable will have the data in it.
  46482. 34:56:03And we're going to go ahead and just
  46483. 34:56:04print the target names. data.target
  46484. 34:56:07names. And let's see what that looks
  46485. 34:56:08like. And you'll see here we have alt
  46486. 34:56:11atheism comp graphics composs
  46487. 34:56:14windows.mmiscellaneous
  46488. 34:56:16and it goes all the way down to talk
  46489. 34:56:18politics.mmiscellaneous talk
  46490. 34:56:20religion.mmiscellaneous. These are the
  46491. 34:56:22categories they've already assigned to
  46492. 34:56:24this news group and it's called fetch 20
  46493. 34:56:26because you'll see there's I believe
  46494. 34:56:27there's 20 different topics in here or
  46495. 34:56:2920 different categories as we scroll
  46496. 34:56:31down. Now, we've gone through the 20
  46497. 34:56:33different categories and we're going to
  46498. 34:56:35go ahead and start defining all the
  46499. 34:56:37categories and set up our data. So,
  46500. 34:56:39we're actually getting here going to go
  46501. 34:56:41ahead and get it get the data all set up
  46502. 34:56:43and take a look at our data. And let's
  46503. 34:56:44move this over to our Jupyter notebook.
  46504. 34:56:47And let's see what this code does.
  46505. 34:56:49First, we're going to set our
  46506. 34:56:51categories. Now, if you noticed up here,
  46507. 34:56:53I could have just as easily set this
  46508. 34:56:54equal to data.target_names target names
  46509. 34:56:58because it's the same thing, but we want
  46510. 34:57:00to kind of spell it out for you so you
  46511. 34:57:01can see the different categories. It
  46512. 34:57:03kind of makes it more visual so you can
  46513. 34:57:04see what your data is looking like in
  46514. 34:57:06the background. Once we've created the
  46515. 34:57:08categories, we're going to open up a
  46516. 34:57:10train set. So this training set of data
  46517. 34:57:13is going to go into fetch 20 news groups
  46518. 34:57:16and it's a subset in there called train
  46519. 34:57:18and categories equals categories. So
  46520. 34:57:20we're pulling out those categories that
  46521. 34:57:22match. And then if you have a train set,
  46522. 34:57:24you should also have the testing set. We
  46523. 34:57:25have test equals fetch 20 news group
  46524. 34:57:27subset equals test and categories equals
  46525. 34:57:29categories. Let's go down one size so it
  46526. 34:57:32all fits on my screen. There we go. And
  46527. 34:57:34just so we can really see what's going
  46528. 34:57:35on, let's see what happens when we print
  46529. 34:57:37out one part of that data. So it creates
  46530. 34:57:42train and under train, it creates train
  46531. 34:57:44data. And we're just going to look at
  46532. 34:57:46data piece number five. And let's go
  46533. 34:57:48ahead and run that and see what that
  46534. 34:57:50looks like. And you can see when I print
  46535. 34:57:52train.data data number five under train.
  46536. 34:57:55It prints out one of the articles. This
  46537. 34:57:57is article number five. You can go
  46538. 34:57:58through and read it on there. And we can
  46539. 34:58:00also go in here and change this to test,
  46540. 34:58:03which should look identical because it's
  46541. 34:58:05splitting the date up into different
  46542. 34:58:06groups. Train and test. And we'll see
  46543. 34:58:08test number five is a a different
  46544. 34:58:10article, but it's another article in
  46545. 34:58:12here. And maybe you're curious and you
  46546. 34:58:14want to see just how many articles are
  46547. 34:58:16in here. We could do length of train.
  46548. 34:58:21data. And if we run that, you'll see
  46549. 34:58:23that the training data has 11,314
  46550. 34:58:27articles. So, we're not going to go
  46551. 34:58:29through all those articles. That's a lot
  46552. 34:58:30of articles, but um we can look at one
  46553. 34:58:32of them just so you can see what kind of
  46554. 34:58:34information is coming out of it and what
  46555. 34:58:35we're looking at. And we'll just look at
  46556. 34:58:37number five for today. And here we have
  46557. 34:58:39it. Rewarding the Second Amendment IDs,
  46558. 34:58:41VTT, line 58, lines 58 in article, uh
  46559. 34:58:45etc. And you can scroll all the way down
  46560. 34:58:47and see all the different parts to
  46561. 34:58:48there. Now, we've looked at it and
  46562. 34:58:50that's pretty complicated when you look
  46563. 34:58:51at one of these articles to try to
  46564. 34:58:52figure out how do you weight this. If
  46565. 34:58:54you look down here, we have different
  46566. 34:58:56words and maybe the word from. Well,
  46567. 34:58:58from is probably in all the articles.
  46568. 34:59:00So, it's not going to have a lot of
  46569. 34:59:02meaning as far as trying to figure out
  46570. 34:59:03whether this article fits one of the
  46571. 34:59:05categories or not. So, trying to figure
  46572. 34:59:06out which category it fits in based on
  46573. 34:59:09these words is where the challenge comes
  46574. 34:59:10in. Now that we've viewed our data,
  46575. 34:59:12we're going to dive in and do the actual
  46576. 34:59:14predictions. This is the actual naive
  46577. 34:59:17bays. And we're going to throw another
  46578. 34:59:18model at you or another module at you
  46579. 34:59:20here in just a second. We can't go into
  46580. 34:59:22too much detail, but it deals
  46581. 34:59:23specifically working with words and text
  46582. 34:59:26and what they call tokenizing those
  46583. 34:59:28words. So, let's take this code and
  46584. 34:59:31let's uh skip on over to our Jupyter
  46585. 34:59:33notebook and walk through it. And here
  46586. 34:59:35we are in our Jupyter notebook. Let's
  46587. 34:59:36paste that in there. And I can run this
  46588. 34:59:38code right off the bat. It's not
  46589. 34:59:39actually going to display anything yet,
  46590. 34:59:41but it has a lot going on in here. So
  46591. 34:59:44the top we had the print module from the
  46592. 34:59:46earlier one. I didn't know why that was
  46593. 34:59:47in there. So we're going to start by
  46594. 34:59:49importing our necessary packages. And
  46595. 34:59:51from the sklearn features
  46596. 34:59:53extraction.ext,
  46597. 34:59:55we're going to import TF IDF vectorzer.
  46598. 34:59:59I told you we're going to throw a module
  46599. 35:00:00at you. We can't go too much into the
  46600. 35:00:03math behind this or how it works. You
  46601. 35:00:04can look it up. The notation for the
  46602. 35:00:06math is usually TF.idf.
  46603. 35:00:09And that's just a way of weighing the
  46604. 35:00:11words. and it weighs the words based on
  46605. 35:00:14how many times are used in a document,
  46606. 35:00:16how many times or how many documents
  46607. 35:00:18they're used in. And it's a well-used
  46608. 35:00:20formula. It's been around for a while.
  46609. 35:00:21It's a little confusing to put this in
  46610. 35:00:23here. Uh, but let's let them know that
  46611. 35:00:25it just goes in there and weights the
  46612. 35:00:27different words in the document for us.
  46613. 35:00:29That way, we don't have to wait. And if
  46614. 35:00:31you put a weight on it, if you remember,
  46615. 35:00:33I was talking about that up here
  46616. 35:00:34earlier. If these are all emails, they
  46617. 35:00:36probably all have the word from in them.
  46618. 35:00:38From probably has a very low weight. It
  46619. 35:00:40has very little value in telling you
  46620. 35:00:41what this document's about. Same with
  46621. 35:00:43words like in an article in articles in
  46622. 35:00:46cost of un maybe cost might or where
  46623. 35:00:50words like criminal weapons destruction
  46624. 35:00:53these might have a heavier weight
  46625. 35:00:55because they describe a little bit more
  46626. 35:00:56what the article is doing. Well, how do
  46627. 35:00:58you figure out all those weights in the
  46628. 35:00:59different articles? That's what this
  46629. 35:01:01module does. That's what the TF
  46630. 35:01:04vectorizer is going to do for us. And
  46631. 35:01:06then we're going to import our
  46632. 35:01:07sklearn.na naive bays and that's our
  46633. 35:01:10multinnomial NB multinnomial naive bay
  46634. 35:01:14pretty easy to understand that where
  46635. 35:01:15that comes from and then finally we have
  46636. 35:01:17the skyarn pipeline import make pipeline
  46637. 35:01:21now the make pipeline is just a cool
  46638. 35:01:23piece of code because we're going to
  46639. 35:01:25take the information we get from the TF
  46640. 35:01:28vectorizer and we're going to pump that
  46641. 35:01:31into the multinnomial NB. So, a pipeline
  46642. 35:01:35is just a way of organizing how things
  46643. 35:01:38flow. It's used commonly. You probably
  46644. 35:01:40already guessed what it is. If you've
  46645. 35:01:41done any businesses, they talk about the
  46646. 35:01:43sales pipeline. If you're on a work crew
  46647. 35:01:46or project manager, you have your
  46648. 35:01:48pipeline of information that's going
  46649. 35:01:49through or your projects and what has to
  46650. 35:01:51be done in what order. That's all this
  46651. 35:01:52pipeline is. We're going to take the
  46652. 35:01:55TFID vectorzer and then we're going to
  46653. 35:01:57push that into the multinnomial inb. Now
  46654. 35:02:00we've designated that as the variable
  46655. 35:02:03model. We have our pipeline model and
  46656. 35:02:05we're going to take that model and this
  46657. 35:02:08is just so elegant. This is done in just
  46658. 35:02:09a couple lines of code. model.fit and
  46659. 35:02:13we're going to fit the data. And first
  46660. 35:02:15the train data and then the train
  46661. 35:02:17target. Now the train data has the
  46662. 35:02:20different articles in it. You can see
  46663. 35:02:22the one we were just looking at and the
  46664. 35:02:24train.target target is what category
  46665. 35:02:26they already categorized that that
  46666. 35:02:28particular article as. And what's
  46667. 35:02:30happening here is the train data is
  46668. 35:02:33going into the TF ID vectorizer. So when
  46669. 35:02:36you have one of these articles, it goes
  46670. 35:02:38in there, it weights all the words in
  46671. 35:02:40there. So there's thousands of words
  46672. 35:02:42with different weights on them. I
  46673. 35:02:43remember once running a model on this
  46674. 35:02:44and I literally had 2.4 million tokens
  46675. 35:02:48go into this. So when you're dealing
  46676. 35:02:50like large document bases, you can have
  46677. 35:02:52a huge number of different words. It
  46678. 35:02:54then takes those words, gives them a
  46679. 35:02:56weight, and then based on that weight,
  46680. 35:02:59based on the words and the weights, and
  46681. 35:03:00then puts that into the multinnomial NB.
  46682. 35:03:03And once we go into our naive bay, we
  46683. 35:03:06want to put the train target in there.
  46684. 35:03:08So the train data that's been mapped to
  46685. 35:03:10the TFID vectorzer is now going through
  46686. 35:03:14the multinnomial NB. And then we're
  46687. 35:03:16telling it, well, these are the answers.
  46688. 35:03:18These are the answers to the different
  46689. 35:03:19documents. So this document that has all
  46690. 35:03:21these words with these different weights
  46691. 35:03:23from the first part is going to be
  46692. 35:03:25whatever category it comes out of. Maybe
  46693. 35:03:27it's the um talk show or the article on
  46694. 35:03:30religion miscellaneous. Once we fit that
  46695. 35:03:33model, we can then take labels and we're
  46696. 35:03:36going to set that equal to
  46697. 35:03:38model.predict. Most of the sklearn use
  46698. 35:03:41the term.predict to let us know that
  46699. 35:03:43we've now trained the model and now we
  46700. 35:03:45want to get some answers. And we're
  46701. 35:03:46going to put our test data in there
  46702. 35:03:48because our test data is the stuff we
  46703. 35:03:50held off to the side. We didn't train it
  46704. 35:03:52on there and we don't know what's going
  46705. 35:03:54to come up out of it and we just want to
  46706. 35:03:55find out how good our labels are. Do
  46707. 35:03:57they match what they should be? Now,
  46708. 35:03:59I've already run this through. There's
  46709. 35:04:01no actual output to it to show. This is
  46710. 35:04:03just setting it all up. This is just
  46711. 35:04:05training our model, creating the labels
  46712. 35:04:07so we can see how good it is, and then
  46713. 35:04:08we move on to the next step to find out
  46714. 35:04:10what happened. To do this, we're going
  46715. 35:04:13to go ahead and create a confusion
  46716. 35:04:15matrix and a heat map. So, the confusion
  46717. 35:04:18matrix, which is confusing just by its
  46718. 35:04:21very name, is basically going to ask how
  46719. 35:04:23confused is our answer. Did it get it
  46720. 35:04:26correct or did it miss some things in
  46721. 35:04:28there or have some missed labels? And
  46722. 35:04:30then we're going to put that on a heat
  46723. 35:04:31map so we have some nice colors to look
  46724. 35:04:33at to see how that plots out. Let's go
  46725. 35:04:35ahead and take this code and see how
  46726. 35:04:37that uh take a walk through it and see
  46727. 35:04:39what that looks like. So, back to our
  46728. 35:04:41Jupyter notebook. I'm going to put the
  46729. 35:04:42code in there and let's go ahead and run
  46730. 35:04:45that code. Take it just a moment. And
  46731. 35:04:48remember, we had the inline. That way,
  46732. 35:04:50my graph shows up on the inline here.
  46733. 35:04:53And let's walk through the code and then
  46734. 35:04:54we'll look at this and see what that
  46735. 35:04:56means. So, make it a little bit bigger.
  46736. 35:04:58There we go. No reason not to use the
  46737. 35:05:00whole screen. Too big. So, we have here
  46738. 35:05:02from sklearn metrics import confusion
  46739. 35:05:05matrix. And that's just going to
  46740. 35:05:07generate a set of data that says I the
  46741. 35:05:10prediction was such the actual truth was
  46742. 35:05:14either agreed with it or was something
  46743. 35:05:16different. And it's going to add up
  46744. 35:05:17those numbers so we can take a look and
  46745. 35:05:18just see how well it worked. And we're
  46746. 35:05:20going to set a variable Matt equal to
  46747. 35:05:22confusion matrix. We have our test
  46748. 35:05:25target, our test data that was not part
  46749. 35:05:27of the training. Very important in data
  46750. 35:05:29science, we always keep our test data
  46751. 35:05:31separate. Otherwise, it's not a valid
  46752. 35:05:34model if we can't properly test it with
  46753. 35:05:36new data. And this is the labels we
  46754. 35:05:38created from that test data. These are
  46755. 35:05:40the ones that we predict it's going to
  46756. 35:05:42be. So, we go in and we create our SN
  46757. 35:05:44heat map. The SNS is our seaborn which
  46758. 35:05:47sits on top of the piplot. So, we create
  46759. 35:05:50a SNS.heet map. We take our confusion
  46760. 35:05:53matrix and it's going to be uh matt.t.
  46761. 35:05:57And then we have other variables that go
  46762. 35:05:59into the SNS heat map. We're not going
  46763. 35:06:02to go into detail what all the variables
  46764. 35:06:04mean. The annotation equals true. That's
  46765. 35:06:06what tells it to put the numbers here.
  46766. 35:06:07So you have the 166, the one, the 00001.
  46767. 35:06:11Format D and C bar equals false have to
  46768. 35:06:13do with the uh format. If you take those
  46769. 35:06:16out, you'll see that some things
  46770. 35:06:17disappear. And then the X tick labels
  46771. 35:06:19and the Y tick labels. Those are our
  46772. 35:06:21target names. And you can see right
  46773. 35:06:23here, that's the alt atheism comp
  46774. 35:06:26graphics composs windows.mmiscellaneous.
  46775. 35:06:29And then finally we have our plt.xl
  46776. 35:06:32label. Remember the SNS or the seabor
  46777. 35:06:35sits on top of our mattplot library our
  46778. 35:06:37plt. And so we want to just tell it x
  46779. 35:06:39label equals a true is is true. The
  46780. 35:06:41labels are true. And then the y label is
  46781. 35:06:44prediction label. So when we say a true,
  46782. 35:06:47this is what it actually is. And the
  46783. 35:06:49prediction is what we predicted. And
  46784. 35:06:51let's look at this graph because that's
  46785. 35:06:53probably a little confusing the way I
  46786. 35:06:54rattled through it. And what I'm going
  46787. 35:06:56to do is I'm going to go ahead and flip
  46788. 35:06:58back to the slides because they have a
  46789. 35:07:00black background they put in there that
  46790. 35:07:02helps it shine a little bit better so
  46791. 35:07:03you can see the graph a little bit
  46792. 35:07:04easier. So in reading this graph, what
  46793. 35:07:07we want to look at is how the color
  46794. 35:07:09scheme has come out. And you'll see a
  46795. 35:07:11line right down the middle diagonally
  46796. 35:07:14from upper left to bottom right. What
  46797. 35:07:16that is is if you look at the labels, we
  46798. 35:07:18have our predicted label on the left and
  46799. 35:07:21our true label on the right. Those are
  46800. 35:07:23the numbers where the prediction and the
  46801. 35:07:25true come together. And this is what we
  46802. 35:07:27want to see is we want to see those lit
  46803. 35:07:28up. That's what that heat map does. As
  46804. 35:07:30you can see that it did a good job of
  46805. 35:07:32finding those data. And you'll notice
  46806. 35:07:34that there's a couple of red spots on
  46807. 35:07:36there where it missed. You know, it it's
  46808. 35:07:38a little confused when we talk about
  46809. 35:07:40talk religion miscellaneous versus talk
  46810. 35:07:42politics miscellaneous, social religion
  46811. 35:07:45Christian versus alt atheism. It
  46812. 35:07:47mislabeled some of those. And those are
  46813. 35:07:49very similar topics. so you could
  46814. 35:07:51understand why it might mislabel them.
  46815. 35:07:53But overall, it did a pretty good job.
  46816. 35:07:55If we're going to create these models,
  46817. 35:07:56we want to go ahead and be able to use
  46818. 35:07:58them. So, let's see what that looks
  46819. 35:08:00like. To do this, let's go ahead and
  46820. 35:08:02create a definition, a function to run.
  46821. 35:08:05And we're going to call this function.
  46822. 35:08:06Let me just expand that just a notch
  46823. 35:08:08here. There we go. I like mine in big
  46824. 35:08:10letters. Predict category. So, we want
  46825. 35:08:12to predict the category. We're going to
  46826. 35:08:14send it as a string. And then we're
  46827. 35:08:17sending it train equals train. We have
  46828. 35:08:19our training model. And then we had our
  46829. 35:08:21pipeline model equals model. This way we
  46830. 35:08:23don't have to resend these variables
  46831. 35:08:24each time. The definition knows that
  46832. 35:08:27because I said train equals train and I
  46833. 35:08:29put the equal for model. And then we're
  46834. 35:08:31going to set the prediction equal to the
  46835. 35:08:32model.predict s. So it's going to send
  46836. 35:08:35whatever string we send to it. It's
  46837. 35:08:38going to push that string through the
  46838. 35:08:39pipeline, the model pipeline. It's going
  46839. 35:08:41to go through and uh tokenize it and put
  46840. 35:08:44it through the TF IDF, convert that into
  46841. 35:08:48numbers and weights for all the
  46842. 35:08:49different documents and words. And then
  46843. 35:08:51it'll put that through our naive bay.
  46844. 35:08:54And from it, we'll go ahead and get our
  46845. 35:08:56prediction. We're going to predict what
  46846. 35:08:58value it is. And so we're going to
  46847. 35:09:00return train.target names predict of
  46848. 35:09:03zero. And remember that the train.target
  46849. 35:09:06names, that's just categories. I could
  46850. 35:09:08have just as easily put uh categories in
  46851. 35:09:10there.predict of zero. So we're taking
  46852. 35:09:13the prediction which is a number and
  46853. 35:09:15we're converting it to an actual
  46854. 35:09:16category. We're converting it from um I
  46855. 35:09:19don't know what the actual numbers are.
  46856. 35:09:20Let's say zero equals alt atheism. So
  46857. 35:09:23we're going to convert that zero to the
  46858. 35:09:24word or uh one maybe it equals comp
  46859. 35:09:27graphics. So we're going to convert
  46860. 35:09:28number one into comp graphics. That's
  46861. 35:09:30all that is. And then we got to go ahead
  46862. 35:09:32and and then we need to go ahead and run
  46863. 35:09:35this. So I load that up. And then once I
  46864. 35:09:38run that, we can start doing some
  46865. 35:09:40predictions. Let me go ahead and type in
  46866. 35:09:42predict category. And let's just do
  46867. 35:09:45predict category, Jesus Christ. And it
  46868. 35:09:47comes back and says it's social,
  46869. 35:09:49religion, Christian. That's pretty good.
  46870. 35:09:52Now note, I didn't put print on this.
  46871. 35:09:54One of the nice things about the Jupiter
  46872. 35:09:56notebook editor and a lot of inline
  46873. 35:09:58editors is if you just put the name of
  46874. 35:10:00the variable out, it's returning the
  46875. 35:10:02variable train.target_ames, target
  46876. 35:10:04names. It'll automatically print that
  46877. 35:10:06for you. In your own IDE, you might have
  46878. 35:10:08to put in print. Let's see where else we
  46879. 35:10:10can take this. And maybe you're a space
  46880. 35:10:12science buff. So, how about sending load
  46881. 35:10:16to international
  46882. 35:10:19space station.
  46883. 35:10:21And if we run that, we get science
  46884. 35:10:24space. Or maybe you're a uh automobile
  46885. 35:10:28buff. And let's do um Oh, they were
  46886. 35:10:31gonna tell me Audi is better than BMW,
  46887. 35:10:32but I'm going to do BMW is better than
  46888. 35:10:36an Audi. So maybe our car buff. And we
  46889. 35:10:39run that. And you'll see it says
  46890. 35:10:41recreational. I'm assuming that's what
  46891. 35:10:42RECC stands for. Autos. So I did a
  46892. 35:10:45pretty good job labeling that one. How
  46893. 35:10:47about uh if we have something like a
  46894. 35:10:49caption running through there, President
  46895. 35:10:51of India. And if we run that, it comes
  46896. 35:10:54up and says talk politics miscellaneous.
  46897. 35:10:58So when we take our definition or our
  46898. 35:11:00function and we run all these things
  46899. 35:11:02through, kudos, we made it. We were able
  46900. 35:11:04to correctly classify texts into
  46901. 35:11:06different groups based on which category
  46902. 35:11:08they belong to using the naive base
  46903. 35:11:11classifier. Now we did throw in the
  46904. 35:11:13pipeline, the TF IDF vectorzer, we threw
  46905. 35:11:17in the graphs. Those are all things that
  46906. 35:11:19you don't necessarily have to know to
  46907. 35:11:21understand the naive base setup or
  46908. 35:11:24classifier, but they're important to
  46909. 35:11:25know. One of the main uses for the naive
  46910. 35:11:27bays is with the TF IDF tokenizer
  46911. 35:11:31vectorzer where it tokenizes a word and
  46912. 35:11:33has labels and we use the pipeline
  46913. 35:11:36because you need to push all that data
  46914. 35:11:38through and it makes it really easy and
  46915. 35:11:39fast. You don't have to know those to
  46916. 35:11:41understand naive bays but they certainly
  46917. 35:11:43help for understanding the industry and
  46918. 35:11:45data science. And we can see our
  46919. 35:11:47categorizer, our naive base classifier.
  46920. 35:11:50We were able to predict the category
  46921. 35:11:52religion, space, motorcycles, autos,
  46922. 35:11:55politics, and properly classify all
  46923. 35:11:57these different things we pushed into
  46924. 35:11:59our prediction and our trained model.
  46925. 35:12:01Before we dive into the SVM, let's take
  46926. 35:12:04a look at applications of the support
  46927. 35:12:06vector machine, at least some general
  46928. 35:12:08ones that are commonly used with it.
  46929. 35:12:10face detection, text and hypertext
  46930. 35:12:12categorization, classification of
  46931. 35:12:14images, and bioinformatics.
  46932. 35:12:17These are only but a few of those that
  46933. 35:12:19are used with this SVM. As we go through
  46934. 35:12:21this lesson, see if you can figure out
  46935. 35:12:23what other ones you could apply it to,
  46936. 35:12:24and also what you would want to use some
  46937. 35:12:26other tools for. So, in this example,
  46938. 35:12:28last week, my son and I visited a fruit
  46939. 35:12:31shop. Dad, is that an apple or a
  46940. 35:12:33strawberry? So, the question comes up,
  46941. 35:12:35what fruit did I just pick up from the
  46942. 35:12:37fruit stand? After a couple of seconds,
  46943. 35:12:39you can figure out that it was a
  46944. 35:12:40strawberry. So, let's take this model a
  46945. 35:12:42step further and let's uh why not build
  46946. 35:12:45a model which can predict an unknown
  46947. 35:12:46data. And in this, we're going to be
  46948. 35:12:48looking at some sweet strawberries or
  46949. 35:12:50crispy apples. We wanted to be able to
  46950. 35:12:52label those two and decide what the
  46951. 35:12:53fruit is. And we do that by having data
  46952. 35:12:56already put in. So, we already have a
  46953. 35:12:58bunch of strawberries. We know our
  46954. 35:12:59strawberries and they're already labeled
  46955. 35:13:00as such. We already have a bunch of
  46956. 35:13:02apples. We know our apples and are
  46957. 35:13:03labeled as such. Then once we train our
  46958. 35:13:05model, that model then can be given the
  46959. 35:13:07new data and the new data is this image.
  46960. 35:13:09In this case, you can see a question
  46961. 35:13:11mark on it and it comes through and goes
  46962. 35:13:12it's a strawberry. In this case, we're
  46963. 35:13:14using the support vector machine model.
  46964. 35:13:17SVM is a supervised learning method that
  46965. 35:13:20looks at data and sorts it into one of
  46966. 35:13:23two categories. And in this case, we're
  46967. 35:13:25sorting the strawberry into the
  46968. 35:13:27strawberry side. At this point, you
  46969. 35:13:29should be asking the question, how does
  46970. 35:13:31the prediction work? Before we dig into
  46971. 35:13:33an example with numbers, let's apply
  46972. 35:13:35this to our fruit scenario. We have our
  46973. 35:13:37support vector machine. We've taken it
  46974. 35:13:39and we've taken labeled sample of data,
  46975. 35:13:42strawberries and apples, and we draw on
  46976. 35:13:44a line down the middle between the two
  46977. 35:13:46groups. This split now allows us to take
  46978. 35:13:49new data, in this case an apple and a
  46979. 35:13:51strawberry, and place them in the
  46980. 35:13:53appropriate group based on which side of
  46981. 35:13:54the line they fall in. And that way we
  46982. 35:13:56can predict the unknown. As colorful and
  46983. 35:13:58tasty as the fruit example is, let's
  46984. 35:14:00take a look at another example with some
  46985. 35:14:02numbers involved. And we can take a
  46986. 35:14:04closer look at how the math works. In
  46987. 35:14:06this example, we're going to be
  46988. 35:14:07classifying men and women. And we're
  46989. 35:14:09going to start with a set of people with
  46990. 35:14:11a different height and a different
  46991. 35:14:13weight. And to make this work, we'll
  46992. 35:14:15have to have a sample data set of female
  46993. 35:14:17where we have their height and weight
  46994. 35:14:19174, 65, 174, 88, and so on. And we'll
  46995. 35:14:23need a sample data set of the male. They
  46996. 35:14:24have a height 179, 90, 180 to 80 and so
  46997. 35:14:28on. Let's go ahead and put this on a
  46998. 35:14:29graph so we have a nice visual. So you
  46999. 35:14:31can see here we have two groups based on
  47000. 35:14:33the height versus the weight. And on the
  47001. 35:14:36left side we're going to have the women,
  47002. 35:14:37on the right side we're going to have
  47003. 35:14:39the men. Now if we're going to create a
  47004. 35:14:40classifier, let's add a new data point
  47005. 35:14:42and figure out if it's male or female.
  47006. 35:14:44So before we can do that, we need to
  47007. 35:14:47split our data first. We can split our
  47008. 35:14:49data by choosing any of these lines. In
  47009. 35:14:52this case, we draw in two lines through
  47010. 35:14:54the data in the middle that separates
  47011. 35:14:56the men from the women. But to predict
  47012. 35:14:57the gender of a new data point, we
  47013. 35:14:59should split the data in the best
  47014. 35:15:01possible way. And we say the best
  47015. 35:15:03possible way because this line has a
  47016. 35:15:06maximum space that separates the two
  47017. 35:15:08classes. Here you can see there's a
  47018. 35:15:10clear split between the two different
  47019. 35:15:12classes. And in this one, there's not so
  47020. 35:15:15much a clear split. This doesn't have
  47021. 35:15:17the maximum space that separates the
  47022. 35:15:18two. That is why this line best splits
  47023. 35:15:21the data. We don't want to just do this
  47024. 35:15:23by eyeballing it. And before we go
  47025. 35:15:25further, we need to add some technical
  47026. 35:15:27terms to this. We can also say that the
  47027. 35:15:30distance between the points in the line
  47028. 35:15:31should be as far as possible. In
  47029. 35:15:33technical terms, we can say the distance
  47030. 35:15:35between the support vector and the hyper
  47031. 35:15:38plane should be as far as possible. And
  47032. 35:15:40this is where the support vectors are
  47033. 35:15:42the extreme points in the data set. And
  47034. 35:15:44if you look at this data set, they have
  47035. 35:15:46circled two points which seem to be
  47036. 35:15:48right on the outskirts of the women and
  47037. 35:15:50one on the outskirts of the men. And
  47038. 35:15:52hyper plane has a maximum distance to
  47039. 35:15:54the support vectors of any class. Now
  47040. 35:15:56you'll see the line down the middle and
  47041. 35:15:58we call this the hyper plane because
  47042. 35:16:00when you're dealing with multiple
  47043. 35:16:01dimensions, it's really not just a line
  47044. 35:16:03but a plane of intersections. And you
  47045. 35:16:05can see here where the support vectors
  47046. 35:16:07have been drawn in dashed lines. The
  47047. 35:16:10math behind this is very simple. We take
  47048. 35:16:12D+ the shortest distance to the closest
  47049. 35:16:15positive point which would be on the
  47050. 35:16:17men's side and D minus is the shortest
  47051. 35:16:19distance to the closest negative point
  47052. 35:16:21which is on the women's side. The sum of
  47053. 35:16:23D plus and D minus is called the
  47054. 35:16:25distance margin or the distance between
  47055. 35:16:27the two support vectors that are shown
  47056. 35:16:29in the dashed lines. And then by finding
  47057. 35:16:32the largest distance margin, we can get
  47058. 35:16:35the optimal hyper plane. Once we've
  47059. 35:16:37created an optimal hyper plane, we can
  47060. 35:16:39easily see which side the new data fits
  47061. 35:16:41in. And based on the hyper plane, we can
  47062. 35:16:43say the new data point belongs to the
  47063. 35:16:44male gender. Hopefully that's clear how
  47064. 35:16:47that works on a visual level. As a data
  47065. 35:16:49scientist, you should also be asking
  47066. 35:16:51what happens if the hyper plane is not
  47067. 35:16:53optimal. If we select a hyper plane
  47068. 35:16:55having low margin, then there is a high
  47069. 35:16:57chance of mclassification. This
  47070. 35:16:59particular SVM model, the one we
  47071. 35:17:02discussed so far, is also called
  47072. 35:17:04referred to as the LSVM.
  47073. 35:17:06So far so clear, but a question should
  47074. 35:17:09be coming up. We have our sample data
  47075. 35:17:11set. But instead of looking like this,
  47076. 35:17:13what if it looked like this where we
  47077. 35:17:16have two sets of data, but one of them
  47078. 35:17:18occurs in the middle of another set. You
  47079. 35:17:20can see here where we have the blue and
  47080. 35:17:22the yellow and then blue again on the
  47081. 35:17:24other side of our data line. In this
  47082. 35:17:26data set, we can't use a hyper plane. So
  47083. 35:17:29when you see data like this, it's
  47084. 35:17:31necessary to move away from a 1D view of
  47085. 35:17:33the data to a two-dimensional view of
  47086. 35:17:35the data. And for the transformation, we
  47087. 35:17:37use what's called a kernel function. The
  47088. 35:17:40kernel function will take the 1D input
  47089. 35:17:42and transfer it to a two-dimensional
  47090. 35:17:44output. As you can see in this picture
  47091. 35:17:47here, the 1D when transferred to a
  47092. 35:17:49two-dimensional makes it very easy to
  47093. 35:17:51draw a line between the two data sets.
  47094. 35:17:53What if we make it even more
  47095. 35:17:55complicated? How do we perform an SVM
  47096. 35:17:57for this type of data set? Here you can
  47097. 35:17:59see we have a two-dimensional data set
  47098. 35:18:01where the data is in the middle
  47099. 35:18:03surrounded by the green data on the
  47100. 35:18:05outside. In this case, we're going to
  47101. 35:18:07segregate the two classes. We have our
  47102. 35:18:09sample data set and if you draw a line
  47103. 35:18:12through, it's obviously not an optimal
  47104. 35:18:13hyper plane in there. So to do that, we
  47105. 35:18:15need to transfer the 2D to a 3D array.
  47106. 35:18:18And when you translate it into a
  47107. 35:18:20three-dimensional array using the
  47108. 35:18:21kernel, you can see where you can place
  47109. 35:18:22a hyper plane right through it and
  47110. 35:18:24easily split the data. Before we start
  47111. 35:18:26looking at a programming example and
  47112. 35:18:28dive into the script, let's look at the
  47113. 35:18:30advantage of the support vector machine.
  47114. 35:18:32We'll start with highdimensional input
  47115. 35:18:34space or sometimes referred to as the
  47116. 35:18:37curse of dimensionality. We looked at
  47117. 35:18:39earlier one dimension, two dimension,
  47118. 35:18:41three dimension. When you get to a
  47119. 35:18:43thousand dimensions, a lot of problems
  47120. 35:18:45start occurring with most algorithms
  47121. 35:18:46that have to be adjusted for. The SVM
  47122. 35:18:49automatically does that in
  47123. 35:18:50highdimensional space. One of the
  47124. 35:18:52highdimensional space, one
  47125. 35:18:54highdimensional space that we work on is
  47126. 35:18:56sparse document vectors. This is where
  47127. 35:18:58we tokenize the words in documents so we
  47128. 35:19:00can run our machine learning algorithms
  47129. 35:19:02over them. I've seen ones get as high as
  47130. 35:19:042.4 million different tokens. That's a
  47131. 35:19:06lot of vectors to look at. And finally,
  47132. 35:19:09we have regularization parameter. The
  47133. 35:19:11realization parameter or lambda is a
  47134. 35:19:14parameter that helps figure out whether
  47135. 35:19:15we're going to have a bias or
  47136. 35:19:17overfitting of the data. Whether it's
  47137. 35:19:19going to be overfitted to very specific
  47138. 35:19:20instance or it's going to be biased to a
  47139. 35:19:22high or low value. With the SVM, it
  47140. 35:19:25naturally avoids the overfitting and
  47141. 35:19:27bias problems that we see in many other
  47142. 35:19:29algorithms. These three advantages of
  47143. 35:19:31the support vector machine make it a
  47144. 35:19:33very powerful tool to add to your
  47145. 35:19:35repertoire of machine learning tools.
  47146. 35:19:37Now, we did promise you a use case
  47147. 35:19:39study. We're actually going to dive in
  47148. 35:19:40to some Python programming. And so we're
  47149. 35:19:42going to go into a problem statement and
  47150. 35:19:44start off with the zoo. So in the zoo
  47151. 35:19:46example, we have um family members going
  47152. 35:19:49to the zoo and we have the young child
  47153. 35:19:50going, "Dad, is that a group of
  47154. 35:19:52crocodiles or alligators?" Well, that's
  47155. 35:19:54hard to differentiate. And zoos are a
  47156. 35:19:56great place to start looking at science
  47157. 35:19:58and understanding how things work,
  47158. 35:20:00especially as a young child. And so we
  47159. 35:20:02can see the parents sitting here
  47160. 35:20:03thinking, well, what is the difference
  47161. 35:20:04between a crocodile and an alligator?
  47162. 35:20:06Well, one, crocodiles are larger in
  47163. 35:20:08size. Alligators are smaller in size.
  47164. 35:20:10Snout width. The crocodiles have a
  47165. 35:20:12narrow snout and alligators have a wider
  47166. 35:20:14snout. And of course, in the modern day
  47167. 35:20:16and age, the father's sitting here is
  47168. 35:20:17thinking, "How can I turn this into a
  47169. 35:20:19lesson for my son?" And he goes, "Let a
  47170. 35:20:21support vector machine segregate the two
  47171. 35:20:23groups." I don't know if my dad ever
  47172. 35:20:25told me that, but that would be funny.
  47173. 35:20:26Now, in this example, we're not going to
  47174. 35:20:29use actual measurements and data. We're
  47175. 35:20:31just using that for imagery. And that's
  47176. 35:20:33very common in a lot of machine learning
  47177. 35:20:34algorithms and setting them up. But
  47178. 35:20:36let's roll up our sleeves and we'll talk
  47179. 35:20:38about that more in just a moment as we
  47180. 35:20:39break into our Python script. So here we
  47181. 35:20:42arrive in our actual coding and I'm
  47182. 35:20:45going to move this into a Python editor
  47183. 35:20:47in just a moment. But let's talk a
  47184. 35:20:49little bit about what we're going to
  47185. 35:20:50cover. First, we're going to cover in
  47186. 35:20:52the code the setup, how to actually
  47187. 35:20:55create our SVM. And you're going to find
  47188. 35:20:57that there's only two lines of code that
  47189. 35:20:58actually create it. And the rest of it
  47190. 35:21:00is done so quick and fast that it's all
  47191. 35:21:02here in the first page. and we'll show
  47192. 35:21:04you what that looks like as far as our
  47193. 35:21:06data because we're going to create some
  47194. 35:21:07data. I talked about creating data just
  47195. 35:21:09a minute ago. And so we'll get into the
  47196. 35:21:10creating data here and you'll see this
  47197. 35:21:12nice correction of our two blobs and
  47198. 35:21:13we'll go through that in just a second.
  47199. 35:21:15And then the second part is we're going
  47200. 35:21:16to take this and we're going to bump it
  47201. 35:21:18up a notch. We're going to show you what
  47202. 35:21:19it looks like behind the scenes. But
  47203. 35:21:21let's start with actually creating our
  47204. 35:21:23setup. I like to use the Anaconda
  47205. 35:21:25Jupyter notebook because it's very easy
  47206. 35:21:27to use, but you can use any of your
  47207. 35:21:29favorite Python editors or setups and go
  47208. 35:21:32in there. But let's go ahead and switch
  47209. 35:21:33over there and see what that looks like.
  47210. 35:21:35So here we are in the Anaconda Python
  47211. 35:21:38notebook or Anaconda Jupyter notebook
  47212. 35:21:40with Python. We're using Python 3. I
  47213. 35:21:43believe this is 3.5, but it should be
  47214. 35:21:45work in any of your 3x versions. And uh
  47215. 35:21:48you'd have to look at the sklearn and
  47216. 35:21:50make sure if you're using a 2x version,
  47217. 35:21:52an earlier version. Let's go and put our
  47218. 35:21:54code in there. And one of the things I
  47219. 35:21:55like about the Jupyter notebook is I can
  47220. 35:21:57go up to view and I'm going to go ahead
  47221. 35:21:59and toggle the line numbers on to make
  47222. 35:22:00it a little bit easier to talk about.
  47223. 35:22:03And we can even increase the size
  47224. 35:22:04because this is edited in in this case
  47225. 35:22:06I'm using Google Chrome explorer and
  47226. 35:22:08that's how it opens up for the editor.
  47227. 35:22:10Although anyone any like I said any
  47228. 35:22:11editor will work. Now the first step is
  47229. 35:22:13going to be our imports and we're going
  47230. 35:22:15to import four different parts. The
  47231. 35:22:18first two I want you to look at are line
  47232. 35:22:20one and line two are numpy as np and
  47233. 35:22:23mapplot library.pipplot
  47234. 35:22:25as plt. Now these are very standardized
  47235. 35:22:29imports when you're doing work. The
  47236. 35:22:30first one is the numbers python. We need
  47237. 35:22:33that because part of the platform we're
  47238. 35:22:35using uses that for the numpy array. And
  47239. 35:22:38I'll talk about that in a minute so you
  47240. 35:22:39can understand why we want to use a
  47241. 35:22:41numpy array versus a standard python
  47242. 35:22:43array. And normally it's pretty standard
  47243. 35:22:45setup to use NP for numpy. The map plot
  47244. 35:22:48library is how we're going to view our
  47245. 35:22:50data. So this has uh you do need the NP
  47246. 35:22:53for the sklearn module, but the map plot
  47247. 35:22:55library is purely for our use for
  47248. 35:22:57visualization. And so you really don't
  47249. 35:22:59need that for the SVM, but we're going
  47250. 35:23:00to put it there so you have a nice
  47251. 35:23:02visual aid and we can show you what it
  47252. 35:23:03looks like. That's really important at
  47253. 35:23:04the end when you finish everything so
  47254. 35:23:06you have a nice display for everybody to
  47255. 35:23:08look at. And then finally, we're going
  47256. 35:23:09to I'm going to jump one ahead to line
  47257. 35:23:12number four. That's the sklearn.datas
  47258. 35:23:14sets.samples generator import make
  47259. 35:23:18blobs. And I told you that we were going
  47260. 35:23:20to make up data. And this is a tool
  47261. 35:23:21that's in the sklearn to make up data. I
  47262. 35:23:23personally don't want to go to the zoo,
  47263. 35:23:25get in trouble for jumping over the
  47264. 35:23:26fence, and probably get eaten by the
  47265. 35:23:28crocodiles or alligators as I work on
  47266. 35:23:30measuring their snouts and width and
  47267. 35:23:32length. Instead, we're just going to
  47268. 35:23:34make up some data. And that's what that
  47269. 35:23:36make blobs is. It's a wonderful tool. If
  47270. 35:23:38you're ready to test your your uh setup
  47271. 35:23:41and you're not sure about what data
  47272. 35:23:42you're going to put in there, you can
  47273. 35:23:43create this blob and it makes it really
  47274. 35:23:45easy to use. And finally, we have our
  47275. 35:23:47actual SVM, the sklearn import SVM on
  47276. 35:23:50line three. So that covers all our
  47277. 35:23:52imports. We're going to create, remember
  47278. 35:23:54I used the make blobs to create data.
  47279. 35:23:56And we're going to create a capital X
  47280. 35:23:58and a lowercase Y equals make blobs in
  47281. 35:24:01samples equals 40. So we're going to
  47282. 35:24:02make 40 lines of data. It's going to
  47283. 35:24:04have two centers with a random state
  47284. 35:24:07equals 20. So each each each group's
  47285. 35:24:09going to have 20 different pieces of
  47286. 35:24:10data in it. And the way that looks is
  47287. 35:24:13that we'll have under X um an XY plane.
  47288. 35:24:16So I have two numbers under X and Y will
  47289. 35:24:18be 01. That's the two different centers.
  47290. 35:24:21So we have yes or no in this case
  47291. 35:24:23alligator crocodile. That's what that
  47292. 35:24:25represents. And then I told you that the
  47293. 35:24:28actual sklearn or the SVM is in two
  47294. 35:24:31lines of code. And we see it right here
  47295. 35:24:33with CLF equals SVM. SVC kernel equals
  47296. 35:24:37linear. And I set C equal to one.
  47297. 35:24:39Although in this example, since we are
  47298. 35:24:41not uh regularizing the data because we
  47299. 35:24:43want it to be very clear and easy to
  47300. 35:24:44see, I went ahead. You can set it to a
  47301. 35:24:46th00and a lot of times when you're not
  47302. 35:24:48doing that. But for this thing linear,
  47303. 35:24:50because it's a very simple linear
  47304. 35:24:51example, we only have the two dimensions
  47305. 35:24:53and it'll be a nice linear hyper plane.
  47306. 35:24:56It'll be a nice linear line instead of a
  47307. 35:24:58full plane. So we're not dealing with a
  47308. 35:25:00huge amount of data. And then all we
  47309. 35:25:01have to do is do clff.fit
  47310. 35:25:04x, y. And that's it. CLF has been
  47311. 35:25:07created. And then we're going to go
  47312. 35:25:08ahead and display it. And I'm going to
  47313. 35:25:10talk about this display here in just a
  47314. 35:25:12second. But let me go ahead and run this
  47315. 35:25:13code. And this is what we've done is
  47316. 35:25:15we've created two blobs. You'll see the
  47317. 35:25:17blue on the side and then kind of an
  47318. 35:25:19orang-ish uh on the other side. That's
  47319. 35:25:21our two sets of data. They represent one
  47320. 35:25:23represents crocodiles and one represents
  47321. 35:25:25alligators. And then we have our
  47322. 35:25:27measurements. In this case, we have like
  47323. 35:25:28the width and length of the snout. And I
  47324. 35:25:31did say I was going to come up here and
  47325. 35:25:32talk just a little bit about our plot.
  47326. 35:25:34And you'll see plt. That's what we
  47327. 35:25:36imported. We're going to do a scatter
  47328. 35:25:38plot. That means we're just putting dots
  47329. 35:25:40on there. And then look at this
  47330. 35:25:41notation. I have the capital X and then
  47331. 35:25:44in brackets I have a colon, 0ero. That's
  47332. 35:25:47from numpy. If you did that in a regular
  47333. 35:25:49array, you'll get an error in a Python
  47334. 35:25:51array. You have to have that in a numpy
  47335. 35:25:53array. It turns out that our make blobs
  47336. 35:25:55returns a numpy array. And this notation
  47337. 35:25:58is great because what it means is the
  47338. 35:26:00first part is the colon means we're
  47339. 35:26:02going to do all the rows. That's all the
  47340. 35:26:04data in our blob we created under
  47341. 35:26:06capital X. And then the second part has
  47342. 35:26:08a comma 0ero. We're only going to take
  47343. 35:26:10the first value. And then if you notice,
  47344. 35:26:13we do the same thing, but we're going to
  47345. 35:26:14take the second value. Remember, we
  47346. 35:26:16always start with zero and then one. So
  47347. 35:26:18we have column zero and column one. And
  47348. 35:26:20you can look at this as our XY plots.
  47349. 35:26:23The first one is the xplot and the
  47350. 35:26:25second one is the y plot. So the first
  47351. 35:26:27one is on the bottom 0 2 4 6 8 and 10.
  47352. 35:26:31And then the second one x of the one is
  47353. 35:26:34the 4 5 6 7 8 9 10 going up the left
  47354. 35:26:37hand side. S= 30 is just the size of the
  47355. 35:26:39dots. We can see them instead of real
  47356. 35:26:41tiny dots. And then cmap equals
  47357. 35:26:44plt.cm.paired.
  47358. 35:26:46And you'll also see the c equals y.
  47359. 35:26:48That's the color. We're using two colors
  47360. 35:26:5101. And that's why we get the nice blue
  47361. 35:26:54and the two different colors for the
  47362. 35:26:55alligator and the crocodile. Now you can
  47363. 35:26:58see here that we did this the actual fit
  47364. 35:27:00was done in two lines of code. A lot of
  47365. 35:27:02times there'll be a third line where we
  47366. 35:27:04regularize the data. We set it between
  47367. 35:27:06like minus one and one and we reshape
  47368. 35:27:08it. But for this it's not necessary and
  47369. 35:27:10it's also kind of nice because you can
  47370. 35:27:12actually see what's going on. And then
  47371. 35:27:14if we wanted to we wanted to actually
  47372. 35:27:16run a prediction. Let's take a look and
  47373. 35:27:17see what that looks like. And to predict
  47374. 35:27:19some new data and we'll show this again
  47375. 35:27:22as we get towards the end of digging in
  47376. 35:27:23deep. You can simply assign your new
  47377. 35:27:26data. In this case I am giving it a uh
  47378. 35:27:29width and length 34 and a width and
  47379. 35:27:31length 56. And note that I put the data
  47380. 35:27:34as a set of brackets and then I have the
  47381. 35:27:36brackets inside. And the reason I do
  47382. 35:27:38that is because when we're looking at
  47383. 35:27:40data it's designed to process a large
  47384. 35:27:43amount of data coming in. We don't want
  47385. 35:27:45to just process one line at a time. And
  47386. 35:27:47so in this case, I'm processing two
  47387. 35:27:48lines. And then I'm just going to print
  47388. 35:27:50and you'll see clf.predict new data. So
  47389. 35:27:53the CLF and the predict part is going to
  47390. 35:27:56give us an answer. And let's see what
  47391. 35:27:57that looks like. And you'll see 01. So
  47392. 35:28:00predicted the first one, the 34 is going
  47393. 35:28:02to be on the one side and the 56 is
  47394. 35:28:04going to be on the other side. So one
  47395. 35:28:06came out as a alligator and one came out
  47396. 35:28:08as a crocodile. Now that's pretty short
  47397. 35:28:10explanation for the setup, but really we
  47398. 35:28:12want to dug in and see what it's going
  47399. 35:28:14on behind the scenes. and let's see what
  47400. 35:28:16that looks like. So, the next step is to
  47401. 35:28:19dig in deep and find out what's going on
  47402. 35:28:22behind the scenes and also put that in a
  47403. 35:28:24nice pretty graph. We're going to spend
  47404. 35:28:26more work on this than we did actually
  47405. 35:28:28generating the original model. And
  47406. 35:28:30you'll see here that we go through a few
  47407. 35:28:32steps and I'm I'll move this over to our
  47408. 35:28:34editor in just a second. We come in, we
  47409. 35:28:36create our original data. It's exactly
  47410. 35:28:38identical to the first part and I'll
  47411. 35:28:40explain why we redid that and show you
  47412. 35:28:41how not to redo that. And then we're
  47413. 35:28:43going to go in there and add in those
  47414. 35:28:45lines. We're going to see what those
  47415. 35:28:47lines look like and how to set those up.
  47416. 35:28:49And finally, we're going to plot all
  47417. 35:28:51that on here and show it. And you'll get
  47418. 35:28:53a nice graph with the what we saw
  47419. 35:28:55earlier when we were going through the
  47420. 35:28:56theory behind this where it shows the
  47421. 35:28:59support vectors and the hyper plane. And
  47422. 35:29:02those are done where you can see the
  47423. 35:29:03support vectors as the dash lines and
  47424. 35:29:05the solid line which is the hyper plane.
  47425. 35:29:07Let's get that into our Jupyter
  47426. 35:29:09notebook. Before I scroll down to a new
  47427. 35:29:12line, I want you to notice line 13. It
  47428. 35:29:15has plot show. And we're going to talk
  47429. 35:29:17about that here in just a second. But
  47430. 35:29:18let's scroll down to a new line down
  47431. 35:29:20here. And I'm going to paste that code
  47432. 35:29:22in. And you'll see that the plot show
  47433. 35:29:24has moved down below. Let's scroll up a
  47434. 35:29:26little bit. And if you look at the top
  47435. 35:29:28here of our new section, 1 2 3 and four
  47436. 35:29:32is the same code we had before. And
  47437. 35:29:34let's go back up here and take a look at
  47438. 35:29:36that. We're going to fit the values on
  47439. 35:29:38our SVM. And then we're going to plot
  47440. 35:29:40scatter it. And then we're going to do a
  47441. 35:29:42plot show. So you should be asking why
  47442. 35:29:44are we redoing the same code. Well, when
  47443. 35:29:47you do the plot show, that blanks out
  47444. 35:29:49what's in the plot. So once I've done
  47445. 35:29:51this plot show, I have to reload that
  47446. 35:29:53data. Now, we could do this simply by
  47447. 35:29:55removing it up here, rerunning it, and
  47448. 35:29:58then coming down here, and then we
  47449. 35:30:00wouldn't have to rerun these first four
  47450. 35:30:01lines of code. Now, in this, it doesn't
  47451. 35:30:04matter too much. And you'll see the plot
  47452. 35:30:05show is down here and then removed right
  47453. 35:30:07there on line five. I'll go ahead and
  47454. 35:30:10just delete that out of there because we
  47455. 35:30:11don't want to blank out our screen. We
  47456. 35:30:13want to move on to the next setup. So,
  47457. 35:30:15we can go ahead and just skip the first
  47458. 35:30:17four lines because we did that before.
  47459. 35:30:19And let's take a look at the ax=
  47460. 35:30:22plt.gca.
  47461. 35:30:24Now, right now, we're actually spending
  47462. 35:30:25a lot of time just graphing. That's all
  47463. 35:30:27we're doing here. Okay. So, this is how
  47464. 35:30:29we display a nice graph with our results
  47465. 35:30:32and our data. AX is very standard not
  47466. 35:30:35used variable when you're talking about
  47467. 35:30:36PLT and it's just setting it to that
  47468. 35:30:39axis the last axis in the PLT. It can
  47469. 35:30:41get very confusing if you're working
  47470. 35:30:43with many different layers of data on
  47471. 35:30:45the same graph and this makes it very
  47472. 35:30:46easy to reference the ax. So this
  47473. 35:30:49reference is looking at the PLT that we
  47474. 35:30:52created and we already mapped out our
  47475. 35:30:54two blobs on. And then we want to know
  47476. 35:30:56the limits. So we want to know how big
  47477. 35:30:58the graph is. And we can find out the x
  47478. 35:31:00limit and the y limit simply with the
  47479. 35:31:02get x limit and get yimit commands which
  47480. 35:31:04is part of our metplot library. And then
  47481. 35:31:07we're going to create a grid. And you'll
  47482. 35:31:09see down here we have we've set the
  47483. 35:31:11variable xx equal to npines space ximit
  47484. 35:31:150 ximit 1a 30. And we've done the same
  47485. 35:31:18thing for the yspace. And then we're
  47486. 35:31:20going to go in here and we create a mesh
  47487. 35:31:22grid. And this is a numpy command. So
  47488. 35:31:25we're back to our numbers python. Let's
  47489. 35:31:28go through what these numpy commands
  47490. 35:31:30mean with the line space in the mesh
  47491. 35:31:32grid. We've taken xx small xx= np line
  47492. 35:31:36space. And we have our x limit zero and
  47493. 35:31:38our x limit one and we're going to
  47494. 35:31:40create 30 points on it. And we're going
  47495. 35:31:43to do the same thing for the y axis. Now
  47496. 35:31:45this has nothing to do with our
  47497. 35:31:46evaluation. It's uh all we're doing is
  47498. 35:31:48we're creating a grid of data. And so
  47499. 35:31:52we're creating a set of points between
  47500. 35:31:53zero and the x limit. We're creating 30
  47501. 35:31:56points. And the same thing with the y.
  47502. 35:31:58And then the mesh grid loops those all
  47503. 35:32:00together. So it forms a nice grid. So if
  47504. 35:32:02we were going to do this say between the
  47505. 35:32:04limit 0 and 10 and do 10 points, we
  47506. 35:32:06would have a 0 0 1 1 0 1 02 03 04 to 10
  47507. 35:32:12and so on. You can just imagine a point
  47508. 35:32:14at each corner one of those boxes. And
  47509. 35:32:16the mesh grid combines them all. So we
  47510. 35:32:18take the y and the xx we created and
  47511. 35:32:20creates the full grid. And we've set
  47512. 35:32:21that grid into the y coordinates and the
  47513. 35:32:24xx coordinates. Now remember, when we're
  47514. 35:32:26working with Numbi in Python, we like to
  47515. 35:32:29separate those. We like to have instead
  47516. 35:32:30of it being x comma 1, you know, x comma
  47517. 35:32:34y and then x2 comma y2 and in the next
  47518. 35:32:38set of data, it would be a column of x's
  47519. 35:32:40and a column of y's. And that's what we
  47520. 35:32:42have here is we have a column of y's. We
  47521. 35:32:44put it as a capital y y and a column of
  47522. 35:32:46x's, capital xx with all those different
  47523. 35:32:49points being listed. And finally, we get
  47524. 35:32:51down to the numpy vstack. Just as we
  47525. 35:32:54created those in the mesh grid, we're
  47526. 35:32:57now going to put them all into one
  47527. 35:32:59array, XY array. Now that we've created
  47528. 35:33:01the stack of data points, we're going to
  47529. 35:33:04do something interesting here. We're
  47530. 35:33:05going to create a value Z. And the Z
  47531. 35:33:08equals the CLF. That's our uh that's our
  47532. 35:33:11support vector machine we created and
  47533. 35:33:13we've already trained. And we have a
  47534. 35:33:16decision function. And we're going to
  47535. 35:33:17put the XY in there. So here we have all
  47536. 35:33:19this data. We're going to put that XY in
  47537. 35:33:22there, that data, and we're going to
  47538. 35:33:23reshape it. And you'll see that we have
  47539. 35:33:25the xx.shape in here. This literally
  47540. 35:33:28takes the xx, resets it up, connected to
  47541. 35:33:31the y, and the zvalue lets us know
  47542. 35:33:34whether it is the left hand side. It's
  47543. 35:33:37going to generate three different
  47544. 35:33:38values. The zvalue does, and it'll tell
  47545. 35:33:40us whether that data is a support vector
  47546. 35:33:43to the left, the hyper plane in the
  47547. 35:33:45middle, or the support vector to the
  47548. 35:33:47right. So it generates three different
  47549. 35:33:49values for each of those points. And
  47550. 35:33:50those points have been reshaped so
  47551. 35:33:52they're right on a line on those three
  47552. 35:33:54different lines. So we've set all of our
  47553. 35:33:56data up. We've labeled it to three
  47554. 35:33:58different areas and we've reshaped it.
  47555. 35:34:00And we've just taken 30 points in each
  47556. 35:34:02direction. If you do the math, you have
  47557. 35:34:0430 * 30. So that's 900 points of data.
  47558. 35:34:07And we separated it between the three
  47559. 35:34:08lines and reshaped it to fit those three
  47560. 35:34:10lines. We can then go back to our map
  47561. 35:34:12plot library where we've created the AX
  47562. 35:34:15and we're going to create a contour. And
  47563. 35:34:16you'll see here where we have contour,
  47564. 35:34:18capital XX, capital Y, Y. These have
  47565. 35:34:21been reshaped to fit those lines. Z is
  47566. 35:34:23the labels. So now we have the three
  47567. 35:34:25different points with the labels in
  47568. 35:34:26there. And we can set the colors equals
  47569. 35:34:28K. And I told you we had three different
  47570. 35:34:30labels, but we have uh three levels of
  47571. 35:34:33data. The alpha is just makes it kind of
  47572. 35:34:36see-through. So it's only uh 0.5 of the
  47573. 35:34:38value in there. So when we graph it, the
  47574. 35:34:40data will show up from behind it,
  47575. 35:34:41wherever the lines go. And finally, the
  47576. 35:34:43line styles. This is where we set the
  47577. 35:34:46two support vectors to be dash dash
  47578. 35:34:49lines and then a single one is just a
  47579. 35:34:51straight line. That's what all that
  47580. 35:34:53setup does. And then finally, we take
  47581. 35:34:55our ax.scatter. We're going to go ahead
  47582. 35:34:58and plot the support vectors, but we've
  47583. 35:35:00programmed it in there so that they look
  47584. 35:35:02nice like the dash dash line and the
  47585. 35:35:04dash line on that grid. And you can see
  47586. 35:35:06here when we do the CLFS support
  47587. 35:35:09vectors, we are looking at column zero
  47588. 35:35:11and column one. And then again we have
  47589. 35:35:14the S equals 100. So we're going to make
  47590. 35:35:16them larger. And the line width equals
  47591. 35:35:181, face colors equals none. Let's take a
  47592. 35:35:20look and see what that looks like when
  47593. 35:35:21we show it. And you can see when we get
  47594. 35:35:23down to our end result, it creates a
  47595. 35:35:25really nice graph. We have our two
  47596. 35:35:28support vectors and dash lines. And they
  47597. 35:35:30have the near data. So you can see those
  47598. 35:35:32two points or in this case the four
  47599. 35:35:34points where those lines nicely cleave
  47600. 35:35:36the data. And then you have your hyper
  47601. 35:35:38plane down the middle which is as far
  47602. 35:35:40from the two different points as
  47603. 35:35:41possible creating the maximum distance.
  47604. 35:35:43So you can see that we have our nice
  47605. 35:35:45output for the size of the body and the
  47606. 35:35:47width of the snout and we've easily
  47607. 35:35:49separated the two groups of crocodile
  47608. 35:35:51and alligator. Congratulations. You've
  47609. 35:35:54done it. We've made it. Of course, these
  47610. 35:35:55are pretend data for our crocodiles and
  47611. 35:35:58alligators. But this hands-on example
  47612. 35:36:00will help you to encounter any support
  47613. 35:36:02vector machine projects in the future.
  47614. 35:36:04And you can see how easy they are to set
  47615. 35:36:06up and look at in depth. We're going to
  47616. 35:36:08cover the K nearest neighbors a lot
  47617. 35:36:10referred to as KNN. And KNN is really a
  47618. 35:36:14fundamental place to start in the
  47619. 35:36:16machine learning. It's a basis of a lot
  47620. 35:36:18of other things and just the logic
  47621. 35:36:19behind it is easy to understand and
  47622. 35:36:21incorporated in other forms of machine
  47623. 35:36:24learning. So today, what's in it for
  47624. 35:36:26you? Why do we need KNN? What is KN&N?
  47625. 35:36:30How do we choose the factor K? When do
  47626. 35:36:33we use KNN? How does KN&N algorithm
  47627. 35:36:37work? And then we'll dive in to my
  47628. 35:36:40favorite part, the use case. Predict
  47629. 35:36:42whether a person will have diabetes or
  47630. 35:36:44not. That is a very common and popular
  47631. 35:36:46used data set as far as testing out
  47632. 35:36:50models and learning how to use the
  47633. 35:36:52different models in machine learning. By
  47634. 35:36:54now, we all know machine learning models
  47635. 35:36:56make predictions by learning from the
  47636. 35:36:58past data available. So we have our
  47637. 35:37:00input values. Our machine learning model
  47638. 35:37:02builds on those inputs of what we
  47639. 35:37:04already know and then we use that to
  47640. 35:37:06create a predicted output. Is that a
  47641. 35:37:09dog? Little kid looking over there and
  47642. 35:37:11watching the black cat cross their path.
  47643. 35:37:14No, dear. You can differentiate between
  47644. 35:37:16a cat and a dog based on their
  47645. 35:37:18characteristics.
  47646. 35:37:20Cats. Cats have sharp claws, uses to
  47647. 35:37:23climb, smaller length of ears, meows and
  47648. 35:37:25purr. Doesn't love to play around. dogs.
  47649. 35:37:29They have dull claws, bigger length of
  47650. 35:37:30ears, barks, loves to run around. You
  47651. 35:37:33usually don't see a cat running around
  47652. 35:37:35people, although I do have a cat that
  47653. 35:37:36does that where dogs do. And we can look
  47654. 35:37:38at these. We can say uh we can evaluate
  47655. 35:37:40the sharpness of the claws. How sharp
  47656. 35:37:42are their claws? And we can evaluate the
  47657. 35:37:45length of the ears. And we can usually
  47658. 35:37:47sort out cats from dogs based on even
  47659. 35:37:49those two characteristics. Now, tell me
  47660. 35:37:52if it is a cat or a dog. Not question.
  47661. 35:37:54Usually little kids know cats and dogs
  47662. 35:37:56by now. unless you live a place where
  47663. 35:37:58there's not many cats or dogs. So, if we
  47664. 35:38:00look at the sharpness of the claws, the
  47665. 35:38:01length of the ears, and we can see that
  47666. 35:38:03the cat has smaller ears and sharper
  47667. 35:38:06claws than the other animals. Its
  47668. 35:38:09features are more like cats. It must be
  47669. 35:38:11a cat. Sharp claws, length of ears, and
  47670. 35:38:14it goes in the cat group. Because KN&N
  47671. 35:38:17is based on feature similarity, we can
  47672. 35:38:19do classification using KN&N classifier.
  47673. 35:38:22So, we have our input value, the picture
  47674. 35:38:24of the black cat. It goes into our
  47675. 35:38:26trained model and it predicts that this
  47676. 35:38:28is a cat coming out. So what is knn?
  47677. 35:38:31What is the kn&n algorithm? K nearest
  47678. 35:38:35neighbors is what that stands for. Is
  47679. 35:38:37one of the simplest supervised machine
  47680. 35:38:39learning algorithms mostly used for
  47681. 35:38:41classification. So we want to know is
  47682. 35:38:44this a dog or it's not a dog? Is it a
  47683. 35:38:46cat or not a cat? It classifies a data
  47684. 35:38:49point based on how its neighbors are
  47685. 35:38:51classified. KN&N stores all available
  47686. 35:38:53cases and classifies new cases based on
  47687. 35:38:56a similarity measure. And here we've
  47688. 35:38:58gone from cats and dogs right into wine.
  47689. 35:39:01Another favorite of mine. KN&N stores
  47690. 35:39:03all available cases and classifies new
  47691. 35:39:05cases based on a similarity measure. And
  47692. 35:39:07here you see we have a measurement of
  47693. 35:39:09sulfur dioxide versus the chloride level
  47694. 35:39:12and then the different wines they've
  47695. 35:39:13tested and where they fall on that graph
  47696. 35:39:15based on how much sulfur dioxide and how
  47697. 35:39:17much chloride. K and K&N is a perimeter
  47698. 35:39:19that refers to the number of nearest
  47699. 35:39:21neighbors to include in the majority of
  47700. 35:39:23the voting process. And so if we add a
  47701. 35:39:25new glass of wine there, red or white,
  47702. 35:39:27we want to know what the neighbors are.
  47703. 35:39:29In this case, we're going to put K
  47704. 35:39:30equals 5. We'll talk about K in just a
  47705. 35:39:33minute. A data point is classified by
  47706. 35:39:35the majority of votes from its five
  47707. 35:39:36nearest neighbors. Here, the unknown
  47708. 35:39:39point would be classified as red since
  47709. 35:39:41four out of five neighbors are red. So,
  47710. 35:39:43how do we choose K? How do we know K
  47711. 35:39:46equals 5? I mean that's was the value we
  47712. 35:39:48put in there. I said we're going to talk
  47713. 35:39:49about it. How do we choose the factor K?
  47714. 35:39:52KN&N algorithm is based on feature
  47715. 35:39:54similarity. Choosing the right value of
  47716. 35:39:56K is a process called parameter tuning
  47717. 35:39:59and is important for better accuracy. So
  47718. 35:40:02at K equals 3, we can classify we have a
  47719. 35:40:04question mark in the middle as either a
  47720. 35:40:06as a square or not. Is it a square or is
  47721. 35:40:08it in this case a triangle? And so if we
  47722. 35:40:10set K equals to three, we're going to
  47723. 35:40:12look at the three nearest neighbors.
  47724. 35:40:14We're going to say this is a square. And
  47725. 35:40:16if we put k equals a 7, we classify as a
  47726. 35:40:19triangle depending on what the other
  47727. 35:40:21data is around it. And you can see as
  47728. 35:40:22the k changes depending on where that
  47729. 35:40:24point is, that drastically changes your
  47730. 35:40:26answer. And uh we jump here. We go, how
  47731. 35:40:29do we choose the factor of k? You'll
  47732. 35:40:31find this in all machine learning.
  47733. 35:40:33Choosing these factors, that's the face
  47734. 35:40:35you get. It's like, oh my gosh, did I
  47735. 35:40:37choose the right K? Did I set it right
  47736. 35:40:39my values in whatever machine learning
  47737. 35:40:41tool you're looking at? so that you
  47738. 35:40:42don't have a huge bias in one direction
  47739. 35:40:45or the other. And in terms of KNN, the
  47740. 35:40:48number of K, if you choose it too low,
  47741. 35:40:50the bias is based on it's just too
  47742. 35:40:52noisy. It's it's right next to a couple
  47743. 35:40:54things and it's going to pick those
  47744. 35:40:56things and you might get a skewed
  47745. 35:40:57answer. And if your K is too big, then
  47746. 35:41:00it's going to take forever to process.
  47747. 35:41:02So you're going to run into processing
  47748. 35:41:03issues and resource issues. So what we
  47749. 35:41:06do the most common use and there's other
  47750. 35:41:08options for choosing k is to use the
  47751. 35:41:11square root of n. So n is a total number
  47752. 35:41:14of values you have you take the square
  47753. 35:41:16root of it. In most cases you also if
  47754. 35:41:18it's an even number so if you're using
  47755. 35:41:20uh like in this case squares and
  47756. 35:41:22triangles if it's even you want to make
  47757. 35:41:24your k value odd. That helps it select
  47758. 35:41:27better. So in other words you're not
  47759. 35:41:28going to have a balance between two
  47760. 35:41:30different factors that are equal. So
  47761. 35:41:32usually take the square root of n and if
  47762. 35:41:34it's even you add one to it or subtract
  47763. 35:41:36one from it and that's where you get the
  47764. 35:41:37k value from that is the most common use
  47765. 35:41:39and it's pretty solid. It works very
  47766. 35:41:41well. When do we use kn? We can use kn
  47767. 35:41:45when data is labeled. So you need a
  47768. 35:41:47label on it. We know we have a group of
  47769. 35:41:49pictures with dogs cats cats. Data is
  47770. 35:41:52noisefree. And so you can see here when
  47771. 35:41:55we have a class and we have like
  47772. 35:41:57underweight 140 23 Hello kitty normal
  47773. 35:42:00that's pretty confusing. We have a a
  47774. 35:42:02high variety of data coming in. So it's
  47775. 35:42:04very noisy and that would cause an
  47776. 35:42:06issue. Data set is small. So we're
  47777. 35:42:08usually working with smaller data sets
  47778. 35:42:10where you might get into gig of data if
  47779. 35:42:13it's really clean. It doesn't have a lot
  47780. 35:42:14of noise because KN&N is a lazy learner.
  47781. 35:42:17I.e. it doesn't learn a discriminative
  47782. 35:42:19function from the training set. So it's
  47783. 35:42:21very lazy. So if you have very
  47784. 35:42:23complicated data and you have a large
  47785. 35:42:25amount of it, you're not going to use
  47786. 35:42:26the kn. But it's really great to get a
  47787. 35:42:28place to start. Even with large data,
  47788. 35:42:30you can sort out a small sample and get
  47789. 35:42:32an idea of what that looks like using
  47790. 35:42:33the KN&N and also just using for smaller
  47791. 35:42:36data sets. KN&N works really good. How
  47792. 35:42:39does the KN&N algorithm work? Consider a
  47793. 35:42:42data set having two variables, height in
  47794. 35:42:44centimeters and weight in kilograms. And
  47795. 35:42:47each point is classified as normal or
  47796. 35:42:49underweight. So we can see right here we
  47797. 35:42:51have two variables, you know, true
  47798. 35:42:53false. They're either normal or they're
  47799. 35:42:55not. They're underweight. On the basis
  47800. 35:42:56of the given data, we have to classify
  47801. 35:42:59the below set as normal or underweight
  47802. 35:43:01using KN&N. So if we have new data
  47803. 35:43:03coming in that says 57 kg and 177 cm, is
  47804. 35:43:08that going to be normal or underweight?
  47805. 35:43:10To find the nearest neighbors, we'll
  47806. 35:43:12calculate the ukitian distance.
  47807. 35:43:14According to the uklitian distance
  47808. 35:43:16formula, the distance between two points
  47809. 35:43:18in the plane with the coordinates xy and
  47810. 35:43:21ab is given by distance d equals the
  47811. 35:43:24square root of x - a^2 + y - b^2. And
  47812. 35:43:29you can remember that from the two edges
  47813. 35:43:31of a triangle. We're computing the third
  47814. 35:43:33edge since we know the x side and the y
  47815. 35:43:36side. Let's calculate it to understand
  47816. 35:43:38clearly. So we have our unknown point
  47817. 35:43:40and we placed it there in red. And we
  47818. 35:43:42have our other points where the data is
  47819. 35:43:44scattered around. The distance d1 is the
  47820. 35:43:47square<unk> of 170 minus 167^ squar + 57
  47821. 35:43:51- 51^ 2ar which is about 6.7 and
  47822. 35:43:55distance 2 is about 13 and distance 3 is
  47823. 35:43:59about 13.4. Similarly, we will calculate
  47824. 35:44:03the ukitian distance of unknown data
  47825. 35:44:05point from all the points in the data
  47826. 35:44:07set. And because we're dealing with
  47827. 35:44:08small amount of data, that's not that
  47828. 35:44:10hard to do and it's actually pretty
  47829. 35:44:11quick for a computer and it's not a
  47830. 35:44:13really complicated math. You can just
  47831. 35:44:14see how close is the data based on the
  47832. 35:44:16uklidian distance. Hence, we have
  47833. 35:44:18calculated the uklidian distance of
  47834. 35:44:20unknown data point from all the points
  47835. 35:44:22as shown where x1 and y1 equal 57 and
  47836. 35:44:26170 whose class we have to classify. So
  47837. 35:44:29now we're looking at that. We're saying
  47838. 35:44:30well here's the ukitian distance. Who's
  47839. 35:44:32going to be their closest neighbors? Now
  47840. 35:44:34let's calculate the nearest neighbor at
  47841. 35:44:36k equals 3. And we can see the three
  47842. 35:44:39closest neighbors puts them at normal.
  47843. 35:44:41And that's pretty self-evident when you
  47844. 35:44:43look at this graph. It's pretty easy to
  47845. 35:44:44say okay what you know we're just voting
  47846. 35:44:46normal normal normal. Three votes for
  47847. 35:44:48normal. This is going to be a normal
  47848. 35:44:49weight. So majority of neighbors are
  47849. 35:44:51pointing towards normal. Hence as per
  47850. 35:44:53KN&N algorithm the class of 571 170
  47851. 35:44:56should be normal. So a recap of KN&N
  47852. 35:44:59positive integer K is specified along
  47853. 35:45:02with a new sample. We select the K
  47854. 35:45:04entries in our database which are
  47855. 35:45:05closest to the new sample. We find the
  47856. 35:45:07most common classification of these
  47857. 35:45:09entries. This is the classification we
  47858. 35:45:11give to the new sample. So, as you can
  47859. 35:45:13see, it's pretty straightforward. We're
  47860. 35:45:14just looking for the closest things that
  47861. 35:45:16match what we got. So, let's take a look
  47862. 35:45:18and see what that looks like in a use
  47863. 35:45:20case in Python. So, let's dive into the
  47864. 35:45:23predict diabetes use case. So, use case,
  47865. 35:45:26predict diabetes. The objective, predict
  47866. 35:45:29whether a person will be diagnosed with
  47867. 35:45:31diabetes or not. We have a data set of
  47868. 35:45:34768 people who were or were not
  47869. 35:45:37diagnosed with diabetes. And let's go
  47870. 35:45:39ahead and open that file and just take a
  47871. 35:45:41look at that data. And this is in a
  47872. 35:45:43simple spreadsheet format. The data
  47873. 35:45:46itself is commaepparated. Very common
  47874. 35:45:48set of data. And it's also a very common
  47875. 35:45:50way to get the data. And you can see
  47876. 35:45:51here we have columns A through I. That's
  47877. 35:45:54what 1 2 3 4 5 6 7 8. um eight columns
  47878. 35:45:59with a particular attribute and then the
  47879. 35:46:02ninth column which is the outcome is
  47880. 35:46:04whether they have diabetes. As a data
  47881. 35:46:06scientist, the first thing you should be
  47882. 35:46:07looking at is insulin. Well, you know,
  47883. 35:46:09if someone has insulin, they have
  47884. 35:46:11diabetes because that's why they're
  47885. 35:46:12taking it. And that could cause issue in
  47886. 35:46:13some of the machine learning packages,
  47887. 35:46:15but for very basic setup, this works
  47888. 35:46:17fine for doing the KNN. And the next
  47889. 35:46:20thing you notice is it it didn't take
  47890. 35:46:22very much to open it up. Um I can scroll
  47891. 35:46:24down to the bottom of the data. There's
  47892. 35:46:25768.
  47893. 35:46:26It's pretty much a small data set. You
  47894. 35:46:28know, at 769, I can easily fit this into
  47895. 35:46:32my RAM on my computer. I can look at it.
  47896. 35:46:35I can manipulate it. And it's not going
  47897. 35:46:37to really tax just a regular desktop
  47898. 35:46:39computer. You don't even need an
  47899. 35:46:40enterprise version to run a lot of this.
  47900. 35:46:42So, let's start with importing all the
  47901. 35:46:44tools we need. And before that, of
  47902. 35:46:46course, we need to discuss what IDE I'm
  47903. 35:46:48using. Certainly, you can use any uh
  47904. 35:46:50particular editor for Python, but I like
  47905. 35:46:52to use for doing uh very basic visual
  47906. 35:46:55stuff. the Anaconda, which is great for
  47907. 35:46:57doing demos with the Jupyter Notebook.
  47908. 35:47:00And just a quick view of the Anaconda
  47909. 35:47:02Navigator, which is the new release out
  47910. 35:47:04there, which is really nice. You can see
  47911. 35:47:06under home, I can choose my application.
  47912. 35:47:09We're going to be using Python 3.6. I
  47913. 35:47:11have a couple different uh versions on
  47914. 35:47:13this particular machine. If I go under
  47915. 35:47:15environments, I can create a unique
  47916. 35:47:16environment for each one, which is nice.
  47917. 35:47:18And there's even a little button there
  47918. 35:47:20where I can install different packages.
  47919. 35:47:21So, if I click on that button and open
  47920. 35:47:23the terminal, I can then use a simple
  47921. 35:47:24pip install to install different
  47922. 35:47:26packages I'm working with. Let's go
  47923. 35:47:28ahead and go back under home and we're
  47924. 35:47:29going to launch our notebook. And I've
  47925. 35:47:31already, you know, kind of like uh the
  47926. 35:47:33old cooking shows, I've already prepared
  47927. 35:47:34a lot of my stuff. So, we don't have to
  47928. 35:47:36wait for it to launch because it takes a
  47929. 35:47:37few minutes for it to open up a browser
  47930. 35:47:40window. In this case, I'm going to it's
  47931. 35:47:42going to open up Chrome because that's
  47932. 35:47:43my default that I use. And since the
  47933. 35:47:45script is pre-done, you'll see I have a
  47934. 35:47:46number of windows open up at the top,
  47935. 35:47:48the one we're working in. And uh since
  47936. 35:47:50we're working on the KN&N predict
  47937. 35:47:53whether a person will have diabetes or
  47938. 35:47:54not. Let's go and put that title in
  47939. 35:47:56there. And I'm also going to go up here
  47940. 35:47:58and click on cell. Actually, we want to
  47941. 35:48:00go ahead and first insert a cell below.
  47942. 35:48:02And then I'm going to go back up to the
  47943. 35:48:04top cell. And I'm going to change the
  47944. 35:48:06cell type to markdown. That means this
  47945. 35:48:09is not going to run as Python. It's a
  47946. 35:48:10markdown language. So if I run this
  47947. 35:48:12first one, it comes up in nice big
  47948. 35:48:13letters, which is kind of nice. Remind
  47949. 35:48:15us what we're working on. And by now you
  47950. 35:48:18should be familiar with doing all of our
  47951. 35:48:19imports. We're going to import the
  47952. 35:48:21pandas as pd import numpy is np. Pandas
  47953. 35:48:25is the uh pandas data frame and numpy is
  47954. 35:48:28a number array. Very powerful tools to
  47955. 35:48:30use in here. So we have our imports. So
  47956. 35:48:33we've brought in our pandas or numpy our
  47957. 35:48:35two general python tools. And then you
  47958. 35:48:38can see over here we have our train test
  47959. 35:48:40split. By now you should be familiar
  47960. 35:48:42with splitting the data. We want to
  47961. 35:48:44split part of it for training our thing
  47962. 35:48:45and then training our particular model
  47963. 35:48:48and then we want to go ahead and test
  47964. 35:48:49the remaining data to see how good it
  47965. 35:48:51is. Pre-processing a standard scaler
  47966. 35:48:54pre-processor so we don't have a bias of
  47967. 35:48:56really large numbers. Remember in the
  47968. 35:48:58data we had like number of pregnancies
  47969. 35:49:00isn't going to get very large where the
  47970. 35:49:02amount of insulin they take and get up
  47971. 35:49:03to 256. So 256 versus six that will skew
  47972. 35:49:08results. So we want to go ahead and
  47973. 35:49:09change that so they're all uniform
  47974. 35:49:11between minus1 and one. And then the
  47975. 35:49:13actual tool. This is the K neighbors
  47976. 35:49:15classifier we're going to use. And
  47977. 35:49:18finally, the last three are three tools
  47978. 35:49:21to test. All about testing our model.
  47979. 35:49:23How good is it? We just put down test on
  47980. 35:49:25there. And we have our confusion matrix,
  47981. 35:49:27our F1 score, and our accuracy. So we
  47982. 35:49:29have our two general Python modules
  47983. 35:49:32we're importing. And then we have our
  47984. 35:49:34six modules specific from the sklearn
  47985. 35:49:37setup. And then we do need to go ahead
  47986. 35:49:39and run this. So these are actually
  47987. 35:49:42imported. There we go. And then move on
  47988. 35:49:44to the next step. And so in this set,
  47989. 35:49:46we're going to go ahead and load the
  47990. 35:49:47database. We're going to use pandas.
  47991. 35:49:49Remember pandas is pd. And we'll take a
  47992. 35:49:51look at the data in Python. We looked at
  47993. 35:49:53it in a simple spreadsheet, but usually
  47994. 35:49:55I like to also pull it up so that we can
  47995. 35:49:57see what we're doing. So here's our data
  47996. 35:49:59set equals PD read CSV. That's a pandas
  47997. 35:50:03command. And the diabetes folder I just
  47998. 35:50:06put in the same folder where my IPython
  47999. 35:50:08script is. If you put in a different
  48000. 35:50:10folder, you'd need the full length on
  48001. 35:50:12there. We can also do a quick length of
  48002. 35:50:15uh the data set. That is a simple Python
  48003. 35:50:18command. Leen for length. We might even
  48004. 35:50:20let's go ahead and print that. We'll go
  48005. 35:50:21print. And if you do it on its own line,
  48006. 35:50:24length data set in the Jupyter notebook,
  48007. 35:50:26it'll automatically print it. But when
  48008. 35:50:28you're in most of your different setups,
  48009. 35:50:30you want to do the print in front of
  48010. 35:50:31there. And then we want to take a look
  48011. 35:50:33at the actual data set. And since we're
  48012. 35:50:34in pandas, we can simply do data set
  48013. 35:50:37head. And again, let's go ahead and add
  48014. 35:50:40the print in there. If you put a bunch
  48015. 35:50:42of these in a row, you know that data
  48016. 35:50:44set one head, data set two head, it only
  48017. 35:50:46prints out the last one. So, I usually
  48018. 35:50:48always like to keep the print statement
  48019. 35:50:49in there. But because most projects only
  48020. 35:50:52use one data frame, Panda's data frame,
  48021. 35:50:54doing it this way doesn't really matter.
  48022. 35:50:56The other way works just fine. And you
  48023. 35:50:58can see when we hit the run button, we
  48024. 35:51:00have the 768 lines, which we knew, and
  48025. 35:51:02we have our pregnancies. It's
  48026. 35:51:04automatically given a label on the left.
  48027. 35:51:06Remember the head only shows the first
  48028. 35:51:08five lines. So we have zero through
  48029. 35:51:11four. And just a quick look at the data.
  48030. 35:51:13You can see it matches what we looked at
  48031. 35:51:15before. We have pregnancy, glucose,
  48032. 35:51:17blood pressure all the way to age. And
  48033. 35:51:20then the outcome on the end. And we're
  48034. 35:51:22going to do a couple things in this next
  48035. 35:51:24step. We're going to create a list of
  48036. 35:51:26columns where we can't have zero.
  48037. 35:51:28There's no such thing as zero skin
  48038. 35:51:30thickness or zero blood pressure, zero
  48039. 35:51:33glucose. Uh any of those, you'd be dead.
  48040. 35:51:36So, not a really good factor if they
  48041. 35:51:37don't if they have a zero in there
  48042. 35:51:39because they didn't have the data. And
  48043. 35:51:40we'll take a look at that because we're
  48044. 35:51:41going to start replacing that
  48045. 35:51:42information with a couple of different
  48046. 35:51:45things. And let's see what that looks
  48047. 35:51:46like. So, first we create a nice list.
  48048. 35:51:49As you can see, we have the values
  48049. 35:51:51talked about glucose, blood pressure,
  48050. 35:51:52skin thickness. Uh, and this is a nice
  48051. 35:51:54way when you're working with columns is
  48052. 35:51:56to list the columns you need to do some
  48053. 35:51:58kind of transformation on. Uh, very
  48054. 35:51:59common thing to do. And then for this
  48055. 35:52:01particular setup, we certainly could use
  48056. 35:52:03the there's some Panda tools that will
  48057. 35:52:06do a lot of this where we can replace
  48058. 35:52:07the NA, but we're going to go ahead and
  48059. 35:52:10do it as a data set column equals data
  48060. 35:52:13set column.replace. This is this is
  48061. 35:52:15still pandas. You can do a direct.
  48062. 35:52:17There's also one that that you look for
  48063. 35:52:19your nan. A lot of different options in
  48064. 35:52:21here. But the nan numpan is what that
  48065. 35:52:24stands for is non doesn't exist. So the
  48066. 35:52:27first thing we're doing here is we're
  48067. 35:52:29replacing the zero with a numpy none.
  48068. 35:52:33There's no data there. That's what that
  48069. 35:52:34says. That's what this is saying right
  48070. 35:52:36here. So put the zero in and we're going
  48071. 35:52:38to replace zeros with no data. So if
  48072. 35:52:41it's a zero, that means the person's
  48073. 35:52:43well hopefully not dead. Hopefully they
  48074. 35:52:44just didn't get the data. The next thing
  48075. 35:52:45we want to do is we're going to create
  48076. 35:52:47the mean which is the in integer from
  48077. 35:52:49the data set from the column mean where
  48078. 35:52:52we skip NAS. We can do that. That is a
  48079. 35:52:55pandas command there, the skip na. So
  48080. 35:52:57we're going to figure out the mean of
  48081. 35:52:59that data set. And then we're going to
  48082. 35:53:00take that data set column and we're
  48083. 35:53:02going to replace all the npnan
  48084. 35:53:06with the means. Why did we do that? And
  48085. 35:53:08we could have actually just uh taken
  48086. 35:53:10this step and gone right down here and
  48087. 35:53:11just replace zero and skip anything
  48088. 35:53:13where except you could actually there's
  48089. 35:53:15a way to skip zeros and then just
  48090. 35:53:16replace all the zeros. But in this case,
  48091. 35:53:18we want to go ahead and do it this way.
  48092. 35:53:19So you could see that we're switching
  48093. 35:53:21this to a non-existent value. Then we're
  48094. 35:53:23going to create the mean. Well, this is
  48095. 35:53:25the average person. So if we don't know
  48096. 35:53:28what it is, if they did not get the data
  48097. 35:53:30and the data is missing, one of the
  48098. 35:53:32tricks is you replace it with the
  48099. 35:53:34average. What is the most common data
  48100. 35:53:36for that? This way you can still use the
  48101. 35:53:39rest of those values to do your
  48102. 35:53:40computation and it kind of just brings
  48103. 35:53:43that particular value or those missing
  48104. 35:53:44values out of the equation. Let's go
  48105. 35:53:46ahead and take this and we'll go ahead
  48106. 35:53:48and run it. Doesn't actually do
  48107. 35:53:50anything. So we're still preparing our
  48108. 35:53:52data. If you want to see what that looks
  48109. 35:53:54like, we don't have anything in the
  48110. 35:53:55first few lines, so it's not going to
  48111. 35:53:57show up. But we certainly could look at
  48112. 35:53:59a row. Let's do that. Let's go into our
  48113. 35:54:01data set. Let's print a data set. And
  48114. 35:54:04let's pick in this case, let's just do
  48115. 35:54:07glucose. And if I run this, this is
  48116. 35:54:10going to print all the different glucose
  48117. 35:54:11levels going down. And we thankfully
  48118. 35:54:14don't see anything in here that looks
  48119. 35:54:16like missing data, at least on the ones
  48120. 35:54:17it shows. You can see it skipped a bunch
  48121. 35:54:19in the middle because that's what it
  48122. 35:54:20does. If you have too many lines in
  48123. 35:54:21Jupyter notebook, it'll skip a few and
  48124. 35:54:23and go on to the next in a data set. Let
  48125. 35:54:25me go and remove this. And we'll just
  48126. 35:54:27zero out that. And of course, before we
  48127. 35:54:30do any processing, before proceeding any
  48128. 35:54:32further, we need to split the data set
  48129. 35:54:34into our train and testing data. That
  48130. 35:54:36way, we have something to train it with
  48131. 35:54:37and something to test it on. And you're
  48132. 35:54:40going to notice we did a little
  48133. 35:54:40something here with the uh pandas
  48134. 35:54:42database code. There we go. My drawing
  48135. 35:54:45tool. We've added in this right here off
  48136. 35:54:47the data set. And what this says is that
  48137. 35:54:50the first one in pandas, this is from
  48138. 35:54:52the PD pandas. It's going to say within
  48139. 35:54:55the data set, we want to look at the eye
  48140. 35:54:57location and it is all rows. That's what
  48141. 35:54:59that says. So we're going to keep all
  48142. 35:55:01the rows, but we're only looking at
  48143. 35:55:02zero, column 0 to 8. Remember column 9.
  48144. 35:55:06Here it is right up here. We printed it
  48145. 35:55:07in here is outcome. Well, that's not
  48146. 35:55:09part of the training data. That's part
  48147. 35:55:10of the answer. Yeah, it's column 9, but
  48148. 35:55:12it's listed as eight. Number eight. So 0
  48149. 35:55:14to eight is nine columns. So uh eight is
  48150. 35:55:17the value. And when you see it in here,
  48151. 35:55:19zero, this is actually 0 to 7. It
  48152. 35:55:22doesn't include the last one. And then
  48153. 35:55:24we go down here to Y, which is our
  48154. 35:55:25answer. And we want just the last one,
  48155. 35:55:29just column 8. And you can do it this
  48156. 35:55:31way with this particular notation. And
  48157. 35:55:33then if you remember, we imported the
  48158. 35:55:34train test split that's part of the
  48159. 35:55:37sklearn right there. And we simply put
  48160. 35:55:39in our X and our Y. We're going to do
  48161. 35:55:42random state equals zero. You don't have
  48162. 35:55:44to necessarily seed it. That's a seed
  48163. 35:55:45number. I think the default is one when
  48164. 35:55:47you seated it. I'd have to look that up.
  48165. 35:55:49And then the test size. Test size is
  48166. 35:55:510.2. That simply means we're going to
  48167. 35:55:53take 20% of the data and put it aside so
  48168. 35:55:55that we can test it later. That's all
  48169. 35:55:57that is. And again, we're going to run
  48170. 35:55:59it. Not very exciting. So far, we
  48171. 35:56:01haven't had any print out other than to
  48172. 35:56:02look at the data. But that is a lot of
  48173. 35:56:04this is prepping this data. Once you
  48174. 35:56:06prep it, the actual lines of code are
  48175. 35:56:08quick and easy. And we're almost there.
  48176. 35:56:10But the actual writing of our KN&N, we
  48177. 35:56:12need to go ahead and do a scale the
  48178. 35:56:14data. If you remember correctly, we're
  48179. 35:56:16fitting the data in a standard scaler,
  48180. 35:56:18which means instead of the data being
  48181. 35:56:20from, you know, five to 303 in one
  48182. 35:56:23column and the next column is 1 to six,
  48183. 35:56:26we're going to set that all so that all
  48184. 35:56:27the data is between minus1 and one.
  48185. 35:56:30That's what that standard scaler does.
  48186. 35:56:32Keeps it standardized. And we only want
  48187. 35:56:34to fit the scaler with the training set,
  48188. 35:56:37but we want to make sure the testing set
  48189. 35:56:39is the X test going in is also
  48190. 35:56:42transformed. So it's processing it the
  48191. 35:56:45same. So here we go with our standard
  48192. 35:56:47scaler. We're going to call it sc__x for
  48193. 35:56:49the scaler. And we're going to import
  48194. 35:56:51the standard scaler into this variable.
  48195. 35:56:53And then our xrain equals sc_x.fit
  48196. 35:56:58transform. So we're creating the scaler
  48197. 35:57:00on the x-ra variable. And then our x
  48198. 35:57:02test, we're also going to transform it.
  48199. 35:57:04So we've trained and transformed the
  48200. 35:57:06x-ra. And then the x test isn't part of
  48201. 35:57:09that training. It isn't part of that of
  48202. 35:57:11training the transformer. it just gets
  48203. 35:57:13transformed. That's all it does. And
  48204. 35:57:15again, we're going to go and run this.
  48205. 35:57:16And if you look at this, we've now gone
  48206. 35:57:18through these steps, all three of them.
  48207. 35:57:21We've taken care of replacing our zeros
  48208. 35:57:24for key columns that shouldn't be zero,
  48209. 35:57:27and we've replaced that with the means
  48210. 35:57:30of those columns. That way, that they
  48211. 35:57:32fit right in with our data models. We've
  48212. 35:57:34come down here, and we split the data.
  48213. 35:57:36So, now we have our test data and our
  48214. 35:57:39training data. And then we've taken and
  48215. 35:57:41we've scaled the data. So all of our
  48216. 35:57:43data going in. No, no, we don't tra we
  48217. 35:57:45don't train the Y part, the Y train and
  48218. 35:57:48Y test that never has to be trained.
  48219. 35:57:51It's only the data going in. That's what
  48220. 35:57:53we want to train in there. Then define
  48221. 35:57:55the model using K neighbors classifier
  48222. 35:57:57and fit the train data in the model. So
  48223. 35:57:59we do all that data prep. And you can
  48224. 35:58:01see down here we're only going to have a
  48225. 35:58:03couple lines of code where we're
  48226. 35:58:05actually building our model and training
  48227. 35:58:07it. That's one of the cool things about
  48228. 35:58:09Python and how far we've come. It's such
  48229. 35:58:11an exciting time to be in machine
  48230. 35:58:12learning because there's so many
  48231. 35:58:13automated tools. Let's see. Before we do
  48232. 35:58:16this, let's do a quick length of and
  48233. 35:58:18let's do y. We want let's just do length
  48234. 35:58:21of y. And we get 768. And if we import
  48235. 35:58:26math, we do math dot square root. Let's
  48236. 35:58:30do y train. There we go. It's actually
  48237. 35:58:33supposed to be x train. Before we do
  48238. 35:58:36this, let's go ahead and do import math
  48239. 35:58:38and do math square root length of y
  48240. 35:58:41test. And when I run that, we get
  48241. 35:58:4312.409.
  48242. 35:58:45I want to see show you where this number
  48243. 35:58:46comes from. We're about to use 12 is an
  48244. 35:58:48even number. So if you know if you're
  48245. 35:58:50ever voting on things, remember the
  48246. 35:58:52neighbors all vote. Don't want to have
  48247. 35:58:53an even number of neighbors voting. So
  48248. 35:58:55we want to do something odd. And let's
  48249. 35:58:57just take one away. We'll make it 11.
  48250. 35:58:59Let me delete this out of here. That's
  48251. 35:59:00one of the reasons I love Jupyter
  48252. 35:59:01Notebook because you can flip around and
  48253. 35:59:03do all kinds of things on the fly. So,
  48254. 35:59:05we'll go ahead and put in our
  48255. 35:59:06classifier. We're creating our
  48256. 35:59:07classifier now and it's going to be the
  48257. 35:59:09K neighbors classifier. In neighbors
  48258. 35:59:11equal 11. Remember, we did 12 - 1 for
  48259. 35:59:1411. So, we have an odd number of
  48260. 35:59:16neighbors. P= 2 because we're looking
  48261. 35:59:18for is it are they diabetic or not? And
  48262. 35:59:21we're using the ukitian metric. There
  48263. 35:59:23are other means of measuring the
  48264. 35:59:25distance. You could do like square
  48265. 35:59:27square means value. There's all kinds of
  48266. 35:59:29measure this, but the uklidian is the
  48267. 35:59:31most common one and it works quite well.
  48268. 35:59:33It's important to evaluate the model.
  48269. 35:59:35Let's use the confusion matrix to do
  48270. 35:59:37that. And we're going to use the
  48271. 35:59:38confusion matrix. Wonderful tool. And
  48272. 35:59:41then we'll jump into the F1 score. And
  48273. 35:59:44finally, accuracy score, which is
  48274. 35:59:46probably the most commonly used quoted
  48275. 35:59:49number when you go into a meeting or
  48276. 35:59:51something like that. So, let's go ahead
  48277. 35:59:52and paste that in there. And we'll set
  48278. 35:59:54the CM equal to confusion matrix. Y
  48279. 35:59:57test, Y predict. So those are the two
  48280. 35:59:59values we're going to put in there. And
  48281. 36:00:01let me go ahead and run that and print
  48282. 36:00:02it out. And the way you interpret this
  48283. 36:00:05is you have the Y predicted, which would
  48284. 36:00:08be your title up here. You can do uh
  48285. 36:00:10let's just do Predicted
  48286. 36:00:13across the top and actual going down.
  48287. 36:00:17Actual. It's always hard to to write in
  48288. 36:00:20here. Actual. That means that this
  48289. 36:00:22column here down the middle, that's the
  48290. 36:00:24important column. And it means that our
  48291. 36:00:26prediction said 94 and prediction in the
  48292. 36:00:30actual agreed on 94 and 32. This number
  48293. 36:00:34here, the 13 and the 15, those are what
  48294. 36:00:38was wrong. So you could have like three
  48295. 36:00:40different if you're looking at this
  48296. 36:00:41across three different variables instead
  48297. 36:00:43of just two. You'd end up with a third
  48298. 36:00:45row down here in the column going down
  48299. 36:00:47the middle. So in the first case, we
  48300. 36:00:48have the the and I believe the zero is a
  48301. 36:00:5294 people who don't have diabetes. The
  48302. 36:00:54prediction said that 13 of those people
  48303. 36:00:56did have diabetes and were at high risk.
  48304. 36:00:58And the 32 that had diabetes had
  48305. 36:01:00correct, but our prediction said another
  48306. 36:01:0315 out of that 15, it classified as
  48307. 36:01:07incorrect. So you can see where that
  48308. 36:01:09classification comes in and how that
  48309. 36:01:12works on the confusion matrix. Then
  48310. 36:01:14we're going to go ahead and print the F1
  48311. 36:01:16score. Let me just run that. And you see
  48312. 36:01:19we get a 69 in our F1 score. The F1
  48313. 36:01:24takes into account both sides of the
  48314. 36:01:26balance of false positives where if we
  48315. 36:01:29go ahead and just do the accuracy
  48316. 36:01:31account and that's what most people
  48317. 36:01:33think of is it looks at just how many we
  48318. 36:01:36got right out of how many we got wrong.
  48319. 36:01:38So a lot of people when you're a data
  48320. 36:01:39scientist and you're talking to other
  48321. 36:01:41data scientists they're going to ask you
  48322. 36:01:43what the F1 score the Fore is. If you're
  48323. 36:01:45talking to the general public or the uh
  48324. 36:01:48decision makers in the business, they're
  48325. 36:01:50going to ask what the accuracy is. And
  48326. 36:01:51the accuracy is always better than the
  48327. 36:01:54F1 score. But the F1 score is more
  48328. 36:01:56telling. It lets us know that there's
  48329. 36:01:58more false positives than we would like
  48330. 36:02:00on here. But 82% not too bad for a quick
  48331. 36:02:03flash look at people's different
  48332. 36:02:06statistics and running an sklearn and
  48333. 36:02:08running the KNN, the K nearest neighbor
  48334. 36:02:10on it. So we have created a model using
  48335. 36:02:13KN&N which can predict whether a person
  48336. 36:02:16will have diabetes or not or at the very
  48337. 36:02:18least whether they should go get a
  48338. 36:02:19checkup and have their glucose checked
  48339. 36:02:21regularly or not. The print accuracy
  48340. 36:02:24score we got the 0818 was pretty close
  48341. 36:02:26to what we got and we can pretty much
  48342. 36:02:28round that off and just say we have an
  48343. 36:02:29accuracy of 80%. Tells us it is a pretty
  48344. 36:02:32fair fit in the model.
  48345. 36:02:34>> So what is game means clustering? C
  48346. 36:02:37means clustering is an unsupervised
  48347. 36:02:40learning algorithm. In this case, you
  48348. 36:02:43don't have labeled data unlike in
  48349. 36:02:45supervised learning. So you have a set
  48350. 36:02:47of data and you want to group them and
  48351. 36:02:50as the name suggests, you want to put
  48352. 36:02:52them into clusters which means objects
  48353. 36:02:55that are similar in nature, similar in
  48354. 36:02:57characteristics need to be put together.
  48355. 36:03:00So that's what K means clustering is all
  48356. 36:03:03about. The term K is basically is a
  48357. 36:03:07number. So we need to tell the system
  48358. 36:03:09how many clusters we need to perform. So
  48359. 36:03:11if K is equal to two, there will be two
  48360. 36:03:13clusters. If K is equal to three, three
  48361. 36:03:15clusters and so on and so forth. That's
  48362. 36:03:17what the K stands for. And of course
  48363. 36:03:19there is a way of finding out what is
  48364. 36:03:21the best or optimum value of K for a
  48365. 36:03:24given data. We will look at that. So
  48366. 36:03:26that is K means clustering. So let's
  48367. 36:03:30take an example. C means clustering is
  48368. 36:03:32used in many many scenarios but let's
  48369. 36:03:35take an example of cricket the game of
  48370. 36:03:37cricket let's say you received data of a
  48371. 36:03:40lot of players from maybe all over the
  48372. 36:03:43country or all over the world and this
  48373. 36:03:46data has information about the runs
  48374. 36:03:49scored by the people or by the player
  48375. 36:03:52and the wickets taken by the player and
  48376. 36:03:55based on this information we need to
  48377. 36:03:58cluster this data into two clusters
  48378. 36:04:02batsmen and bowlers. So this is an
  48379. 36:04:04interesting example. Let's see how we
  48380. 36:04:06can perform this. So we have the data
  48381. 36:04:09which consists of primarily two
  48382. 36:04:12characteristics which is the runs and
  48383. 36:04:15the wickets. So the bowlers basically
  48384. 36:04:17take wickets and the batsmen score runs.
  48385. 36:04:20There will be of course a few bowlers
  48386. 36:04:22who can score some runs and similarly
  48387. 36:04:25there will be some batsmen who will who
  48388. 36:04:27would have taken a few wickets. But with
  48389. 36:04:29this information, we want to cluster
  48390. 36:04:31this players into batsmen and bowlers.
  48391. 36:04:34So how does this work? Let's say this is
  48392. 36:04:36how the data is. So there are
  48393. 36:04:39information there is information on the
  48394. 36:04:41y-axis about the run scored and on the
  48395. 36:04:44x-axis about the wickets taken by the
  48396. 36:04:46players. So if we do a quick plot, this
  48397. 36:04:49is how it would look. And um when we do
  48398. 36:04:53the clustering, we need to have the
  48399. 36:04:55clusters like shown in the third diagram
  48400. 36:04:59out here. We need to have a cluster
  48401. 36:05:01which consists of people who have scored
  48402. 36:05:04high runs which is basically the
  48403. 36:05:06batsmen. And then we need a cluster with
  48404. 36:05:08people who have taken a lot of wickets
  48405. 36:05:11which is typically the bowlers. There
  48406. 36:05:13may be a certain amount of overlap but
  48407. 36:05:15we will not talk about it right now. So
  48408. 36:05:18with K means clustering we will have
  48409. 36:05:20here that means K is equal to two and we
  48410. 36:05:23will have two clusters which is batsmen
  48411. 36:05:25and bowlers. So how does this work? The
  48412. 36:05:27way it works is the first step in K
  48413. 36:05:30means clustering is the allocation of
  48414. 36:05:33two centroidids randomly. So two points
  48415. 36:05:37are assigned as so-called centrids. So
  48416. 36:05:41in this case we want two clusters which
  48417. 36:05:44means K is equal to two. So two points
  48418. 36:05:47have been randomly assigned as centrids.
  48419. 36:05:51Keep in mind these points can be
  48420. 36:05:54anywhere. There are random points. They
  48421. 36:05:56are not initially they are not really
  48422. 36:05:59the centroidids. Centr means it's a
  48423. 36:06:02central point of a given data set. But
  48424. 36:06:04in this case when it starts off it's not
  48425. 36:06:07really the centroid. Okay. So these
  48426. 36:06:09points though in our presentation here
  48427. 36:06:12we have shown them one point closer to
  48428. 36:06:14these data points and another closer to
  48429. 36:06:16these data points. They can be assigned
  48430. 36:06:18randomly anywhere. Okay. So that's the
  48431. 36:06:21first step. The next step is to
  48432. 36:06:23determine the distance of each of the
  48433. 36:06:26data points from each of the randomly
  48434. 36:06:30assigned centrids. So for example we
  48435. 36:06:32take this point and find the distance
  48436. 36:06:35from this centr and the distance from
  48437. 36:06:38this cent. This point is taken and the
  48438. 36:06:40distance is found from this centroid and
  48439. 36:06:42this c and so on and so forth. So for
  48440. 36:06:45every point the distance is measured
  48441. 36:06:48from both the centroids and then
  48442. 36:06:51whichever distance is less that point is
  48443. 36:06:54assigned to that centroid. So for
  48444. 36:06:56example in this case visually it is very
  48445. 36:06:59obvious that all these data points are
  48446. 36:07:01assigned to this centroid and all these
  48447. 36:07:04data points are assigned to this
  48448. 36:07:05centroid and that's what is represented
  48449. 36:07:07here in blue color and in this yellow
  48450. 36:07:10color. The next step is to actually
  48451. 36:07:12determine the central point or the
  48452. 36:07:15actual centrid for these two clusters.
  48453. 36:07:18So we have this one initial cluster,
  48454. 36:07:21this one initial cluster. But as you can
  48455. 36:07:23see these points are not really the
  48456. 36:07:25centroid. Centroid means it should be
  48457. 36:07:27the central position of this data set.
  48458. 36:07:30Central position of this data set. So
  48459. 36:07:32that is what needs to be determined as
  48460. 36:07:35the next step. So the central point of
  48461. 36:07:38the actual centrid is determined and the
  48462. 36:07:41original randomly allocated centr is
  48463. 36:07:43repositioned to the actual centroid of
  48464. 36:07:46this new clusters and this process is
  48465. 36:07:50actually repeated. Now what might happen
  48466. 36:07:52is some of these points may get
  48467. 36:07:55reallocated. In our example that is not
  48468. 36:07:57happening probably but it may so happen
  48469. 36:07:59that the distance is found between each
  48470. 36:08:02of these data points once again with
  48471. 36:08:04these centroidids. And if there is if it
  48472. 36:08:06is required some points may be
  48473. 36:08:08reallocated. We will see that in a later
  48474. 36:08:10example but for now we will keep it
  48475. 36:08:12simple. So this process is continued
  48476. 36:08:16till the centrid repositioning stops and
  48477. 36:08:20that is our final cluster. So this is
  48478. 36:08:23our so after iteration we come to this
  48479. 36:08:26position this situation where the
  48480. 36:08:28centroid doesn't need any more
  48481. 36:08:30repositioning and that means our
  48482. 36:08:33algorithm has converged convergence has
  48483. 36:08:36occurred and we have the cluster two
  48484. 36:08:38clusters we have the clusters with a
  48485. 36:08:41centroid. So this process is repeated.
  48486. 36:08:45The process of calculating the distance
  48487. 36:08:47and repositioning the centrid is
  48488. 36:08:50repeated till the repositioning stops
  48489. 36:08:53which means that the algorithm has
  48490. 36:08:56converged and we have the final cluster
  48491. 36:09:00with the data points and the
  48492. 36:09:01centroidids. So this is what you're
  48493. 36:09:03going to learn from this session. We
  48494. 36:09:05will talk about the types of clustering.
  48495. 36:09:08What is K means clustering? application
  48496. 36:09:10of K means clustering. C means
  48497. 36:09:12clustering is done using distance
  48498. 36:09:14measure. So we will talk about the
  48499. 36:09:17common distance measures and then we
  48500. 36:09:19will talk about how K means clustering
  48501. 36:09:22works and go into the details of K means
  48502. 36:09:24clustering algorithm and then we will
  48503. 36:09:27end with a demo and a use case for K
  48504. 36:09:29means clustering. So let's begin. First
  48505. 36:09:32of all, what are the types of
  48506. 36:09:33clustering? There are primarily two
  48507. 36:09:35categories of clustering. hierarchical
  48508. 36:09:38clustering and then partitional
  48509. 36:09:40clustering and each of these categories
  48510. 36:09:42are further subdivided into elomerative
  48511. 36:09:45and divisive clustering and K means and
  48512. 36:09:48fuzzy C means clustering. Let's take a
  48513. 36:09:50quick look at what each of these types
  48514. 36:09:52of clustering are. In hierarchical
  48515. 36:09:55clustering, the clusters have a treelike
  48516. 36:09:58structure and hierarchical clustering is
  48517. 36:10:01further divided into elomerative and
  48518. 36:10:04divisive. Elomemerative clustering is a
  48519. 36:10:07bottomup approach. We begin with each
  48520. 36:10:09element as a separate cluster and merge
  48521. 36:10:12them into successively larger clusters.
  48522. 36:10:14So for example, we have A B CDE E F. We
  48523. 36:10:18start by combining B and C form one
  48524. 36:10:20cluster. D and E form one more. Then we
  48525. 36:10:23combine D, E and F one more bigger
  48526. 36:10:25cluster and then add BC to that and then
  48527. 36:10:27finally A to it. Compared to that
  48528. 36:10:30divisive clustering or divisive
  48529. 36:10:32clustering is a top- down approach. We
  48530. 36:10:34begin with the whole set and proceed to
  48531. 36:10:36divide it into successively smaller
  48532. 36:10:39clusters. So we have ABCDE E F. We first
  48533. 36:10:42take that as a single cluster and then
  48534. 36:10:44break it down into A B C D E and F. Then
  48535. 36:10:49we have partitional clustering split
  48536. 36:10:51into two subtypes. K means clustering
  48537. 36:10:54and fuzzy C means. In K means clustering
  48538. 36:10:58the objects are divided into the number
  48539. 36:11:00of clusters mentioned by the number K.
  48540. 36:11:04That's where the K comes from. So if we
  48541. 36:11:06say K is equal to two, the objects are
  48542. 36:11:08divided into two clusters C1 and C2. And
  48543. 36:11:12the way it is done is the features or
  48544. 36:11:14characteristics are compared and all
  48545. 36:11:17objects having similar characteristics
  48546. 36:11:19are clubed together. So that's how K
  48547. 36:11:22means clustering is done. We will see it
  48548. 36:11:24in more detail as we move forward. And
  48549. 36:11:26fuzzy C means is very similar to K means
  48550. 36:11:29in the sense that it clubs objects that
  48551. 36:11:32have similar characteristics together.
  48552. 36:11:34But while in K means clustering two
  48553. 36:11:37objects cannot belong to or any object a
  48554. 36:11:40single object cannot belong to two
  48555. 36:11:41different clusters in C means objects
  48556. 36:11:44can belong to more than one cluster. So
  48557. 36:11:46that is the primary difference between K
  48558. 36:11:49means and fuzzy C means. So what are
  48559. 36:11:51some of the applications of K means
  48560. 36:11:54clustering? C means clustering is used
  48561. 36:11:56in a variety of examples or variety of
  48562. 36:11:59business cases in real life starting
  48563. 36:12:02from academic performance, diagnostic
  48564. 36:12:04systems, search engines and wireless
  48565. 36:12:07sensor networks and many more. So let us
  48566. 36:12:10take a little deeper look at each of
  48567. 36:12:11these examples. Academic performance. So
  48568. 36:12:14based on the scores of the students,
  48569. 36:12:16students are categorized into A, B, C
  48570. 36:12:19and so on. Clustering forms a backbone
  48571. 36:12:21of search engines. When a search is
  48572. 36:12:24performed, the search results need to be
  48573. 36:12:26grouped together. The search engines
  48574. 36:12:28very often use clustering to do this.
  48575. 36:12:31And similarly, in case of wireless
  48576. 36:12:34sensor networks, the clustering
  48577. 36:12:36algorithm plays the role of finding the
  48578. 36:12:38cluster heads which collects all the
  48579. 36:12:41data in its respective cluster. So
  48580. 36:12:44clustering especially K means clustering
  48581. 36:12:46uses distance measure. So let's take a
  48582. 36:12:49look at what is distance measure. So
  48583. 36:12:51while these are the different types of
  48584. 36:12:53clustering in this video we will focus
  48585. 36:12:56on K means clustering. So distance
  48586. 36:12:59measure tells how similar some objects
  48587. 36:13:03are. So the similarity is measured using
  48588. 36:13:07what is known as distance measure and
  48589. 36:13:09what are the various types of distance
  48590. 36:13:11measures. There is ukidian distance.
  48591. 36:13:14There is Manhattan distance. Then we
  48592. 36:13:17have squared ukitian distance measure
  48593. 36:13:19and cosine distance measure. These are
  48594. 36:13:22some of the distance measures supported
  48595. 36:13:24by k means clustering. Let's take a look
  48596. 36:13:26at each of these. What is ukidian
  48597. 36:13:29distance measure? This is nothing but
  48598. 36:13:31the distance between two points. So we
  48599. 36:13:33have learned in high school how to find
  48600. 36:13:35the distance between two points. This is
  48601. 36:13:38a little sophisticated formula for that.
  48602. 36:13:40But we know a simpler one is square
  48603. 36:13:43roo<unk> of y2 - y1 square + x2 - x1
  48604. 36:13:48square. So this is an extension of that
  48605. 36:13:50formula. So that is the ukidian distance
  48606. 36:13:53between two points. What is the squared
  48607. 36:13:56ukidian distance measure? It's nothing
  48608. 36:13:58but the square of the ukidian distance
  48609. 36:14:02as the name suggests. So instead of
  48610. 36:14:04taking the square root, we leave the
  48611. 36:14:08square as it is. And then we have
  48612. 36:14:10Manhattan distance measure. In case of
  48613. 36:14:12Manhattan distance, it is the sum of the
  48614. 36:14:16distances across the x-axis and the
  48615. 36:14:19y-axis. And note that we are taking the
  48616. 36:14:22absolute value so that the negative
  48617. 36:14:24values don't come into play. So that is
  48618. 36:14:25the Manhattan distance measure. Then we
  48619. 36:14:27have cosine distance measure. In this
  48620. 36:14:30case, we take the angle between the two
  48621. 36:14:33vectors formed by joining the points
  48622. 36:14:35from the origin. So that is the cosine
  48623. 36:14:38distance measure. Okay. So that was a
  48624. 36:14:40quick overview about the various
  48625. 36:14:42distance measures that are supported by
  48626. 36:14:45K means. Now let's go and check how
  48627. 36:14:48exactly K means clustering works. Okay.
  48628. 36:14:51So this is how K means clustering works.
  48629. 36:14:55This is like a flowchart of the whole
  48630. 36:14:57process. There is a starting point and
  48631. 36:15:00then we specify the number of clusters
  48632. 36:15:03that we want. Now there are couple of
  48633. 36:15:05ways of doing this. We can do by trial
  48634. 36:15:08and error. So we specify a certain
  48635. 36:15:11number maybe k is equal to 3 or four or
  48636. 36:15:13five to start with and then as we
  48637. 36:15:15progress we keep changing until we get
  48638. 36:15:18the best clusters or there is a
  48639. 36:15:20technique called elbow technique whereby
  48640. 36:15:23we can determine the value of k. What
  48641. 36:15:26should be the best value of k? How many
  48642. 36:15:28clusters should be formed? So once we
  48643. 36:15:30have the value of K we specify that and
  48644. 36:15:33then the system will assign that many
  48645. 36:15:37centrids. So it picks randomly that to
  48646. 36:15:39start with randomly that many points
  48647. 36:15:42that are considered to be the centrids
  48648. 36:15:44of these clusters and then it measures
  48649. 36:15:47the distance of each of the data points
  48650. 36:15:50from these centroidids and assigns those
  48651. 36:15:54points to the corresponding centr from
  48652. 36:15:57which the distance is minimum. So each
  48653. 36:15:59data point will be assigned to the
  48654. 36:16:02centrid which is closest to it and
  48655. 36:16:05thereby we have k number of initial
  48656. 36:16:09clusters. However this is not the final
  48657. 36:16:12clusters. The next step it does is for
  48658. 36:16:16the new groups for the clusters that
  48659. 36:16:19have been formed it calculates the mean
  48660. 36:16:21position thereby calculates the new
  48661. 36:16:25centroid position. the position of the
  48662. 36:16:27centrid moves compared to the randomly
  48663. 36:16:30allocated one. So it's an iterative
  48664. 36:16:31process. Once again the distance of each
  48665. 36:16:34point is measured from this new centroid
  48666. 36:16:36point and if required the data points
  48667. 36:16:40are reallocated to the new centroidids
  48668. 36:16:43and the mean position or the new centrid
  48669. 36:16:46is calculated once again. If the centrid
  48670. 36:16:48moves then the iteration continues which
  48671. 36:16:51means the convergence has not happened.
  48672. 36:16:53The clustering has not converged. So as
  48673. 36:16:56long as there is a movement of the
  48674. 36:16:58centrid this iteration keeps happening.
  48675. 36:17:01But once the centrid stops moving which
  48676. 36:17:04means that the cluster has converged or
  48677. 36:17:07the clustering process has converged
  48678. 36:17:10that will be the end result. So now we
  48679. 36:17:12have the final position of the centroid
  48680. 36:17:15and the data points are allocated
  48681. 36:17:18accordingly to the closest centrid. I
  48682. 36:17:21know it's a little difficult to
  48683. 36:17:22understand from this simple flowchart.
  48684. 36:17:25So let's do a little bit of
  48685. 36:17:27visualization and see if we can explain
  48686. 36:17:30it better. Let's take an example. If we
  48687. 36:17:32have a data set for a grocery shop. So
  48688. 36:17:35let's say we have a data set for a
  48689. 36:17:37grocery shop and now we want to find out
  48690. 36:17:41how many clusters this has to be spread
  48691. 36:17:44across. So how do we find the optimum
  48692. 36:17:47number of clusters? There is a technique
  48693. 36:17:49called the elbow method. So when these
  48694. 36:17:52clusters are formed, there is a
  48695. 36:17:54parameter called within sum of squares.
  48696. 36:17:57And the lower this value is, the better
  48697. 36:18:01the cluster is. That means all these
  48698. 36:18:04points are very close to each other. So
  48699. 36:18:07we use this within sum of squares as a
  48700. 36:18:10measure to find the optimum number of
  48701. 36:18:15clusters that can be formed for a given
  48702. 36:18:17data set. So we create clusters or we
  48703. 36:18:20let the system create clusters of a
  48704. 36:18:23variety of numbers maybe of 10 10
  48705. 36:18:26clusters and for each value of K the
  48706. 36:18:30within SS is measured and the value of K
  48707. 36:18:34which has the least amount of within SS
  48708. 36:18:37or WSS that is taken as the optimum
  48709. 36:18:41value of K. So this is the diagrammatic
  48710. 36:18:44representation. So we have on the y-axis
  48711. 36:18:47the within sum of squares or wss and on
  48712. 36:18:51the x-axis we have the number of
  48713. 36:18:53clusters. So as you can imagine if you
  48714. 36:18:56have k is equal to one which means all
  48715. 36:18:58the data points are in a single cluster
  48716. 36:19:00the within ss value will be very high
  48717. 36:19:03because they are probably scattered all
  48718. 36:19:05over. The moment you split it into two
  48719. 36:19:08there will be a drastic fall in the
  48720. 36:19:10within ss value and that's what is
  48721. 36:19:13represented here. But then as the value
  48722. 36:19:15of K increases the decrease the rate of
  48723. 36:19:19decrease will not be so high. It will
  48724. 36:19:22continue to decrease but probably the
  48725. 36:19:24rate of decrease will not be high. So
  48726. 36:19:26that gives us an idea. So from here we
  48727. 36:19:29get an idea for example the optimum
  48728. 36:19:32value of K should be either two or three
  48729. 36:19:35or at the most four but beyond that
  48730. 36:19:38increasing the number of clusters is not
  48731. 36:19:41dramatically changing the value in WSS
  48732. 36:19:45because that pretty much gets
  48733. 36:19:46stabilized. Okay. Now that we have got
  48734. 36:19:49the value of K and let's assume that
  48735. 36:19:52these are our delivery points. The next
  48736. 36:19:54step is basically to assign two centrids
  48737. 36:19:59randomly. So let's say C1 and C2 are the
  48738. 36:20:02centrids assigned randomly. Now the
  48739. 36:20:04distance of each location from the
  48740. 36:20:07centrid is measured and each point is
  48741. 36:20:10assigned to the centrid which is closest
  48742. 36:20:14to it. So for example these points are
  48743. 36:20:17very obvious that these are closest to
  48744. 36:20:19C1 whereas this point is far away from
  48745. 36:20:22C2. So these points will be assigned
  48746. 36:20:26which are close to C1 will be assigned
  48747. 36:20:27to C1 and these points or locations
  48748. 36:20:30which are close to C2 will be assigned
  48749. 36:20:32to C2. And then so this is the how the
  48750. 36:20:35initial grouping is done. This is part
  48751. 36:20:38of C1 and this is part of C2. Then the
  48752. 36:20:41next step is to calculate the actual
  48753. 36:20:44centrid of this data because remember C1
  48754. 36:20:47and C2 are not the centrids. They've
  48755. 36:20:49been randomly assigned points and only
  48756. 36:20:53thing that has been done was the data
  48757. 36:20:55points which are closest to them have
  48758. 36:20:57been assigned to them. But now in this
  48759. 36:21:00step the actual centroid will be
  48760. 36:21:02calculated which may be for each of
  48761. 36:21:04these data sets somewhere in the middle.
  48762. 36:21:06So that's like the main point that will
  48763. 36:21:08be calculated and the centr will
  48764. 36:21:10actually be positioned or repositioned
  48765. 36:21:13there. Same with C2. So the new centroid
  48766. 36:21:17for this group is C2. this new position
  48767. 36:21:19and C1 is in this new position. Once
  48768. 36:21:22again, the distance of each of the data
  48769. 36:21:24points is calculated from these
  48770. 36:21:26centroids. Now remember, it's not
  48771. 36:21:28necessary that the distance still
  48772. 36:21:30remains the or each of these data points
  48773. 36:21:32still remain in the same group. By
  48774. 36:21:34recalculating the distance, it may be
  48775. 36:21:36possible that some points get
  48776. 36:21:38reallocated like so. You see this? So
  48777. 36:21:41this point earlier was closer to C2
  48778. 36:21:45because C2 was here. But after
  48779. 36:21:48recalculating repositioning it is
  48780. 36:21:51observed that this is closer to C1 than
  48781. 36:21:55C2. So this is the new grouping. So some
  48782. 36:21:58points will be reassigned. And again the
  48783. 36:22:01centrid will be calculated and if the
  48784. 36:22:04centroid doesn't change so that is a
  48785. 36:22:06repetative process, iterative process.
  48786. 36:22:08And if the centroid doesn't change once
  48787. 36:22:10the centroid stops changing that means
  48788. 36:22:13the algorithm has converged and this is
  48789. 36:22:16our final cluster with this as the
  48790. 36:22:18centroid C1 and C2 as the centroids
  48791. 36:22:21these data points as a part of each
  48792. 36:22:23cluster. So I hope this helps in
  48793. 36:22:25understanding the whole process
  48794. 36:22:27iterative process of K means clustering.
  48795. 36:22:30So let's take a look at the K means
  48796. 36:22:33clustering algorithm. Let's say we have
  48797. 36:22:36x1, x2, x3, n number of points as our
  48798. 36:22:40inputs and we want to split this into k
  48799. 36:22:44clusters or we want to create k
  48800. 36:22:46clusters. So the first step is to
  48801. 36:22:48randomly pick k points and call them
  48802. 36:22:52centroidids. They are not real centrids
  48803. 36:22:55because centr is supposed to be a center
  48804. 36:22:57point but they are just called centrids.
  48805. 36:23:00And we calculate the distance of each
  48806. 36:23:03and every input point from each of the
  48807. 36:23:07centroidids. So the distance of X1 from
  48808. 36:23:12C1 from C2 C3 each of the distances we
  48809. 36:23:16calculate and then find out which
  48810. 36:23:19distance is the lowest and assign X1 to
  48811. 36:23:22that particular random centroid. Repeat
  48812. 36:23:24that process for X2. calculate its
  48813. 36:23:28distance from each of the centroid C1,
  48814. 36:23:31C2, C3 up to CK and find which is the
  48815. 36:23:34lowest distance and assign X2 to that
  48816. 36:23:36particular centroid. Same with X3 and so
  48817. 36:23:38on. So that is the first round of
  48818. 36:23:41assignment that is done. Now we have K
  48819. 36:23:45groups because there are we have
  48820. 36:23:47assigned the value of K. So there are K
  48821. 36:23:50centroids and uh so there are K groups.
  48822. 36:23:54All these inputs have been split into K
  48823. 36:23:56groups. However, remember we picked the
  48824. 36:23:59centrids randomly. So they are not real
  48825. 36:24:02centrids. So now what we have to do, we
  48826. 36:24:05have to calculate the actual centroids
  48827. 36:24:08for each of these groups which is like
  48828. 36:24:10the mean position which means that the
  48829. 36:24:13position of the randomly selected
  48830. 36:24:15centrids will now change and they will
  48831. 36:24:18be the main positions of these newly
  48832. 36:24:21formed K groups. And once that is done,
  48833. 36:24:25we once again repeat this process of
  48834. 36:24:28calculating the distance. Right? So this
  48835. 36:24:31is what we are doing as a part of step
  48836. 36:24:33four. We repeat step two and three. So
  48837. 36:24:36we again calculate the distance of X1
  48838. 36:24:40from the centroid C1, C2, C3 and then
  48839. 36:24:44see which is the lowest value and assign
  48840. 36:24:47X1 to that. Calculate the distance of X2
  48841. 36:24:50from C1, C2, C3 or whatever up to CK and
  48842. 36:24:54find whichever is the lowest distance
  48843. 36:24:56and assign X2 to that centroid and so
  48844. 36:24:59on. In this process there may be some
  48845. 36:25:01reassignment. X1 was probably assigned
  48846. 36:25:04to cluster C2 and after doing this
  48847. 36:25:07calculation maybe now X1 is assigned to
  48848. 36:25:09C1. So that kind of reallocation may
  48849. 36:25:12happen. So we repeat the steps two and
  48850. 36:25:14three till the position of the centrids
  48851. 36:25:18don't change or stop changing and that's
  48852. 36:25:21when we have convergence. So let's take
  48853. 36:25:24a detailed look at at each of these
  48854. 36:25:26steps. So we randomly pick K cluster
  48855. 36:25:29centers. We call them centroidids
  48856. 36:25:31because they are not initially they are
  48857. 36:25:34not really the centrids. So we let us
  48858. 36:25:36name them C1 C2 up to CK. And then step
  48859. 36:25:39two, we assign each data point to the
  48860. 36:25:42closest center. So what we do, we
  48861. 36:25:45calculate the distance of each X value
  48862. 36:25:48from each C value. So the distance
  48863. 36:25:51between X1 C1 distance between X1 C2 X1
  48864. 36:25:57C3 and then we find which is the lowest
  48865. 36:26:00value. Right? That's the minimum value
  48866. 36:26:02we find and assign X1 to that particular
  48867. 36:26:06centroid. Then we go next to x2. Find
  48868. 36:26:09the distance of x2 from c1, x2 from c2,
  48869. 36:26:13x2 from c3 and so on up to ck. And then
  48870. 36:26:16assign it to the point or to the
  48871. 36:26:19centroid which has the lowest value and
  48872. 36:26:21so on. So that is step number two. In
  48873. 36:26:24step number three, we now find the
  48874. 36:26:27actual centr for each group. So what has
  48875. 36:26:30happened as a part of step number two?
  48876. 36:26:33We now have all the points, all the data
  48877. 36:26:36points grouped into K groups because we
  48878. 36:26:40we wanted to create K clusters, right?
  48879. 36:26:42So we have K groups. Each one may be
  48880. 36:26:45having a certain number of input values.
  48881. 36:26:47They need not be equally distributed. By
  48882. 36:26:50the way, based on the distance, we will
  48883. 36:26:52have K groups. But remember the initial
  48884. 36:26:56values of the C1 C2 were not really the
  48885. 36:26:59centrids of these groups, right? we
  48886. 36:27:01assign them randomly. So now in step
  48887. 36:27:04three, we actually calculate the centr
  48888. 36:27:08of each group which means the original
  48889. 36:27:11point which we thought was the centrid
  48890. 36:27:13will shift to the new position which is
  48891. 36:27:16the actual centrid for each of these
  48892. 36:27:18groups. Okay? And we again calculate the
  48893. 36:27:22distance. So we go back to step two
  48894. 36:27:25which is what we calculate again the
  48895. 36:27:27distance of each of these points from
  48896. 36:27:30the newly positioned centroidids and if
  48897. 36:27:34required we reassign these points to the
  48898. 36:27:38new centroidids. So as I said earlier
  48899. 36:27:41there may be a reallocation. So we now
  48900. 36:27:43have a new set or a new group. We still
  48901. 36:27:47have K groups but the number of items
  48902. 36:27:50and the actual assignment may be
  48903. 36:27:52different from what was in step two
  48904. 36:27:56here. Okay, so that might change. Then
  48905. 36:27:59we perform step three once again to find
  48906. 36:28:02the new centroid of this new group. So
  48907. 36:28:05we have again a new set of clusters, new
  48908. 36:28:09centroidids and new assignments. We
  48909. 36:28:12repeat this step two again. Once again
  48910. 36:28:14we find and then it is possible that
  48911. 36:28:17after iterating through three or four or
  48912. 36:28:20five times the centrid will stop moving
  48913. 36:28:24in the sense that when you calculate the
  48914. 36:28:26new value of the centrid that will be
  48915. 36:28:29same as the original value or there will
  48916. 36:28:31be very marginal change. So that is when
  48917. 36:28:34we say convergence has occurred and that
  48918. 36:28:37is our final cluster. That's the
  48919. 36:28:41formation of the final cluster. All
  48920. 36:28:43right. So let's see a couple of demos of
  48921. 36:28:47uh K means clustering. We will actually
  48922. 36:28:50see some live demos in uh Python
  48923. 36:28:52notebook using Python notebook. But
  48924. 36:28:54before that let's find out what's the
  48925. 36:28:57problem that we are trying to solve. The
  48926. 36:28:59problem statement is let's say Walmart
  48927. 36:29:01wants to open a chain of stores across
  48928. 36:29:04the state of Florida and uh it wants to
  48929. 36:29:08find the optimal store locations. Now
  48930. 36:29:11the issue here is if they open too many
  48931. 36:29:14stores close to each other obviously the
  48932. 36:29:16they will not make profit but if they if
  48933. 36:29:20the stores are too far apart then they
  48934. 36:29:22will not have enough sales. So how do
  48935. 36:29:24they optimize this? Now for an
  48936. 36:29:27organization like Walmart which is an
  48937. 36:29:30e-commerce giant they already have the
  48938. 36:29:33addresses of their customers in their
  48939. 36:29:36database. So they can actually use this
  48940. 36:29:39information or this data and use K means
  48941. 36:29:42clustering to find the optimal location.
  48942. 36:29:46Now before we go into the Python
  48943. 36:29:49notebook and show you the live code, I
  48944. 36:29:52wanted to take you through very quickly
  48945. 36:29:54a summary of the code in the slides and
  48946. 36:29:57then we will go into the Python
  48947. 36:29:59notebook. So in this block we are
  48948. 36:30:02basically importing all the required
  48949. 36:30:05libraries like numpy, mattplot lib and
  48950. 36:30:09so on and we are loading the data that
  48951. 36:30:13is available in the form of let's say
  48952. 36:30:15the addresses for simplicity sake we
  48953. 36:30:18will just take them as some data points.
  48954. 36:30:21Then the next thing we do is quickly do
  48955. 36:30:23a scatter plot to see how they are
  48956. 36:30:27related to each other with respect to
  48957. 36:30:29each other. So in the scatter plot we
  48958. 36:30:31see that there are a few distinct groups
  48959. 36:30:35already being formed. So you can
  48960. 36:30:37actually get an idea about how the
  48961. 36:30:39cluster would look and how many clusters
  48962. 36:30:41what is the optimal number of clusters
  48963. 36:30:44and then starts the actual K means
  48964. 36:30:46clustering process. So we will assign
  48965. 36:30:50each of these points to the centrids and
  48966. 36:30:53then check whether they are the optimal
  48967. 36:30:56distance which is the shortest distance
  48968. 36:30:58and assign each of the points data
  48969. 36:31:01points to the centroidids and then go
  48970. 36:31:05through this iterative process till the
  48971. 36:31:07whole process converges and finally we
  48972. 36:31:10get an output like this. So we have four
  48973. 36:31:13distinct clusters and um which is we can
  48974. 36:31:17say that this is how the population is
  48975. 36:31:20probably distributed across Florida
  48976. 36:31:22state and uh these centroidids are like
  48977. 36:31:26the location where the store should be
  48978. 36:31:29the optimum location where the store
  48979. 36:31:31should be. So that's the way we
  48980. 36:31:34determine the best locations for the
  48981. 36:31:37store and that's how we can help Walmart
  48982. 36:31:40find the best locations for their stores
  48983. 36:31:42in Florida. So now let's take this into
  48984. 36:31:46Python notebook. Let's see how this
  48985. 36:31:48looks when we are learning running the
  48986. 36:31:50code live. All right. So this is the
  48987. 36:31:52code for K means clustering in Jupyter
  48988. 36:31:56notebook. We have a few examples here
  48989. 36:31:59which we will demonstrate how K means
  48990. 36:32:02clustering is used and even there is a
  48991. 36:32:04small implementation of K means
  48992. 36:32:06clustering as well. Okay. So let's get
  48993. 36:32:08started. Okay. So this block is
  48994. 36:32:11basically importing the various
  48995. 36:32:13libraries that are required like
  48996. 36:32:14mattplot lib and numpy and so on and so
  48997. 36:32:17forth which would be used as a part of
  48998. 36:32:20the code. Then we are going and creating
  48999. 36:32:23blobs which are similar to clusters. Now
  49000. 36:32:26this is a very neat feature which is
  49001. 36:32:28available in scikitlearn. Make blobs is
  49002. 36:32:31a nice feature which creates clusters of
  49003. 36:32:34data sets. So that's a wonderful
  49004. 36:32:37functionality that is readily available
  49005. 36:32:38for us to create some test data kind of
  49006. 36:32:41thing. Okay. So that's exactly what we
  49007. 36:32:44are doing here. We are using make blobs
  49008. 36:32:46and we can specify how many clusters we
  49009. 36:32:50want. So centers we are mentioning here.
  49010. 36:32:53So it will go ahead and so we just
  49011. 36:32:54mentioned four. So it will go ahead and
  49012. 36:32:56create some test data for us. And this
  49013. 36:33:00is how it looks. As you can see visually
  49014. 36:33:02also we can figure out that there are
  49015. 36:33:05four distinct classes or clusters in
  49016. 36:33:08this data set. And that is what make
  49017. 36:33:12blobs actually provides. Now from here
  49018. 36:33:16onwards we will basically run the
  49019. 36:33:18standard K means functionality that is
  49020. 36:33:21readily available. So we really don't
  49021. 36:33:23have to implement K means itself. The C
  49022. 36:33:26means functionality or the the function
  49023. 36:33:28is readily available. You just need to
  49024. 36:33:31feed the data and we'll create the
  49025. 36:33:33clusters. So this is the code for that.
  49026. 36:33:36We import k means and then we create an
  49027. 36:33:39instance of k means and we specify the
  49028. 36:33:42value of k. This n_clusters is the value
  49029. 36:33:46of k. Remember K means in K means K is
  49030. 36:33:49basically the number of clusters that
  49031. 36:33:51you want to create and it is a integer
  49032. 36:33:54value. So this is where we are
  49033. 36:33:55specifying that. So we have K is equal
  49034. 36:33:57to four and so that instance is created.
  49035. 36:34:01We take that instance and as with any
  49036. 36:34:04other machine learning functionality fit
  49037. 36:34:06is what we use the function or the
  49038. 36:34:09method rather fit is what we use to
  49039. 36:34:12train the model. Here there is no real
  49040. 36:34:14training uh kind of thing but that's the
  49041. 36:34:17call. Okay. So we are calling fit and
  49042. 36:34:19what we are doing here we are just
  49043. 36:34:21passing the data. So x has these values
  49044. 36:34:23the data that has been created right. So
  49045. 36:34:26that is what we are passing here and uh
  49046. 36:34:30this will go ahead and create the
  49047. 36:34:33clusters and uh then we are using
  49048. 36:34:38after doing uh fit we run the predict
  49049. 36:34:42which basically assigns for each of
  49050. 36:34:44these observations which cluster it
  49051. 36:34:47belongs to. All right. So it will name
  49052. 36:34:50the clusters. Maybe this is cluster one.
  49053. 36:34:52This is two, three and so on. Or will
  49054. 36:34:55actually start from zero, cluster 0, 1,
  49055. 36:34:582 and 3 maybe. And then for each of the
  49056. 36:35:02observations it will assign based on
  49057. 36:35:04which cluster it belongs to it will
  49058. 36:35:06assign a value. So that is stored in y_k
  49059. 36:35:10means when we call predict that is what
  49060. 36:35:12it does. And we can take a quick look at
  49061. 36:35:15these uh y_k means or the cluster
  49062. 36:35:19numbers that have been assigned for each
  49063. 36:35:21observation. So this is the cluster
  49064. 36:35:23number assigned for observation one.
  49065. 36:35:25Maybe this is for observation two,
  49066. 36:35:27observation three and so on. So we have
  49067. 36:35:30how many about I think 300 samples
  49068. 36:35:32right? So all the 300 samples there are
  49069. 36:35:34300 values here. Each of them the
  49070. 36:35:37cluster number is given and the cluster
  49071. 36:35:39number goes from 0 to three. So there
  49072. 36:35:42are four clusters. So the numbers go
  49073. 36:35:44from 0 1 2 3. So that's what is seen
  49074. 36:35:48here. Okay. Now, so this was a quick
  49075. 36:35:50example of generating some dummy data
  49076. 36:35:53and then clustering that. Okay. And this
  49077. 36:35:56can be applied if you have proper data.
  49078. 36:35:57You can just load it up into X for
  49079. 36:36:00example here and then run the K. So this
  49080. 36:36:03is the central part of the K means
  49081. 36:36:05clustering program example. So you
  49082. 36:36:07basically create an instance and you
  49083. 36:36:10mention how many clusters you want by
  49084. 36:36:12specifying this parameter n_clusters and
  49085. 36:36:15that is also the value of k and then
  49086. 36:36:17pass the data to get the values. Now the
  49087. 36:36:20next section of this code is the
  49088. 36:36:23implementation of a k means. Now this is
  49089. 36:36:26kind of a a rough implementation of the
  49090. 36:36:29k means algorithm. So we will just walk
  49091. 36:36:32you through I will walk you through the
  49092. 36:36:33code uh at each step what it is doing
  49093. 36:36:36and then we will see a couple of more
  49094. 36:36:38examples of how K means clustering can
  49095. 36:36:41be used in maybe some real life examples
  49096. 36:36:44real life use cases. All right. So in
  49097. 36:36:46this case here what we're doing is
  49098. 36:36:48basically implementing K means
  49099. 36:36:51clustering and there is a function or a
  49100. 36:36:55library calculates for a given two pairs
  49101. 36:36:58of points it will calculate the the
  49102. 36:37:01distance between them and see which one
  49103. 36:37:03is the closest and so on. So this is
  49104. 36:37:05like this is pretty much like what K
  49105. 36:37:07means does right. So it calculates the
  49106. 36:37:09distance of each point or each data set
  49107. 36:37:12from predefined centroid and then based
  49108. 36:37:15on whichever is the lowest this
  49109. 36:37:17particular data point is assigned to
  49110. 36:37:19that centroid. So that is basically
  49111. 36:37:21available as a standard function and we
  49112. 36:37:23will be using that here. So as explained
  49113. 36:37:26in the slides the first step that is
  49114. 36:37:28done in case of C means clustering is to
  49115. 36:37:32randomly assign some centrides. So as a
  49116. 36:37:37first step we randomly allocate a couple
  49117. 36:37:40of centrids which we call here we're
  49118. 36:37:42calling as centers
  49119. 36:37:44and then we put this in a loop and we
  49120. 36:37:47take it through an iterative process.
  49121. 36:37:50For each of the data points, we first
  49122. 36:37:53find out using this function pair-wise
  49123. 36:37:55distance argument. For each of the
  49124. 36:37:57points, we find out which one which
  49125. 36:38:00center or which uh randomly selected
  49126. 36:38:02centrid is the closest and accordingly
  49127. 36:38:06we assign that data or the data point to
  49128. 36:38:10that particular centrid or cluster. And
  49129. 36:38:13once that is done for all the data
  49130. 36:38:16points, we calculate the new centr by
  49131. 36:38:20finding out the mean position with the
  49132. 36:38:22the center position. Right? So we
  49133. 36:38:24calculate the new centroid and then we
  49134. 36:38:26check if the new centroid is the
  49135. 36:38:29coordinates or the position is the same
  49136. 36:38:32as the previous centroid. The positions
  49137. 36:38:35we will compare and if it is the same
  49138. 36:38:39that means the process has converged. So
  49139. 36:38:42remember we do this process till the
  49140. 36:38:44centroidids or the centrid doesn't move
  49141. 36:38:46anymore right so the centroid gets
  49142. 36:38:48relocated each time this reallocation is
  49143. 36:38:52done so the moment it doesn't change
  49144. 36:38:55anymore the position of the cent doesn't
  49145. 36:38:57change anymore we know that convergence
  49146. 36:38:59has occurred so till then so you see
  49147. 36:39:01here this is like an infinite loop while
  49148. 36:39:03true is an infinite loop it only breaks
  49149. 36:39:06when the centers are the same the new
  49150. 36:39:09center and old center positions are the
  49151. 36:39:11name and once that is uh done uh we
  49152. 36:39:14return the centers and the labels. Now
  49153. 36:39:16of course as explained this is not a
  49154. 36:39:18very sophisticated and advanced
  49155. 36:39:20implementation very basic implementation
  49156. 36:39:22because one of the flaws in this is that
  49157. 36:39:24sometimes what happens is the centroid
  49158. 36:39:27the position will keep moving but in the
  49159. 36:39:31change will be very minor. So in that
  49160. 36:39:33case also that is actually convergence
  49161. 36:39:36right. So for example the change is 0.1
  49162. 36:39:39we can consider that as convergence
  49163. 36:39:41otherwise what will happen is this will
  49164. 36:39:43either take forever or it will be never
  49165. 36:39:46ending. So that's a small flaw here. So
  49166. 36:39:48that is something additional checks may
  49167. 36:39:50have to be added here. But again as
  49168. 36:39:52mentioned this is not the most
  49169. 36:39:54sophisticated uh implementation. This is
  49170. 36:39:57like a kind of a rough implementation of
  49171. 36:39:59the k means clustering. Okay. So if we
  49172. 36:40:02execute this code this is what we get as
  49173. 36:40:05the output. So this is the definition of
  49174. 36:40:07this particular function and then we
  49175. 36:40:10call that find_clusters and we pass our
  49176. 36:40:13data x and the number of clusters which
  49177. 36:40:15is four and if we run that and plot it
  49178. 36:40:19this is the output that we get. So this
  49179. 36:40:21is of course each cluster is represented
  49180. 36:40:23by a different color. So we have a
  49181. 36:40:25cluster in green color, yellow color and
  49182. 36:40:27so on and so forth. And these big points
  49183. 36:40:30here these are the centroidids is the
  49184. 36:40:32final position of the centroidids. And
  49185. 36:40:33as you can see visually also this
  49186. 36:40:35appears like a kind of a center of all
  49187. 36:40:38these points here. Right? Similarly this
  49188. 36:40:41is like the center of all these points
  49189. 36:40:43here and so on. So this is the example
  49190. 36:40:46or this is an example of a
  49191. 36:40:47implementation of K means clustering and
  49192. 36:40:52uh next we will move on to see a couple
  49193. 36:40:54of examples of how K means clustering is
  49194. 36:40:58used in maybe some real life scenarios
  49195. 36:41:01or use cases. In the next example or
  49196. 36:41:04demo, we are going to see how we can use
  49197. 36:41:07K means clustering to perform color
  49198. 36:41:10compression. We will take a couple of
  49199. 36:41:12images. So there will be two examples
  49200. 36:41:15and uh we will try to use C means
  49201. 36:41:18clustering to compress the colors. This
  49202. 36:41:21is a common situation in image
  49203. 36:41:23processing when you have an image with
  49204. 36:41:26millions of uh colors but then you
  49205. 36:41:29cannot render it on some devices which
  49206. 36:41:32may not have enough memory. Uh so that
  49207. 36:41:34is the scenario where where something
  49208. 36:41:36like this can be used. So before again
  49209. 36:41:40we go into the Python notebook let's
  49210. 36:41:43take a look at quickly the the code. As
  49211. 36:41:46usual we import the libraries and then
  49212. 36:41:49we import the image and uh then we will
  49213. 36:41:53flatten it. So the reshaping is
  49214. 36:41:55basically we have the image information
  49215. 36:41:58is stored in the form of pixels and uh
  49216. 36:42:02if the image is like for example 427x
  49217. 36:42:05640 and it has three colors. So that's
  49218. 36:42:08the overall dimension of the of the
  49219. 36:42:12initial image. we just reshape it and um
  49220. 36:42:16then feed this to our algorithm and this
  49221. 36:42:20will then create clusters of only 16
  49222. 36:42:23clusters. So this this colors there are
  49223. 36:42:26millions of colors and now we need to
  49224. 36:42:29bring it down to 16 colors. So we use k
  49225. 36:42:32is equal to 16 and u this is how when we
  49226. 36:42:36visualize this is how it looks. There
  49227. 36:42:37are these are all about 16 million
  49228. 36:42:40possible colors. The input color space
  49229. 36:42:42has 16 million possible colors and we
  49230. 36:42:46just sub compress it to 16 colors. So
  49231. 36:42:49this is how it would look when we
  49232. 36:42:51compress it to 16 colors. And this is
  49233. 36:42:54how the original image looks. And after
  49234. 36:42:57compression to 16 colors, this is how
  49235. 36:43:00the new image looks. As you can see,
  49236. 36:43:02there is not a lot of information that
  49237. 36:43:06has been lost. though the image quality
  49238. 36:43:10is definitely reduced a little bit. So
  49239. 36:43:14this is an example which we are going to
  49240. 36:43:16now see in Python notebook. Let's go
  49241. 36:43:19into the Python and once again as always
  49242. 36:43:22we will import some libraries and load
  49243. 36:43:25this image called flower.jpg.
  49244. 36:43:29Okay. So let we'll load that and this is
  49245. 36:43:31how it looks. This is the original image
  49246. 36:43:34which has I think 16 million colors and
  49247. 36:43:38uh this is the shape of this image which
  49248. 36:43:40is basically what is the shape is
  49249. 36:43:42nothing but the overall size right so
  49250. 36:43:45this is 427 pixel by 640 pixel and then
  49251. 36:43:49there are three layers which is this
  49252. 36:43:51three basically is for RGB which is red
  49253. 36:43:53green blue so color image will have that
  49254. 36:43:56right so that is the shape of this now
  49255. 36:43:58what we need to do is data let's take a
  49256. 36:44:01look at how data is looking. So let me
  49257. 36:44:03just create a new cell and show you what
  49258. 36:44:07is in data. Basically we have captured
  49259. 36:44:09this information.
  49260. 36:44:12So data is what? Let me just show you
  49261. 36:44:14here.
  49262. 36:44:18All right. So let's take a look at
  49263. 36:44:21China. What are the values in China? And
  49264. 36:44:25uh if you see here, this is how the data
  49265. 36:44:27is stored. This is nothing but the pixel
  49266. 36:44:29values. Okay? So this is like a matrix
  49267. 36:44:32and each one has about for for this 427x
  49268. 36:44:36640 pixels. All right. So this is how it
  49269. 36:44:39looks. Now the issue here is these
  49270. 36:44:40values are large. The numbers are large.
  49271. 36:44:44So we need to normalize them to between
  49272. 36:44:460 and one. Right? So that's why we will
  49273. 36:44:49basically create one more variable which
  49274. 36:44:52is data which will contain the values
  49275. 36:44:54between 0 and one. And the way to do
  49276. 36:44:56that is divide by 255. So we divide
  49277. 36:44:59China by 255 and we get the new values
  49278. 36:45:02in data. So let's just run this uh piece
  49279. 36:45:05of code and this is the shape. So we now
  49280. 36:45:08have also yeah what we have done is we
  49281. 36:45:10changed using reshape we converted into
  49282. 36:45:13the three-dimensional into a
  49283. 36:45:15two-dimensional data set. And let us
  49284. 36:45:18also take a look at how
  49285. 36:45:21let me just insert
  49286. 36:45:24probably a cell here and take a look at
  49287. 36:45:27how data is looking. All right. So this
  49288. 36:45:29is how data is looking and now you see
  49289. 36:45:31this is the values are between 0 and
  49290. 36:45:34one. Right? So if you earlier noticed in
  49291. 36:45:36case of China the values were large
  49292. 36:45:38numbers. Now everything is between 0 and
  49293. 36:45:40one. This is one of the things we need
  49294. 36:45:42to do. All right. So after that the next
  49295. 36:45:45thing that we need to do is to visualize
  49296. 36:45:47this and uh we can take random set of
  49297. 36:45:51maybe 10,000 points and plot it and
  49298. 36:45:54check and see how this looks. So let us
  49299. 36:45:57just plot this and uh so this is how the
  49300. 36:46:00original the color the pixel
  49301. 36:46:02distribution is. These are two plots one
  49302. 36:46:05is red against green and another is red
  49303. 36:46:07against blue and this is the original
  49304. 36:46:09distribution of the color. So then what
  49305. 36:46:11we will do is we will use K means
  49306. 36:46:13clustering to create just 16 clusters
  49307. 36:46:17for the various colors and then apply
  49308. 36:46:20that to the image. Now what will happen
  49309. 36:46:23is since the data is large because there
  49310. 36:46:25are millions of colors using regular K
  49311. 36:46:27means may be a little time consuming. So
  49312. 36:46:30there is another version of K means
  49313. 36:46:32which is called mini batch K means. So
  49314. 36:46:34we will use that which is which
  49315. 36:46:36processes in the overall concept remains
  49316. 36:46:39the same but this basically processes it
  49317. 36:46:41in smaller batches. That's the only
  49318. 36:46:43thing. Okay. So the results will pretty
  49319. 36:46:45much be the same. So let's go ahead and
  49320. 36:46:48execute this piece of code and also
  49321. 36:46:51visualize this so that we can see that
  49322. 36:46:53there are the this is how the 16 colors
  49323. 36:46:56uh would look. So this is red against
  49324. 36:46:58green and this is red against blue.
  49325. 36:47:01there is uh quite a bit of similarity
  49326. 36:47:03between this original color schema and
  49327. 36:47:06the new one. Right? So it doesn't look
  49328. 36:47:08very very completely different or
  49329. 36:47:10anything like that. Now we apply this
  49330. 36:47:12the newly created colors to the image
  49331. 36:47:16and uh we can take a look uh how this is
  49332. 36:47:18uh looking. Now we can compare both the
  49333. 36:47:20images. So this is our original image
  49334. 36:47:23and this is our new image. So as you can
  49335. 36:47:25see there is not a lot of information
  49336. 36:47:28that has been lost. uh it pretty much
  49337. 36:47:31looks like the original image. Yes, we
  49338. 36:47:34can see that for example here there is a
  49339. 36:47:36little bit uh it appears a little
  49340. 36:47:39dullish compared to this one right
  49341. 36:47:41because uh we kind of took off some of
  49342. 36:47:43the finer details of the color but
  49343. 36:47:47overall the highle information has been
  49344. 36:47:50maintained. At the same time, the main
  49345. 36:47:52advantage is that now this can be this
  49346. 36:47:54is an image which can be rendered on a
  49347. 36:47:57device which may not be that very
  49348. 36:47:59sophisticated. Now let's take one more
  49349. 36:48:01example with a different image. In the
  49350. 36:48:04second example, we will take an image of
  49351. 36:48:06the summer palace in China and we repeat
  49352. 36:48:10the same process. This is a high
  49353. 36:48:12definition color image with millions of
  49354. 36:48:15colors and also uh three-dimensional.
  49355. 36:48:19Now we will reduce that to 16 colors
  49356. 36:48:22using K means clustering. And um we do
  49357. 36:48:26the same process like before. We reshape
  49358. 36:48:28it and then we cluster the colors to 16
  49359. 36:48:32and then we render the image once again.
  49360. 36:48:36And we will see that the color the
  49361. 36:48:38quality of the image is slightly
  49362. 36:48:41deteriorates. As you can see here, this
  49363. 36:48:43has much finer details in this which are
  49364. 36:48:46probably missing here. But then that's
  49365. 36:48:48the compromise because there are some
  49366. 36:48:50devices which may not be able to handle
  49367. 36:48:53this kind of a high density images. So
  49368. 36:48:56let's run this code in Python notebook.
  49369. 36:49:00All right. So let's apply the same
  49370. 36:49:02technique for another picture which is
  49371. 36:49:04uh even more intricate and has probably
  49372. 36:49:08much complicated color schema. So this
  49373. 36:49:11is the image. Now once again uh we can
  49374. 36:49:14take a look at the shape which is 427x
  49375. 36:49:17640x3
  49376. 36:49:19and this is the new data would look
  49377. 36:49:22somewhat like this compared to the
  49378. 36:49:24flower image. So we have some new values
  49379. 36:49:27here and we will also bring this as you
  49380. 36:49:31can see the numbers are much big. So we
  49381. 36:49:33will much bigger so we will now have to
  49382. 36:49:36scale them down to values between 0 and
  49383. 36:49:39one. And that is done by dividing by
  49384. 36:49:41255. So let's go ahead and uh do that
  49385. 36:49:46and reshape it. Okay. So we get a
  49386. 36:49:49two-dimensional matrix and uh we will
  49387. 36:49:53then as the next step we will go ahead
  49388. 36:49:55and visualize this how it looks the the
  49389. 36:49:5816 colors and this is basically how it
  49390. 36:50:01would look 16 million colors. And now we
  49391. 36:50:05can create the clusters out of this. The
  49392. 36:50:0916 K means clusters we will create. So
  49393. 36:50:13this is how the distribution of the
  49394. 36:50:15pixels would look with 16 colors. And
  49395. 36:50:19then we go ahead and uh apply this and
  49396. 36:50:23visualize how it is looking for with the
  49397. 36:50:26with the new just the 16 color. So once
  49398. 36:50:29again, as you can see, this looks much
  49399. 36:50:32richer in color, but at the same time,
  49400. 36:50:35and this probably doesn't have, as we
  49401. 36:50:38can see, it doesn't look as rich as this
  49402. 36:50:40one, but nevertheless, the information
  49403. 36:50:42is not lost, the shape and all that
  49404. 36:50:44stuff. And this can be also rendered on
  49405. 36:50:47a slightly a device which is probably
  49406. 36:50:51not that sophisticated. Okay, so that's
  49407. 36:50:54pretty much it. So we have seen two
  49408. 36:50:56examples of how color compression can be
  49409. 36:50:59done uh using K means clustering and we
  49410. 36:51:03have also seen in the previous examples
  49411. 36:51:05of how to implement C means the code to
  49412. 36:51:08roughly how to implement C means
  49413. 36:51:10clustering and we use some sample data
  49414. 36:51:13using blob to just execute the C means
  49415. 36:51:17cluster that takes place after data
  49416. 36:51:20collection and before statistical
  49417. 36:51:21analysis. So before you conduct any
  49418. 36:51:24statistical formulas and analysis on the
  49419. 36:51:26data and squeeze the data to extract
  49420. 36:51:30some valuable insights, the process
  49421. 36:51:32which you perform is called as initial
  49422. 36:51:34data analysis. Like taking the data from
  49423. 36:51:36the source, cleaning the data,
  49424. 36:51:38transforming the data into a readable
  49425. 36:51:40format and using that readable data to
  49426. 36:51:44build some basic charts what exactly is
  49427. 36:51:46happening with this particular company,
  49428. 36:51:49brand or anything. Let's say I give you
  49429. 36:51:51some data from the company. Then you get
  49430. 36:51:53some insights of it. How many number of
  49431. 36:51:55traffic you received? How many number of
  49432. 36:51:57orders you received? What's the sale
  49433. 36:51:59that you made in a specific month,
  49434. 36:52:01specific quarter or specific year? And
  49435. 36:52:04what was the profit? So basic
  49436. 36:52:06information which you convert from the
  49437. 36:52:08data and create a dashboard. Right? That
  49438. 36:52:10is called as initial data analysis. So a
  49439. 36:52:14step beyond initial data analysis is
  49440. 36:52:16known as the exploratory data analysis.
  49441. 36:52:19This is where you perform some
  49442. 36:52:20statistics and probability and predict
  49443. 36:52:23the future. Right? So let's dive deep
  49444. 36:52:25and learn what exactly is exploratory
  49445. 36:52:28data analysis. So a simple definition
  49446. 36:52:30for exploratory data analysis is as
  49447. 36:52:32follows. Exploratory data analysis is a
  49448. 36:52:35key step in data analysis process that
  49449. 36:52:38helps you identify patterns, outliners
  49450. 36:52:40and relationships between variables
  49451. 36:52:43before making assumptions. It is not
  49452. 36:52:46like you just create a dashboard out of
  49453. 36:52:47the initial data analysis and you can
  49454. 36:52:49predict the future. No, it's not like
  49455. 36:52:51that. You might have to last. You might
  49456. 36:52:53have to go through some permutations and
  49457. 36:52:55combinations. You might have to check
  49458. 36:52:57the seasons. You might have to check the
  49459. 36:52:59possibilities, right? During some
  49460. 36:53:02particular seasons in the year, let's
  49461. 36:53:04say it's Christmas, then you can expect
  49462. 36:53:06some good sales. Let's say it's some
  49463. 36:53:09festival, it's some special occasion,
  49464. 36:53:11you can expect some good sales on the
  49465. 36:53:13product, right? and maybe a part of the
  49466. 36:53:16year, maybe a part of 10 years, a
  49467. 36:53:19decade, right? In a certain period of
  49468. 36:53:21time, there might be some reason due to
  49469. 36:53:23which the sales of a certain product
  49470. 36:53:26were high. So to make sure that your
  49471. 36:53:29assumptions to make sure that your
  49472. 36:53:31projection of the sales is 100% accurate
  49473. 36:53:34or at least 90 to 95% accurate then you
  49474. 36:53:37might have to go through the exploratory
  49475. 36:53:39data analysis where you make use of
  49476. 36:53:42statistics in your data analysis. Now
  49477. 36:53:45there are certain steps that you might
  49478. 36:53:46have to follow while going through
  49479. 36:53:48exploratory data analysis. So following
  49480. 36:53:51are the steps. So a first few steps
  49481. 36:53:54might be slightly relevant to initial
  49482. 36:53:57data analysis like connecting data,
  49483. 36:53:59cleaning it, transforming it and loading
  49484. 36:54:01it. After that you will import certain
  49485. 36:54:04libraries from Python and after that you
  49486. 36:54:07read the data what exactly you have in
  49487. 36:54:10your data. The number of columns, the
  49488. 36:54:12number of rows and if there are any null
  49489. 36:54:14values, if there are any uh entries
  49490. 36:54:16which are invalid, you might have to
  49491. 36:54:18check that, read that and you might have
  49492. 36:54:20to check for duplicate entries. It is
  49493. 36:54:22possible that one entry might have been
  49494. 36:54:25entered by two different people, right?
  49495. 36:54:27There might be a duplication. So you
  49496. 36:54:28might have to eliminate those
  49497. 36:54:30duplications. You might have to check
  49498. 36:54:32for some missing values. You might have
  49499. 36:54:34to calculate the total number of missing
  49500. 36:54:35values from the data set and try to
  49501. 36:54:37eliminate them from the calculation
  49502. 36:54:40during your exploratory data analysis.
  49503. 36:54:42Followed by that you have to do some
  49504. 36:54:43model engineering. Followed by that you
  49505. 36:54:45might have to do some feature
  49506. 36:54:46engineering creating features and then
  49507. 36:54:49you will get started with exploratory
  49508. 36:54:51data analysis and the at the end you
  49509. 36:54:53will generate a projection or a
  49510. 36:54:55prediction or give your assumption that
  49511. 36:54:57this might happen in the future and you
  49512. 36:54:59might have to take action to avoid it or
  49513. 36:55:02you might have to take action to
  49514. 36:55:04improvise it right so this is how the
  49515. 36:55:06steps in exploratory data analysis take
  49516. 36:55:08part now let's proceed and start with
  49517. 36:55:12our demo on Python's exploratory data
  49518. 36:55:15analysis and in this session we will be
  49519. 36:55:18using the use case that we discussed
  49520. 36:55:19before which happens to be the students
  49521. 36:55:22performance data set. So in this
  49522. 36:55:23particular data set we will be having
  49523. 36:55:25some columns based on physical activity
  49524. 36:55:28the distance from home parental
  49525. 36:55:30education the subjects the marks they
  49526. 36:55:32have scored in the previous exam. The
  49527. 36:55:34marks that they have scored in the
  49528. 36:55:36previous exam and if they have any
  49529. 36:55:38disabilities if they are having any
  49530. 36:55:40resources that they require to write the
  49531. 36:55:42exams right. So a few parameters the
  49532. 36:55:45important parameters that we will be
  49533. 36:55:46discussing in this session and followed
  49534. 36:55:48by that we will project the future that
  49535. 36:55:52how they will you know improve in their
  49536. 36:55:54exams and if there is a problem and if
  49537. 36:55:58there is a solution to it then we can
  49538. 36:56:00implement that solution and help
  49539. 36:56:01students to gain better marks in their
  49540. 36:56:03exam. So that's the overall use case for
  49541. 36:56:05this demonstration. Now let's get
  49542. 36:56:07started with our Jupyter notebook. Now
  49543. 36:56:10we are on Jupiter notebook. Now let's
  49544. 36:56:13get started. So I would like to have a
  49545. 36:56:15title for my um notebook. So I'll write
  49546. 36:56:18an HTML code for that.
  49547. 36:56:22So HTML code uh
  49548. 36:56:26and I have uh three apostrophes
  49549. 36:56:32and here I would like to write something
  49550. 36:56:34in H1. So I want my title to be in H1.
  49551. 36:56:37So
  49552. 36:56:39style will be
  49553. 36:56:43background color
  49554. 36:56:51name will be dark blue.
  49555. 36:56:56So we let's uh proceed with the simply
  49556. 36:56:58dance background which will be t usually
  49557. 36:57:02and the color
  49558. 36:57:04of text will be orange
  49559. 36:57:09and I also want to have the font let's
  49560. 36:57:13let me give the font size as 30.
  49561. 36:57:24And in the next line, I'd like to have
  49562. 36:57:26border
  49563. 36:57:30radius. I can give the border radius as
  49564. 36:57:3320 pixels
  49565. 36:57:37and padding
  49566. 36:57:39to be 16 pixels.
  49567. 36:57:48Text alignment, I'd like to keep it
  49568. 36:57:50center.
  49569. 36:57:55There you go. Let's code the Let's close
  49570. 36:57:58the H1. And now let's code the uh
  49571. 36:58:02background color or border color. So
  49572. 36:58:06B style.
  49573. 36:58:09So what we can do is we can basically
  49574. 36:58:12have this uh code here and what we will
  49575. 36:58:14do is we will reuse this particular code
  49576. 36:58:17segment because we will be having
  49577. 36:58:19multiple HTML codes in this particular
  49578. 36:58:23uh workbook which will explain the
  49579. 36:58:26results of the analysis that we are
  49580. 36:58:29doing. So the way we just run the code
  49581. 36:58:31and after that we will be getting some
  49582. 36:58:33visualizations and I will be writing
  49583. 36:58:36some textual content in XT in an HTML
  49584. 36:58:39page so that it will be easier for the
  49585. 36:58:42people to understand what's what exactly
  49586. 36:58:44is happening here right so the color
  49587. 36:58:47will be light blue and now comes the
  49588. 36:58:50text we will close this and here we will
  49589. 36:58:53write the text as
  49590. 36:58:56Python
  49591. 36:58:58explorate
  49592. 36:59:00data analysis
  49593. 36:59:04and we will
  49594. 36:59:07break here
  49595. 36:59:10the student performance
  49596. 36:59:19and here we will close the H1
  49597. 36:59:24and lastly we will display this in HTML
  49598. 36:59:27code. So basically we missed this
  49599. 36:59:29library. So we will be importing from
  49600. 36:59:32ipython display import html to display
  49601. 36:59:34this particular code. So without this
  49602. 36:59:36particular library we cannot display any
  49603. 36:59:39HTML codes in our notebook. So we will
  49604. 36:59:41quickly run that and we have a title
  49605. 36:59:43over here. Now let's proceed with the
  49606. 36:59:46next part. Now we will uh use some
  49607. 36:59:49libraries like CAD boost and light bgm.
  49608. 36:59:53So for that we might have to install the
  49609. 36:59:55these libraries. So we will be using pip
  49610. 36:59:58install here
  49611. 37:00:00pip install cat boost
  49612. 37:00:04and control enter to run this particular
  49613. 37:00:06code segment or you can also use run and
  49614. 37:00:09it's installed and after that we will
  49615. 37:00:11also install
  49616. 37:00:15light bgm
  49617. 37:00:19g sorry not bgm
  49618. 37:00:24so it's already installed
  49619. 37:00:26Now we will start importing the
  49620. 37:00:29libraries that we need. So we will be
  49621. 37:00:32needing numpy, panda, seabon, mattplot
  49622. 37:00:35lib and we will also import another
  49623. 37:00:37special library which is for ignoring
  49624. 37:00:40warnings. So we will import warnings and
  49625. 37:00:43after that from we will be importing
  49626. 37:00:45that warnings from IPython display
  49627. 37:00:47import clear output and after that we
  49628. 37:00:50will uh tell the jupyter notebook to
  49629. 37:00:52import if there are any warnings. So
  49630. 37:00:54that code will be warning dot filter
  49631. 37:00:57warnings and in the uh brackets we will
  49632. 37:01:00write ignore. So uh the basic libraries
  49633. 37:01:04which we will be needing are as follows.
  49634. 37:01:07import
  49635. 37:01:09numpy
  49636. 37:01:11as np. Let's quickly copy this and
  49637. 37:01:16proceed. Enter. And now we will be
  49638. 37:01:19needing pandas as pd.
  49639. 37:01:23Enter. And now we will be needing
  49640. 37:01:26seabbone
  49641. 37:01:30SNS.
  49642. 37:01:32And we will also use mattplot lib
  49643. 37:01:36py plot
  49644. 37:01:43as plt.
  49645. 37:01:46And after that we will import warnings
  49646. 37:01:51from
  49647. 37:01:54I Python
  49648. 37:01:56dot display
  49649. 37:01:59import
  49650. 37:02:00clear
  49651. 37:02:02output
  49652. 37:02:05warnings
  49653. 37:02:08dot filter warnings
  49654. 37:02:12ignore.
  49655. 37:02:15Now we'll just quickly run this query.
  49656. 37:02:18Run. And now we will proceed with
  49657. 37:02:21feature engineering.
  49658. 37:02:23So you can use a hashtag to ignore that
  49659. 37:02:26particular line from execution for
  49660. 37:02:29Jupyter notebook.
  49661. 37:02:32And here we will be importing import
  49662. 37:02:37from skarn
  49663. 37:02:40dot impute import
  49664. 37:02:45simple computer
  49665. 37:02:50from skarn
  49666. 37:02:52dot model
  49667. 37:02:54selection
  49668. 37:02:57import kf fold. There you go. Now let's
  49669. 37:03:01quickly run this query.
  49670. 37:03:04There you go. Now we will proceed with
  49671. 37:03:06modeling and model evaluation. Once we
  49672. 37:03:09are done with this then we will directly
  49673. 37:03:12input the data into our notebook. So for
  49674. 37:03:15modeling the data we will again use a
  49675. 37:03:17hash code so that this particular line
  49676. 37:03:20will not execute
  49677. 37:03:24and import some libraries lit
  49678. 37:03:28GBM
  49679. 37:03:31as LGB
  49680. 37:03:34from light GBM library.
  49681. 37:03:41import
  49682. 37:03:42LGBM regressor
  49683. 37:03:47from CAD boost import
  49684. 37:03:51boost regressor. There you go. Let's
  49685. 37:03:54quickly run this command.
  49686. 37:03:57There you go. Now, lastly, we have one
  49687. 37:04:00more task before importing the data that
  49688. 37:04:03is model evaluation. Once that is done,
  49689. 37:04:06we can proceed with importing the data.
  49690. 37:04:09There you go. Let's quickly run it. And
  49691. 37:04:11now so far so good. We have uh done the
  49692. 37:04:14basic library imports and feature
  49693. 37:04:17engineering is done, modeling is done
  49694. 37:04:19and model evaluation is also done. Now
  49695. 37:04:21we can begin with importing the data. So
  49696. 37:04:25we let's also add um the HTML code where
  49697. 37:04:29we have done this importing. So here I
  49698. 37:04:32will try to add another segment and I
  49699. 37:04:35will import the HTML code here which
  49700. 37:04:38displays a similar HTML format which
  49701. 37:04:41explains what exactly is happening here.
  49702. 37:04:43Just a moment I have the code ready.
  49703. 37:04:46I'll just paste it here. There you go.
  49704. 37:04:48Let's quickly run this so that we have a
  49705. 37:04:51HTML code here page here which explains
  49706. 37:04:54what exactly is happening here. So we're
  49707. 37:04:55importing libraries and also student
  49708. 37:04:58data. Now let's import the student data.
  49709. 37:05:01For that let's write the query. So we
  49710. 37:05:04are importing the data as data frame and
  49711. 37:05:07after that we will write pandas read CSV
  49712. 37:05:13right and here we will add the location
  49713. 37:05:17of the file. Right? So the file is
  49714. 37:05:20located in my downloads section. So
  49715. 37:05:24let's quickly copy that location and
  49716. 37:05:25paste it here. So this is the location
  49717. 37:05:28of my file. Let's quickly run it. So
  49718. 37:05:31there might be some error. It's okay. We
  49719. 37:05:33can resolve it. So in such scenarios
  49720. 37:05:37don't have to worry either you can add
  49721. 37:05:38an R but even if that doesn't work you
  49722. 37:05:41might have to change uh in this case it
  49723. 37:05:44worked but in case if it doesn't work
  49724. 37:05:45what you can do is uh you can eliminate
  49725. 37:05:48uh the r and you can just change this
  49726. 37:05:51from uh forward slash to backlash. This
  49727. 37:05:55will also help. So this could be worth
  49728. 37:05:58it. So this can also work. So these are
  49729. 37:06:00the situations where you can use this.
  49730. 37:06:02Now let's proceed with some more uh
  49731. 37:06:06interesting facts. Let's try to
  49732. 37:06:08understand what's going on with our uh
  49733. 37:06:10data set. Right? So what you can do is
  49734. 37:06:13now the data is stored in df variable as
  49735. 37:06:17a data frame. So what you can do is read
  49736. 37:06:19this particular data df.shape shape so
  49737. 37:06:23that you can understand what's the uh
  49738. 37:06:26what's happening with this data. Right?
  49739. 37:06:28So it can tell you that there are uh 6
  49740. 37:06:32sorry 6,67
  49741. 37:06:34rows and 20 columns. Now let's add
  49742. 37:06:36another query part and here you can uh
  49743. 37:06:40try to see the head right what head
  49744. 37:06:42means basically head means uh the column
  49745. 37:06:45headers. So what you can do is just
  49746. 37:06:47write df dot head
  49747. 37:06:51and run. Now you have the column headers
  49748. 37:06:54and couple of sample columns
  49749. 37:06:57sorry couple of sample rows. Now let's
  49750. 37:07:01proceed with uh checking the duplicates
  49751. 37:07:04and identifying the total number of
  49752. 37:07:06duplicate entries in this particular
  49753. 37:07:07data set. So you can write down df dot
  49754. 37:07:11duplicated
  49755. 37:07:14dot
  49756. 37:07:16sum and you will get the number of
  49757. 37:07:18duplicates present in this particular
  49758. 37:07:21file. So we have zero duplicate entries.
  49759. 37:07:23Now let's see if there are any null uh
  49760. 37:07:26elements in this particular data frame.
  49761. 37:07:28So df dot is null
  49762. 37:07:33dot sum.
  49763. 37:07:35So these are all functions. Control
  49764. 37:07:36enter. There you go. So in the column
  49765. 37:07:40teacher quality there are 78 null
  49766. 37:07:43entries and in parental educational
  49767. 37:07:45level there are 90 null entries and
  49768. 37:07:47distance from home there are u 67 null
  49769. 37:07:51entries. Now what we can do is uh from
  49770. 37:07:55the okay so from df shape we can add a
  49771. 37:07:59new cell here and we can write a HTML
  49772. 37:08:02code so that the viewer can understand
  49773. 37:08:04that we are trying to understand our
  49774. 37:08:05data. So let's write a quick code for
  49775. 37:08:08that. So let's not waste much time in
  49776. 37:08:10just writing the HTML code. So I've got
  49777. 37:08:12that HTML code written in a notepad
  49778. 37:08:14already. So I'll just paste it over here
  49779. 37:08:16and let's quickly run it so that we have
  49780. 37:08:18a HTML visibility here. So this was
  49781. 37:08:22supposed to be the result. So we will
  49782. 37:08:23add it here. So what we will do is
  49783. 37:08:26quickly edit this particular content.
  49784. 37:08:28We'll just quickly copy this code and
  49785. 37:08:32paste it over here.
  49786. 37:08:36Change the content from importing
  49787. 37:08:40libraries to
  49788. 37:08:43reading
  49789. 37:08:45student data and we will cut this code
  49790. 37:08:48from here and we will paste it in here
  49791. 37:08:55so that it get give us some information.
  49792. 37:08:58Let's also run this particular code
  49793. 37:09:00segment
  49794. 37:09:02so that we have trading student data.
  49795. 37:09:04There you go. Now the next part of this
  49796. 37:09:06session will be about creating a target
  49797. 37:09:09variable. So overall target of this
  49798. 37:09:11particular data analysis is about exam
  49799. 37:09:14score. Right? So we can name our target
  49800. 37:09:16variable as exam score. And let's
  49801. 37:09:18understand the distribution of this
  49802. 37:09:20particular exam score with uh the
  49803. 37:09:25variables we have. Now let's write down
  49804. 37:09:28plot dot figure.
  49805. 37:09:33Figure size should be around 15 comma 9
  49806. 37:09:41equals to let's add a bracket here 15
  49807. 37:09:44comma 9 or let's keep it as six 9 would
  49808. 37:09:48be a little bigger. Now enter now we
  49809. 37:09:52will use seabbond library here. SNS dot
  49810. 37:09:55count plot
  49811. 37:09:58x is equals to data frame cleaned
  49812. 37:10:01target. Okay. Uh before cleaned target
  49813. 37:10:04we might have to run a few more. Okay.
  49814. 37:10:07We did not perform data cleaning so far,
  49815. 37:10:09right? So let's proceed with data
  49816. 37:10:12cleaning so far. So we found some empty
  49817. 37:10:14entries, right? Null entries and we also
  49818. 37:10:17found some So here we have 299 rows
  49819. 37:10:21which have missing values. So we will
  49820. 37:10:23have to remove that. For that uh we
  49821. 37:10:26might have to create a new column which
  49822. 37:10:28has to be named as not um assigned right
  49823. 37:10:31df data frame not assigned which is
  49824. 37:10:34equals to df dot drop na. So we will be
  49825. 37:10:39dropping the null values here. It's a
  49826. 37:10:41function. And here let's print the
  49827. 37:10:44values. Print df dot df
  49828. 37:10:49na dot shape. So how many number of rows
  49829. 37:10:53and columns we have right and after that
  49830. 37:10:57let's also try to eliminate the null
  49831. 37:10:59values as well dot is null so we don't
  49832. 37:11:02have basically we don't have null values
  49833. 37:11:04but we have u some illegal entries maybe
  49834. 37:11:09some there you go now let's quickly run
  49835. 37:11:11this query so there you go so the new
  49836. 37:11:14data is about 678
  49837. 37:11:1820
  49838. 37:11:19so there is Um, okay. We did uh some
  49839. 37:11:23mistake here. So, we supposed to add it
  49840. 37:11:25as null. N U N L N N N N N N N N N N N N
  49841. 37:11:27N N N N N N L N N N N N N N N N N N N N
  49842. 37:11:27N N N N N N N N N N N N N N N N N N N N
  49843. 37:11:27N N N N N N N U L. Now quickly run this.
  49844. 37:11:30So we should not get any errors this
  49845. 37:11:31time.
  49846. 37:11:34There you go. No errors. So far so good.
  49847. 37:11:37Now let's describe the new data set. df
  49848. 37:11:40na dot describe. So these are the new
  49849. 37:11:45columns and rows that we have. And we
  49850. 37:11:48have uh mean, standard variation,
  49851. 37:11:50standard deviation, minimum, maximum. So
  49852. 37:11:53the scores are split into 25%, 50% and
  49853. 37:11:5775%. Which could be based on hours
  49854. 37:12:00studied which could be based on
  49855. 37:12:01attendance, sleep hours, previous
  49856. 37:12:03course, due training sessions,
  49857. 37:12:04everything. So uh minimum sleep hours,
  49858. 37:12:07maximum sleep hours, 25% of that, 50% of
  49859. 37:12:10that, 75% of that. So that is supposed
  49860. 37:12:12to be the u describe. Now what we will
  49861. 37:12:16do is we have a target variable which is
  49862. 37:12:18exam score. Right? Now we will do some
  49863. 37:12:22changes to it. We already know in an
  49864. 37:12:25exam there will be a threshold value. It
  49865. 37:12:28can be 25 marks per exam. It can be 50
  49866. 37:12:31marks per exam and it can be 100 marks
  49867. 37:12:34per exam. Right? In our situation let's
  49868. 37:12:36consider the threshold value is 100
  49869. 37:12:38marks. Right? If there is a situation
  49870. 37:12:41where marks is entered in a wrong way
  49871. 37:12:45right it can if they add if they wanted
  49872. 37:12:47to add 11 but by mistake if they added
  49873. 37:12:49uh another one right triple one it's not
  49874. 37:12:52a right entry right so we will try to
  49875. 37:12:55eliminate those kind of uh data
  49876. 37:12:59so we will create a new uh data frame
  49877. 37:13:01here which is dataf frame cleaned is
  49878. 37:13:04equals to dataf frame
  49879. 37:13:08not null so We have eliminated the null
  49880. 37:13:10values. BF NA. Now we will add our
  49881. 37:13:15target variable which is exam score
  49882. 37:13:19should be. So let's come out of this and
  49883. 37:13:23here we will add it as should be less
  49884. 37:13:26than or equal to 100 but not more than
  49885. 37:13:28100. Let's use square brackets.
  49886. 37:13:32Here we also the format is square
  49887. 37:13:35brackets. ing action now enter and
  49888. 37:13:41we will describe this particular
  49889. 37:13:44data set instead of the FNA we will copy
  49890. 37:13:47paste this here now let's run this
  49891. 37:13:51okay exam score is not identified let's
  49892. 37:13:55quickly check the error and resolve it
  49893. 37:13:56yeah so we missed out to add colons here
  49894. 37:14:00it's okay not a problem so this was
  49895. 37:14:03supposed to be how it is now let's run
  49896. 37:14:05this and we will have the answer over
  49897. 37:14:07here. So we have the output. Now let's
  49898. 37:14:10check the uh head of this particular
  49899. 37:14:12clean data set. So we can make use of
  49900. 37:14:14the same code here and paste it right
  49901. 37:14:17here and instead of describe let's write
  49902. 37:14:20head so that we have the header of uh
  49903. 37:14:23this particular data set. So we have our
  49904. 37:14:25study and everything normal and we will
  49905. 37:14:28categorize the data right. So we will
  49906. 37:14:31make use of three columns our study
  49907. 37:14:33attendance and previous scores and uh
  49908. 37:14:36after that we will also make use of
  49909. 37:14:39other columns in this particular data
  49910. 37:14:41set which happens to be the parental
  49911. 37:14:43involvement access to resources sleep
  49912. 37:14:45hours ting sessions etc. And now our
  49913. 37:14:49target will be the exam score that we
  49914. 37:14:51created over here. Right? This exam
  49915. 37:14:53score will be our target. And using this
  49916. 37:14:55particular exam score target, we will
  49917. 37:14:58categorize the data. Okay? We will
  49918. 37:15:00categorize the data in terms of uh let's
  49919. 37:15:03say uh first class uh second class and
  49920. 37:15:06uh pass or something like that. Right?
  49921. 37:15:09So if if a student is uh scoring below
  49922. 37:15:1264 and uh that is a separate category.
  49923. 37:15:15If the score student is scoring equals
  49924. 37:15:17to or above 65, that is a different
  49925. 37:15:19category. And if the student is scoring
  49926. 37:15:21beyond 70, that's a different category.
  49927. 37:15:25And uh before we proceed with that,
  49928. 37:15:28let's try to add a HTML code before this
  49929. 37:15:31particular data set so that we have u an
  49930. 37:15:34understanding of what exactly happened
  49931. 37:15:36here. So I let's uh I'll just quickly
  49932. 37:15:38copy paste this particular code here. So
  49933. 37:15:40we will run it and now next we will
  49934. 37:15:43describe check the data type column
  49935. 37:15:44separate and everything and we will you
  49936. 37:15:46know create data type category variables
  49937. 37:15:50right now so far so good. Now we will
  49938. 37:15:54create the categories.
  49939. 37:15:56So num call
  49940. 37:15:58equals to
  49941. 37:16:02hours studied
  49942. 37:16:05comma attendance.
  49943. 37:16:10So let's quickly add the data
  49944. 37:16:14previous exam scores.
  49945. 37:16:17Just a minute. Let's quickly add the
  49946. 37:16:19columns. Let me take a while. There you
  49947. 37:16:22go. I've added the columns. So we are
  49948. 37:16:23considering three different columns. uh
  49949. 37:16:26our study attendance and previous scores
  49950. 37:16:28for num call and cat call. We are
  49951. 37:16:30considering the other columns apart from
  49952. 37:16:32the first three and our target value is
  49953. 37:16:35exam score. Let's quickly run this.
  49954. 37:16:36There you go. And now we will try to
  49955. 37:16:40build some visualizations and before
  49956. 37:16:42that let's try to uh add some data uh
  49957. 37:16:46from in HTML. Let's try to create a
  49958. 37:16:50Okay, what we can do is simply copy this
  49959. 37:16:53particular HTML file here. We can take
  49960. 37:16:57this
  49961. 37:16:59and add it here so that we will
  49962. 37:17:01understand what exactly is happening
  49963. 37:17:03next. And in place of reading student
  49964. 37:17:06data, we will write data visualization
  49965. 37:17:10for student data.
  49966. 37:17:13And we will keep the colors same dark
  49967. 37:17:15blue background and u the color for data
  49968. 37:17:19visualization will be light blue and
  49969. 37:17:20student data will be orange. Let's
  49970. 37:17:23quickly run and there we have it. Now
  49971. 37:17:25our target variable is exam score. Right
  49972. 37:17:28now we will compare this particular
  49973. 37:17:31target variable with three other
  49974. 37:17:34parameters. So our parameters will be
  49975. 37:17:37the following uh as we discussed uh
  49976. 37:17:40creating the segregation in data set
  49977. 37:17:42right. So first will be distribution of
  49978. 37:17:45target variable with other parameters.
  49979. 37:17:48So we will create another HTML file for
  49980. 37:17:51that right here. Just a minute while I
  49981. 37:17:53paste the code for um HTML quickly run
  49982. 37:17:57this. There you go. So distribution of
  49983. 37:18:00target variable exam score against some
  49984. 37:18:03parameters. Now we will write the plot
  49985. 37:18:06for that
  49986. 37:18:08plot dot figure. So we are going to
  49987. 37:18:12consider the size equals to 15 6
  49988. 37:18:17big size
  49989. 37:18:20is equals to 15 6 the same one that we
  49990. 37:18:24considered before. And we will be using
  49991. 37:18:26Cbond SNS dot count plot
  49992. 37:18:32open bracket. This is a function x is
  49993. 37:18:34equals to df
  49994. 37:18:37clean. Okay. Uh what we can do is just
  49995. 37:18:39quickly take the column name so that we
  49996. 37:18:42don't create any mistakes here and we
  49997. 37:18:46will paste it over here instead of df.
  49998. 37:18:49There you go. Or target.
  49999. 37:18:53So our target is exam score,
  50000. 37:18:57and we will use the pellet as green
  50001. 37:19:02and the plot title will be distribution
  50002. 37:19:04of target variable exam score. We can
  50003. 37:19:07copy this. Okay, just a minute before
  50004. 37:19:10that plt do.
  50005. 37:19:14Should be let's use double quotes now.
  50006. 37:19:16Copy this and paste it here.
  50007. 37:19:20I think semicolons went off. Okay, not a
  50008. 37:19:23problem.
  50009. 37:19:24It's right here. Let's add a dot as a
  50010. 37:19:27full stop. And the next line, if you
  50011. 37:19:31need, you can add the full stop. If not
  50012. 37:19:32you can ignore plt dot grid true
  50013. 37:19:38which equals to major
  50014. 37:19:43comma
  50015. 37:19:44access is equals to y
  50016. 37:19:49comma line style equals to so I want
  50017. 37:19:55lines to be hyphen hyphen in this way I
  50018. 37:19:57want the lines and comma line width um
  50019. 37:20:01let's say 0.5 or 0.7. Let's go with 0.7
  50020. 37:20:07line width equals to 0.9 mm. There you
  50021. 37:20:11go. Now let's quickly run this query.
  50022. 37:20:13There you go. Done. And we have the
  50023. 37:20:16first visualization. So here you can see
  50024. 37:20:19there are some students which are
  50025. 37:20:21scoring 58 59 and you can see maximum
  50026. 37:20:24number of students are already scoring
  50027. 37:20:25good marks which is under 65 and uh
  50028. 37:20:28sorry which is under 70 and above 65 and
  50029. 37:20:32there is a good number of students uh
  50030. 37:20:34which are also scoring uh above 65 as
  50031. 37:20:37well right so sorry 70 70 as well. So we
  50032. 37:20:40have now less than or equal to 70 71.
  50033. 37:20:43And highest scorer in some situations
  50034. 37:20:46there is also 100. If you can see there
  50035. 37:20:48is slight growth here. There are a few
  50036. 37:20:50students toppers maybe which have
  50037. 37:20:52already scored 100 as well. Now we have
  50038. 37:20:55the list here. Now we what we need to do
  50039. 37:20:56is we need to segregate that is part one
  50040. 37:20:59which is less than or equal to 64 which
  50041. 37:21:02falls under 65 and another category
  50042. 37:21:05which falls in between 65 to 70 and
  50043. 37:21:08above 70. So we need to categorize these
  50044. 37:21:11three uh datas and segregate them as
  50045. 37:21:13bottom 65, top which is above 70 and mid
  50046. 37:21:17between 70 to 65. Right now before that
  50047. 37:21:21if you want to add an HTML document uh
  50048. 37:21:24sorry segment here which explains about
  50049. 37:21:26the distribution of target you can also
  50050. 37:21:28do that. It's already added here. Now
  50051. 37:21:31let's continue. But in case if you want
  50052. 37:21:33to uh add uh some data which explains
  50053. 37:21:37that we're trying to segregate, you can
  50054. 37:21:39also do that. I would like to do that.
  50055. 37:21:42Let's quickly uh add that HTML code
  50056. 37:21:44here. So what this particular code will
  50057. 37:21:47do is it will tell the percentage of
  50058. 37:21:49students which are scoring less than 65.
  50059. 37:21:52Number of stu uh percentage of students
  50060. 37:21:55uh scoring in between 65 and 70 and the
  50061. 37:21:58percent of students which are scoring
  50062. 37:22:00beyond 70. Right? Let's run this. And
  50063. 37:22:02here we have the result. Bottom 21.81%
  50064. 37:22:05scores under 64 while 24% scores over 70
  50065. 37:22:09and 50% are in between 65 to 69. Right
  50066. 37:22:13now let's uh continue with the
  50067. 37:22:15segregation part of the data. So for
  50068. 37:22:17segregation we will create three
  50069. 37:22:19different variables A, B, C. So first A
  50070. 37:22:22is equals to length of DF claimed. So
  50071. 37:22:25let's copy the column name sorry data
  50072. 37:22:29frame name length of DF cleaned inside
  50073. 37:22:34the square brackets we'll again add DF
  50074. 37:22:36cleaned of target variable which is exam
  50075. 37:22:40score
  50076. 37:22:42let's also add uh single quotes here
  50077. 37:22:46who are scoring in between or um less
  50078. 37:22:50than let's start with less than or equal
  50079. 37:22:53to 64 we'll not consider is 65 we'll
  50080. 37:22:56consider 64 divided by len of df cleaned
  50081. 37:23:02target variable exam score single quotes
  50082. 37:23:06into 100
  50083. 37:23:09which will give us the percentage now
  50084. 37:23:11similarly let's just copy and paste this
  50085. 37:23:14three more times for B and C. So here
  50086. 37:23:18instead of minus we're supposed to add
  50087. 37:23:19equals to and another one. So here
  50088. 37:23:24equals to and instead of A I will write
  50089. 37:23:26B and the last one is C. And instead of
  50090. 37:23:3064 we will add 70 here for top and here
  50091. 37:23:35we will make some changes. It should be
  50092. 37:23:38greater than or equal to
  50093. 37:23:4265. So this is the third category A B C
  50094. 37:23:45and then we will proceed with printing
  50095. 37:23:50the files. So print
  50096. 37:23:54the bottom
  50097. 37:23:57for the first one which is f of
  50098. 37:24:01a
  50099. 37:24:04is to do 2f
  50100. 37:24:07and we will add the percentage symbol
  50101. 37:24:08over here r under
  50102. 37:24:1264.
  50103. 37:24:14Now we can copy paste the same here and
  50104. 37:24:17we can change the variables.
  50105. 37:24:21So here we will be adding under over 70.
  50106. 37:24:27Lastly in between the ones in between
  50107. 37:24:3365 and 70.
  50108. 37:24:36There you go. Here we will change the
  50109. 37:24:39values from A to B and here A to C.
  50110. 37:24:44There you go. And we can quickly run
  50111. 37:24:46this query. So it's not 70. It was
  50112. 37:24:49supposed to be 69.
  50113. 37:24:51There you go. So we forgot to mention
  50114. 37:24:54this particular one. Now let's run this.
  50115. 37:24:58There you go. So we have 21%
  50116. 37:25:01of people who are scoring under 64, 24%
  50117. 37:25:04over 70 and 53% are in between average.
  50118. 37:25:08So I think the school is focusing on
  50119. 37:25:10improving this particular percentage,
  50120. 37:25:13reducing this particular percentage and
  50121. 37:25:15increasing this particular percentage
  50122. 37:25:17and try to eliminate if possible this
  50123. 37:25:19particular one which are under 64. So
  50124. 37:25:22that is the overall moto I guess. Now so
  50125. 37:25:25far so good. Let's now try to remove
  50126. 37:25:28infinite values from HTML, right? So
  50127. 37:25:30before that, let's add uh this
  50128. 37:25:32particular HTML code here. So which
  50129. 37:25:34explains what we are trying to do. So we
  50130. 37:25:37will first implement the code that
  50131. 37:25:39prevents warning about infinite values
  50132. 37:25:41during data visualization. And now let's
  50133. 37:25:43add the code which will try to eliminate
  50134. 37:25:45the uh infinite values. Let's copy this
  50135. 37:25:48particular data frame cleaned uh data
  50136. 37:25:52frame name here. Now df cleaned
  50137. 37:25:56dotreplace
  50138. 37:26:01square brackets np dot info
  50139. 37:26:07comma
  50140. 37:26:08minus np
  50141. 37:26:11dot info out of these square brackets
  50142. 37:26:14dot np
  50143. 37:26:16na
  50144. 37:26:18non na values we're trying to eliminate
  50145. 37:26:20na values in place of those values you
  50146. 37:26:24can write true and after that we will
  50147. 37:26:27try to eliminate the null. So if it is
  50148. 37:26:31null
  50149. 37:26:33dot sum give me the total number of null
  50150. 37:26:36values after this. Right? So let's try
  50151. 37:26:38to execute that. There you go. So all
  50152. 37:26:40the null entries have been removed here.
  50153. 37:26:43Now let's see the distribution of
  50154. 37:26:45numerical values here. So before that
  50155. 37:26:48let's add the HTML code for that. So
  50156. 37:26:51let's quickly run this. So distribution
  50157. 37:26:54of numerical values. So we will be
  50158. 37:26:56considering these three parameters. So
  50159. 37:26:59if you go back here you can see our
  50160. 37:27:02studied attendance and previous scores.
  50161. 37:27:05So we will be considering these three
  50162. 37:27:07values or these three columns and check
  50163. 37:27:10the distribution of these variables
  50164. 37:27:13against the exam score. So uh is it
  50165. 37:27:16making any um you know kind of variation
  50166. 37:27:19if the u attendance is high? If if the
  50167. 37:27:22attendance is high is the exam score
  50168. 37:27:24high and uh apart from that we have if
  50169. 37:27:28our studies is high is the uh mark score
  50170. 37:27:31is high and if the previous scores are
  50171. 37:27:34high is there a chance to get better
  50172. 37:27:36scores in this particular exam. So what
  50173. 37:27:38we are trying to do is we are trying to
  50174. 37:27:40see if there is any direct involvement
  50175. 37:27:42of number of study hours and number of
  50176. 37:27:45days attended and number of uh or the
  50177. 37:27:48number of marks they received in the
  50178. 37:27:49previous course and we'll try to build a
  50179. 37:27:52visualization on that front.
  50180. 37:27:55So let's go and build that. So we'll try
  50181. 37:27:57the try to write the code here.
  50182. 37:28:00Figure axis
  50183. 37:28:05equals tot
  50184. 37:28:08dot
  50185. 37:28:10subplots. So we will be having three
  50186. 37:28:12different plots here since we're
  50187. 37:28:13considering three different u
  50188. 37:28:16categories.
  50189. 37:28:18And the fixed size should be equal to
  50190. 37:28:2212A 4. There you go. Now
  50191. 37:28:28access
  50192. 37:28:30is equals to access dot variable
  50193. 37:28:36per idx
  50194. 37:28:39comma call
  50195. 37:28:41in enumerate
  50196. 37:28:52and we will try to import the seaborn
  50197. 37:28:54library here. We will try to create
  50198. 37:28:57histo plots here. Histograms here
  50199. 37:29:01plot.
  50200. 37:29:22So line style we will be selecting this
  50201. 37:29:24one
  50202. 37:29:25and the comma
  50203. 37:29:28line width will be 0.7.
  50204. 37:29:34There you go. Next will be access
  50205. 37:29:41dot set
  50206. 37:29:44title. So for this we will be uh setting
  50207. 37:29:47the title as distribution of columns. So
  50208. 37:29:50the columns will be the three uh ones
  50209. 37:29:52attendance, hours studied and uh what
  50210. 37:29:55was the third one that we considered
  50211. 37:29:59previous course. Right? So instead of
  50212. 37:30:02mentioning them specifically, what we
  50213. 37:30:03can do is we can just write columns
  50214. 37:30:05here. C L and close.
  50215. 37:30:10There you go. And lastly,
  50216. 37:30:15plt.tight Right.
  50217. 37:30:18And show the plot. There you go. Let's
  50218. 37:30:21quickly run this query. Run. And now we
  50219. 37:30:24will be having the visualizations here.
  50220. 37:30:27So um you can directly see the
  50221. 37:30:29involvement of these three parameters
  50222. 37:30:30here. If uh they are trying to help if
  50223. 37:30:34the number of hours are increased then
  50224. 37:30:36you can see if there is a better
  50225. 37:30:37improvement in scores. If the attendance
  50226. 37:30:39is increased, if there is a betterment
  50227. 37:30:40in scores or if the previous uh scores
  50228. 37:30:44are helping then you can find it out how
  50229. 37:30:46it is. There you go. Now we can write a
  50230. 37:30:49result here in the form of HTML page.
  50231. 37:30:52And if we run this, it will give you the
  50232. 37:30:54result. The breaks or gaps in the hour
  50233. 37:30:56study variable may be due to the
  50234. 37:30:58respondents answering appropriately. The
  50235. 37:31:00variables attendance and previous scores
  50236. 37:31:02which exhibit a uniform distribution
  50237. 37:31:05have a normal impact on exam scores
  50238. 37:31:07variable which is our target variable.
  50239. 37:31:09Now let's proceed with another part of
  50240. 37:31:11this session which will be about the
  50241. 37:31:13relationship between the numerical
  50242. 37:31:15values and the target variables. Now we
  50243. 37:31:17will copy paste the same code and make
  50244. 37:31:19some minute changes to it. So the only
  50245. 37:31:22change that we did to it is we're trying
  50246. 37:31:24to uh build a scatter plot. So we will
  50247. 37:31:27be getting a scatter plot here. But
  50248. 37:31:28before that let's try to add another
  50249. 37:31:33column here and try to add an HTML code
  50250. 37:31:35which explains why we are doing it.
  50251. 37:31:38There you go a scatter plot. So
  50252. 37:31:40basically these two are one and the
  50253. 37:31:41same. Here we use some column graphs. So
  50254. 37:31:44here we did the same using the scatter
  50255. 37:31:46plot which will help for a better
  50256. 37:31:48understanding. Now we will try to build
  50257. 37:31:50some correlations.
  50258. 37:31:52So basically a list is called
  50259. 37:31:54correlation is created containing the
  50260. 37:31:56names of the columns for which you want
  50261. 37:31:58to calculate the correlation. So here in
  50262. 37:32:00our situation it is the df c r which is
  50263. 37:32:04a data frame and it is created by
  50264. 37:32:06selecting only these columns for df
  50265. 37:32:08cleaned data frame effectively created a
  50266. 37:32:11new data frame containing only the
  50267. 37:32:12specified columns. Now the second one
  50268. 37:32:14which is the co r which is a calculate
  50269. 37:32:17correlation. So this method computes the
  50270. 37:32:20correlation matrix for selected columns
  50271. 37:32:22which is n df c r. The one indicates
  50272. 37:32:26perfect positive correlation minus one
  50273. 37:32:28indicates the perfect negative
  50274. 37:32:30correlation and zero indicates no
  50275. 37:32:31correlation. And lastly the setup of
  50276. 37:32:34plot. This line sets up the figure size
  50277. 37:32:37of the plot. In this particular
  50278. 37:32:39situation we are choosing five and four.
  50279. 37:32:42Right now let's quickly try to execute
  50280. 37:32:44this query and see the answer. And we
  50281. 37:32:47will also add the HTML code for this so
  50282. 37:32:50that we have a better understanding for
  50283. 37:32:52this. So we will be adding that HTML
  50284. 37:32:56uh box here which will explain what
  50285. 37:32:58exactly happened here. So this is our
  50286. 37:33:01plot and this is the correlation. Now
  50287. 37:33:03let's try to add that HTML code right
  50288. 37:33:06here. The result of this particular data
  50289. 37:33:09visualization will be maintained here.
  50290. 37:33:12So the hours studying and attendance
  50291. 37:33:14shows a positive correlation with the
  50292. 37:33:16target variable which is exams hour.
  50293. 37:33:18However, previous course appears to have
  50294. 37:33:20no or little relationship with the
  50295. 37:33:22target variable. Right now let's
  50296. 37:33:24continue with our next uh part of this
  50297. 37:33:27session. So now we will try to identify
  50298. 37:33:30the relationship between studies
  50299. 37:33:34hours and attendance and extracurricular
  50300. 37:33:37scores. Right? So we have other u
  50301. 37:33:40columns to consider which is
  50302. 37:33:41extracurricular activities. So there is
  50303. 37:33:43a belief that extracurricular activities
  50304. 37:33:45will also help students to study better.
  50305. 37:33:47So we will find if there is a relation
  50306. 37:33:49between the target variable and this
  50307. 37:33:51extracurricular activities attendance
  50308. 37:33:53and study hours. Now let's quickly add
  50309. 37:33:56the code here. Now let's quickly execute
  50310. 37:33:58the code. Now we have the visualization
  50311. 37:34:02which explains the relationship between
  50312. 37:34:04the number of hours studied
  50313. 37:34:06extracurricular activities etc. So here
  50314. 37:34:08it is and now let's add an HTML code
  50315. 37:34:10which explains about this result
  50316. 37:34:13in this particular code was supposed to
  50317. 37:34:15be added here.
  50318. 37:34:20So this uh is the resultant column here.
  50319. 37:34:24Now here it explains about the
  50320. 37:34:26influences that it performs. So the
  50321. 37:34:28extracal activities, parental income and
  50322. 37:34:30extra things that influence the scores
  50323. 37:34:34and there you go.
  50324. 37:34:37Now let us also consider other columns
  50325. 37:34:40right the other parameters like
  50326. 37:34:43resources are available or not parental
  50327. 37:34:45education and other things which also
  50328. 37:34:47might have influenced the exam scores of
  50329. 37:34:50students. So for that let's add an HTML
  50330. 37:34:52code so that we have an HTML page here
  50331. 37:34:54which explains what is the next
  50332. 37:34:55procedure that we are following. Right
  50333. 37:34:57now let's add the query here. So here we
  50334. 37:35:00are considering the other parameters
  50335. 37:35:02like family income, peer influence,
  50336. 37:35:04motivation level, gender, parental
  50337. 37:35:06involvement, parental educational level
  50338. 37:35:08and extracurricular activities. And we
  50339. 37:35:10are considering them against the target
  50340. 37:35:13variable which is exam score. And now
  50341. 37:35:16let's execute this query. There you go.
  50342. 37:35:19Now we have generated a graph which
  50343. 37:35:20explains about this particular
  50344. 37:35:23parameters against the target variable.
  50345. 37:35:25And now let's add the HTML page here
  50346. 37:35:27which explains about these results.
  50347. 37:35:31Let's quickly run it. And there you go.
  50348. 37:35:34So when certain factors affect Q1 and Q2
  50349. 37:35:36but not Q2, it can be understood that
  50350. 37:35:38individual has overcome challenges
  50351. 37:35:40through personal effort. Right? So if
  50352. 37:35:42government policies and corporate social
  50353. 37:35:45contributors are focused on addressing
  50354. 37:35:47these aspect, it seems that we could
  50355. 37:35:49create dynamic country with greater
  50356. 37:35:52social mobility and open opportunities
  50357. 37:35:53for all. Right? So if extracurricular
  50358. 37:35:56activities can outweigh the influence of
  50359. 37:35:58other variables in academic performance
  50360. 37:35:59then we should foster that kind of
  50361. 37:36:01environment right. So this is how u you
  50362. 37:36:04can get extract some statistical
  50363. 37:36:07analysis on this particular data set.
  50364. 37:36:09Now let's quickly rename this uh python
  50365. 37:36:14eda
  50366. 37:36:16students
  50367. 37:36:18performance
  50368. 37:36:20and you can quickly rename and save it.
  50369. 37:36:23Welcome to math refresher probability
  50370. 37:36:26and statistics.
  50371. 37:36:28In this lesson, we are going to explain
  50372. 37:36:30the concepts of statistics and
  50373. 37:36:33probability.
  50374. 37:36:34Describe conditional probability. Define
  50375. 37:36:37the chain rule of probability. Discuss
  50376. 37:36:40the measure of variance. Identify the
  50377. 37:36:42types of gshian distribution.
  50378. 37:36:45Basic of statistics and probability.
  50379. 37:36:48Probability and statistics. Data science
  50380. 37:36:51relies heavily on estimates and
  50381. 37:36:53predictions. A significant portion of
  50382. 37:36:56data science is made up of evaluations
  50383. 37:36:58and forecast.
  50384. 37:37:00Statistical methods are used to make
  50385. 37:37:02estimates for further analysis.
  50386. 37:37:05Probability theory is helpful for making
  50387. 37:37:07predictions. Statistical methods are
  50388. 37:37:10highly dependent on probability theory
  50389. 37:37:13and all probability and statistics are
  50390. 37:37:16dependent on data. Data is information
  50391. 37:37:19acquired for reference or research via
  50392. 37:37:22observations, facts, and measurements.
  50393. 37:37:26Data is a set of facts structured in the
  50394. 37:37:29form that computers can interpret such
  50395. 37:37:31as numbers, words, estimations, and
  50396. 37:37:34views. Importance of data. Data aids in
  50397. 37:37:38seeing more about the information by
  50398. 37:37:40identifying possible connections between
  50399. 37:37:42two features. Data assists in the
  50400. 37:37:45detection of distortion by uncovering
  50401. 37:37:48hidden patterns based on prior
  50402. 37:37:50information patterns. Data may be
  50403. 37:37:53utilized to anticipate the future or
  50404. 37:37:55predict the current state of affairs.
  50405. 37:37:58Also, data aids in determining whether
  50406. 37:38:00two pieces of information have any
  50407. 37:38:02instance in common or not. Types of
  50408. 37:38:05data. Data might be quantitative. That
  50409. 37:38:09is data that can be measured or counted
  50410. 37:38:11in numbers. Or it may be qualitative
  50411. 37:38:14which is data which is generally divided
  50412. 37:38:16into groups or in simpler words which
  50413. 37:38:19cannot be counted or measured in
  50414. 37:38:21numbers. Let's consider an example. A
  50415. 37:38:24customer information data of a bank may
  50416. 37:38:27contain quantitative and qualitative
  50417. 37:38:29data. Consider this snapshot where we
  50418. 37:38:32have customer ID, surname, geography,
  50419. 37:38:36gender, age, balance, has C or card is
  50420. 37:38:39active member. Amongst these variables
  50421. 37:38:42we can see surname is mostly qualitative
  50422. 37:38:45as it cannot be counted and measured in
  50423. 37:38:47numbers. Geography and gender are also
  50424. 37:38:51qualitative as they cannot be counted in
  50425. 37:38:53numbers and are mostly groups. has C or
  50426. 37:38:57card that is has credit card and is
  50427. 37:39:00active member although are containing
  50428. 37:39:02numerical in form but these are
  50429. 37:39:05categorical that means these have been
  50430. 37:39:07divided into groups of one and zero that
  50431. 37:39:11represent yes and no as an answer hence
  50432. 37:39:14these two variables are also qualitative
  50433. 37:39:18customer ID is again although a
  50434. 37:39:21numerical data however the significance
  50435. 37:39:24or intuition behind Customer ID is
  50436. 37:39:27categorical.
  50437. 37:39:28Hence, it may be kept in the qualitative
  50438. 37:39:31data also. However, age and balance
  50439. 37:39:35these are numerical information which
  50440. 37:39:37have been measured or counted and
  50441. 37:39:39numerical operations can be performed on
  50442. 37:39:42them. Hence, these are under
  50443. 37:39:44quantitative data categories.
  50444. 37:39:46Introduction to descriptive statistics.
  50445. 37:39:49Descriptive statistics. A descriptive
  50446. 37:39:52measurement is summary measure that
  50447. 37:39:54quantitatively portrays the most
  50448. 37:39:56important features of a set of data
  50449. 37:39:59allowing for a better comprehension of
  50450. 37:40:01the information. Data can be measured as
  50451. 37:40:04different levels. The levels of
  50452. 37:40:06measurement describe the nature of
  50453. 37:40:08information stored in the data assigned
  50454. 37:40:10to the variables. Qualitative data can
  50455. 37:40:13be measured as nominal or ordinal.
  50456. 37:40:15Quantitative data can be measured in
  50457. 37:40:17terms of interval and ratio type.
  50458. 37:40:20Nominal data. The data is categorized
  50459. 37:40:23using names, labels or qualities. For
  50460. 37:40:26example, brand name, zip code, and
  50461. 37:40:28gender. Ordinal data can be arranged in
  50462. 37:40:31order or ranked, and can be compared.
  50463. 37:40:34Examples include grades, star reviews,
  50464. 37:40:38position, and race, and date. Interval
  50465. 37:40:41data is the data that is ordered and has
  50466. 37:40:43meaningful differences between the data
  50467. 37:40:46points. Example, temperature in Celsius
  50468. 37:40:49and year of birth. Ratio data is similar
  50469. 37:40:52to the interval level with the added
  50470. 37:40:55property of inherent zero. Mathematical
  50471. 37:40:58calculations can be performed on both
  50472. 37:41:00interval as well as ratio data. For
  50473. 37:41:03example, height, age, and weight.
  50474. 37:41:06Population versus sample. Before
  50475. 37:41:09analyzing the data, it's important to
  50476. 37:41:11figure out if it's from a population or
  50477. 37:41:13a sample. Population is a collection of
  50478. 37:41:17all available items as well as each unit
  50479. 37:41:19in our study. Sample is a subset of the
  50480. 37:41:22population that contains only a few
  50481. 37:41:25units of the population. Population data
  50482. 37:41:28is used for study when the data pool is
  50483. 37:41:31very small and can give all the required
  50484. 37:41:33information. Samples are collected
  50485. 37:41:36randomly and represent the entire
  50486. 37:41:39population in the best possible way.
  50487. 37:41:42Measures of central tendency. The
  50488. 37:41:45central tendency is a single value that
  50489. 37:41:48aids in the description of the data by
  50490. 37:41:50determining its center position.
  50491. 37:41:53Measures of central tendency are
  50492. 37:41:55sometimes known as summary statistics or
  50493. 37:41:58measures of central location. The most
  50494. 37:42:01popular measurements of central tendency
  50495. 37:42:04are mean, median, and mode. The normal
  50496. 37:42:07distribution is a bell-shaped
  50497. 37:42:10symmetrical distribution in which mean,
  50498. 37:42:12median, and mode all are equal. The
  50499. 37:42:15curve over here shows the bell-shaped
  50500. 37:42:17curve or the normal distribution of
  50501. 37:42:19variable X. The point over here that is
  50502. 37:42:22X1 is the point which represents the
  50503. 37:42:26mean, median and mode of this
  50504. 37:42:28distribution. Mean mean is calculated by
  50505. 37:42:31dividing these sum of all data values by
  50506. 37:42:34the total number of data values. It gets
  50507. 37:42:38affected when there are unusual or
  50508. 37:42:40extreme values. It is sensitive to the
  50509. 37:42:43outliers. Mean can be calculated as
  50510. 37:42:46summation over all the values of X in a
  50511. 37:42:49collection divided by the size of the
  50512. 37:42:51collection.
  50513. 37:42:53For example, we have a collection where
  50514. 37:42:55we have values as 7 3 4 1 6 and 7.
  50515. 37:43:01We find out the sum of these values
  50516. 37:43:03which is 28 and there are total of six
  50517. 37:43:07values. So 28 / 6 gives us a mean value
  50518. 37:43:11of 4.66.
  50519. 37:43:14Median,
  50520. 37:43:16it is the middle value in the set of the
  50521. 37:43:18data that has been sorted in ascending
  50522. 37:43:20order.
  50523. 37:43:22It is a better alternative to mean since
  50524. 37:43:24it is less impacted by outliers and
  50525. 37:43:27skewess.
  50526. 37:43:28It is closer to the actual central
  50527. 37:43:31value.
  50528. 37:43:32Median is calculated differently for
  50529. 37:43:35different sizes of data.
  50530. 37:43:37Differentiated as if the total number of
  50531. 37:43:39values is odd or if the total number of
  50532. 37:43:43values is even. If the size of the data
  50533. 37:43:46is odd. For example, in this case we
  50534. 37:43:50have five elements.
  50535. 37:43:54After sorting whatever middle value we
  50536. 37:43:56get
  50537. 37:43:58that means n + 1 by 2 term in this case
  50538. 37:44:035 + 1 / 2
  50539. 37:44:06that is the third term which is four is
  50540. 37:44:09the median value.
  50541. 37:44:12In case when the total number of values
  50542. 37:44:14is even like here there are six values.
  50543. 37:44:18The average or the mean of the two
  50544. 37:44:20central values is considered as the
  50545. 37:44:22median. In this case the median is the
  50546. 37:44:25mean of 6 and four which is five. Mode.
  50547. 37:44:30Mode represents the most common value in
  50548. 37:44:33the data set. It is not at all affected
  50549. 37:44:36by extreme observations.
  50550. 37:44:40It is the best measure of central
  50551. 37:44:42tendency for highly skewed or non-normal
  50552. 37:44:45distribution.
  50553. 37:44:46Mode for categorical data is determined
  50554. 37:44:49by estimating the frequencies for each
  50555. 37:44:51categories
  50556. 37:44:53and then the category with the highest
  50557. 37:44:55frequency is considered to be mode.
  50558. 37:44:58Like in this case 7 has the highest
  50559. 37:45:01frequency. Hence seven becomes the mode
  50560. 37:45:03value. However, in case of continuous
  50561. 37:45:07data or quantitative data, the
  50562. 37:45:09calculation of mode is slightly
  50563. 37:45:11different. The first step in calculation
  50564. 37:45:14of mode is dividing the data into
  50565. 37:45:16classes which are equal with then
  50566. 37:45:18getting the frequency of data points
  50567. 37:45:20lying in within that range of classes
  50568. 37:45:23and finally selecting the class with the
  50569. 37:45:26highest frequency.
  50570. 37:45:29Using the range of that class and the
  50571. 37:45:31frequencies, we can get the final mode
  50572. 37:45:33value.
  50573. 37:45:35Using the formula L+
  50574. 37:45:39minus F_sub_1 multiplied to H divided by
  50575. 37:45:42FM minus F_sub_1 plus FM minus F_sub_2.
  50576. 37:45:47Here L is the lower limit or the lower
  50577. 37:45:50observation of the mode class.
  50578. 37:45:53H is the size of the mode class.
  50579. 37:45:57FM is the frequency of the mode class.
  50580. 37:46:00F_sub_1 is the frequency of the class
  50581. 37:46:03proceeding to mode. And F_sub_2 is the
  50582. 37:46:06frequency of the class succeeding to
  50583. 37:46:09mode. This gives us the final mode
  50584. 37:46:11value.
  50585. 37:46:13Mean versus expectation.
  50586. 37:46:16Now let's talk about mean versus
  50587. 37:46:18expectation.
  50588. 37:46:19So in general we use the expected value
  50589. 37:46:22or expectation when we want to calculate
  50590. 37:46:25the mean of a probability distribution
  50591. 37:46:28that represents the average value we
  50592. 37:46:31expect to occur before collecting any
  50593. 37:46:33data. And mean on the other hand mean is
  50594. 37:46:36basically used when we want to calculate
  50595. 37:46:39the average value of a given sample.
  50596. 37:46:42This represents the average value of raw
  50597. 37:46:45data that we may have already collected.
  50598. 37:46:48We can understand this by using a simple
  50599. 37:46:51example.
  50600. 37:46:53Now to calculate the expected value of
  50601. 37:46:56this probability distribution, we can
  50602. 37:46:59use a specific formula from the previous
  50603. 37:47:01discussion.
  50604. 37:47:03This is going to be the expected value
  50605. 37:47:05where X is going to be the data value
  50606. 37:47:08and this PX is the probability of value.
  50607. 37:47:13For example, we could calculate the
  50608. 37:47:15expected value for this probability
  50609. 37:47:17distribution to be as shown.
  50610. 37:47:22So here it will be 1.45 goals.
  50611. 37:47:27So this represents the expected number
  50612. 37:47:29of goals that the team will score in any
  50613. 37:47:32given game.
  50614. 37:47:34And then if you talk about calculating
  50615. 37:47:36mean, so we typically calculate the mean
  50616. 37:47:39after we have actually collected raw
  50617. 37:47:41data.
  50618. 37:47:44For example, suppose we record the
  50619. 37:47:46number of goals that a soccer team will
  50620. 37:47:48score in 15 different games.
  50621. 37:47:53Now to calculate the mean number of
  50622. 37:47:55goals scored per game,
  50623. 37:47:58we can use the following formula
  50624. 37:48:01where sum of x is basically the sum of
  50625. 37:48:04all the goals divided by n and the
  50626. 37:48:06number of records or we can say the
  50627. 37:48:08sample size.
  50628. 37:48:11It is as shown on the screen.
  50629. 37:48:29So this represents the mean number of
  50630. 37:48:31goals scored per game by the team.
  50631. 37:48:34Measures of asymmetry.
  50632. 37:48:37The difference between the three
  50633. 37:48:38distinct curves can be studied in this
  50634. 37:48:41image.
  50635. 37:48:42The central curve is the normal or no
  50636. 37:48:45skeus curve. Here mean, median and mode
  50637. 37:48:48all lie on the same point. This normal
  50638. 37:48:51curve is symmetrical about its mean,
  50639. 37:48:53median and mode.
  50640. 37:48:57That means the left hand side of the
  50641. 37:48:59curve is a mirror image of the right
  50642. 37:49:01hand side of the curve.
  50643. 37:49:04However, in case of negatively skewed
  50644. 37:49:07data, the tail is elongated on the left
  50645. 37:49:11hand side
  50646. 37:49:12and the mean is smaller than the mode
  50647. 37:49:15and the median values or is on the left
  50648. 37:49:18hand side of the mode.
  50649. 37:49:21Hence indicating that the outliers are
  50650. 37:49:23in the negative direction.
  50651. 37:49:26On the other hand, in case of positively
  50652. 37:49:29skewed, the data is concentrated on the
  50653. 37:49:31left hand side of the curve.
  50654. 37:49:35While the tail is elongated or longer on
  50655. 37:49:37the right hand side of the curve,
  50656. 37:49:41the mean is greater than the mode and
  50657. 37:49:43median
  50658. 37:49:44or is on the right hand side of the mode
  50659. 37:49:47and median indicating that the outliers
  50660. 37:49:49are in the positive direction.
  50661. 37:49:56Let's consider an example.
  50662. 37:49:59The graph here shows the global income
  50663. 37:50:02distribution for the year 2003 2013 and
  50664. 37:50:06a projection for 2035.
  50665. 37:50:09If we see the global income distribution
  50666. 37:50:12statistics for 2003 it is highly right
  50667. 37:50:15skewed.
  50668. 37:50:18We can observe in the previous graph
  50669. 37:50:20that in 2003
  50670. 37:50:24the mean of $3,451
  50671. 37:50:29was higher than the median of $1090.
  50672. 37:50:33The global income is definitely not
  50673. 37:50:36evenly distributed. The majority of
  50674. 37:50:38people make less than $2,000 each year.
  50675. 37:50:44while only a small percentage of the
  50676. 37:50:46population earns more than $14,000.
  50677. 37:50:51Measures of variability.
  50678. 37:50:56Measures of variability.
  50679. 37:50:58Dispersion. The measure of central
  50680. 37:51:01tendencies provide a single value that
  50681. 37:51:03addresses the full worth. However, the
  50682. 37:51:06central tendency cannot depict the
  50683. 37:51:08viewpoint entirely. The metric of
  50684. 37:51:11dispersion helps us focus on the
  50685. 37:51:13inconsistency in the data spread.
  50686. 37:51:16Measures of dispersion describe the
  50687. 37:51:18spread of the data.
  50688. 37:51:21The range, intercortile range, standard
  50689. 37:51:24deviation and variance are examples of
  50690. 37:51:27dispersion measures.
  50691. 37:51:30Range.
  50692. 37:51:32The range of distribution is the
  50693. 37:51:34difference between the largest and the
  50694. 37:51:36smallest amount of data.
  50695. 37:51:39The range, for example, does not include
  50696. 37:51:42all of a series positive aspects.
  50697. 37:51:46It concentrates on the most shocking
  50698. 37:51:48aspects and ignores that aren't
  50699. 37:51:50considered critical. For example, for a
  50700. 37:51:52set 13, 33, 45, 67, 70.
  50701. 37:51:58The range is 57. That is the maximum of
  50702. 37:52:03this which is 70 minus the minimum over
  50703. 37:52:05here which is 13.
  50704. 37:52:10Variance.
  50705. 37:52:12Variance is the average of all squared
  50706. 37:52:14deviations.
  50707. 37:52:17It is defined as the sum of squared
  50708. 37:52:19distance between each point and the mean
  50709. 37:52:22or the dispersion around the mean.
  50710. 37:52:25The standard deviation is used as
  50711. 37:52:28variance suffers from a unit difference.
  50712. 37:52:32Variance can be computed as sigma square
  50713. 37:52:35summation over x - mu^ 2
  50714. 37:52:39divided by n
  50715. 37:52:41where mu is the mean of the data, x is
  50716. 37:52:45the individual data point
  50717. 37:52:48and n is the size of the data.
  50718. 37:52:52This representation is for a population
  50719. 37:52:54data.
  50720. 37:52:56for a sample data variance can be
  50721. 37:52:58computed as X minus
  50722. 37:53:01Xar whole square summation
  50723. 37:53:04over it divided by n minus one.
  50724. 37:53:08Here Xbar is the mean of these sample
  50725. 37:53:11data and n is the sample size.
  50726. 37:53:16The units of values and variance are not
  50727. 37:53:18equal.
  50728. 37:53:20So another variability measure is used.
  50729. 37:53:24Standard deviation.
  50730. 37:53:28Standard deviation is a statistical term
  50731. 37:53:30used to measure the amount of
  50732. 37:53:32variability or dispersion around a mean.
  50733. 37:53:37The standard deviation is calculated as
  50734. 37:53:40the square root of variance. It depicts
  50735. 37:53:43the concentration of the data around the
  50736. 37:53:46mean of the data set.
  50737. 37:53:50Standard deviation as indicated
  50738. 37:53:52previously can be computed as square
  50739. 37:53:54root of variance
  50740. 37:53:56for a population data. Standard
  50741. 37:53:59deviation sigma can be computed as
  50742. 37:54:02square root of summation over x i minus
  50743. 37:54:05mu^ square / n
  50744. 37:54:09where mu is the mean of the data x i are
  50745. 37:54:13the data points and n is the size. Let's
  50746. 37:54:16consider an example.
  50747. 37:54:19Let's find out the mean, variance, and
  50748. 37:54:22standard deviation for this data. The
  50749. 37:54:25data values are 3, 5, 6, 9, and 10. To
  50750. 37:54:30find out the mean, we first find the sum
  50751. 37:54:32of all these data values
  50752. 37:54:36that is 33 and divide it by the count,
  50753. 37:54:39which is five.
  50754. 37:54:42We get the mean of 6.6. To compute the
  50755. 37:54:45variance, we start by computing the
  50756. 37:54:48deviation.
  50757. 37:54:49That is X minus the mean of X. Here 3 is
  50758. 37:54:54one of the values of the data and 6.6 is
  50759. 37:54:57the mean.
  50760. 37:54:59So 3 - 6.6 squared and we do that
  50761. 37:55:05to find out sum of all the deviations
  50762. 37:55:07divided by the count
  50763. 37:55:10which is five.
  50764. 37:55:12We end up getting an overall variance of
  50765. 37:55:146.64.
  50766. 37:55:18Standard deviation as we know is
  50767. 37:55:21measured at square root of variance that
  50768. 37:55:23is square<unk> of 6.64
  50769. 37:55:27which amounts to 2.576.
  50770. 37:55:31Measures of relationship.
  50771. 37:55:34Measures of relationship coariance.
  50772. 37:55:37Covariance is the measure of joint
  50773. 37:55:39variability of two variables.
  50774. 37:55:43It measures the direction of the
  50775. 37:55:44relationship between the variables. It
  50776. 37:55:47determines if one variable will cause
  50777. 37:55:50the other to alter in the same way.
  50778. 37:55:54Coariance between variable X and Y can
  50779. 37:55:57be computed as summation over the
  50780. 37:56:00product of X I - XR
  50781. 37:56:03and Y I - Y bar the whole divided by N
  50782. 37:56:07minus one.
  50783. 37:56:10Here Xar and Y bar are the mean of X and
  50784. 37:56:13Y respectively. The value of covariance
  50785. 37:56:16can range from minus infinity to a plus
  50786. 37:56:19infinity.
  50787. 37:56:22Correlation. Correlation is normalized
  50788. 37:56:25coariance.
  50789. 37:56:28It measures the strength of association
  50790. 37:56:31between two variables. The most common
  50791. 37:56:33measure for correlation is the Pearson
  50792. 37:56:36correlation coefficient.
  50793. 37:56:38Correlation between two variables
  50794. 37:56:42X and Y can be measured with respect to
  50795. 37:56:44coariance as coariance between X
  50796. 37:56:48and Y divided by the standard deviation
  50797. 37:56:51of X and standard deviation of Y.
  50798. 37:56:55The value of correlation ranges from a
  50799. 37:56:58negative 1 to positive 1.
  50800. 37:57:02Types of correlation.
  50801. 37:57:06Correlation can be either a positive
  50802. 37:57:08correlation,
  50803. 37:57:10zero correlation or a negative
  50804. 37:57:12correlation.
  50805. 37:57:16The first picture over here represents a
  50806. 37:57:18perfect positive correlation
  50807. 37:57:22wherein a straight line with a positive
  50808. 37:57:24slope
  50809. 37:57:26is representing the relationship between
  50810. 37:57:28the two variables.
  50811. 37:57:31Zero correlation means that the line
  50812. 37:57:33representing the relationship between
  50813. 37:57:35the two variables is horizontal to the
  50814. 37:57:38xaxis.
  50815. 37:57:41Perfect negative correlation can be
  50816. 37:57:44represented by a straight line with a
  50817. 37:57:46negative slope.
  50818. 37:57:49Correlation equals to 1 implies a
  50819. 37:57:52positive relationship. That is when one
  50820. 37:57:55variable increases the other variable
  50821. 37:57:57also increases. A correlation value of
  50822. 37:58:00negative one implies a negative
  50823. 37:58:02relationship. That is when one variable
  50824. 37:58:05increases the other decreases.
  50825. 37:58:09The correlation coefficient of zero
  50826. 37:58:12shows that the variables are completely
  50827. 37:58:14independent of each other.
  50828. 37:58:17Let's consider an example.
  50829. 37:58:21Here we have two variables height and
  50830. 37:58:24weight.
  50831. 37:58:27To compute the correlation between
  50832. 37:58:29height and weight,
  50833. 37:58:31we use the correlation formula as
  50834. 37:58:33covariance of X
  50835. 37:58:36and Y divided by standard deviation of X
  50836. 37:58:39and standard deviation of Y.
  50837. 37:58:42Here height is the X variable and weight
  50838. 37:58:45is the Y variable.
  50839. 37:58:48First to compute coariance we compute
  50840. 37:58:51the x - xar and y - y bar values and
  50841. 37:58:56then the product of them.
  50842. 37:58:59We then compute x - xr²
  50843. 37:59:04and y - y bar square values to compute
  50844. 37:59:07the standard deviations of height and
  50845. 37:59:09weight respectively. Correlation as we
  50846. 37:59:12know has been defined as covariance of X
  50847. 37:59:16and I and Y divided by standard
  50848. 37:59:18deviations of X and Y.
  50849. 37:59:22This can also be represented as
  50850. 37:59:24summation over x - xr multiplied to y -
  50851. 37:59:29y bar
  50852. 37:59:31divided by square root of summation over
  50853. 37:59:33sum of squared deviations that is x - xr
  50854. 37:59:38square multiplied to square root of
  50855. 37:59:40summation over y - y bar square that is
  50856. 37:59:45sum of square deviations for y.
  50857. 37:59:49Now let's find out values to put into
  50858. 37:59:52this formula.
  50859. 37:59:55First we find out the overall sum of
  50860. 37:59:58height to get the mean of height which
  50861. 38:00:01is 5.14.
  50862. 38:00:03Similarly we get the sum of weight to
  50863. 38:00:05get the mean of weight as 50. We now get
  50864. 38:00:08the summation over x - xr multiplied to
  50865. 38:00:12y - y bar to get the numerator for the
  50866. 38:00:15formula. Then we compute x - xr square
  50867. 38:00:19summation
  50868. 38:00:21and y - y bar square that is sum of
  50869. 38:00:24squared deviation of x and y
  50870. 38:00:27respectively.
  50871. 38:00:29Now we put in the values in this final
  50872. 38:00:31correlation formula to get a correlation
  50873. 38:00:33value of 0.889.
  50874. 38:00:38This indicates that height and weight
  50875. 38:00:40have a positive relationship.
  50876. 38:00:43It is evident that as height grows,
  50877. 38:00:46weight also increases.
  50878. 38:00:50In this module, we will be talking about
  50879. 38:00:52expectation and variance.
  50880. 38:00:56So the expected value or we can say mean
  50881. 38:00:58of a given variable that we can denote
  50882. 38:01:01by X is a discrete random variable where
  50883. 38:01:04it is a weighted average of the possible
  50884. 38:01:06values that X can take and each value is
  50885. 38:01:09going to be according to the probability
  50886. 38:01:11of that specific event occurring.
  50887. 38:01:15So usually the expected value of X is
  50888. 38:01:17denoted by a simple formula where we can
  50889. 38:01:21define the expectation based on the X
  50890. 38:01:23parameter
  50891. 38:01:26which is going to be the sum of each
  50892. 38:01:28possible outcome multiplied by the
  50893. 38:01:30probability of the outcome occurring.
  50894. 38:01:34So in more concrete terms, the
  50895. 38:01:37expectation is what we would expect the
  50896. 38:01:39outcome of an experiment to be on
  50897. 38:01:41average.
  50898. 38:01:45We can take an example for the coin. If
  50899. 38:01:48a coin is being tossed 10 times, then
  50900. 38:01:51one is most likely to get five heads and
  50901. 38:01:54five tails.
  50902. 38:01:57Same logic can be discussed if we talk
  50903. 38:02:00about another example of rolling a
  50904. 38:02:02dieice. So there are six possible
  50905. 38:02:04outcomes when you roll a dieice. 1 2 3 4
  50906. 38:02:085 6. And each of these has a probability
  50907. 38:02:11of 1x 6 of occurring. So we can say that
  50908. 38:02:15the expectation is going to be 1
  50909. 38:02:17multiplied by the probability of that
  50910. 38:02:19happening which is going to be 1x 6 + 2x
  50911. 38:02:236 + 3x 6 + 4x 6 + 5x 6 + 6x 6 and that
  50912. 38:02:31is going to give us 3.5 as an output.
  50913. 38:02:34The expected value is 3.5.
  50914. 38:02:38So if you think about it, 3.5 is halfway
  50915. 38:02:41between the possible values that I can
  50916. 38:02:43take and this is what we should have
  50917. 38:02:46expected.
  50918. 38:02:47Next we talk about the concept of
  50919. 38:02:49variance. So variance of a random
  50920. 38:02:52variable allows us to know something
  50921. 38:02:54about the spread of the possible values
  50922. 38:02:57of the variable.
  50923. 38:02:59So for a discrete random variable X, the
  50924. 38:03:02variances of X is going to be denoted by
  50925. 38:03:04using a simple formula that is going to
  50926. 38:03:06be var=
  50927. 38:03:08E X - M the whole square where M is
  50928. 38:03:12basically the expected value of the
  50929. 38:03:14expectation of X. So this is more like a
  50930. 38:03:17standard deviation of X which can also
  50931. 38:03:20be represented by using this formula. So
  50932. 38:03:23the variance does not behave in the same
  50933. 38:03:25way as expectation when we multiply and
  50934. 38:03:28add constants to random variables.
  50935. 38:03:33So now there are two different type of
  50936. 38:03:35variance that we can have a fair
  50937. 38:03:37understanding on. First of all we have
  50938. 38:03:40low variance and then we have high
  50939. 38:03:43variance.
  50940. 38:03:45So low variance simply means that there
  50941. 38:03:48is a small variation in the production
  50942. 38:03:50of the target function with changes in
  50943. 38:03:53the trading data set and at the same
  50944. 38:03:55time high variance as we can see here
  50945. 38:03:58high variance shows a large variation in
  50946. 38:04:01prediction of the target function with
  50947. 38:04:03changes in the trading data set. So a
  50948. 38:04:06model that shows high variance learns a
  50949. 38:04:08lot and perform well with the training
  50950. 38:04:10data set and it does not generalize well
  50951. 38:04:13with the unseen data set and that's why
  50952. 38:04:16as a result such a model gives good
  50953. 38:04:19results with training data set but shows
  50954. 38:04:21high error rates on the test data set
  50955. 38:04:24and since the high variance a model
  50956. 38:04:26learns too much from the data set it
  50957. 38:04:29leads to an overfitting of the model. So
  50958. 38:04:32model with high variance will be having
  50959. 38:04:34couple of issues like it may lead to
  50960. 38:04:36overfitting or it may also lead to
  50961. 38:04:39increase in model complexities.
  50962. 38:04:43Next we have skewess.
  50963. 38:04:46So skewess in simple terms is basically
  50964. 38:04:49a measure of asymmetry of a
  50965. 38:04:51distribution. So distribution is
  50966. 38:04:54asymmetrical when its left and right
  50967. 38:04:56sides are not the mirror images.
  50968. 38:04:59Right now this is a mirrored image and a
  50969. 38:05:02distribution can have right positive or
  50970. 38:05:04we can say negative or it can have zero
  50971. 38:05:07skewess.
  50972. 38:05:10So right skewed in this scenario is
  50973. 38:05:13basically the distribution is longer on
  50974. 38:05:15the right side of its peak
  50975. 38:05:18and a left skew distribution is going to
  50976. 38:05:20be we can say where it is longer on the
  50977. 38:05:23left side.
  50978. 38:05:25So we can see we have this one as a part
  50979. 38:05:28of right side. It is more elongated
  50980. 38:05:30towards the right side and this one is
  50981. 38:05:33more elongated towards the left side. So
  50982. 38:05:36we can think of skewess in terms of
  50983. 38:05:38tails. A tail is long tampering and the
  50984. 38:05:41end of a distribution. So it simply
  50985. 38:05:44indicates that they are observations at
  50986. 38:05:46one end of the distribution but that
  50987. 38:05:48they are relatively infrequent. So a
  50988. 38:05:51right skew distribution has a long tail
  50989. 38:05:54on the right side as you can see here.
  50990. 38:05:56So the number supports observed. Let's
  50991. 38:05:59say we have a data on a per year basis.
  50992. 38:06:02So again we can have a more skewess
  50993. 38:06:04towards the right side where data is
  50994. 38:06:06being dropping as we continue to
  50995. 38:06:09increase the number of years. For
  50996. 38:06:11example we may have a high sales towards
  50997. 38:06:14the beginning of year suppose in 2022
  50998. 38:06:17but again as we proceed to 2023 second
  50999. 38:06:21half we are seeing the dip in
  51000. 38:06:23performance. So that is rightly skewed
  51001. 38:06:26and same way let's suppose if we started
  51002. 38:06:28with the sales figure it was really less
  51003. 38:06:31in suppose 2002
  51004. 38:06:34but again as we proceeded to 2023 now
  51005. 38:06:37our sales have been gradually
  51006. 38:06:39increasing. So it's more like skew
  51007. 38:06:42towards the left section as a part of
  51008. 38:06:44negative skew. Next we have curtosis.
  51009. 38:06:49So curtosis is basically a measure of
  51010. 38:06:52the tailness of a distribution.
  51011. 38:06:55So taeness is how often the outliers
  51012. 38:06:58occur and act as curtis is the tailness
  51013. 38:07:01of the distribution related to a normal
  51014. 38:07:04distribution. So a distribution with
  51015. 38:07:07medium curttosis is called as meocurtic.
  51016. 38:07:10A distribution with low curtosis like
  51017. 38:07:12this one. This is called as the
  51018. 38:07:15platicurtic and then distribution with
  51019. 38:07:17high curtosis like this one. This is
  51020. 38:07:19called as the leptoccuric.
  51021. 38:07:23So tails here they are tapering ends on
  51022. 38:07:25either side of a distribution like this.
  51023. 38:07:28So they represent the probability or the
  51024. 38:07:31frequency of values that are extremely
  51025. 38:07:33high or extremely low to the mean.
  51026. 38:07:36In other words, tails here represents
  51027. 38:07:39how often the outliers occur.
  51028. 38:07:43So there are three type of curtis. We
  51029. 38:07:46have platocurtic which is negative,
  51030. 38:07:48leptocortic which is a positive towards
  51031. 38:07:50the upper end and then we have messertic
  51032. 38:07:53which is a normal distribution. So
  51033. 38:07:56messertic is the medium tail. So normal
  51034. 38:07:59distributions they have a curtosis of
  51035. 38:08:01three. So any distribution with a curtis
  51036. 38:08:04of a prox value of three is going to be
  51037. 38:08:07messertic. And curtosis is described in
  51038. 38:08:10terms of excess curttosis which is
  51039. 38:08:13curtosis minus3. And since normal
  51040. 38:08:16distribution they have a curtosis of
  51041. 38:08:19three axis curtises makes comparing a
  51042. 38:08:22distribution curtosis to a normal
  51043. 38:08:24distribution even easier. Introduction
  51044. 38:08:27to probability.
  51045. 38:08:30Probability theory. Probability is a
  51046. 38:08:33measure of the likelihood that an event
  51047. 38:08:35will occur.
  51048. 38:08:38Let's consider an example of coin toss
  51049. 38:08:42where the chances of getting heads on a
  51050. 38:08:44coin are 1 by two or 50%.
  51051. 38:08:48The probability of each given event is
  51052. 38:08:50between zero and one both inclusive.
  51053. 38:08:54Sum of an events cumulative probability
  51054. 38:08:57cannot be greater than one.
  51055. 38:09:00Hence the probability of an event x lies
  51056. 38:09:03between zero and one. This means that
  51057. 38:09:06the integral of probability of
  51058. 38:09:08distribution over x equals to 1.
  51059. 38:09:14Conditional probability. Conditional
  51060. 38:09:16probability of any event A is defined as
  51061. 38:09:20the probability of occurrence of A given
  51062. 38:09:23that event B has previously occurred.
  51063. 38:09:28Condition probability of event A given B
  51064. 38:09:31can be estimated as probability of A
  51065. 38:09:34intersection B that is probability of
  51066. 38:09:37both A and B happening together
  51067. 38:09:40divided by the probability of B.
  51068. 38:09:45It is also written as that probability
  51069. 38:09:47of A intersection B equals to
  51070. 38:09:50probability of A given B multiplied to
  51071. 38:09:54probability of B.
  51072. 38:09:59Let's consider an example.
  51073. 38:10:01In a coin, we are doing a two coin flip.
  51074. 38:10:04Coin one gets heads, tails, heads, and
  51075. 38:10:07tails in subsequent flips.
  51076. 38:10:12while coin 2 gets tails, heads, heads,
  51077. 38:10:15and tails in the subsequent flips. Now,
  51078. 38:10:18the probability that coin one will get a
  51079. 38:10:20head is 2 out of four. While the
  51080. 38:10:23probability that coin two will get heads
  51081. 38:10:26is again two out of four.
  51082. 38:10:29The probability that both coin one and
  51083. 38:10:31coin two will have a heads is just one
  51084. 38:10:34out of the four flips.
  51085. 38:10:38Hence the probability that coin one will
  51086. 38:10:40get heads given that coin 2 is already
  51087. 38:10:43heads can be computed as probability of
  51088. 38:10:46coin one edge intersection coin 2 edge
  51089. 38:10:50that is 1x4 divided by probability of
  51090. 38:10:54coin 2 edge
  51091. 38:10:57that's a given that is 2x 4 which is
  51092. 38:11:00going to be 0.5 or 50% based
  51093. 38:11:05base theorem Base theorem calculates the
  51094. 38:11:08conditional probability of an event
  51095. 38:11:10based on its prior probabilities.
  51096. 38:11:14Basically base theorem incorporates the
  51097. 38:11:17prior probability distribution to
  51098. 38:11:19predict the posterior probabilities base
  51099. 38:11:22theorem for conditional probability
  51100. 38:11:25can be expressed as probability of A
  51101. 38:11:28given B equals probability of B given A
  51102. 38:11:32divided by probability of B multiplied
  51103. 38:11:35to probability of A.
  51104. 38:11:38Base theorem allows updating the
  51105. 38:11:40probability values by using new
  51106. 38:11:42information or evidence. Here
  51107. 38:11:45probability of A is known as prior
  51108. 38:11:48probability. That is the probability of
  51109. 38:11:50event that before any new data is
  51110. 38:11:52collected. Probability of A given B is
  51111. 38:11:56known as the posterior probability. It
  51112. 38:11:59is the revised probability of an event
  51113. 38:12:01occurring after taking into
  51114. 38:12:03consideration the new information
  51115. 38:12:05probability of B given A is known as the
  51116. 38:12:09likelihood and probability of B is
  51117. 38:12:11probability of observing an evidence B
  51118. 38:12:14model. An example consider an example
  51119. 38:12:18for calculating the likelihood of having
  51120. 38:12:20diabetes based on frequency of fast food
  51121. 38:12:23consumption. Here is the observed data.
  51122. 38:12:26Let's say the fast food audience is 20%.
  51123. 38:12:30Diabetes prevalence is 10% and 5% is
  51124. 38:12:34fast food and diabetes.
  51125. 38:12:37The chances of diabetes given fast food
  51126. 38:12:39that is the conditional probability of D
  51127. 38:12:42given B can be calculated as probability
  51128. 38:12:45of diabetes and fast food together
  51129. 38:12:48divided by probability of fast food.
  51130. 38:12:51That means 5% divided by 20%. that
  51131. 38:12:55equals 25%.
  51132. 38:12:57Define an analysis can state eating fast
  51133. 38:13:00food increases the chance of having
  51134. 38:13:02diabetes by 25%.
  51135. 38:13:05The multiplication rule of probability
  51136. 38:13:08if events A and B are statistically
  51137. 38:13:11independent and probability of A
  51138. 38:13:14intersection B can be given as
  51139. 38:13:16probability of A given B multiplied to
  51140. 38:13:20probability of B. However, probability
  51141. 38:13:23of A intersection B is also given as
  51142. 38:13:26probability of A multiplied to
  51143. 38:13:29probability of B. Here probability of A
  51144. 38:13:33given B equals to probability of A when
  51145. 38:13:37we assume that probability of B is non
  51146. 38:13:40zero. Similarly, probability of B equals
  51147. 38:13:43probability of B given A assuming
  51148. 38:13:46probability of A is non zero. Chain rule
  51149. 38:13:50of probability joint probability
  51150. 38:13:52distributions over many random variables
  51151. 38:13:55can be reduced into conditional
  51152. 38:13:57distributions over a single variable. It
  51153. 38:14:00can be expressed as probability of X1 X2
  51154. 38:14:04so on until XN equals probability of X1
  51155. 38:14:08intersection probability of X I given
  51156. 38:14:11probability of X1 till X I minus one.
  51157. 38:14:16For example, the joint probability of A,
  51158. 38:14:19B and C can be given as probability of A
  51159. 38:14:23given B. C multiplied to probability of
  51160. 38:14:27B given C multiply to probability of C.
  51161. 38:14:32Logistic sigmoid.
  51162. 38:14:35The logistics function is a type of
  51163. 38:14:37sigmoid function that aims to predict
  51164. 38:14:39the class to which a particular sample
  51165. 38:14:42belongs. Its outcome is discrete binary
  51166. 38:14:45value. a probability between zero and
  51167. 38:14:48one. The logistic sigmoid is a useful
  51168. 38:14:51function that follows the yes curve. It
  51169. 38:14:54saturates when the input is very large
  51170. 38:14:56or very small. Logistic sigmoid is
  51171. 38:15:00expressed as sigma of x= 1 upon 1 + e to
  51172. 38:15:04the power minus x.
  51173. 38:15:07The logistic sigmoid can be expressed as
  51174. 38:15:10sigmoid function of x is given as 1 upon
  51175. 38:15:131 + e ^ minus x where e is the ooler's
  51176. 38:15:17number.
  51177. 38:15:19Gshian distribution.
  51178. 38:15:22The gossian distribution is a type of
  51179. 38:15:24distribution in which data tends to
  51180. 38:15:26cluster around a central value with
  51181. 38:15:29little or no bias to the left or right.
  51182. 38:15:33It is often referred to as normal
  51183. 38:15:35distribution.
  51184. 38:15:37In absence of prior information, the
  51185. 38:15:40normal distribution is frequently a fair
  51186. 38:15:42assumption in machine learning
  51187. 38:15:45equation.
  51188. 38:15:47The formula for calculating Gaussian
  51189. 38:15:49distribution is described as the normal
  51190. 38:15:52distribution of X.
  51191. 38:15:55That is the function of x given mean as
  51192. 38:15:57mu and variance is sigma square can be
  51193. 38:16:00calculated as 1 upon sigma square
  51194. 38:16:03roo<unk> of 2 pi e to the power -/ x -
  51195. 38:16:07mood / sigma square
  51196. 38:16:11where mu is the mean or peak value which
  51197. 38:16:14also is the expected value of x.
  51198. 38:16:18Sigma is the standard deviation. Sigma
  51199. 38:16:21square is the variance.
  51200. 38:16:23A standard normal distribution has a
  51201. 38:16:26mean of zero and a standard deviation of
  51202. 38:16:28one.
  51203. 38:16:31Goshan distribution can be univariate
  51204. 38:16:35which describes the distribution of a
  51205. 38:16:37single variable X.
  51206. 38:16:39It can also be multivariat where it can
  51207. 38:16:42just use to describe the distribution of
  51208. 38:16:44several variables.
  51209. 38:16:47It is represented in 3D of ND formats.
  51210. 38:16:53Law of large numbers.
  51211. 38:16:57Now let's talk about law of large
  51212. 38:16:59numbers. The law of large numbers states
  51213. 38:17:02that an observed sample average from a
  51214. 38:17:05large sample will be close to the true
  51215. 38:17:07population average and that it will get
  51216. 38:17:10closer in the larger sample. So the law
  51217. 38:17:12of large number does not guarantee that
  51218. 38:17:15a given sample spatially a small sample
  51219. 38:17:17will reflect the true population
  51220. 38:17:19characteristics or that a sample does
  51221. 38:17:22not reflect the true population will be
  51222. 38:17:24balanced by a subsequent sample. This is
  51223. 38:17:27for the law of large numbers to express
  51224. 38:17:30the relationship between scale and
  51225. 38:17:32growth rate.
  51226. 38:17:35So there are multiple examples through
  51227. 38:17:37which we can understand
  51228. 38:17:41and it is widely used in statistical
  51229. 38:17:43analysis in working with the central
  51230. 38:17:45limit theorem in terms of the business
  51231. 38:17:47growth. So there are multiple real time
  51232. 38:17:50setup in which these are going to be
  51233. 38:17:52used. So if you talk about tossing a
  51234. 38:17:55coin so tossing a coin in a number of
  51235. 38:17:58times will give us two different type of
  51236. 38:18:00outcomes.
  51237. 38:18:03the result will spread evenly between
  51238. 38:18:05head and tails and the expected average
  51239. 38:18:08value is going to be half.
  51240. 38:18:10That means 50 times tails and 30 times
  51241. 38:18:13heads. But again, if you toss a coin
  51242. 38:18:161,000 times, then the result can be in
  51243. 38:18:19different manners because out of 1,000,
  51244. 38:18:22let's say 850 times it has been head and
  51245. 38:18:26only 150 times it has been tails and so
  51246. 38:18:30on. So that's why the possibility of one
  51247. 38:18:32event occurring is going to be changed
  51248. 38:18:35in large sample sets as compared to a
  51249. 38:18:37small sample sets as in let's say 10
  51250. 38:18:40times. So the number of heads and tails
  51251. 38:18:43unbalanced for lower number of trials.
  51252. 38:18:45So we can see it is unbalanced.
  51253. 38:18:49But again as soon as we toss more number
  51254. 38:18:51of coins more leans towards the balance
  51255. 38:18:54value or we can see the observed
  51256. 38:18:56averages.
  51257. 38:18:58Next we have p value.
  51258. 38:19:01So p value is basically a number
  51259. 38:19:04calculated from the statistical test
  51260. 38:19:07that describes how likely we are to have
  51261. 38:19:09found a particular set of observations
  51262. 38:19:11if the null hypothesis were true. So p
  51263. 38:19:15values are used in hypothesis testing to
  51264. 38:19:18help decide whether to reject the null
  51265. 38:19:20hypothesis.
  51266. 38:19:22And the smaller the p value, the more
  51267. 38:19:24likely we are to reject the null
  51268. 38:19:26hypothesis.
  51269. 38:19:27So we have a term called as null
  51270. 38:19:30hypothesis. So all statistical tests
  51271. 38:19:33they have null hypothesis. So for most
  51272. 38:19:36tests the null hypothesis is that there
  51273. 38:19:38is no relationship between our variables
  51274. 38:19:41of in first or that there is no
  51275. 38:19:43difference among groups. For example in
  51276. 38:19:46a two-taile t test the non-hypothesis is
  51277. 38:19:49that the difference between two groups
  51278. 38:19:51is going to be zero.
  51279. 38:19:54So p value is going to tell us how
  51280. 38:19:56likely it is that our data could have
  51281. 38:19:58occurred under the null hypothesis.
  51282. 38:20:02It is done by calculating the likelihood
  51283. 38:20:04of a test statistic
  51284. 38:20:06which is the number calculated by a
  51285. 38:20:08statistical test using our data. So p
  51286. 38:20:12value tell us how often we would expect
  51287. 38:20:14to see a test statistic as extreme or
  51288. 38:20:17more extreme
  51289. 38:20:18than one calculated by a statistical
  51290. 38:20:21test. if the null hypothesis of the test
  51291. 38:20:24was true.
  51292. 38:20:26So there are multiple limitations as
  51293. 38:20:28well. So first one is the results can be
  51294. 38:20:31significant but again they are they may
  51295. 38:20:34not be practical as we have compared it
  51296. 38:20:37can be based on multiple hypothesis for
  51297. 38:20:39a game for the healthcare test. If the
  51298. 38:20:42test is going to be positive or not it
  51299. 38:20:45may show even values of the effect of a
  51300. 38:20:47variable but not the magnitude in real
  51301. 38:20:50life. What exactly is going to be the
  51302. 38:20:52application of a drug test being failed
  51303. 38:20:55in pharma company? Therefore, it is
  51304. 38:20:58recommended to use confidence and levels
  51305. 38:21:00in addition to the p values to quantify
  51306. 38:21:03or we can say to give a solid figure to
  51307. 38:21:05the reserve which we are going to get.
  51308. 38:21:08The p values they are interpreted as
  51309. 38:21:11supporting or we can say refuting the
  51310. 38:21:13alternative hypothesis.
  51311. 38:21:15So p value can only tell you whether or
  51312. 38:21:18not the null hypothesis is supported. It
  51313. 38:21:21cannot tell us whether our alternative
  51314. 38:21:23hypothesis is true or why. So the risk
  51315. 38:21:27of rejecting the null hypothesis is
  51316. 38:21:30often higher than the p value. So
  51317. 38:21:33especially when we are looking at a
  51318. 38:21:34single study or when using small sample
  51319. 38:21:37sizes. So this is because the smaller
  51320. 38:21:40frame of reference, the greater are the
  51321. 38:21:42chance that as we stumble across a
  51322. 38:21:45statistically significant pattern
  51323. 38:21:47completely by accident.
  51324. 38:21:49Key takeaways.
  51325. 38:21:51Key takeaways. Probability and
  51326. 38:21:54statistics structure the premise of the
  51327. 38:21:56data. The data helps in anticipating the
  51328. 38:22:00future or gauging in view of the past
  51329. 38:22:03patterns of information.
  51330. 38:22:05The central tendency is a single value
  51331. 38:22:08that helps to describe the data by
  51332. 38:22:10identifying these central positions. The
  51333. 38:22:12mean, median and mode are the measures
  51334. 38:22:15of central tendencies.
  51335. 38:22:18The distribution where the data tends to
  51336. 38:22:20be around a central value with a lack of
  51337. 38:22:23bias or minimal bias towards the left or
  51338. 38:22:26right is called as gshian distribution.
  51339. 38:22:30So now let's dive into the definition of
  51340. 38:22:32the probability distribution function.
  51341. 38:22:36What is probability distribution
  51342. 38:22:38function? A function which defines the
  51343. 38:22:41relationship between a random variable
  51344. 38:22:43and its probability such that you can
  51345. 38:22:45find the probability of the variable
  51346. 38:22:47using the function is called a
  51347. 38:22:49probability density function.
  51348. 38:22:52In simple words, probability density is
  51349. 38:22:55the relationship between an observation
  51350. 38:22:58and the probability. Some outcomes of a
  51351. 38:23:00random variable will have low
  51352. 38:23:02probability density and other outcomes
  51353. 38:23:04will have a very high probability
  51354. 38:23:06density. Basically, the probability of a
  51355. 38:23:09variable X happening or occurring will
  51356. 38:23:12vary and it can sometimes take on a
  51357. 38:23:15lower value or it can take on a way
  51358. 38:23:16higher value.
  51359. 38:23:18The overall shape of the probability
  51360. 38:23:20density is referred to as probability
  51361. 38:23:22distribution. And the calculation of
  51362. 38:23:24probabilities for specific outcomes of a
  51363. 38:23:27random variable is performed by a
  51364. 38:23:29probability density function or PDF for
  51365. 38:23:32short. Now consider a variable X with a
  51366. 38:23:35PDF of f ofx.
  51367. 38:23:39This is what your probability density
  51368. 38:23:41function will look like. There might be
  51369. 38:23:42a point where the probability of X
  51370. 38:23:45occurring is very high. Hence your
  51371. 38:23:49probability distribution function or f
  51372. 38:23:51ofx will also be very high. At other
  51373. 38:23:54points the distribution or the
  51374. 38:23:55probability of X happening or occurring
  51375. 38:23:58is going to be very low. Hence your f
  51376. 38:24:00ofx is also going to have a very small
  51377. 38:24:03value. Basically given the random sample
  51378. 38:24:06of a variable we might want to know
  51379. 38:24:08things like the shape of the probability
  51380. 38:24:10distribution. This here is something
  51381. 38:24:12called a normal distribution where a
  51382. 38:24:14probability distribution function takes
  51383. 38:24:16on a bell shape.
  51384. 38:24:19However, this is not the probability
  51385. 38:24:22density function that might always
  51386. 38:24:23occur. There are different probability
  51387. 38:24:25distribution functions and all of their
  51388. 38:24:27graphs look very different from each
  51389. 38:24:29other. Knowing the probability
  51390. 38:24:31distribution for a random variable can
  51391. 38:24:33help you calculate movements of the
  51392. 38:24:35distribution like the mean and variance.
  51393. 38:24:37But it can also be useful for other more
  51394. 38:24:40general considerations like determining
  51395. 38:24:42whether an observation is unlikely or
  51396. 38:24:44very unlikely and might be an outlier or
  51397. 38:24:47an anomaly like consider this graph
  51398. 38:24:50itself. In this graph, these points over
  51399. 38:24:54here which have very less probability
  51400. 38:24:57distribution
  51401. 38:24:58are outliers which means that the chance
  51402. 38:25:01of them occurring is very low. And
  51403. 38:25:04basically this is not something that
  51404. 38:25:06you're going to see in your regular
  51405. 38:25:08scenario for your variable X. Now let's
  51406. 38:25:11consider two points A and B which are
  51407. 38:25:13values that a variable X can take. P of
  51408. 38:25:17A and P of B just represent the
  51409. 38:25:19probability of A and the probability of
  51410. 38:25:22B which can be found out by drawing a
  51411. 38:25:25straight line and coinciding it with our
  51412. 38:25:27graphs. The area under the graph over
  51413. 38:25:30here which is going to give you your
  51414. 38:25:33probability of this region occurring can
  51415. 38:25:36be written as probability of A less than
  51416. 38:25:40equal to X which is a probability that
  51417. 38:25:42we're searching for here less than equal
  51418. 38:25:44to probability of B. What does this mean
  51419. 38:25:47exactly? This means that this area is
  51420. 38:25:51always going to be greater than or equal
  51421. 38:25:53to the probability of A but less than or
  51422. 38:25:56equal to the probability of B. This
  51423. 38:25:58gives us the narrow region
  51424. 38:26:01of the probability which is present over
  51425. 38:26:03here. And doing this we can find the
  51426. 38:26:06probability of occurrence for any value
  51427. 38:26:09of X. Suppose you want to find the
  51428. 38:26:12probability of B happening. For a
  51429. 38:26:15probability distribution function, the
  51430. 38:26:18probability of B happening is not simply
  51431. 38:26:21this point here, but the entire area of
  51432. 38:26:24the graph which is taking place before
  51433. 38:26:27this point itself. So if you want to
  51434. 38:26:30find the probability between these
  51435. 38:26:32regions, you're going to have to find
  51436. 38:26:33the entire area and not simply the
  51437. 38:26:36probability at one point.
  51438. 38:26:39Now, so far we've been talking about
  51439. 38:26:41different types of variables which is
  51440. 38:26:42discrete random variables and continuous
  51441. 38:26:44variables. What exactly do these mean? A
  51442. 38:26:48variable which can only take a value
  51443. 38:26:50within a certain range is called a
  51444. 38:26:52discrete random variable. The value is
  51445. 38:26:55usually within a certain distance of
  51446. 38:26:57another finite value. An example of this
  51447. 38:27:00would be the sum of two dices. Basically
  51448. 38:27:03values which are well defined are called
  51449. 38:27:05discrete values or and a variable which
  51450. 38:27:08has well- definfined values will be
  51451. 38:27:10called a discrete random variable.
  51452. 38:27:13This variable can only take values which
  51453. 38:27:15fall within a certain set of values.
  51454. 38:27:19Let's say you roll a dice. The dice can
  51455. 38:27:21only give you specific outcomes which
  51456. 38:27:23range from 1 to six. This is what you
  51457. 38:27:26would call a discrete output.
  51458. 38:27:29On the other hand, a continuous random
  51459. 38:27:31variable can take on infinite different
  51460. 38:27:33values within a range of values. For
  51461. 38:27:36example, the height of a student. The
  51462. 38:27:38height of a student is not fixed. Even
  51463. 38:27:40if the height is 1.7 m, in reality, the
  51464. 38:27:46height can be 1.77 or 1.765
  51465. 38:27:50or 1.789.
  51466. 38:27:52The exact height is very hard to
  51467. 38:27:54determine because it's not easy for us
  51468. 38:27:56to find the precise value of the height
  51469. 38:28:00of a student. So basically the height
  51470. 38:28:02can take on an infinite different range
  51471. 38:28:04of values. When we're trying to define
  51472. 38:28:06the values that a continuous random
  51473. 38:28:08variable can take, we usually say it in
  51474. 38:28:11the form of a range of values which
  51475. 38:28:13means that the value can fall in that
  51476. 38:28:15range and can take on any value in that
  51477. 38:28:18range. It's not like a discrete random
  51478. 38:28:20variable where you can define definitive
  51479. 38:28:23values.
  51480. 38:28:26Now let's understand a probability
  51481. 38:28:27density function with the help of a
  51482. 38:28:29graph. Consider the graph below which
  51483. 38:28:31shows the rainfall distribution in an
  51484. 38:28:33year in a city. The x-axis has the
  51485. 38:28:36rainfall in inches or the amount of rain
  51486. 38:28:38that we're getting and the y-axis has
  51487. 38:28:40the probability density function of
  51488. 38:28:42getting that amount of rain.
  51489. 38:28:45The probability of some amount of
  51490. 38:28:47rainfall is obtained by finding the area
  51491. 38:28:50of the curve to the left of it. So let's
  51492. 38:28:53say we have a 3.
  51493. 38:28:56If you want to find the probability of 3
  51494. 38:28:59in of rainfall occurring, we would have
  51495. 38:29:01to find the area of the curve
  51496. 38:29:05which falls to the left of three. When
  51497. 38:29:08we draw a line from three which
  51498. 38:29:10intercepts the graph and further extend
  51499. 38:29:12it onto the yaxis, we get a value of
  51500. 38:29:160.5.
  51501. 38:29:18Simply put, this means that the
  51502. 38:29:20probability of 3 in of rainfall
  51503. 38:29:23occurring is going to be lesser than or
  51504. 38:29:25equal to 0.5. The exact probability can
  51505. 38:29:29be found out by finding the area of the
  51506. 38:29:31curve
  51507. 38:29:34which falls to the left of three.
  51508. 38:29:38How do we find the probability
  51509. 38:29:40distribution function?
  51510. 38:29:43The first step is to summarize your
  51511. 38:29:45density with the help of a histogram.
  51512. 38:29:47The first step in a density estimation
  51513. 38:29:49is to create a histogram of the
  51514. 38:29:51observations
  51515. 38:29:53in the random sample.
  51516. 38:29:56Now what is a histogram? A histogram is
  51517. 38:29:58a plot which involves first grouping the
  51518. 38:30:00observation into bins and counting the
  51519. 38:30:03number of events that fall in each bin.
  51520. 38:30:06The counts or frequency of observation
  51521. 38:30:08in each bins are then plotted as a bar
  51522. 38:30:10graph with the bins on the x-axis and
  51523. 38:30:13the frequency on the y-axis. The choice
  51524. 38:30:16of the number of bins is important as it
  51525. 38:30:18controls the coarseness of the
  51526. 38:30:19distribution and in turn how well the
  51527. 38:30:22density of the observation is plotted.
  51528. 38:30:24It is a good idea to experiment with
  51529. 38:30:26different bin sizes for a given data
  51530. 38:30:28sample to get multiple perspectives or
  51531. 38:30:31views on the same data.
  51532. 38:30:35At the same time, the number of bins is
  51533. 38:30:37important as it determines how many bars
  51534. 38:30:40the histogram will have and their
  51535. 38:30:41widths. This will change not only the
  51536. 38:30:44shape of the graph but also how the
  51537. 38:30:46graph is read. This will also determine
  51538. 38:30:49how our density is plotted. Now let's
  51539. 38:30:51see how we can summarize our density
  51540. 38:30:53with histograms using Python. First
  51541. 38:30:56let's import all of our necessary
  51542. 38:30:58modules which we're going to require.
  51543. 38:31:00We're going to require Mattplot lib to
  51544. 38:31:02plot graphs. We're going to need the
  51545. 38:31:05normal random function so that we can
  51546. 38:31:07get a normal distribution. We're going
  51547. 38:31:09to import mean and standard deviation
  51548. 38:31:12from numpy to use on our graphs and also
  51549. 38:31:15going to normalize our uh data. So we're
  51550. 38:31:18going to import the nom function from
  51551. 38:31:20sci.
  51552. 38:31:25We finished importing all of our
  51553. 38:31:27necessary modules. Now let's generate a
  51554. 38:31:30sample
  51555. 38:31:32which has a size of thousand and it's
  51556. 38:31:34going to be a normal distribution. And
  51557. 38:31:36we're going to also plot this with the
  51558. 38:31:38help of a histogram in bins of 10.
  51559. 38:31:44So as you can see here you get a normal
  51560. 38:31:46distribution which is nothing but a
  51561. 38:31:48almost bell-shaped curve and we have 10
  51562. 38:31:52bins here which are centered at zero and
  51563. 38:31:54which extend from minus3 to 3.
  51564. 38:31:59How will our graph look if we change the
  51565. 38:32:02number of bins though?
  51566. 38:32:04Let's run it and see. So you still have
  51567. 38:32:07a normal distribution but it's not as
  51568. 38:32:09well defined because of how less the
  51569. 38:32:11number of bins are. you lose a majority
  51570. 38:32:13of the data which will contribute to
  51571. 38:32:15your normal distribution. It doesn't
  51572. 38:32:18look like a proper normal distribution
  51573. 38:32:19but it looks more like a discrete data
  51574. 38:32:21at this point. Now let's take a look at
  51575. 38:32:24the next step of finding a probability
  51576. 38:32:27distribution function.
  51577. 38:32:29The next step is called parametric
  51578. 38:32:31density estimation. What exactly is
  51579. 38:32:34parametric density estimation? The
  51580. 38:32:37probability density function is of many
  51581. 38:32:39types. The shape of your histogram will
  51582. 38:32:42help you determine what type of a
  51583. 38:32:43function it is. We can also calculate
  51584. 38:32:46the parameters associated with the
  51585. 38:32:48function to get our density. Now
  51586. 38:32:51different probability distribution
  51587. 38:32:54functions will have
  51588. 38:32:57different graphs which will have
  51589. 38:33:00different shapes and which will also
  51590. 38:33:01have different parameters like mean,
  51591. 38:33:03standard deviation etc associated with
  51592. 38:33:06them. Using these parameters, we can
  51593. 38:33:09find important points of our data.
  51594. 38:33:13Hence, it's very important for us to
  51595. 38:33:15recognize what type of a distribution it
  51596. 38:33:17is. Common distributions will occur
  51597. 38:33:20again and again in different and
  51598. 38:33:21sometimes unexpected domains.
  51599. 38:33:24Getting familiar with common probability
  51600. 38:33:26distributions will help you identify a
  51601. 38:33:29distribution from a histogram. And once
  51602. 38:33:31identified, you can attempt to estimate
  51603. 38:33:33the density of the random variable with
  51604. 38:33:36a chosen probability distribution. This
  51605. 38:33:38can be achieved by estimating the
  51606. 38:33:40parameters of the distribution from a
  51607. 38:33:42random sample of data. Now, an example
  51608. 38:33:45of this would be a normal distribution
  51609. 38:33:46which has two main parameters, the mean
  51610. 38:33:49and standard deviation. Given these two
  51611. 38:33:51parameters, we will now know the
  51612. 38:33:53probability distribution function. These
  51613. 38:33:56parameters can be estimated from data by
  51614. 38:33:58calculating the sample mean and sample
  51615. 38:34:01standard deviation. This entire process
  51616. 38:34:03is known as parametric density
  51617. 38:34:05estimation and it includes identifying
  51618. 38:34:08your probability distribution function
  51619. 38:34:11and getting the parameters which are
  51620. 38:34:12associated with it.
  51621. 38:34:16Now once we have estimated the density,
  51622. 38:34:18we can check if it's a good fit.
  51623. 38:34:24This can be done in three different
  51624. 38:34:26ways. One is plotting the density of the
  51625. 38:34:28function and comparing the shape to the
  51626. 38:34:30histogram. The next is sampling the
  51627. 38:34:33density function and comparing the
  51628. 38:34:35generated sample to the real sample. And
  51629. 38:34:38the last one is using a statistical test
  51630. 38:34:40to confirm if the data fits the
  51631. 38:34:42distribution. Now over here as you can
  51632. 38:34:44see all we've done is taken our data and
  51633. 38:34:48plotted the density function on top of
  51634. 38:34:51our histogram and we've compared the
  51635. 38:34:53shape. So the distribution so the
  51636. 38:34:55density function that we're actually
  51637. 38:34:56considering here is a normal
  51638. 38:34:58distribution and from this graph we can
  51639. 38:35:01see that it's almost an exact fit to our
  51640. 38:35:03histogram. Now let's see how we can
  51641. 38:35:06perform parametric density estimation
  51642. 38:35:08using Python. To begin with,
  51643. 38:35:12let's generate a random sample of
  51644. 38:35:14thousand observations from a normal
  51645. 38:35:16distribution with a mean of 50 which is
  51646. 38:35:19determined by the LOC parameter and a
  51647. 38:35:22standard deviation of five which is
  51648. 38:35:24determined by the scale parameter.
  51649. 38:35:29Now just to show you what the
  51650. 38:35:31distribution looks like, we're going to
  51651. 38:35:33plot it in the form of a histogram. So
  51652. 38:35:35this is what the histogram looks like.
  51653. 38:35:38But this is just to give you a basic
  51654. 38:35:40idea of our data and what it looks like
  51655. 38:35:42once plotted. But let's assume that we
  51656. 38:35:45don't know the probability distribution
  51657. 38:35:49and and we don't know what it looks like
  51658. 38:35:51as a histogram and we don't know that
  51659. 38:35:53that it's normal. So now if we just
  51660. 38:35:56assume that it's normal, we can
  51661. 38:35:58calculate the parameters of the
  51662. 38:35:59distribution specifically the mean and
  51663. 38:36:02the standard deviation.
  51664. 38:36:04We would not expect the mean and
  51665. 38:36:05standard deviation to be 50 and five.
  51666. 38:36:07Exactly given the small sample size and
  51667. 38:36:10the noise in the sampling data.
  51668. 38:36:12So because of this noise and the small
  51669. 38:36:15sample size, we have a mean of almost 50
  51670. 38:36:18and a standard deviation of a little
  51671. 38:36:20more than five.
  51672. 38:36:22Now let's define the distribution as
  51673. 38:36:25normal. So now using this we've defined
  51674. 38:36:28a normal distribution. We've used the
  51675. 38:36:30norm method of the sci-fi uh library and
  51676. 38:36:34uh we're doing this with the mean and
  51677. 38:36:37the standard deviation that we've
  51678. 38:36:39obtained from our samples. So up until
  51679. 38:36:42now we're just assuming that it's a
  51680. 38:36:44normal distribution and because of that
  51681. 38:36:46the parameters that we've calculated is
  51682. 38:36:48the mean and standard deviation
  51683. 38:36:51and using the calculated mean and
  51684. 38:36:53standard deviation we've gotten a normal
  51685. 38:36:55distribution.
  51686. 38:36:57And up until now again keep in mind we
  51687. 38:36:59do not know for certain that it is a
  51688. 38:37:02normal distribution. So far all we have
  51689. 38:37:04is this data.
  51690. 38:37:06So the next thing that we're going to do
  51691. 38:37:08is fit the distribution with these
  51692. 38:37:10parameters
  51693. 38:37:12and then sample the probabilities for a
  51694. 38:37:14distribution for a range of values in
  51695. 38:37:17our domain which in this case is 30 and
  51696. 38:37:1970. So all we're doing is we're
  51697. 38:37:22calculating probabilities for a range of
  51698. 38:37:24outcomes. And in this case we've taken
  51699. 38:37:2730 and 70 as our domain.
  51700. 38:37:32So these are the probability
  51701. 38:37:33distribution values for the normal
  51702. 38:37:36distribution that we've defined over
  51703. 38:37:38here. And this is going to uh this is
  51704. 38:37:41basically going to give you the outline
  51705. 38:37:42of your normal distribution.
  51706. 38:37:45These are the points at which your
  51707. 38:37:46normal distribution will be plotted. Uh
  51708. 38:37:48now what we're basically going to do is
  51709. 38:37:50we're going to plot our histograms using
  51710. 38:37:52the samples that we've already generated
  51711. 38:37:55along with the values and probabilities
  51712. 38:37:58of the normal function that we defined
  51713. 38:38:00over here.
  51714. 38:38:05So as you can see it's an all it's
  51715. 38:38:07almost a complete fit. The normal
  51716. 38:38:09distribution that we have here is made
  51717. 38:38:13using the mean and the standard
  51718. 38:38:14deviation of our actual samples.
  51719. 38:38:18The reason we took mean and standard
  51720. 38:38:20deviation was because we assumed it was
  51721. 38:38:22a normal distribution and the parameters
  51722. 38:38:25associated with the normal distribution
  51723. 38:38:28are mean and standard deviation. Using
  51724. 38:38:30the mean and standard deviation, we got
  51725. 38:38:32the normal distribution. We calculated
  51726. 38:38:36probabilities for this normal
  51727. 38:38:38distribution using a random domain of 30
  51728. 38:38:40and 70
  51729. 38:38:42and we plotted the probabilities and the
  51730. 38:38:45values on top of our histogram to see if
  51731. 38:38:48the normal distribution was a fit to our
  51732. 38:38:50histogram. If it was not a fit, you
  51733. 38:38:53would have to go and do the same
  51734. 38:38:55procedure with other common probability
  51735. 38:38:58density functions
  51736. 38:39:01until you found a function which was a
  51737. 38:39:03proper fit to your histogram. Now let's
  51738. 38:39:06move on to the final step which is used
  51739. 38:39:08in the calculation of a PDF. This final
  51740. 38:39:11step is called nonparametric density
  51741. 38:39:13estimation and it's only used when the
  51742. 38:39:16shape of a histogram doesn't match a
  51743. 38:39:18common probability density function or
  51744. 38:39:21it cannot be made to fit one. In this
  51745. 38:39:23case, we will calculate the density
  51746. 38:39:25using all samples in our data using
  51747. 38:39:28certain algorithms.
  51748. 38:39:30This is only done when a data sample
  51749. 38:39:31does not resemble a common probability
  51750. 38:39:34distribution or it cannot be easily made
  51751. 38:39:36to fit the distribution. And this is
  51752. 38:39:38often the case when the data has two
  51753. 38:39:40peaks. This is also called a biodal
  51754. 38:39:43distribution or it has many peaks which
  51755. 38:39:46is also called a multimodal
  51756. 38:39:48distribution. In this case, the
  51757. 38:39:50parametric density estimation will not
  51758. 38:39:52be feasible and alternative methods can
  51759. 38:39:55be used that do not use a common
  51760. 38:39:57distribution. Instead, you will use an
  51761. 38:39:59algorithm which is used to approximate
  51762. 38:40:01the probability distribution of the data
  51763. 38:40:04without a predefined distribution which
  51764. 38:40:06is also referred to as a non-parametric
  51765. 38:40:09method because we're not using any
  51766. 38:40:12predefined parameters. The distribution
  51767. 38:40:14will still have parameters but these are
  51768. 38:40:17not controllable in the same way as a
  51769. 38:40:19simple probability distribution. For
  51770. 38:40:22example, a non-parametric method might
  51771. 38:40:24estimate the density using all
  51772. 38:40:26observations in a random sample in
  51773. 38:40:29effect making all observations in the
  51774. 38:40:31sample parameters.
  51775. 38:40:37Now consider this graph which has two
  51776. 38:40:39peaks. You this is not a normal
  51777. 38:40:42distribution or any other sort of
  51778. 38:40:44distribution that we are familiar with.
  51779. 38:40:46So for this we're not going to use a
  51780. 38:40:48parametric estimation method but we're
  51781. 38:40:52just going to calculate the parameters
  51782. 38:40:54for every single sample point in this.
  51783. 38:40:57Perhaps the most common nonparametric
  51784. 38:41:00approach for estimating the probability
  51785. 38:41:02density function of a continuous random
  51786. 38:41:04variable is called kernel smoothing or
  51787. 38:41:07kernel density estimation or KDE for
  51788. 38:41:09short. Kernal density estimation is a
  51789. 38:41:13nonparametric method for using a data
  51790. 38:41:15set to estimate probabilities for new
  51791. 38:41:17points.
  51792. 38:41:19It uses a mathematical function and
  51793. 38:41:21smoothing probabilities. So the so the
  51794. 38:41:23sum of the resultant probabilities is
  51795. 38:41:26always one. Now in this case a kernel is
  51796. 38:41:28a mathematical function that returns a
  51797. 38:41:31probability for a given value of a
  51798. 38:41:33random variable. The kernel effectively
  51799. 38:41:35smooths or interpolates the
  51800. 38:41:37probabilities across a range of outcomes
  51801. 38:41:40for a random variable such that the sum
  51802. 38:41:42of probabilities always equals one. A
  51803. 38:41:45requirement of well- behaved
  51804. 38:41:47probabilities. You also have a parameter
  51805. 38:41:50called the smoothing parameter which
  51806. 38:41:52controls the scope or the window of
  51807. 38:41:54observations from the data samples that
  51808. 38:41:57contributes to estimating the
  51809. 38:41:59probability for a given sample. As such
  51810. 38:42:02the kernel density estimation is s is
  51811. 38:42:05sometimes referred to as your parsen
  51812. 38:42:07rosenbalt window. Now at the end you
  51813. 38:42:10also have a basis function which is a
  51814. 38:42:12function which is chosen to control the
  51815. 38:42:14contribution of samples in the data set
  51816. 38:42:16towards estimating the probability of a
  51817. 38:42:18new point. This is only done to make
  51818. 38:42:21sure that you're not learning from a lot
  51819. 38:42:23of noise and that you're not using a lot
  51820. 38:42:26of the outliers. Again let's see how we
  51821. 38:42:29can perform non-parametric density
  51822. 38:42:31estimation with the help of Python.
  51823. 38:42:35So first we'll start by importing all
  51824. 38:42:37the necessary modules along with the
  51825. 38:42:39kernel density estimation which can be
  51826. 38:42:42imported from skarn.
  51827. 38:42:47Now let's create a biodial distribution
  51828. 38:42:51by combining two different samples.
  51829. 38:42:53Sample one and sample two. Sample one
  51830. 38:42:56has 300 examples with a mean of 20 and a
  51831. 38:43:00standard deviation of five. While sample
  51832. 38:43:02two has 700 examples with a mean of 40
  51833. 38:43:06and a standard deviation of five.
  51834. 38:43:10We're then going to use it stack to com
  51835. 38:43:13to merge both of them together to get a
  51836. 38:43:15final sample.
  51837. 38:43:18The means that we've chosen which is 20
  51838. 38:43:20and 40 are chosen close together to
  51839. 38:43:23ensure that the distributions overlap in
  51840. 38:43:25the combined sample.
  51841. 38:43:28So this is what our distribution is.
  51842. 38:43:30Let's just plot it so you get a basic
  51843. 38:43:32idea of what our graph looks like.
  51844. 38:43:35So this is what our graph looks like.
  51845. 38:43:39Now we already know that none of the
  51846. 38:43:41various different uh uh probability
  51847. 38:43:44distribution functions fit these graphs.
  51848. 38:43:47So now we're going to perform
  51849. 38:43:48nonparametric estimations.
  51850. 38:43:52To perform nonparametric estimations,
  51851. 38:43:57we're going to use the scikitlearn
  51852. 38:43:59machine learning library which provides
  51853. 38:44:01the kernel density class that implements
  51854. 38:44:04kernel density sorry that implements
  51855. 38:44:07kernel density estimation. First the
  51856. 38:44:10class is constructed with the desired
  51857. 38:44:12bandwidth or window size of two
  51858. 38:44:17and your basis function
  51859. 38:44:20which in this case is a gshian function.
  51860. 38:44:24It's a good idea to at least test
  51861. 38:44:25different configurations to your data.
  51862. 38:44:28And in this case we're only going to try
  51863. 38:44:29a bandwidth of two and a gshian kernel.
  51864. 38:44:32Uh but usually there are multiple
  51865. 38:44:34different kernels that you can uh you
  51866. 38:44:37know like uh that you can play around
  51867. 38:44:39with and you can also tweak your
  51868. 38:44:40bandwidth to exactly fit the
  51869. 38:44:43distribution that you have.
  51870. 38:44:47Now let's run this.
  51871. 38:44:50Uh so now we've gotten our kernel
  51872. 38:44:52density estimation. Uh now we can
  51873. 38:44:54evaluate how well the density estimates
  51874. 38:44:57matches our data by calculating
  51875. 38:45:00probabilities
  51876. 38:45:01for a range of observations and
  51877. 38:45:03comparing shapes to the histogram just
  51878. 38:45:05like we did for the parametric case
  51879. 38:45:07before. So again we're going to just
  51880. 38:45:10calculate different probabilities using
  51881. 38:45:12the kernel density function
  51882. 38:45:15and we're just going to plot it on top
  51883. 38:45:16of a histogram to see how well this the
  51884. 38:45:19kernel density function is estimating
  51885. 38:45:22for our data.
  51886. 38:45:24So these are the probabilities that
  51887. 38:45:25we've gotten finally with the con uh
  51888. 38:45:27with the kernel density estimation. And
  51889. 38:45:30now we're going to plot it on top of our
  51890. 38:45:32histogram.
  51891. 38:45:34So as you can see it's almost a complete
  51892. 38:45:37fit. It's just left out some of these
  51893. 38:45:39outlier values which again are ranging
  51894. 38:45:42very high. But overall we have a pretty
  51895. 38:45:44good fit.
  51896. 38:45:47Uh the only problem is it's not very
  51897. 38:45:49smooth and you can uh again try tweaking
  51898. 38:45:52the bandwidth to different values. Uh so
  51899. 38:45:55let's just in this case try tweaking it
  51900. 38:45:57to three and see how well it runs. Okay.
  51901. 38:46:02So now we've got a new probabilities and
  51902. 38:46:04let's run it on top of our bandwidth. Uh
  51903. 38:46:08so again we using a bandwidth of three.
  51904. 38:46:10You can see that we're fitting our data
  51905. 38:46:12even better and we're again
  51906. 38:46:16ignoring a lot of the outliers which are
  51907. 38:46:18out there. So this is going to give us a
  51908. 38:46:20better estimation. The first question
  51909. 38:46:22that is probably in your mind is what's
  51910. 38:46:24in it for you? What can you expect from
  51911. 38:46:27this video?
  51912. 38:46:29First we will explain the concept of
  51913. 38:46:31regression a machine learning algorithm
  51914. 38:46:34to you.
  51915. 38:46:36Next we will take a look at the R squar
  51916. 38:46:38error which can be used to calculate the
  51917. 38:46:41error in regression models.
  51918. 38:46:43Next
  51919. 38:46:45we will teach you how to calculate the R
  51920. 38:46:47squar error and finally we will
  51921. 38:46:50implement the R squared error with the
  51922. 38:46:52help of Python.
  51923. 38:46:54So what is regression?
  51924. 38:46:58Regression is nothing but a machine
  51925. 38:47:00learning algorithm that helps us
  51926. 38:47:02determine the relationship between two
  51927. 38:47:04or more variables. It uses input or
  51928. 38:47:07independent variables to find the value
  51929. 38:47:09of the output or dependent variables.
  51930. 38:47:13Regression is a prediction algorithm
  51931. 38:47:16which means that given some variables we
  51932. 38:47:19can predict the value of an output
  51933. 38:47:21variable.
  51934. 38:47:23The predicted value is not going to be
  51935. 38:47:25from a set of values but it's going to
  51936. 38:47:27be a unique value in itself.
  51937. 38:47:31Now let's understand what exactly
  51938. 38:47:33regression is with the help of a few
  51939. 38:47:35independent input variables. In this
  51940. 38:47:38case, the variables that we'll be
  51941. 38:47:40looking at is rainwater, fertilizer, and
  51942. 38:47:43seeds. When we pass these independent
  51943. 38:47:46input variables through regression
  51944. 38:47:48model, we're going to get a predicted
  51945. 38:47:50output.
  51946. 38:47:52The output predicted is that a crop will
  51947. 38:47:55germinate when all three of these
  51948. 38:47:56components are put together in certain
  51949. 38:47:59quantities. So with the help of
  51950. 38:48:01regression given raw input data we can
  51951. 38:48:04find out the dependent output variable
  51952. 38:48:07that we'll get
  51953. 38:48:09when all of these input variables are
  51954. 38:48:12more or less combined. Regression is
  51955. 38:48:14nothing but a statistical method which
  51956. 38:48:17is used to determine the strength and
  51957. 38:48:19character of the relationship between a
  51958. 38:48:21dependent variable and a series of other
  51959. 38:48:24variables.
  51960. 38:48:26Now using regression if you have a set
  51961. 38:48:28of data points we can use a regression
  51962. 38:48:31model to fit a line which passes through
  51963. 38:48:33most of these data points and use it to
  51964. 38:48:35predict the outcome for new data points.
  51965. 38:48:38The line is fit using equation of a
  51966. 38:48:40straight line or a polomial equation.
  51967. 38:48:43Now in this graph consider that we have
  51968. 38:48:46our dependent variable or the output y
  51969. 38:48:48and our independent variable or the
  51970. 38:48:50output x. for a certain value of our
  51971. 38:48:54independent variable. We are going to
  51972. 38:48:56get a certain value of our dependent
  51973. 38:48:58variable or our output. Using the data
  51974. 38:49:01given here, we can see how our output
  51975. 38:49:03varies when our input varies. To predict
  51976. 38:49:06the value of our outcome Y, we're going
  51977. 38:49:09to need to find a relationship between
  51978. 38:49:11all of these data points. To do this,
  51979. 38:49:14we're going to plot a straight line
  51980. 38:49:15through it. Because, as you can see, all
  51981. 38:49:17the data points lie more or less along a
  51982. 38:49:20given straight line. Now using the
  51983. 38:49:22straight line for any value of x we can
  51984. 38:49:25find the approximate value of y. So
  51985. 38:49:29suppose you want to find the value of y
  51986. 38:49:32at a point x which is given here say
  51987. 38:49:35then all you have to do is extend a line
  51988. 38:49:37from this point onto a predicted model
  51989. 38:49:40which is this line here and then we can
  51990. 38:49:43see where this point on the line
  51991. 38:49:46coincides with the y-axis and get the
  51992. 38:49:48approximate value of the outcome.
  51993. 38:49:52The equation of a straight line is given
  51994. 38:49:55as y = b + b1 x + e. Where b is a
  51995. 38:50:01constant given by the y intercept of our
  51996. 38:50:04line or basically where the line
  51997. 38:50:06intersects on the y-axis.
  51998. 38:50:11B1 is the slope of our line and x is the
  51999. 38:50:14point for which we want to find the
  52000. 38:50:16output. E in this case is nothing but an
  52001. 38:50:19error correcting term.
  52002. 38:50:22So this is basically how regression
  52003. 38:50:24works and this is how prediction takes
  52004. 38:50:27place in a regression model. Next we
  52005. 38:50:30will explain the concept of the R squar
  52006. 38:50:32error to you. So what is R squar error?
  52007. 38:50:36R squared error is nothing but an error
  52008. 38:50:39measurement term which calculates how
  52009. 38:50:41well a regression model fits the data.
  52010. 38:50:44It determines the amount of variance in
  52011. 38:50:46a model caused by the input variables.
  52012. 38:50:49Now, R squar is a statistical measure
  52013. 38:50:52that represents the portion of variance
  52014. 38:50:54for a dependent variable that's
  52015. 38:50:56explained by an independent variable or
  52016. 38:50:58variables in a regression model. R 2 is
  52017. 38:51:01used to explain to what extent the
  52018. 38:51:04variance of one variable affects the
  52019. 38:51:06variance of a second variable.
  52020. 38:51:09So if the R square of a model is 0.5
  52021. 38:51:12then approximately half of the observed
  52022. 38:51:14variation can be explained by the
  52023. 38:51:16model's inputs.
  52024. 38:51:20In other words, an R squar of 60%
  52025. 38:51:23reveals that 60% of our data fits our
  52026. 38:51:26regression model exactly.
  52027. 38:51:30Now in this case in this graph if we
  52028. 38:51:34have a variance of 60% it means that 60%
  52029. 38:51:38of our data points fall exactly on our
  52030. 38:51:41regression line. However it is not
  52031. 38:51:44always the case that a high R squar is
  52032. 38:51:46good for a regression model. The quality
  52033. 38:51:49of the statistical measure depends on
  52034. 38:51:51many factors such as the nature of the
  52035. 38:51:53variables employed in the model, the
  52036. 38:51:56unit of measure of variables and the
  52037. 38:51:58applied data transformation.
  52038. 38:52:00Thus, sometimes a high R squar can
  52039. 38:52:03indicate the problems with the
  52040. 38:52:05regression model. A low R squar figure
  52041. 38:52:08is generally a bad sign for predictive
  52042. 38:52:10models. Now, all this time we've been
  52043. 38:52:13talking about variance in our model.
  52044. 38:52:15What exactly is variance? Variance is
  52045. 38:52:18nothing but a statistical term which
  52046. 38:52:20determines how spread out our data is
  52047. 38:52:22and tells us how many outliers are
  52048. 38:52:25present in it. Basically, it's a measure
  52049. 38:52:27of how far a set of numbers is spread
  52050. 38:52:30out from their average value.
  52051. 38:52:33Using variance, you can basically figure
  52052. 38:52:35out where your data is centered
  52053. 38:52:38and how spread it is from the mean and
  52054. 38:52:41also you can find out how many outliers
  52055. 38:52:44it has.
  52056. 38:52:46Now how can you calculate the R squar
  52057. 38:52:49error?
  52058. 38:52:52Let's start by considering the data that
  52059. 38:52:55we have been given as shown below. Now
  52060. 38:52:57we can find the relationship between the
  52061. 38:52:59input and the output variables by
  52062. 38:53:01plotting a straight line or a regression
  52063. 38:53:03model that passes through most of the
  52064. 38:53:06data. To get the perfect fit for a
  52065. 38:53:08model, we don't necessarily need to have
  52066. 38:53:11the line passing through as many data
  52067. 38:53:13points as possible.
  52068. 38:53:15A true measure of a good model is that
  52069. 38:53:17we reduce the error which is present in
  52070. 38:53:19our model. Now how do you find this
  52071. 38:53:22error? The error present in our model is
  52072. 38:53:25given by nothing but the distance
  52073. 38:53:27between our predicted line and the data
  52074. 38:53:29points which do not fall on the line.
  52075. 38:53:32This is what we have to minimize.
  52076. 38:53:35The distance between our data points and
  52077. 38:53:37our line can be calculated by
  52078. 38:53:39subtracting
  52079. 38:53:41our data point from the point at which
  52080. 38:53:44it coincides on our regression line. We
  52081. 38:53:47square this just to get rid of any
  52082. 38:53:49negative coefficients that may occur due
  52083. 38:53:51to finding the difference between the
  52084. 38:53:52two points. Now to find the variance in
  52085. 38:53:55our data, we're going to find the mean
  52086. 38:53:57and subtract the data points from the
  52087. 38:53:59mean. This will basically tell us how
  52088. 38:54:02spread out our data point is from the
  52089. 38:54:04average value. We can then square these
  52090. 38:54:08differences and add up the result to get
  52091. 38:54:10our total variance. This is also known
  52092. 38:54:12as the sum of squares total. And using
  52093. 38:54:14this we can find the total variance in
  52094. 38:54:16our data. The mean is nothing but the
  52095. 38:54:19average of our data. And using the mean
  52096. 38:54:22we can find the center of our data.
  52097. 38:54:26So this is exactly where our data is
  52098. 38:54:28centered. This is the average value that
  52099. 38:54:30occurs in our data. Now, the variance is
  52100. 38:54:33nothing but the distance of all of our
  52101. 38:54:35data points from the mean. If we do
  52102. 38:54:39this, we're basically going to find out
  52103. 38:54:40how spread apart our data is from the
  52104. 38:54:43average value or how far all of our data
  52105. 38:54:46points lie from each other.
  52106. 38:54:48When we subtract the position of our
  52107. 38:54:50data points from our mean and square it
  52108. 38:54:53and add all of that up, we get something
  52109. 38:54:55known as the sum of squares total.
  52110. 38:54:59Now the R squ error is the total
  52111. 38:55:01variance in our input data. It can be
  52112. 38:55:03obtained by dividing the SSR by the SST
  52113. 38:55:06and subtracting the results from one.
  52114. 38:55:10So the R squar error totally becomes 1
  52115. 38:55:13minus the sum of our squared errors
  52116. 38:55:16divided by the sum of the squared
  52117. 38:55:19difference between our data points and
  52118. 38:55:22the mean. Now this value gives us the
  52119. 38:55:25variance and this is why we can say that
  52120. 38:55:29R squ is used to find the portion of
  52121. 38:55:32variation in our data.
  52122. 38:55:35Now how can you implement R squ error
  52123. 38:55:37with Python?
  52124. 38:55:39To calculate R squared error with
  52125. 38:55:41Python, we're going to look at the data
  52126. 38:55:44which depicts the weather conditions
  52127. 38:55:46which were present during World War II.
  52128. 38:55:48And using the variables which are
  52129. 38:55:50present in our data, we're going to
  52130. 38:55:51create a model which predicts the daily
  52131. 38:55:54weather forecast during World War II.
  52132. 38:55:57And then we're going to use R squared
  52133. 38:55:58error to find the accuracy of our model.
  52134. 38:56:02So this is our R square uh model. So
  52135. 38:56:05we're going to start off by importing
  52136. 38:56:07all of our necessary modules. We're
  52137. 38:56:09going to use the model numpy to perform
  52138. 38:56:12numerical calculations on our database
  52139. 38:56:15and arrays and we're going to use
  52140. 38:56:17seaborn and mattplot lib to plot our
  52141. 38:56:20data.
  52142. 38:56:22So now we've managed to import all of
  52143. 38:56:24our data sets.
  52144. 38:56:27Let's also load our data by reading in
  52145. 38:56:30the CSV file that it is stored as in the
  52146. 38:56:33form of a data frame.
  52147. 38:56:36So over here as you can see we've read
  52148. 38:56:38in the CSV file into uh a variable
  52149. 38:56:42called weather and then after that we're
  52150. 38:56:44changing weather into a panda's data
  52151. 38:56:46frame called climate. This is what our
  52152. 38:56:48data frame finally looks like.
  52153. 38:56:53So as you can see in our data frame we
  52154. 38:56:55have five rows because we're only
  52155. 38:56:58looking at the top five rows and we have
  52156. 38:57:0031 columns. So these are values which
  52157. 38:57:03are not required and which are basically
  52158. 38:57:05going to increase our error value. So
  52159. 38:57:08let's drop them and get rid of them.
  52160. 38:57:13Now let's also drop any
  52161. 38:57:17empty values which may occur in the
  52162. 38:57:19remaining columns of our data set and
  52163. 38:57:21see what the final data set looks like.
  52164. 38:57:24So this is our final data set. We have
  52165. 38:57:27at the end we're only left with max
  52166. 38:57:29temperature, minimum temperature, and
  52167. 38:57:30mean temperature.
  52168. 38:57:33Now let's plot a count plot of a max
  52169. 38:57:36temperature.
  52170. 38:57:41A count plot is basically going to go
  52171. 38:57:44through the entire max temperature
  52172. 38:57:46column and figure out how many times
  52173. 38:57:49every temperature value occurs. So it's
  52174. 38:57:53going to figure out how many it's going
  52175. 38:57:54to count how many times 29.44444
  52176. 38:57:58has occurred and it's going to plot that
  52177. 38:58:00on this graph and it is going to do that
  52178. 38:58:03for every unique temperature value which
  52179. 38:58:05is present in our column. So finally
  52180. 38:58:08this is the value that we get. Uh so
  52181. 38:58:11over here as you can see the majority of
  52182. 38:58:14our temperature values are concentrated
  52183. 38:58:16within this range. This means that these
  52184. 38:58:19temperature values are the ones which
  52185. 38:58:22occur most frequently. The other ones
  52186. 38:58:24can be considered as outliers because
  52187. 38:58:27they rarely are seen in our data
  52188. 38:58:31and they can further skew the output
  52189. 38:58:33that we're going to get. Now let's do
  52190. 38:58:35the same with minimum temperature. Let's
  52191. 38:58:38plot a count plot for minimum
  52192. 38:58:40temperature.
  52193. 38:58:42So for the minimum temperature we can
  52194. 38:58:44see a very similar plot to the one that
  52195. 38:58:46we got for a maximum temperature. Most
  52196. 38:58:49of the values are concentrated around
  52197. 38:58:51this region but the outlier values here
  52198. 38:58:54are way fewer.
  52199. 38:58:59Now let's plot a regression plot between
  52200. 38:59:01our maximum and our minimum temperature.
  52201. 38:59:04So using the regression plot we can plot
  52202. 38:59:07a regression line for the two variables
  52203. 38:59:10in our x and y axis. So this is our
  52204. 38:59:13x-axis and this is our y-axis. This is
  52205. 38:59:15basically going to plot a straight line
  52206. 38:59:19which best fits the data
  52207. 38:59:24that we are getting here. So over here
  52208. 38:59:27as you can see this is a regression
  52209. 38:59:28line. This thin blue line is a
  52210. 38:59:30regression line which intersects our
  52211. 38:59:33minimum temperature at a value which is
  52212. 38:59:37between -30 and -40 and it passes
  52213. 38:59:40through our entire data. So uh for a
  52214. 38:59:44value of maximum temperature which is
  52215. 38:59:47zero using this we can predict the
  52216. 38:59:50minimum temperature that would have
  52217. 38:59:53occurred on the same day.
  52218. 38:59:56So for zero it'll be somewhere around -
  52219. 38:59:5910°C. So if we saw maximum temperature
  52220. 39:00:02of 0° on that day, we would have seen a
  52221. 39:00:04minimum temperature of minus 10 on the
  52222. 39:00:06same day. Now let's plot a heat map to
  52223. 39:00:09see how these values are correlated with
  52224. 39:00:12each other. So the correlation is
  52225. 39:00:14basically
  52226. 39:00:15used to find which values affect each
  52227. 39:00:19other linearly
  52228. 39:00:22or which values
  52229. 39:00:25when changed will also affect the change
  52230. 39:00:27in other values. So over here as you can
  52231. 39:00:30see minimum temperature
  52232. 39:00:39and maximum temperature have a
  52233. 39:00:41correlation of 0.8. 88 which means if
  52234. 39:00:44minimum temperature changes then the
  52235. 39:00:47maximum temperature will also change to
  52236. 39:00:510.88.
  52237. 39:00:55Now the best correlation is obviously
  52238. 39:00:57going to be between the mean
  52239. 39:00:58temperatures and the minimum and maximum
  52240. 39:01:00temperatures. The mean temperature is
  52241. 39:01:02nothing but the average temperature
  52242. 39:01:04value that we have. So this this is
  52243. 39:01:07basically going to lie in the middle of
  52244. 39:01:10all of our temperature values which is
  52245. 39:01:12why we going to have a better
  52246. 39:01:14correlation for mean temperature. But
  52247. 39:01:16minimum temperature and maximum
  52248. 39:01:17temperature are also pretty well
  52249. 39:01:19correlated with a correlation value of
  52250. 39:01:210.88.
  52251. 39:01:22This means that if our minimum
  52252. 39:01:25temperature fluctuates, our maximum
  52253. 39:01:28temperature will also fluctuate
  52254. 39:01:29proportionately.
  52255. 39:01:31Now let's separate our input and output
  52256. 39:01:33values. We're going to predict the
  52257. 39:01:36maximum temperature given our minimum
  52258. 39:01:38temperature. Here x is our input
  52259. 39:01:40variable and y is our output variable.
  52260. 39:01:46So now we're basically just going to get
  52261. 39:01:48all the important values in our x and y
  52262. 39:01:51data sets. So after that this is what
  52263. 39:01:53our x and y data sets are going to look
  52264. 39:01:55like. Now let's split our data set into
  52265. 39:01:58training and testing sets.
  52266. 39:02:01The training set will be used to train
  52267. 39:02:03our regression model and the testing set
  52268. 39:02:06will be used to predict how well our
  52269. 39:02:08regression model is performing.
  52270. 39:02:11The training data is the data which will
  52271. 39:02:14be visible to our model or the data
  52272. 39:02:17which a model is allowed to have access
  52273. 39:02:19to. Testing data will be data which the
  52274. 39:02:21model has never seen before or which it
  52275. 39:02:24doesn't have access to and hence it will
  52276. 39:02:27be made to work on completely new data
  52277. 39:02:29to better test how well we've fitted to
  52278. 39:02:32our data set. Now we can split our data
  52279. 39:02:35set into training and testing sets by
  52280. 39:02:36using the train test split functionality
  52281. 39:02:39from our skarn.mmodel selection library.
  52282. 39:02:44Now finally from our scikitlearn library
  52283. 39:02:47let's import a linear regression model.
  52284. 39:02:51We're going to initialize a linear
  52285. 39:02:53regression model to a variable called
  52286. 39:02:55regressor and then we're going to fit a
  52287. 39:02:57linear regression model to our training
  52288. 39:03:00data set.
  52289. 39:03:02So we finally got our trained linear
  52290. 39:03:05regression model. Now let's use this
  52291. 39:03:07model to perform predictions on our
  52292. 39:03:09testing data set. So these are the
  52293. 39:03:12values that we've gotten after running
  52294. 39:03:14our linear regression model on our
  52295. 39:03:17testing data set. Let's see how well
  52296. 39:03:20we've performed.
  52297. 39:03:22We are going to import the R2 score from
  52298. 39:03:25our skarn metric.
  52299. 39:03:28Now the R2 score will directly perform R
  52300. 39:03:31squared error on our prediction and
  52301. 39:03:34testing data set and see how well our
  52302. 39:03:36testing data set matches a prediction
  52303. 39:03:39data set.
  52304. 39:03:42So now we've gotten an R squared error
  52305. 39:03:44of 0.9345
  52306. 39:03:46which basically means that 93% of our
  52307. 39:03:53output values are influenced by our
  52308. 39:03:55input values. This also means that a
  52309. 39:03:58model is 93% accurate
  52310. 39:04:02and that approximately
  52311. 39:04:0493% of our observed variation can be
  52312. 39:04:08explained by the model's inputs. Ever
  52313. 39:04:10wondered how to build an AI project that
  52314. 39:04:13actually gets noticed by Google, OpenAI,
  52315. 39:04:15or top startups, not just a chartboard
  52316. 39:04:18or recycled homework. Today I'm going to
  52317. 39:04:21walk you through 10 AI project ideas for
  52318. 39:04:2426 that are practical, futuristic and
  52319. 39:04:28portfolio ready. I'll tell you exactly
  52320. 39:04:30which models, framework and data sets to
  52321. 39:04:33you so that you can start coding
  52322. 39:04:35immediately. Now before we jump into hit
  52323. 39:04:38that like button, share and subscribe
  52324. 39:04:40because keeping up with future proof AI
  52325. 39:04:43projects is going to give you a massive
  52326. 39:04:45edge. Let's start with the AI shopping
  52327. 39:04:48buddy. This project acts like a personal
  52328. 39:04:51stylist and interior designer. Users
  52329. 39:04:54upload photos of the room, outfit, or
  52330. 39:04:56even face, and the AI suggest products
  52331. 39:04:59that match color, style, and
  52332. 39:05:01preferences. This isn't just about
  52333. 39:05:04throwing recommendations at someone.
  52334. 39:05:06It's about computer vision to understand
  52335. 39:05:08images, generative AI to create style
  52336. 39:05:11suggestions, and recommendation
  52337. 39:05:13algorithms to find the perfect products.
  52338. 39:05:15Personalized recommendation systems
  52339. 39:05:17drive massive engagement and conversions
  52340. 39:05:20which is why companies like Amazon,
  52341. 39:05:22Flipkart, Myntra or Urban Ladder would
  52342. 39:05:26be thrilled to hit someone who can build
  52343. 39:05:28this. Completing a project like this
  52344. 39:05:30demonstrates skills in deep learning,
  52345. 39:05:33computer vision, generative AI and full
  52346. 39:05:35stack deployment for web or mobile app.
  52347. 39:05:38While shopping and lifestyle AI is
  52348. 39:05:40exciting, the next project takes up to
  52349. 39:05:43our health and wellness. The smart
  52350. 39:05:46health analyzer predicts stress burnout
  52351. 39:05:48or sleep issues by analyzing voice,
  52352. 39:05:51facial microp expressions and variable
  52353. 39:05:54data. It uses multimodel AI that
  52354. 39:05:57integrates time series analysis for
  52355. 39:05:59variable data. NLP for voice and text
  52356. 39:06:02and computer vision for micro
  52357. 39:06:04expressions. Health tech startups in
  52358. 39:06:06India and globally like healthy, cure
  52359. 39:06:08fit, Fitbit and Apple Health are looking
  52360. 39:06:11for engineers who can make predictive
  52361. 39:06:14wellness tools. Building this project
  52362. 39:06:17demonstrates your ability to work with
  52363. 39:06:19multimodel AI, pre-process complex data
  52364. 39:06:22sets, train models, and visualize
  52365. 39:06:24result. Moving from personal health to
  52366. 39:06:27professional efficiency, the AI
  52367. 39:06:29productivity agent automates your daily
  52368. 39:06:32workflow. It reads emails, scans your
  52369. 39:06:34calendar, understands priorities, and
  52370. 39:06:37builds an optimized schedule. It uses
  52371. 39:06:39NLP to parse emails, API integration
  52372. 39:06:43with other tools like Gmail, Slack, and
  52373. 39:06:45Notion, and optimization algorithms to
  52374. 39:06:48prioritize task efficiently.
  52375. 39:06:50Productivity loss is a major issue for
  52376. 39:06:52companies which is why tech giants like
  52377. 39:06:55Google Workspace, Microsoft 365 and
  52378. 39:06:58startups in workflow automation would be
  52379. 39:07:00very interested in this project. It is a
  52380. 39:07:03great way to demonstrate automation, NLP
  52381. 39:07:06API integration and practical problem
  52382. 39:07:08solving skills. Taking automation to the
  52383. 39:07:11next level, the voice toaction system
  52384. 39:07:13allows users to speak commands and have
  52385. 39:07:16the AI perform multi-step action such as
  52386. 39:07:19booking flights, organizing files or
  52387. 39:07:21generating reports. It relies on
  52388. 39:07:24speechtoext models, intent
  52389. 39:07:26classification using NLP and task
  52390. 39:07:29automation pipelines. You can train
  52391. 39:07:32intent classification models using data
  52392. 39:07:34set into the snips NLU data set. This
  52393. 39:07:38project builds directly on productivity
  52394. 39:07:40AI and is exactly the kind of work that
  52395. 39:07:43Amazon openai or Apple would notice for
  52396. 39:07:46voiced driven automation solutions. Once
  52397. 39:07:49we have automated task, why not explore
  52398. 39:07:52creativity? Generative AI story maker
  52399. 39:07:54allows you to create full stories
  52400. 39:07:56including scripts, characters and
  52401. 39:07:59visuals based on just a few keywords.
  52402. 39:08:02Now it uses large language models for
  52403. 39:08:05text generation and image generation
  52404. 39:08:07models like stable diffusion deli3 for
  52405. 39:08:10visuals and you can also train
  52406. 39:08:12fine-tuned models on data sets like CMU
  52407. 39:08:15book summary corpus on writing prompts
  52408. 39:08:18text to speech library such as scope TTS
  52409. 39:08:21or GTTS can add narration media
  52410. 39:08:25companies like Netflix, Ubisoft and
  52411. 39:08:27Adobe are actively investing in
  52412. 39:08:30generative AI and a project like this
  52413. 39:08:32would definitely stand out. Building on
  52414. 39:08:34the idea of multiple AI capabilities
  52415. 39:08:37working together, the multi- aent AI
  52416. 39:08:40team project introduces collaboration
  52417. 39:08:42between AI agents. Multiple AI agents
  52418. 39:08:45are assigned specialized roles such as
  52419. 39:08:48researching, writing, criticing, and
  52420. 39:08:50summarizing. They communicate and
  52421. 39:08:53coordinate to complete complex task
  52422. 39:08:55using multi- aent reinforcement,
  52423. 39:08:57learning and communication protocols.
  52424. 39:09:00Enterprise AI and automation platforms
  52425. 39:09:03are investing heavily in this approach
  52426. 39:09:05and companies like Enthropic, OpenAI, AI
  52427. 39:09:08workflow startups are actively seeking
  52428. 39:09:11engineers who can build collaborative AI
  52429. 39:09:14systems from collaboration to
  52430. 39:09:16observation. The AI body language reader
  52431. 39:09:18analyzes micro expressions, tone of
  52432. 39:09:21voice and posture to provide feedback on
  52433. 39:09:24communication skills. It can be applied
  52434. 39:09:26in interviews, public speaking or remote
  52435. 39:09:29coaching. Computer vision tools such as
  52436. 39:09:32open pose or media pipe pose track
  52437. 39:09:34gestures and posture. While audio
  52438. 39:09:37processing libraries such as librosa
  52439. 39:09:40analyze tone models can be trained using
  52440. 39:09:43data sets like Raves for audio and CK
  52441. 39:09:46plus for facial expressions. HR tech
  52442. 39:09:49companies like High View, Pytrics and AI
  52443. 39:09:52coaching startups would highly value
  52444. 39:09:54this type of project as it help bridge
  52445. 39:09:57human behavior and AI analysis. Nucation
  52446. 39:10:01is also another area being transformed
  52447. 39:10:03by AI. The personalized tutor with
  52448. 39:10:06adaptive difficulty creates an AI tutor
  52449. 39:10:09that adjusts lessons in real time based
  52450. 39:10:12on student performance. Knowledge
  52451. 39:10:14tracing models like deep knowledge
  52452. 39:10:16tracing combined with transformers for
  52453. 39:10:18content generation allow the AI to adapt
  52454. 39:10:21to each learner. Data sets such as
  52455. 39:10:24assessments or edn nets can be used for
  52456. 39:10:27training. Now this AI can generate new
  52457. 39:10:29questions, explanations and motivational
  52458. 39:10:32feedback based on learning pace.
  52459. 39:10:35Companies like Baiju, Vidanto, Corsera
  52460. 39:10:38and Udemy are constantly looking for
  52461. 39:10:40talent that can build adaptive learning
  52462. 39:10:43platforms. Next, we move into research
  52463. 39:10:45augmentation with the autonomous
  52464. 39:10:48research agent. This AI can answer
  52465. 39:10:50research questions by reading academic
  52466. 39:10:52papers, extracting insights, summarizing
  52467. 39:10:55information, and citing sources
  52468. 39:10:57automatically. It uses the S2 or data
  52469. 39:11:01set for academic papers. Cybboard for
  52470. 39:11:04scientific text embeddings and hugging
  52471. 39:11:06face transformers or lang chain for
  52472. 39:11:09reasoning and summarization. Citation
  52473. 39:11:12extraction can be done with NLP passers
  52474. 39:11:15or rejects academic platforms. AI labs
  52475. 39:11:18and companies like Google research, open
  52476. 39:11:20AI, research gate and LCV would hire
  52477. 39:11:23engineers who can build this system.
  52478. 39:11:26This project connects perfectly with
  52479. 39:11:28education focused AI extending learning
  52480. 39:11:31into automated research capabilities.
  52481. 39:11:35Finally, we arrive at realworld robotics
  52482. 39:11:38control with the AI. The ultimate
  52483. 39:11:40demonstration of cuttingedge skill. This
  52484. 39:11:43project trains AI to control robotic
  52485. 39:11:46arms or humanoids based on goals rather
  52486. 39:11:49than just rigid instructions. It uses pi
  52487. 39:11:52bullet or vbots for simulation. Stable
  52488. 39:11:55baseline 3 or R lib for reinforcement
  52489. 39:11:58learning and open CV or media pipe for
  52490. 39:12:01vision input. Sim to real transfer
  52491. 39:12:04techniques bring simulations into realw
  52492. 39:12:06world scenarios. Robotics companies like
  52493. 39:12:09Boston Dynamics, Appronic, Amazon
  52494. 39:12:12Robotics and Agibot are seeking
  52495. 39:12:15engineers capable of endtoend AIdriven
  52496. 39:12:18robotic systems. After exploring
  52497. 39:12:20softwarebased AI projects, robotics is
  52498. 39:12:23the next step to show mastery of AI
  52499. 39:12:26applied in the physical world. These 10
  52500. 39:12:28AI projects are more than just ideas.
  52501. 39:12:31>> We will learn about some of the machine
  52502. 39:12:33learning and deep learning interview
  52503. 39:12:34questions.
  52504. 39:12:36So let's begin with our first question.
  52505. 39:12:39The first question is how to detect
  52506. 39:12:41outliers in data. So in data analytics
  52507. 39:12:45and machine learning, you often find
  52508. 39:12:47data points that lie at an abnormal
  52509. 39:12:49distance from other points in a random
  52510. 39:12:51sample from a population. Those are
  52511. 39:12:54called outliers. Now outliers in data
  52512. 39:12:56can significantly impact any prediction
  52513. 39:12:58analysis. There are majorly three
  52514. 39:13:00different methods to treat outliers.
  52515. 39:13:03First we have the univariate method. It
  52516. 39:13:06is one of the simplest methods for
  52517. 39:13:07detecting outliers. The univariate
  52518. 39:13:10method uses box plots. A box plot is a
  52519. 39:13:13graphical display for describing the
  52520. 39:13:14distributions of the data. Box plots use
  52521. 39:13:18the median and the lower and upper
  52522. 39:13:19quartiles.
  52523. 39:13:21This method looks for data points with
  52524. 39:13:23extreme values on one variable. Next, we
  52525. 39:13:26have the multivariate method. So, the
  52526. 39:13:28multivariate outliers can be found in an
  52527. 39:13:30n- dimensional space having n features.
  52528. 39:13:33We look for unusual combinations of all
  52529. 39:13:35the variables in this method. Finally,
  52530. 39:13:37we have Minowski error. This method
  52531. 39:13:40reduces the contribution of potential
  52532. 39:13:42outliers in the training process. The
  52533. 39:13:44Minowski error is a loss index that is
  52534. 39:13:46more insensitive to outliers than the
  52535. 39:13:48standard mean squared error. Now moving
  52536. 39:13:50on to the second question. What is a
  52537. 39:13:53confusion matrix? So a confusion matrix
  52538. 39:13:56is a table that is used to describe the
  52539. 39:13:58performance of a classification model on
  52540. 39:14:00a set of test data for which the true
  52541. 39:14:02values are already known. The target
  52542. 39:14:05variable has two values positive or
  52543. 39:14:07negative. The columns represent the
  52544. 39:14:10actual values of the target variable
  52545. 39:14:12which you can see here. The rows
  52546. 39:14:14represent the predicted values of the
  52547. 39:14:15target variable which you can see here.
  52548. 39:14:18Now there are four important terms that
  52549. 39:14:20are related to confusion matrix. First
  52550. 39:14:23we have true positive which is this one.
  52551. 39:14:27So in true positive the predicted value
  52552. 39:14:29matches the actual value. So the actual
  52553. 39:14:31value was positive and the model also
  52554. 39:14:33predicted a positive value. Then we have
  52555. 39:14:36true negative which is also represented
  52556. 39:14:39as tn. The true negative depicts the
  52557. 39:14:42predicted value matches the actual
  52558. 39:14:44value. Now the actual value was negative
  52559. 39:14:46and the model predicted a negative
  52560. 39:14:48value. Next we have false positive. Now
  52561. 39:14:52false positive is also known as a type
  52562. 39:14:54one error. In false positive the
  52563. 39:14:57predicted value was falsely predicted.
  52564. 39:14:59The actual value was negative but the
  52565. 39:15:01model predicted a positive value.
  52566. 39:15:03Finally we have false negative. A false
  52567. 39:15:06negative is also known as type two
  52568. 39:15:08error. So in false negative the
  52569. 39:15:10predicted value was falsely predicted.
  52570. 39:15:13The actual value was positive but the
  52571. 39:15:15model predicted a negative value. Now
  52572. 39:15:17moving to our third question which is
  52573. 39:15:20explain the ROC curve. Now the ROC curve
  52574. 39:15:23is one of the most important evaluation
  52575. 39:15:25metrics for checking the performance of
  52576. 39:15:27any classification model. ROC stands for
  52577. 39:15:30receiver operating characteristic.
  52578. 39:15:33Receiver operating characteristic or ROC
  52579. 39:15:35curve is a method to compare the
  52580. 39:15:37diagnostic tests. The ROC curve is
  52581. 39:15:39created by plotting the true positive
  52582. 39:15:41rate against the false positive rate at
  52583. 39:15:43various threshold settings. So here on
  52584. 39:15:46the y-axis you have the true positive
  52585. 39:15:47rate. On the x-axis we have the false
  52586. 39:15:50positive rate. The true positive rate
  52587. 39:15:52indicates the proportion of observations
  52588. 39:15:54that were correctly predicted to be
  52589. 39:15:55positive out of all positive
  52590. 39:15:57observations. Similarly, the false
  52591. 39:16:00positive rate is the proportion of
  52592. 39:16:02observations that are incorrectly
  52593. 39:16:03predicted to be positive out of all
  52594. 39:16:05negative observations.
  52595. 39:16:08You can take an example. Suppose in
  52596. 39:16:10medical testing, the true positive rate
  52597. 39:16:12is the rate in which people are
  52598. 39:16:14correctly identified to test positive
  52599. 39:16:16for the disease in question. Let's say
  52600. 39:16:18the corona virus testing. ROC does not
  52601. 39:16:21depend on any class distribution. This
  52602. 39:16:23makes it useful for evaluating
  52603. 39:16:25classifiers predicting rare events such
  52604. 39:16:27as diseases or disasters. Now moving to
  52605. 39:16:29the fourth question we have what are the
  52606. 39:16:32assumptions for linear regression. So
  52607. 39:16:34linear regression analysis is used for
  52608. 39:16:36modeling the relationship between a
  52609. 39:16:38single dependent variable Y and one or
  52610. 39:16:40more feature or predictor variables.
  52611. 39:16:43Some of the important assumptions for
  52612. 39:16:45linear regression are so first they
  52613. 39:16:48should have linearity. So linear
  52614. 39:16:50regression needs the relationship
  52615. 39:16:52between the independent and the
  52616. 39:16:53dependent variables to be linear. It is
  52617. 39:16:56also crucial to check for outliers since
  52618. 39:16:58linear regression is sensitive to
  52619. 39:17:00outlier effects. Next we have
  52620. 39:17:02homocyasticity.
  52621. 39:17:04Homoscadasticity
  52622. 39:17:05illustrates a situation in which the
  52623. 39:17:07error term that is the noise or random
  52624. 39:17:10disturbance in the relationship between
  52625. 39:17:11the features and the target variable is
  52626. 39:17:13the same across all levels of the
  52627. 39:17:15dependent variables. Third we have
  52628. 39:17:17independence. So observations should be
  52629. 39:17:19independent of each other. Finally, we
  52630. 39:17:22have no multi-olinearity.
  52631. 39:17:25So there should be little or no
  52632. 39:17:26multi-olinearity.
  52633. 39:17:28Independent variables should not be too
  52634. 39:17:29highly correlated. Now moving to our
  52635. 39:17:32fifth question in our list of interview
  52636. 39:17:33questions.
  52637. 39:17:35The question is what is regularization
  52638. 39:17:38in machine learning? Explain the L2
  52639. 39:17:40regularization.
  52640. 39:17:43So regularization is a machine learning
  52641. 39:17:45technique that is used to reduce the
  52642. 39:17:46errors by fitting the function
  52643. 39:17:48appropriately on the training set in
  52644. 39:17:50order to avoid overfitting of data. So
  52645. 39:17:52overfitting happens when a model learns
  52646. 39:17:54the detail and noise in the training
  52647. 39:17:56data to the extent that it negatively
  52648. 39:17:58impacts the performance of the model on
  52649. 39:18:00new data. So here you can see we have a
  52650. 39:18:03nice plot which shows how overfitting of
  52651. 39:18:06data can be visualized
  52652. 39:18:08and here we have a good fit line over
  52653. 39:18:13the same data points. So this is also
  52654. 39:18:15known as the regression line. Now L2
  52655. 39:18:18regularization is also known as ridge
  52656. 39:18:20regression. So ridge regression modifies
  52657. 39:18:23the overfitted model by adding the
  52658. 39:18:24squared magnitude of coefficient as a
  52659. 39:18:26penalty term to the loss function. So on
  52660. 39:18:29the right you can see a set of data
  52661. 39:18:32points plotted and we have our linear
  52662. 39:18:34regression line and here we are
  52663. 39:18:37calculating the cost function for the
  52664. 39:18:39ridge regression line. So our cost
  52665. 39:18:41function is actually loss plus lambda
  52666. 39:18:45into summation of w ^ 2 where loss is
  52667. 39:18:48actually the sum of squared errors or
  52668. 39:18:50squared residuals. Lambda stands for
  52669. 39:18:53penalty for the errors. W is called the
  52670. 39:18:55slope of the curve or line. Okay. Now
  52671. 39:18:59consider a case where there are two
  52672. 39:19:02points passing through the linear
  52673. 39:19:04regression line. Now if you calculate
  52674. 39:19:06the cost function, we get the value as
  52675. 39:19:091.69. So here we have assumed that loss
  52676. 39:19:12is zero since the two points lie
  52677. 39:19:14directly on the line. We have taken
  52678. 39:19:17lambda to be 1 and w is 1.3. So if you
  52679. 39:19:20use this function or this formula, you
  52680. 39:19:24get the cost function as 1.69.
  52681. 39:19:28Now moving ahead, let's consider another
  52682. 39:19:31situation where we'll calculate the same
  52683. 39:19:33cost function for the ridge regression
  52684. 39:19:35line. There is some loss for both the
  52685. 39:19:38points as they are not on the same line.
  52686. 39:19:41So here you can see the sum of squared
  52687. 39:19:43residuals is 0.05 which actually is the
  52688. 39:19:47sum of 2 squared. I'm assuming this as 2
  52689. 39:19:50and 0.1 for this one.
  52690. 39:19:53So if you square both and add it the
  52691. 39:19:56value is 05 a lambda is again 1 and w we
  52692. 39:20:00have assumed to be 6. Now if you find
  52693. 39:20:03the cost function the value is 41. Let's
  52694. 39:20:06draw the linear regression line and the
  52695. 39:20:08ridge regression line with all the
  52696. 39:20:10points we find that the ridge regression
  52697. 39:20:12line as the best fit since its cost
  52698. 39:20:14function is less. Now coming to the
  52699. 39:20:17sixth question.
  52700. 39:20:19What are the different methods to split
  52701. 39:20:21a tree in a decision tree algorithm?
  52702. 39:20:25So there are three methods to split a
  52703. 39:20:27decision tree. First we have variance.
  52704. 39:20:30So reduction in variance is an algorithm
  52705. 39:20:33that is used for continuous target
  52706. 39:20:34variables. This algorithm uses the
  52707. 39:20:37standard formula variance to choose the
  52708. 39:20:39best split. So here you can see the
  52709. 39:20:41standard formula variance which is
  52710. 39:20:42summation of x that is all the
  52711. 39:20:45individual points minus xar which is the
  52712. 39:20:48mean squared divided by the total number
  52713. 39:20:51of observations. Now the split with
  52714. 39:20:53lower variance is selected as the
  52715. 39:20:55criteria to split the population.
  52716. 39:20:58Now the steps to calculate variance is
  52717. 39:21:00you need to calculate variance for each
  52718. 39:21:02node and then you need to calculate for
  52719. 39:21:05each split as the weighted average of
  52720. 39:21:07each node variance.
  52721. 39:21:10Moving ahead, the second method we have
  52722. 39:21:12is information gain. So information gain
  52723. 39:21:15is used for splitting the nodes when the
  52724. 39:21:17target variable is categorical.
  52725. 39:21:20It works on the concept of entropy. Now
  52726. 39:21:23the degree of disorganization in a
  52727. 39:21:25system is known as entropy. So here you
  52728. 39:21:27can see the formula for information gain
  52729. 39:21:29which is 1 minus entropy.
  52730. 39:21:32Finally we have genie impurity. So genie
  52731. 39:21:36impurity is the probability of
  52732. 39:21:37incorrectly classifying a randomly
  52733. 39:21:39chosen element in the data set if it
  52734. 39:21:42were randomly labeled according to the
  52735. 39:21:43class distribution in the data set. So
  52736. 39:21:45below you can see the formula for gen
  52737. 39:21:48impurity. So we have 1 minus summation
  52738. 39:21:51of pi whole square where n represents
  52739. 39:21:54the number of classes and p of i
  52740. 39:21:56represents the probability of randomly
  52741. 39:21:58picking an element of class i.
  52742. 39:22:02Now moving to the seventh question.
  52743. 39:22:06So the question is how do we find the
  52744. 39:22:08optimum cluster value in K means
  52745. 39:22:09clustering algorithm. Now there are two
  52746. 39:22:12methods to find the optimum cluster
  52747. 39:22:14value. So first we have the elbow method
  52748. 39:22:17which is one of the most wellknown for
  52749. 39:22:18finding the optimum number of clusters.
  52750. 39:22:22So in this method you need to calculate
  52751. 39:22:23the within cluster sum of squared errors
  52752. 39:22:26for different values of K and choose the
  52753. 39:22:28K for which within cluster sum of
  52754. 39:22:30squared errors first starts to diminish.
  52755. 39:22:34So in the below plot of squared errors
  52756. 39:22:36versus the number of clusters K you can
  52757. 39:22:39see at K is equal to 4 the squared error
  52758. 39:22:42starts to diminish. So hence our optimum
  52759. 39:22:46K value is four. Next we have the siloid
  52760. 39:22:50method.
  52761. 39:22:51So the celloid method measures how
  52762. 39:22:53similar a point is to its own cluster
  52763. 39:22:56compared to other clusters. The average
  52764. 39:22:59seloid method computes the average of
  52765. 39:23:01observations for different values of K.
  52766. 39:23:04The optimum number of clusters K is the
  52767. 39:23:06one that maximizes the average over a
  52768. 39:23:09range of possible values for K. The
  52769. 39:23:12seloid score reaches its global maximum
  52770. 39:23:14at the optimal K. So in our case the
  52771. 39:23:17average seloid reaches maximum at k is
  52772. 39:23:20equal to two which you can see here. So
  52773. 39:23:22our optimum cluster value will be two
  52774. 39:23:25here. Moving ahead the eighth question
  52775. 39:23:28in our list is how does the pooling
  52776. 39:23:31layer work in a convolutional neural
  52777. 39:23:32network. So the pooling layer performs a
  52778. 39:23:35downsampling operation in order to
  52779. 39:23:37reduce the dimensionality of the feature
  52780. 39:23:39map. So in the pooling operation, you
  52781. 39:23:42slide a two-dimensional filter over each
  52782. 39:23:45channel of feature map and summarize the
  52783. 39:23:47features lying within the region covered
  52784. 39:23:49by the filter.
  52785. 39:23:51It is a common practice to periodically
  52786. 39:23:53insert a pooling layer in between
  52787. 39:23:54successive convolutional layers in a
  52788. 39:23:56convolutional neural network
  52789. 39:23:58architecture.
  52790. 39:23:59So the pooling layer operates
  52791. 39:24:01independently on every depth slice in
  52792. 39:24:04the input and resizes it specially using
  52793. 39:24:07the max operation. So in the diagram
  52794. 39:24:10shown here you can see we have a
  52795. 39:24:12rectified feature map. We are using a 2
  52796. 39:24:16+2 filter and performing a max pooling
  52797. 39:24:19operation. So consider this as the
  52798. 39:24:21filter. If you perform the max operation
  52799. 39:24:23over the values
  52800. 39:24:27let's say 0 5 3 and 1. So considering
  52801. 39:24:30this one our pool feature map maximum
  52802. 39:24:33value will be five. Similarly for this
  52803. 39:24:35chunk of data it is going to be seven.
  52804. 39:24:38Next, if you slide the filter over this
  52805. 39:24:40square frame, you get eight. And
  52806. 39:24:42similarly here you get six. So this is
  52807. 39:24:45also known as a pulled feature map.
  52808. 39:24:48Moving ahead, the ninth question in our
  52809. 39:24:50list is how does LSTM network work? So
  52810. 39:24:53long short-term memory networks are a
  52811. 39:24:55type of recurrent neural networks that
  52812. 39:24:57are capable of learning order dependence
  52813. 39:24:59and sequence prediction problems. So
  52814. 39:25:01remembering information for long periods
  52815. 39:25:03of time is practically their default
  52816. 39:25:05behavior. Now, LSTMs also have this
  52817. 39:25:08chain-like structure which you can see
  52818. 39:25:10here.
  52819. 39:25:12But the repeating module has a different
  52820. 39:25:14structure. So, instead of having a
  52821. 39:25:16single neural network layer, there are
  52822. 39:25:18four interacting in a very special way.
  52823. 39:25:21Now, you can see these are called as
  52824. 39:25:23gates. These gates contain sigmoid
  52825. 39:25:25activations. A sigmoid activation is
  52826. 39:25:28similar to the tanh activation. Instead
  52827. 39:25:30of squishing values between minus1 and +
  52828. 39:25:34one, it squishes values between 0 and
  52829. 39:25:36one. An LSTDM has four gates.
  52830. 39:25:41Now these are called forget, remember,
  52831. 39:25:44learn and use or output. So if you see
  52832. 39:25:47this in the first step, we use the
  52833. 39:25:49forget gate that decides what
  52834. 39:25:51information should be thrown away or
  52835. 39:25:53kept. the information from the previous
  52836. 39:25:55hidden state and the information from
  52837. 39:25:58the current input is passed through the
  52838. 39:25:59sigmoid function. Values come out
  52839. 39:26:02between zero and one.
  52840. 39:26:04So if the value is closer to zero, it
  52841. 39:26:06means you need to forget that
  52842. 39:26:07information and if the value is closer
  52843. 39:26:09to one, it means you need to keep that
  52844. 39:26:11information.
  52845. 39:26:13Next we have the input gate. So the
  52846. 39:26:16input gate is used to update the cell
  52847. 39:26:18state. First, we pass the previous
  52848. 39:26:20hidden state and the current input into
  52849. 39:26:22a sigmoid function
  52850. 39:26:25that decides which values will be
  52851. 39:26:26updated by transforming the values to be
  52852. 39:26:29between 0 and 1. Zero means not
  52853. 39:26:31important and one means important. You
  52854. 39:26:34also pass the hidden state and current
  52855. 39:26:36input into the tan function to flatten
  52856. 39:26:38the values between minus1 and + one.
  52857. 39:26:41This helps to regulate the network.
  52858. 39:26:44Then you multiply the tan output with
  52859. 39:26:46the sigmoid output. The sigmoid output
  52860. 39:26:49will decide which information is
  52861. 39:26:50important to keep from the tanh output.
  52862. 39:26:54And finally in step three we have the
  52863. 39:26:56output gate. This output gate is used to
  52864. 39:26:59decide what the next hidden state should
  52865. 39:27:02be. First we pass the previous hidden
  52866. 39:27:04state and the current input into a
  52867. 39:27:06sigmoid function. Then we pass the newly
  52868. 39:27:09modified cell state into the tanage
  52869. 39:27:11function. We then multiply the tanage
  52870. 39:27:14output with the sigmoid output to decide
  52871. 39:27:16what information the hidden state should
  52872. 39:27:18carry. The output is the hidden state.
  52873. 39:27:21The new cell state and the new hidden
  52874. 39:27:23state is then carried over to the next
  52875. 39:27:25time step. Finally, talking about the
  52876. 39:27:29last question in our list of interview
  52877. 39:27:30questions, we have explained the concept
  52878. 39:27:33of gradient descent in deep learning.
  52879. 39:27:36Now gradient descent is an optimization
  52880. 39:27:38algorithm which is mainly used to find
  52881. 39:27:40the minimum of a function in machine
  52882. 39:27:42learning. Gradient descent is used to
  52883. 39:27:44update the parameters in a model.
  52884. 39:27:47Parameters can vary according to the
  52885. 39:27:49algorithms such as coefficients in
  52886. 39:27:51linear regression and weights in neural
  52887. 39:27:53networks.
  52888. 39:27:55You can see we have these maps and on
  52889. 39:27:59the y-axis we have the loss. On the
  52890. 39:28:02x-axis we have the weight and here we
  52891. 39:28:05are trying to find the local minimum or
  52892. 39:28:08the global minimum. Now this gradient
  52893. 39:28:11descent method is used to minimize the
  52894. 39:28:13cost function and update the parameters
  52895. 39:28:14of the learning model. The gradient
  52896. 39:28:17always points in the direction of the
  52897. 39:28:19steepest increase in the loss function.
  52898. 39:28:22The gradient descent algorithm takes a
  52899. 39:28:24step in the direction of the negative
  52900. 39:28:25gradient in order to reduce the loss as
  52901. 39:28:28quickly as possible. To determine the
  52902. 39:28:30next point along the loss function
  52903. 39:28:32curve, the gradient descent algorithm
  52904. 39:28:34adds some fraction of the gradient's
  52905. 39:28:36magnitude to the starting point. Now
  52906. 39:28:38this process is repeated to find the
  52907. 39:28:40global minimum.
  52908. 39:28:42>> Welcome to math refresher probability
  52909. 39:28:45and statistics.
  52910. 39:28:47In this lesson, we are going to explain
  52911. 39:28:49the concepts of statistics and
  52912. 39:28:52probability.
  52913. 39:28:53Describe conditional probability. Define
  52914. 39:28:56the chain rule of probability. Discuss
  52915. 39:28:59the measure of variance. Identify the
  52916. 39:29:01types of gshian distribution.
  52917. 39:29:04Basic of statistics and probability.
  52918. 39:29:07Probability and statistics. Data science
  52919. 39:29:10relies heavily on estimates and
  52920. 39:29:12predictions. A significant portion of
  52921. 39:29:15data science is made up of evaluations
  52922. 39:29:17and forecast.
  52923. 39:29:19Statistical methods are used to make
  52924. 39:29:21estimates for further analysis.
  52925. 39:29:24Probability theory is helpful for making
  52926. 39:29:26predictions. Statistical methods are
  52927. 39:29:29highly dependent on probability theory
  52928. 39:29:32and all probability and statistics are
  52929. 39:29:35dependent on data.
  52930. 39:29:37Data is information acquired for
  52931. 39:29:39reference or research via observations,
  52932. 39:29:43facts, and measurements. Data is a set
  52933. 39:29:46of facts structured in the form that
  52934. 39:29:48computers can interpret such as numbers,
  52935. 39:29:51words, estimations, and views.
  52936. 39:29:54Importance of data. Data aids in seeing
  52937. 39:29:57more about the information by
  52938. 39:29:59identifying possible connections between
  52939. 39:30:01two features. Data assists in the
  52940. 39:30:04detection of distortion by uncovering
  52941. 39:30:06hidden patterns based on prior
  52942. 39:30:09information patterns. Data may be
  52943. 39:30:12utilized to anticipate the future or
  52944. 39:30:14predict the current state of affairs.
  52945. 39:30:17Also, data aids in determining whether
  52946. 39:30:19two pieces of information have any
  52947. 39:30:21instance in common or not. Types of
  52948. 39:30:24data. Data might be quantitative. That
  52949. 39:30:28is data that can be measured or counted
  52950. 39:30:30in numbers or it may be qualitative
  52951. 39:30:33which is data which is generally divided
  52952. 39:30:35into groups or in simpler words which
  52953. 39:30:38cannot be counted or measured in
  52954. 39:30:40numbers. Let's consider an example a
  52955. 39:30:43customer information data of a bank may
  52956. 39:30:46contain quantitative and qualitative
  52957. 39:30:48data. Consider this snapshot where we
  52958. 39:30:51have customer ID, surname, geography,
  52959. 39:30:55gender, age, balance, has C or card is
  52960. 39:30:58active member. Amongst these variables
  52961. 39:31:01we can see surname is mostly qualitative
  52962. 39:31:04as it cannot be counted and measured in
  52963. 39:31:06numbers. Geography and gender are also
  52964. 39:31:10qualitative as they cannot be counted in
  52965. 39:31:12numbers and are mostly groups. has C or
  52966. 39:31:16card that is has credit card and is
  52967. 39:31:19active member although are containing
  52968. 39:31:21numerical in form but these are
  52969. 39:31:24categorical that means these have been
  52970. 39:31:26divided into groups of one and zero that
  52971. 39:31:30represent yes and no as an answer hence
  52972. 39:31:33these two variables are also qualitative
  52973. 39:31:37customer ID is again although a
  52974. 39:31:40numerical data however the significance
  52975. 39:31:43or intuition behind Customer ID is
  52976. 39:31:46categorical.
  52977. 39:31:47Hence, it may be kept in the qualitative
  52978. 39:31:50data also. However, age and balance
  52979. 39:31:54these are numerical information which
  52980. 39:31:56have been measured or counted and
  52981. 39:31:58numerical operations can be performed on
  52982. 39:32:01them. Hence, these are under
  52983. 39:32:03quantitative data categories.
  52984. 39:32:05Introduction to descriptive statistics.
  52985. 39:32:08Descriptive statistics. A descriptive
  52986. 39:32:11measurement is summary measure that
  52987. 39:32:13quantitatively portrays the most
  52988. 39:32:15important features of a set of data
  52989. 39:32:18allowing for a better comprehension of
  52990. 39:32:20the information. Data can be measured as
  52991. 39:32:22different levels. The levels of
  52992. 39:32:25measurement describe the nature of
  52993. 39:32:26information stored in the data assigned
  52994. 39:32:29to the variables. Qualitative data can
  52995. 39:32:32be measured as nominal or ordinal.
  52996. 39:32:34Quantitative data can be measured in
  52997. 39:32:36terms of interval and ratio type.
  52998. 39:32:39Nominal data. The data is categorized
  52999. 39:32:42using names, labels or qualities. For
  53000. 39:32:45example, brand name, zip code, and
  53001. 39:32:47gender. Ordinal data can be arranged in
  53002. 39:32:50order or ranked and can be compared.
  53003. 39:32:53Examples include grades, star reviews,
  53004. 39:32:57position, and race, and date. Interval
  53005. 39:33:00data is the data that is ordered and has
  53006. 39:33:02meaningful differences between the data
  53007. 39:33:05points. Example temperature in Celsius
  53008. 39:33:08and year of birth. Ratio data is similar
  53009. 39:33:11to the interval level with the added
  53010. 39:33:14property of inherent zero. Mathematical
  53011. 39:33:17calculations can be performed on both
  53012. 39:33:19interval as well as ratio data. For
  53013. 39:33:22example, height, age, and weight.
  53014. 39:33:25Population versus sample. Before
  53015. 39:33:28analyzing the data, it's important to
  53016. 39:33:30figure out if it's from a population or
  53017. 39:33:32a sample. Population is a collection of
  53018. 39:33:36all available items as well as each unit
  53019. 39:33:38in our study. Sample is a subset of the
  53020. 39:33:41population that contains only a few
  53021. 39:33:44units of the population. Population data
  53022. 39:33:47is used for study when the data pool is
  53023. 39:33:50very small and can give all the required
  53024. 39:33:52information. Samples are collected
  53025. 39:33:55randomly and represent the entire
  53026. 39:33:58population in the best possible way.
  53027. 39:34:01Measures of central tendency.
  53028. 39:34:04The central tendency is a single value
  53029. 39:34:06that aids in the description of the data
  53030. 39:34:09by determining its center position.
  53031. 39:34:12Measures of central tendency are
  53032. 39:34:14sometimes known as summary statistics or
  53033. 39:34:17measures of central location.
  53034. 39:34:19The most popular measurements of central
  53035. 39:34:22tendency are mean, median, and mode. The
  53036. 39:34:26normal distribution is a bell-shaped
  53037. 39:34:28symmetrical distribution in which mean,
  53038. 39:34:31median, and mode all are equal. The
  53039. 39:34:34curve over here shows the bell-shaped
  53040. 39:34:36curve or the normal distribution of
  53041. 39:34:38variable X. The point over here that is
  53042. 39:34:41X1 is the point which represents the
  53043. 39:34:45mean, median and mode of this
  53044. 39:34:47distribution. Mean mean is calculated by
  53045. 39:34:50dividing these sum of all data values by
  53046. 39:34:53the total number of data values. It gets
  53047. 39:34:57affected when there are unusual or
  53048. 39:34:59extreme values. It is sensitive to the
  53049. 39:35:02outliers. Mean can be calculated as
  53050. 39:35:05summation over all the values of X in a
  53051. 39:35:08collection divided by the size of the
  53052. 39:35:10collection.
  53053. 39:35:12For example, we have a collection where
  53054. 39:35:14we have values as 7 3 4 1 6 and 7.
  53055. 39:35:20We find out the sum of these values
  53056. 39:35:22which is 28 and there are total of six
  53057. 39:35:26values. So 28 / 6 gives us a mean value
  53058. 39:35:30of 4.66.
  53059. 39:35:33Median,
  53060. 39:35:35it is the middle value in the set of the
  53061. 39:35:37data that has been sorted in ascending
  53062. 39:35:39order.
  53063. 39:35:41It is a better alternative to mean since
  53064. 39:35:43it is less impacted by outliers and
  53065. 39:35:46skewess.
  53066. 39:35:47It is closer to the actual central
  53067. 39:35:50value.
  53068. 39:35:51Median is calculated differently for
  53069. 39:35:54different sizes of data.
  53070. 39:35:56Differentiated as if the total number of
  53071. 39:35:58values is odd or if the total number of
  53072. 39:36:02values is even. If the size of the data
  53073. 39:36:05is odd. For example, in this case we
  53074. 39:36:09have five elements.
  53075. 39:36:12After sorting whatever middle value we
  53076. 39:36:15get
  53077. 39:36:17that means n + 1 by 2 term in this case
  53078. 39:36:225 + 1 / 2
  53079. 39:36:25that is the third term which is four is
  53080. 39:36:28the median value.
  53081. 39:36:31In case when the total number of values
  53082. 39:36:33is even like here there are six values.
  53083. 39:36:37The average or the mean of the two
  53084. 39:36:39central values is considered as the
  53085. 39:36:41median. In this case the median is the
  53086. 39:36:44mean of 6 and four which is five. Mode.
  53087. 39:36:50Mode represents the most common value in
  53088. 39:36:52the data set. It is not at all affected
  53089. 39:36:55by extreme observations.
  53090. 39:36:59It is the best measure of central
  53091. 39:37:01tendency for highly skewed or non-normal
  53092. 39:37:04distribution.
  53093. 39:37:05Mode for categorical data is determined
  53094. 39:37:08by estimating the frequencies for each
  53095. 39:37:10categories
  53096. 39:37:12and then the category with the highest
  53097. 39:37:14frequency is considered to be mode.
  53098. 39:37:17Like in this case 7 has the highest
  53099. 39:37:20frequency. Hence seven becomes the mode
  53100. 39:37:22value. However, in case of continuous
  53101. 39:37:26data or quantitative data, the
  53102. 39:37:28calculation of mode is slightly
  53103. 39:37:30different. The first step in calculation
  53104. 39:37:33of mode is dividing the data into
  53105. 39:37:35classes which are equal with then
  53106. 39:37:37getting the frequency of data points
  53107. 39:37:39lying in within that range of classes
  53108. 39:37:42and finally selecting the class with the
  53109. 39:37:45highest frequency.
  53110. 39:37:48Using the range of that class and the
  53111. 39:37:50frequencies, we can get the final mode
  53112. 39:37:52value.
  53113. 39:37:54Using the formula L+
  53114. 39:37:58minus F_sub_1 multiplied to H / FM minus
  53115. 39:38:02F_sub_1 plus FM minus F_sub_2.
  53116. 39:38:06Here L is the lower limit or the lower
  53117. 39:38:09observation of the mode class.
  53118. 39:38:12H is the size of the mode class.
  53119. 39:38:16FM is the frequency of the mode class.
  53120. 39:38:19F_sub_1 is the frequency of the class
  53121. 39:38:22proceeding to mode and F_sub_2 is the
  53122. 39:38:25frequency of the class succeeding to
  53123. 39:38:28mode. This gives us the final mode
  53124. 39:38:30value,
  53125. 39:38:32mean versus expectation.
  53126. 39:38:35Now let's talk about mean versus
  53127. 39:38:37expectation.
  53128. 39:38:38So in general we use the expected value
  53129. 39:38:41or expectation when we want to calculate
  53130. 39:38:44the mean of a probability distribution
  53131. 39:38:47that represents the average value we
  53132. 39:38:50expect to occur before collecting any
  53133. 39:38:52data. And mean on the other hand mean is
  53134. 39:38:55basically used when we want to calculate
  53135. 39:38:58the average value of a given sample.
  53136. 39:39:01This represents the average value of raw
  53137. 39:39:04data that we may have already collected.
  53138. 39:39:07We can understand this by using a simple
  53139. 39:39:10example.
  53140. 39:39:12Now to calculate the expected value of
  53141. 39:39:15this probability distribution, we can
  53142. 39:39:18use a specific formula from the previous
  53143. 39:39:20discussion.
  53144. 39:39:22This is going to be the expected value
  53145. 39:39:24where X is going to be the data value
  53146. 39:39:27and this PX is the probability of value.
  53147. 39:39:32For example, we could calculate the
  53148. 39:39:34expected value for this probability
  53149. 39:39:36distribution to be as shown.
  53150. 39:39:41So here it will be 1.45 goals.
  53151. 39:39:46So this represents the expected number
  53152. 39:39:48of goals that the team will score in any
  53153. 39:39:50given game.
  53154. 39:39:53And then if you talk about calculating
  53155. 39:39:55mean, so we typically calculate the mean
  53156. 39:39:58after we have actually collected raw
  53157. 39:40:00data.
  53158. 39:40:03For example, suppose we record the
  53159. 39:40:05number of goals that a soccer team will
  53160. 39:40:07score in 15 different games.
  53161. 39:40:12Now to calculate the mean number of
  53162. 39:40:14goals scored per game,
  53163. 39:40:17we can use the following formula
  53164. 39:40:20where sum of x is basically the sum of
  53165. 39:40:23all the goals divided by n and the
  53166. 39:40:26number of records or we can say the
  53167. 39:40:27sample size.
  53168. 39:40:30It is as shown on the screen.
  53169. 39:40:48So this represents the mean number of
  53170. 39:40:50goals scored per game by the team.
  53171. 39:40:53Measures of asymmetry.
  53172. 39:40:56The difference between the three
  53173. 39:40:57distinct curves can be studied in this
  53174. 39:41:00image.
  53175. 39:41:01The central curve is the normal or no
  53176. 39:41:04skewess curve. Here mean, median and
  53177. 39:41:07mode all lie on the same point. This
  53178. 39:41:10normal curve is symmetrical about its
  53179. 39:41:12mean, median and mode.
  53180. 39:41:16That means the left hand side of the
  53181. 39:41:18curve is a mirror image of the right
  53182. 39:41:20hand side of the curve.
  53183. 39:41:23However, in case of negatively skewed
  53184. 39:41:26data, the tail is elongated on the left
  53185. 39:41:30hand side
  53186. 39:41:32and the mean is smaller than the mode
  53187. 39:41:34and the median values or is on the left
  53188. 39:41:37hand side of the mode.
  53189. 39:41:40Hence indicating that the outliers are
  53190. 39:41:42in the negative direction.
  53191. 39:41:45On the other hand, in case of positively
  53192. 39:41:48skewed, the data is concentrated on the
  53193. 39:41:50left hand side of the curve.
  53194. 39:41:54While the tail is elongated or longer on
  53195. 39:41:56the right hand side of the curve,
  53196. 39:42:00the mean is greater than the mode and
  53197. 39:42:02median
  53198. 39:42:03or is on the right hand side of the mode
  53199. 39:42:06and median indicating that the outliers
  53200. 39:42:08are in the positive direction.
  53201. 39:42:15Let's consider an example.
  53202. 39:42:18The graph here shows the global income
  53203. 39:42:20distribution for the year 2003 2013 and
  53204. 39:42:25a projection for 2035.
  53205. 39:42:28If we see the global income distribution
  53206. 39:42:30statistics for 2003 it is highly right
  53207. 39:42:34skewed.
  53208. 39:42:37We can observe in the previous graph
  53209. 39:42:39that in 2003
  53210. 39:42:43the mean of $3,451
  53211. 39:42:48was higher than the median of $1090.
  53212. 39:42:52The global income is definitely not
  53213. 39:42:55evenly distributed. The majority of
  53214. 39:42:57people make less than $2,000 each year.
  53215. 39:43:03while only a small percentage of the
  53216. 39:43:05population earns more than $14,000.
  53217. 39:43:10Measures of variability.
  53218. 39:43:15Measures of variability.
  53219. 39:43:17Dispersion. The measure of central
  53220. 39:43:20tendencies provide a single value that
  53221. 39:43:22addresses the full worth. However, the
  53222. 39:43:25central tendency cannot depict the
  53223. 39:43:27viewpoint entirely. The metric of
  53224. 39:43:30dispersion helps us focus on the
  53225. 39:43:32inconsistency in the data spread.
  53226. 39:43:35Measures of dispersion describe the
  53227. 39:43:37spread of the data.
  53228. 39:43:40The range, intercortile range, standard
  53229. 39:43:43deviation and variance are examples of
  53230. 39:43:46dispersion measures.
  53231. 39:43:49Range.
  53232. 39:43:51The range of distribution is the
  53233. 39:43:53difference between the largest and the
  53234. 39:43:55smallest amount of data.
  53235. 39:43:58The range, for example, does not include
  53236. 39:44:01all of a series positive aspects.
  53237. 39:44:04It concentrates on the most shocking
  53238. 39:44:07aspects and ignores that aren't
  53239. 39:44:09considered critical. For example, for a
  53240. 39:44:11set 13, 33, 45, 67, 70.
  53241. 39:44:17The range is 57. That is the maximum of
  53242. 39:44:22this which is 70 minus the minimum over
  53243. 39:44:24here which is 13.
  53244. 39:44:28Variance.
  53245. 39:44:31Variance is the average of all squared
  53246. 39:44:33deviations.
  53247. 39:44:36It is defined as the sum of squared
  53248. 39:44:38distance between each point and the mean
  53249. 39:44:41or the dispersion around the mean.
  53250. 39:44:44The standard deviation is used as
  53251. 39:44:46variance suffers from a unit difference.
  53252. 39:44:51Variance can be computed as sigma square
  53253. 39:44:54summation over x - mu^ 2
  53254. 39:44:58divided by n
  53255. 39:45:00where mu is the mean of the data, x is
  53256. 39:45:03the individual data point
  53257. 39:45:06and n is the size of the data.
  53258. 39:45:11This representation is for a population
  53259. 39:45:13data.
  53260. 39:45:15For a sample data variance can be
  53261. 39:45:18computed as X minus
  53262. 39:45:20Xar whole square summation
  53263. 39:45:23over it divided by n minus one.
  53264. 39:45:27Here Xbar is the mean of these sample
  53265. 39:45:30data and n is the sample size.
  53266. 39:45:34The units of values and variance are not
  53267. 39:45:37equal.
  53268. 39:45:39So another variability measure is used.
  53269. 39:45:43Standard deviation.
  53270. 39:45:47Standard deviation is a statistical term
  53271. 39:45:49used to measure the amount of
  53272. 39:45:51variability or dispersion around a mean.
  53273. 39:45:56The standard deviation is calculated as
  53274. 39:45:59the square root of variance. It depicts
  53275. 39:46:02the concentration of the data around the
  53276. 39:46:05mean of the data set.
  53277. 39:46:08Standard deviation as indicated
  53278. 39:46:11previously can be computed as square
  53279. 39:46:13root of variance
  53280. 39:46:15for a population data. Standard
  53281. 39:46:18deviation sigma can be computed as
  53282. 39:46:21square root of summation over x i minus
  53283. 39:46:24mu^ square / n
  53284. 39:46:28where mu is the mean of the data x i are
  53285. 39:46:32the data points and n is the size. Let's
  53286. 39:46:35consider an example.
  53287. 39:46:38Let's find out the mean, variance, and
  53288. 39:46:41standard deviation for this data. The
  53289. 39:46:44data values are three, 5, 6, 9, and 10.
  53290. 39:46:49To find out the mean, we first find the
  53291. 39:46:51sum of all these data values
  53292. 39:46:55that is 33 and divide it by the count,
  53293. 39:46:58which is five.
  53294. 39:47:01We get the mean of 6.6. To compute the
  53295. 39:47:04variance, we start by computing the
  53296. 39:47:07deviation.
  53297. 39:47:08That is X minus the mean of X. Here
  53298. 39:47:12three is one of the values of the data
  53299. 39:47:14and 6.6 is the mean.
  53300. 39:47:18So 3 - 6.6 squared and we do that
  53301. 39:47:24to find out sum of all the deviations
  53302. 39:47:26divided by the count
  53303. 39:47:29which is five.
  53304. 39:47:31we end up getting an overall variance of
  53305. 39:47:336.64.
  53306. 39:47:37Standard deviation as we know is
  53307. 39:47:39measured at square root of variance that
  53308. 39:47:42is square<unk> of 6.64
  53309. 39:47:46which amounts to 2.576.
  53310. 39:47:50Measures of relationship.
  53311. 39:47:53Measures of relationship. Coariance.
  53312. 39:47:56Coariance is the measure of joint
  53313. 39:47:58variability of two variables.
  53314. 39:48:01It measures the direction of the
  53315. 39:48:03relationship between the variables. It
  53316. 39:48:06determines if one variable will cause
  53317. 39:48:08the other to alter in the same way.
  53318. 39:48:13Coariance between variable X and Y can
  53319. 39:48:16be computed as summation over the
  53320. 39:48:19product of X I - XR
  53321. 39:48:22and Y I - Y bar the whole divided by N
  53322. 39:48:26minus one.
  53323. 39:48:29Here Xar and Y bar are the mean of X and
  53324. 39:48:32Y respectively. The value of covariance
  53325. 39:48:36can range from minus infinity to a plus
  53326. 39:48:38infinity.
  53327. 39:48:41Correlation.
  53328. 39:48:43Correlation is normalized coariance.
  53329. 39:48:47It measures the strength of association
  53330. 39:48:50between two variables. The most common
  53331. 39:48:52measure for correlation is the Pearson
  53332. 39:48:55correlation coefficient.
  53333. 39:48:57Correlation between two variables
  53334. 39:49:01X and Y can be measured with respect to
  53335. 39:49:03coariance as coariance between X
  53336. 39:49:07and Y divided by the standard deviation
  53337. 39:49:10of X and standard deviation of Y.
  53338. 39:49:14The value of correlation ranges from a
  53339. 39:49:17negative 1 to positive 1.
  53340. 39:49:21Types of correlation.
  53341. 39:49:24Correlation can be either a positive
  53342. 39:49:27correlation,
  53343. 39:49:28zero correlation or a negative
  53344. 39:49:31correlation.
  53345. 39:49:35The first picture over here represents a
  53346. 39:49:37perfect positive correlation
  53347. 39:49:41wherein a straight line with a positive
  53348. 39:49:43slope
  53349. 39:49:45is representing the relationship between
  53350. 39:49:47the two variables.
  53351. 39:49:50Zero correlation means that the line
  53352. 39:49:52representing the relationship between
  53353. 39:49:54the two variables is horizontal to the
  53354. 39:49:57xaxis.
  53355. 39:50:00Perfect negative correlation can be
  53356. 39:50:02represented by a straight line with a
  53357. 39:50:05negative slope.
  53358. 39:50:08Correlation equals to 1 implies a
  53359. 39:50:11positive relationship. That is when one
  53360. 39:50:14variable increases the other variable
  53361. 39:50:16also increases. A correlation value of
  53362. 39:50:19negative 1 implies a negative
  53363. 39:50:21relationship. That is when one variable
  53364. 39:50:24increases the other decreases.
  53365. 39:50:28The correlation coefficient of zero
  53366. 39:50:31shows that the variables are completely
  53367. 39:50:33independent of each other.
  53368. 39:50:36Let's consider an example.
  53369. 39:50:40Here we have two variables height and
  53370. 39:50:43weight.
  53371. 39:50:46To compute the correlation between
  53372. 39:50:48height and weight,
  53373. 39:50:50we use the correlation formula as
  53374. 39:50:52covariance of X
  53375. 39:50:55and Y divided by standard deviation of X
  53376. 39:50:58and standard deviation of Y.
  53377. 39:51:01Here height is the X variable and weight
  53378. 39:51:04is the Y variable.
  53379. 39:51:07First to compute coariance we compute
  53380. 39:51:10the x - xar and y - y bar values and
  53381. 39:51:14then the product of them.
  53382. 39:51:18We then compute x - xr²
  53383. 39:51:23and y - y bar square values to compute
  53384. 39:51:26the standard deviations of height and
  53385. 39:51:28weight respectively. Correlation as we
  53386. 39:51:31know has been defined as covariance of x
  53387. 39:51:34and i and y divided by standard
  53388. 39:51:37deviations of x and y.
  53389. 39:51:41This can also be represented as
  53390. 39:51:43summation over x - xr multiplied to y -
  53391. 39:51:48y bar
  53392. 39:51:50divided by square root of summation over
  53393. 39:51:52sum of squared deviations
  53394. 39:51:55that is x - xr square multiplied to
  53395. 39:51:58square root of summation over y - yar
  53396. 39:52:02whole square that is sum of square
  53397. 39:52:05deviations for y.
  53398. 39:52:08Now let's find out values to put into
  53399. 39:52:11this formula.
  53400. 39:52:14First we find out the overall sum of
  53401. 39:52:17height to get the mean of height which
  53402. 39:52:20is 5.14.
  53403. 39:52:22Similarly we get the sum of weight to
  53404. 39:52:24get the mean of weight as 50. We now get
  53405. 39:52:27the summation over x - xr multiplied to
  53406. 39:52:31y - y bar to get the numerator for the
  53407. 39:52:34formula. Then we compute x - xr square
  53408. 39:52:38summation
  53409. 39:52:40and y - y bar square that is sum of
  53410. 39:52:43squared deviation of x and y
  53411. 39:52:46respectively.
  53412. 39:52:48Now we put in the values in this final
  53413. 39:52:50correlation formula to get a correlation
  53414. 39:52:52value of 0.889.
  53415. 39:52:57This indicates that height and weight
  53416. 39:52:59have a positive relationship.
  53417. 39:53:02It is evident that as height grows,
  53418. 39:53:05weight also increases.
  53419. 39:53:09In this module, we will be talking about
  53420. 39:53:11expectation and variance.
  53421. 39:53:14So the expected value or we can say mean
  53422. 39:53:17of a given variable that we can denote
  53423. 39:53:20by X is a discrete random variable where
  53424. 39:53:22it is a weighted average of the possible
  53425. 39:53:25values that X can take and each value is
  53426. 39:53:28going to be according to the probability
  53427. 39:53:30of that specific event occurring.
  53428. 39:53:34So usually the expected value of X is
  53429. 39:53:36denoted by a simple formula where we can
  53430. 39:53:40define the expectation based on the X
  53431. 39:53:42parameter.
  53432. 39:53:45which is going to be the sum of each
  53433. 39:53:47possible outcome multiplied by the
  53434. 39:53:49probability of the outcome occurring.
  53435. 39:53:53So in more concrete terms, the
  53436. 39:53:56expectation is what we would expect the
  53437. 39:53:58outcome of an experiment to be on
  53438. 39:54:00average.
  53439. 39:54:04We can take an example for the coin. If
  53440. 39:54:07a coin is being tossed 10 times, then
  53441. 39:54:10one is most likely to get five heads and
  53442. 39:54:13five tails.
  53443. 39:54:16Same logic can be discussed if we talk
  53444. 39:54:19about another example of rolling a die.
  53445. 39:54:22So there are six possible outcomes when
  53446. 39:54:24you roll a dieice 1 2 3 4 5 6. And each
  53447. 39:54:28of these has a probability of 1 by 6 of
  53448. 39:54:31occurring. So we can say that the
  53449. 39:54:34expectation is going to be 1 multiplied
  53450. 39:54:36by the probability of that happening
  53451. 39:54:39which is going to be 1x 6 + 2x 6 + 3x 6
  53452. 39:54:44+ 4x 6 + 5x 6 + 6x 6 and that is going
  53453. 39:54:50to give us 3.5 as an output. The
  53454. 39:54:53expected value is 3.5.
  53455. 39:54:57So if you think about it, 3.5 is halfway
  53456. 39:55:00between the possible values that I can
  53457. 39:55:02take and this is what we should have
  53458. 39:55:05expected.
  53459. 39:55:06Next we talk about the concept of
  53460. 39:55:08variance. So variance of a random
  53461. 39:55:11variable allows us to know something
  53462. 39:55:13about the spread of the possible values
  53463. 39:55:16of the variable. So for a discrete
  53464. 39:55:19random variable X the variances of X is
  53465. 39:55:22going to be denoted by using a simple
  53466. 39:55:24formula that is going to be var=
  53467. 39:55:27E X - M the whole square where M is
  53468. 39:55:31basically the expected value of the
  53469. 39:55:33expectation of X. So this is more like a
  53470. 39:55:36standard deviation of X which can also
  53471. 39:55:39be represented by using this formula. So
  53472. 39:55:42the variance does not behave in the same
  53473. 39:55:44way as expectation when we multiply and
  53474. 39:55:47add constants to random variables.
  53475. 39:55:52So now there are two different type of
  53476. 39:55:54variance that we can have a fair
  53477. 39:55:56understanding on. First of all we have
  53478. 39:55:59low variance and then we have high
  53479. 39:56:02variance.
  53480. 39:56:04So low variance simply means that there
  53481. 39:56:06is a small variation in the production
  53482. 39:56:09of the target function with changes in
  53483. 39:56:12the trading data set and at the same
  53484. 39:56:14time high variance as we can see here
  53485. 39:56:17high variance shows a large variation in
  53486. 39:56:20prediction of the target function with
  53487. 39:56:22changes in the trading data set. So a
  53488. 39:56:25model that shows high variance learns a
  53489. 39:56:27lot and perform well with the training
  53490. 39:56:29data set and it does not generalize well
  53491. 39:56:32with the unseen data set and that's why
  53492. 39:56:35as a result such a model gives good
  53493. 39:56:38results with training data set but shows
  53494. 39:56:40high error rates on the test data set
  53495. 39:56:43and since the high variance a model
  53496. 39:56:45learns too much from the data set it
  53497. 39:56:48leads to an overfitting of the model. So
  53498. 39:56:51model with high variance will be having
  53499. 39:56:53couple of issues like it may lead to
  53500. 39:56:55overfitting or it may also lead to
  53501. 39:56:57increase in model complexities.
  53502. 39:57:02Next we have skewess.
  53503. 39:57:05So skewess in simple terms is basically
  53504. 39:57:08a measure of asymmetry of a
  53505. 39:57:10distribution. So distribution is
  53506. 39:57:13asymmetrical when its left and right
  53507. 39:57:15sides are not the mirror images.
  53508. 39:57:18Right now this is a mirrored image and a
  53509. 39:57:21distribution can have right positive or
  53510. 39:57:23we can say negative or it can have zero
  53511. 39:57:26skewess.
  53512. 39:57:29So right skewed in this scenario is
  53513. 39:57:32basically the distribution is longer on
  53514. 39:57:34the right side of its peak
  53515. 39:57:37and a left skew distribution is going to
  53516. 39:57:39be we can say where it is longer on the
  53517. 39:57:42left side.
  53518. 39:57:44So we can see we have this one as a part
  53519. 39:57:46of right side. It is more elongated
  53520. 39:57:49towards the right side and this one is
  53521. 39:57:52more elongated towards the left side. So
  53522. 39:57:55we can think of skewess in terms of
  53523. 39:57:57tails. A tail is long tampering and the
  53524. 39:58:00end of a distribution. So it simply
  53525. 39:58:03indicates that they are observations at
  53526. 39:58:05one end of the distribution but that
  53527. 39:58:07they are relatively infrequent. So a
  53528. 39:58:10right skew distribution has a long tail
  53529. 39:58:13on the right side as you can see here.
  53530. 39:58:15So the number supports observed. Let's
  53531. 39:58:18say we have a data on a per year basis.
  53532. 39:58:21So again we can have a more skewess
  53533. 39:58:23towards the right side where data is
  53534. 39:58:25being dropping as we continue to
  53535. 39:58:28increase the number of years. For
  53536. 39:58:30example we may have a high sales towards
  53537. 39:58:32the beginning of year suppose in 2022
  53538. 39:58:36but again as we proceed to 2023 second
  53539. 39:58:40half we are seeing the dip in
  53540. 39:58:42performance. So that is rightly skewed
  53541. 39:58:45and same way let's suppose if we started
  53542. 39:58:47with the sales figure it was really less
  53543. 39:58:50in suppose 2002
  53544. 39:58:53but again as we proceeded to 2023 now
  53545. 39:58:56our sales have been gradually
  53546. 39:58:58increasing. So it's more like skew
  53547. 39:59:01towards the left section as a part of
  53548. 39:59:03negative skew. Next we have curtosis.
  53549. 39:59:08So curtosis is basically a measure of
  53550. 39:59:11the tailness of a distribution.
  53551. 39:59:14So taeness is how often the outliers
  53552. 39:59:17occur and act as curtis is the tailness
  53553. 39:59:20of the distribution related to a normal
  53554. 39:59:23distribution. So a distribution with
  53555. 39:59:26medium curttosis is called as messortic.
  53556. 39:59:29A distribution with low curtosis like
  53557. 39:59:31this one. This is called as the
  53558. 39:59:33platicurtic and then distribution with
  53559. 39:59:36high curtosis like this one. This is
  53560. 39:59:38called as the leptocortic.
  53561. 39:59:42So tails here they are tapering ends on
  53562. 39:59:44either side of a distribution like this.
  53563. 39:59:47So they represent the probability or the
  53564. 39:59:49frequency of values that are extremely
  53565. 39:59:52high or extremely low to the mean.
  53566. 39:59:55In other words, tails here represents
  53567. 39:59:58how often the outliers occur.
  53568. 40:00:02So there are three type of curtis. We
  53569. 40:00:05have platicurtic which is negative,
  53570. 40:00:07leptocortic which is a positive towards
  53571. 40:00:09the upper end and then we have messertic
  53572. 40:00:12which is a normal distribution. So
  53573. 40:00:15messertic is the medium tail. So normal
  53574. 40:00:18distributions they have a curtosis of
  53575. 40:00:20three. So any distribution with a
  53576. 40:00:23kurtosis of approx value of three is
  53577. 40:00:26going to be messertic. And curtosis is
  53578. 40:00:29described in terms of excess curttosis
  53579. 40:00:32which is curtosis minus3. And since
  53580. 40:00:35normal distribution they have a curtosis
  53581. 40:00:38of three axis curtises makes comparing a
  53582. 40:00:41distribution curtosis to a normal
  53583. 40:00:43distribution even easier. Introduction
  53584. 40:00:46to probability.
  53585. 40:00:49Probability theory. Probability is a
  53586. 40:00:52measure of the likelihood that an event
  53587. 40:00:54will occur.
  53588. 40:00:57Let's consider an example of coin toss
  53589. 40:01:01where the chances of getting heads on a
  53590. 40:01:03coin are 1 by two or 50%.
  53591. 40:01:07The probability of each given event is
  53592. 40:01:09between zero and one both inclusive. Sum
  53593. 40:01:13of an events cumulative probability
  53594. 40:01:16cannot be greater than one.
  53595. 40:01:19Hence the probability of an event X lies
  53596. 40:01:22between zero and one. This means that
  53597. 40:01:25the integral of probability of
  53598. 40:01:27distribution over x equals to 1.
  53599. 40:01:33Conditional probability. Conditional
  53600. 40:01:35probability of any event A is defined as
  53601. 40:01:38the probability of occurrence of A given
  53602. 40:01:42that event B has previously occurred.
  53603. 40:01:47Condition probability of event A given B
  53604. 40:01:50can be estimated as probability of A
  53605. 40:01:53intersection B that is probability of
  53606. 40:01:56both A and B happening together
  53607. 40:01:59divided by the probability of B.
  53608. 40:02:04It is also written as that probability
  53609. 40:02:06of A intersection B equals to
  53610. 40:02:09probability of A given B multiplied to
  53611. 40:02:13probability of B.
  53612. 40:02:18Let's consider an example.
  53613. 40:02:20In a coin, we are doing a two coin flip.
  53614. 40:02:23Coin one gets heads, tails, heads, and
  53615. 40:02:26tails in subsequent flips.
  53616. 40:02:30while coin two gets tails, heads, heads,
  53617. 40:02:33and tails in the subsequent flips. Now,
  53618. 40:02:37the probability that coin one will get a
  53619. 40:02:39head is 2 out of four. While the
  53620. 40:02:42probability that coin two will get heads
  53621. 40:02:45is again two out of four.
  53622. 40:02:48The probability that both coin one and
  53623. 40:02:50coin two will have a heads is just one
  53624. 40:02:53out of the four flips.
  53625. 40:02:57Hence the probability that coin one will
  53626. 40:02:59get heads given that coin 2 is already
  53627. 40:03:02heads can be computed as probability of
  53628. 40:03:05coin one edge intersection coin 2 edge
  53629. 40:03:09that is 1x4 divided by probability of
  53630. 40:03:12coin 2 edge
  53631. 40:03:16that's a given that is 2x 4 which is
  53632. 40:03:19going to be 0.5 or 50% based
  53633. 40:03:24base theorem Base theorem calculates the
  53634. 40:03:27conditional probability of an event
  53635. 40:03:29based on its prior probabilities.
  53636. 40:03:33Basically base theorem incorporates the
  53637. 40:03:36prior probability distribution to
  53638. 40:03:38predict the posterior probabilities.
  53639. 40:03:40Base theorem for conditional probability
  53640. 40:03:44can be expressed as probability of A
  53641. 40:03:47given B equals probability of B given A
  53642. 40:03:51divided by probability of B multiplied
  53643. 40:03:54to probability of A.
  53644. 40:03:57Base theorem allows updating the
  53645. 40:03:59probability values by using new
  53646. 40:04:01information or evidence. Here
  53647. 40:04:04probability of A is known as prior
  53648. 40:04:06probability. That is the probability of
  53649. 40:04:09event before any new data is collected.
  53650. 40:04:12Probability of A given B is known as the
  53651. 40:04:16posterior probability. It is the revised
  53652. 40:04:19probability of an event occurring after
  53653. 40:04:21taking into consideration the new
  53654. 40:04:24information probability of B given A is
  53655. 40:04:27known as the likelihood and probability
  53656. 40:04:29of B is probability of observing an
  53657. 40:04:32evidence B model. An example consider an
  53658. 40:04:36example for calculating the likelihood
  53659. 40:04:38of having diabetes based on frequency of
  53660. 40:04:41fast food consumption. Here is the
  53661. 40:04:44observed data. Let's say the fast food
  53662. 40:04:47audience is 20%. Diabetes prevalence is
  53663. 40:04:5110% and 5% is fast food and diabetes.
  53664. 40:04:56The chances of diabetes given fast food
  53665. 40:04:58that is the conditional probability of D
  53666. 40:05:01given B can be calculated as probability
  53667. 40:05:04of diabetes and fast food together
  53668. 40:05:07divided by probability of fast food.
  53669. 40:05:10That means 5% divided by 20%. that
  53670. 40:05:14equals 25%.
  53671. 40:05:16Define an analysis can state eating fast
  53672. 40:05:19food increases the chance of having
  53673. 40:05:21diabetes by 25%.
  53674. 40:05:24The multiplication rule of probability
  53675. 40:05:27if events A and B are statistically
  53676. 40:05:30independent and probability of A
  53677. 40:05:33intersection B can be given as
  53678. 40:05:35probability of A given B multiplied to
  53679. 40:05:39probability of B. However, probability
  53680. 40:05:42of A intersection B is also given as
  53681. 40:05:45probability of A multiplied to
  53682. 40:05:48probability of B. Here probability of A
  53683. 40:05:52given B equals to probability of A when
  53684. 40:05:56we assume that probability of B is non
  53685. 40:05:59zero. Similarly, probability of B equals
  53686. 40:06:02probability of B given A assuming
  53687. 40:06:05probability of A is non zero.
  53688. 40:06:08Chain rule of probability joint
  53689. 40:06:10probability distributions over many
  53690. 40:06:13random variables can be reduced into
  53691. 40:06:16conditional distributions over a single
  53692. 40:06:18variable. It can be expressed as
  53693. 40:06:21probability of X1 X2 so on until Xn
  53694. 40:06:25equals probability of X1 intersection
  53695. 40:06:28probability of X I given probability of
  53696. 40:06:31X1 till X I minus one.
  53697. 40:06:36For example, the joint probability of A,
  53698. 40:06:38B and C can be given as probability of A
  53699. 40:06:42given B. C multiplied to probability of
  53700. 40:06:46B given C multiply to probability of C.
  53701. 40:06:50Logistic sigmoid.
  53702. 40:06:54The logistics function is a type of
  53703. 40:06:56sigmoid function that aims to predict
  53704. 40:06:58the class to which a particular sample
  53705. 40:07:01belongs. Its outcome is discrete binary
  53706. 40:07:04value. a probability between zero and
  53707. 40:07:07one. The logistic sigmoid is a useful
  53708. 40:07:10function that follows the yes curve. It
  53709. 40:07:13saturates when the input is very large
  53710. 40:07:15or very small. Logistic sigmoid is
  53711. 40:07:19expressed as sigma of x= 1 upon 1 + e to
  53712. 40:07:23the power minus x.
  53713. 40:07:26The logistic sigmoid can be expressed as
  53714. 40:07:29sigmoid function of x is given as 1 upon
  53715. 40:07:321 + e ^ minus x where e is the ooler's
  53716. 40:07:36number.
  53717. 40:07:38Gshian distribution.
  53718. 40:07:41The gossian distribution is a type of
  53719. 40:07:43distribution in which data tends to
  53720. 40:07:45cluster around a central value with
  53721. 40:07:48little or no bias to the left or right.
  53722. 40:07:52It is often referred to as normal
  53723. 40:07:54distribution.
  53724. 40:07:56In absence of prior information, the
  53725. 40:07:59normal distribution is frequently a fair
  53726. 40:08:01assumption in machine learning
  53727. 40:08:04equation.
  53728. 40:08:06The formula for calculating Gaussian
  53729. 40:08:08distribution is described as the normal
  53730. 40:08:11distribution of X.
  53731. 40:08:14That is the function of x given mean as
  53732. 40:08:16mu and variance is sigma square can be
  53733. 40:08:19calculated as 1 upon sigma square
  53734. 40:08:22roo<unk> of 2 pi e to the power -/ x -
  53735. 40:08:26mood / sigma square
  53736. 40:08:30where mu is the mean or peak value which
  53737. 40:08:33also is the expected value of x.
  53738. 40:08:37Sigma is the standard deviation. Sigma
  53739. 40:08:40square is the variance.
  53740. 40:08:42A standard normal distribution has a
  53741. 40:08:44mean of zero and a standard deviation of
  53742. 40:08:47one.
  53743. 40:08:50Gshian distribution can be univariate
  53744. 40:08:54which describes the distribution of a
  53745. 40:08:56single variable X.
  53746. 40:08:58It can also be multivariate where it can
  53747. 40:09:01just use to describe the distribution of
  53748. 40:09:03several variables.
  53749. 40:09:06It is represented in 3D of ND formats.
  53750. 40:09:12Law of large numbers.
  53751. 40:09:16Now let's talk about law of large
  53752. 40:09:18numbers. The law of large numbers states
  53753. 40:09:21that an observed sample average from a
  53754. 40:09:24large sample will be close to the true
  53755. 40:09:26population average and that it will get
  53756. 40:09:28closer in the larger sample. So the law
  53757. 40:09:32of large number does not guarantee that
  53758. 40:09:34a given sample spatially a small sample
  53759. 40:09:36will reflect the true population
  53760. 40:09:38characteristics or that a sample does
  53761. 40:09:41not reflect the true population will be
  53762. 40:09:43balanced by a subsequent sample. This is
  53763. 40:09:46for the law of large numbers to express
  53764. 40:09:49the relationship between scale and
  53765. 40:09:51growth rate.
  53766. 40:09:54So there are multiple examples through
  53767. 40:09:56which we can understand
  53768. 40:10:00and it is widely used in statistical
  53769. 40:10:02analysis in working with the central
  53770. 40:10:04limit theorem in terms of the business
  53771. 40:10:06growth. So there are multiple real time
  53772. 40:10:09setup in which these are going to be
  53773. 40:10:11used. So if you talk about tossing a
  53774. 40:10:14coin so tossing a coin in a number of
  53775. 40:10:17times will give us two different type of
  53776. 40:10:19outcomes.
  53777. 40:10:22the result will spread evenly between
  53778. 40:10:24head and tails and the expected average
  53779. 40:10:27value is going to be half.
  53780. 40:10:29That means 50 times tails and 30 times
  53781. 40:10:32heads. But again, if you toss a coin
  53782. 40:10:351,000 times, then the result can be in
  53783. 40:10:38different manners because out of 1,000,
  53784. 40:10:41let's say 850 times it has been head and
  53785. 40:10:45only 150 times it has been tails and so
  53786. 40:10:49on. So that's why the possibility of one
  53787. 40:10:51event occurring is going to be changed
  53788. 40:10:54in large sample sets as compared to a
  53789. 40:10:56small sample sets as in let's say 10
  53790. 40:10:59times. So the number of heads and tails
  53791. 40:11:02unbalanced for lower number of trials.
  53792. 40:11:04So we can see it is unbalanced.
  53793. 40:11:08But again as soon as we toss more number
  53794. 40:11:10of coins more leans towards the balance
  53795. 40:11:13value or we can see the observed
  53796. 40:11:15averages.
  53797. 40:11:17Next we have p value.
  53798. 40:11:20So p value is basically a number
  53799. 40:11:23calculated from the statistical test
  53800. 40:11:26that describes how likely we are to have
  53801. 40:11:28found a particular set of observations
  53802. 40:11:30if the null hypothesis were true. So p
  53803. 40:11:34values are used in hypothesis testing to
  53804. 40:11:37help decide whether to reject the null
  53805. 40:11:39hypothesis. And the smaller the p value,
  53806. 40:11:42the more likely we are to reject the
  53807. 40:11:45null hypothesis.
  53808. 40:11:46So we have a term called as null
  53809. 40:11:48hypothesis. So all statistical tests
  53810. 40:11:52they have null hypothesis. So for most
  53811. 40:11:55tests the null hypothesis is that there
  53812. 40:11:57is no relationship between our variables
  53813. 40:12:00of in first or that there is no
  53814. 40:12:02difference among groups. For example in
  53815. 40:12:05a two-tail t test the non-hypothesis is
  53816. 40:12:08that the difference between two groups
  53817. 40:12:10is going to be zero.
  53818. 40:12:13So p value is going to tell us how
  53819. 40:12:15likely it is that our data could have
  53820. 40:12:17occurred under the null hypothesis.
  53821. 40:12:21It is done by calculating the likelihood
  53822. 40:12:23of a test statistic
  53823. 40:12:25which is the number calculated by a
  53824. 40:12:27statistical test using our data. So p
  53825. 40:12:30value tell us how often we would expect
  53826. 40:12:33to see a test statistic as extreme or
  53827. 40:12:36more extreme
  53828. 40:12:37than one calculated by a statistical
  53829. 40:12:40test. if the null hypothesis of the test
  53830. 40:12:43was true.
  53831. 40:12:45So there are multiple limitations as
  53832. 40:12:47well. So first one is the results can be
  53833. 40:12:50significant but again they are they may
  53834. 40:12:53not be practical as we have compared it
  53835. 40:12:56can be based on multiple hypothesis for
  53836. 40:12:58a game for the healthcare test. If the
  53837. 40:13:01test is going to be positive or not it
  53838. 40:13:04may show even values of the effect of a
  53839. 40:13:06variable but not the magnitude in real
  53840. 40:13:09life. What exactly is going to be the
  53841. 40:13:11application of a drug test being failed
  53842. 40:13:14in pharma company? Therefore, it is
  53843. 40:13:17recommended to use confidence and levels
  53844. 40:13:19in addition to the p values to quantify
  53845. 40:13:22or we can say to give a solid figure to
  53846. 40:13:24the reserve which we are going to get.
  53847. 40:13:27The p values they are interpreted as
  53848. 40:13:30supporting or we can say refuting the
  53849. 40:13:32alternative hypothesis.
  53850. 40:13:34So p value can only tell you whether or
  53851. 40:13:37not the null hypothesis is supported. It
  53852. 40:13:40cannot tell us whether our alternative
  53853. 40:13:42hypothesis is true or why. So the risk
  53854. 40:13:46of rejecting the null hypothesis is
  53855. 40:13:49often higher than the p value. So
  53856. 40:13:52especially when we are looking at a
  53857. 40:13:53single study or when using small sample
  53858. 40:13:56sizes. So this is because the smaller
  53859. 40:13:59frame of reference, the greater are the
  53860. 40:14:01chance that as we stumble across a
  53861. 40:14:03statistically significant pattern
  53862. 40:14:06completely by accident.
  53863. 40:14:08Key takeaways.
  53864. 40:14:10Key takeaways. Probability and
  53865. 40:14:13statistics structure the premise of the
  53866. 40:14:15data. The data helps in anticipating the
  53867. 40:14:18future or gauging in view of the past
  53868. 40:14:21patterns of information.
  53869. 40:14:24The central tendency is a single value
  53870. 40:14:27that helps to describe the data by
  53871. 40:14:29identifying these central positions. The
  53872. 40:14:31mean, median, and mode are the measures
  53873. 40:14:34of central tendencies.
  53874. 40:14:37The distribution where the data tends to
  53875. 40:14:39be around a central value with a lack of
  53876. 40:14:42bias or minimal bias towards the left or
  53877. 40:14:45right is called as gshian distribution.
  53878. 40:14:49>> My name is Richard Kersner with the
  53879. 40:14:50simply learn team. That's get certified,
  53880. 40:14:53get ahead. We're going to cover
  53881. 40:14:54mathematics for machine learning. So
  53882. 40:14:57today's agenda is going to cover data
  53883. 40:14:59and its types. Then we're going to dive
  53884. 40:15:01into linear algebra and its concepts,
  53885. 40:15:04calculus, statistics for machine
  53886. 40:15:06learning, probability for machine
  53887. 40:15:08learning, hands-on demos, and of course
  53888. 40:15:12throwing in there in the middle is going
  53889. 40:15:13to be your matrixes and a few other
  53890. 40:15:15things to go along with all this.
  53891. 40:15:18Data and its types. Data denotes the
  53892. 40:15:20individual pieces of factual information
  53893. 40:15:22collected from various sources. It is
  53894. 40:15:25stored, processed and later used for
  53895. 40:15:26analysis.
  53896. 40:15:28And so we see here uh just a huge
  53897. 40:15:30grouping of information, a lot of tech
  53898. 40:15:32stuff, money, dollar signs, numbers
  53899. 40:15:36uh and then you have your performing
  53900. 40:15:38analytics to drive insights and
  53901. 40:15:40hopefully you have a nice share your
  53902. 40:15:41shareholders gathered at the meeting and
  53903. 40:15:43you're able to explain it in something
  53904. 40:15:44they can understand. So we talk about
  53905. 40:15:47datas types of data we have in our types
  53906. 40:15:50of data we have a qualitative
  53907. 40:15:52categorical
  53908. 40:15:54you think nominal or ordinal and then
  53909. 40:15:56you have your quantitative or numerical
  53910. 40:15:58which is discrete or continuous
  53911. 40:16:01and let's look a little closer at those
  53912. 40:16:03data type vocabulary always people's
  53913. 40:16:06favorite is the vocabulary words okay
  53914. 40:16:09not mine uh but let's dive into this
  53915. 40:16:11what we mean by nominal nominal they are
  53916. 40:16:14used to label various just uh label our
  53917. 40:16:17variables without providing any
  53918. 40:16:19measurable value. Uh country, gender,
  53919. 40:16:22race, hair, color, etc. It's something
  53920. 40:16:26that you either mark true or false. This
  53921. 40:16:28is a label. It's on or off. Either they
  53922. 40:16:30have a red hat on or they do not. Uh so
  53923. 40:16:33a lot of times when you're thinking
  53924. 40:16:34nominal data labels, uh think of it as a
  53925. 40:16:38true false kind of setup. And we look at
  53926. 40:16:40ordinal. This is categorical data with a
  53927. 40:16:42set order or a scale to it. Uh and you
  53928. 40:16:45can think of salary range is a great
  53929. 40:16:47one. Uh movie ratings etc. You see here
  53930. 40:16:50the salary range if you have 10,000 to
  53931. 40:16:5220,000 number of employees earning that
  53932. 40:16:55rate is 150 20,000 to 30,000 100 and so
  53933. 40:16:59forth. Some of the terms you'll hear is
  53934. 40:17:02bucket. Uh this is where you have 10
  53935. 40:17:04different buckets and you want to
  53936. 40:17:05separate it into something that makes
  53937. 40:17:07sense into those 10 buckets. And so when
  53938. 40:17:10we start talking about ordinal, a lot of
  53939. 40:17:12times when you get down to the brass
  53940. 40:17:14bones, again, we're talking true false.
  53941. 40:17:17Uh so if you're a member of the 10 to
  53942. 40:17:1820k range, uh so forth, those would each
  53943. 40:17:22be either part of that group or you're
  53944. 40:17:24not. But now we're talking about buckets
  53945. 40:17:26and we want to count how many people are
  53946. 40:17:27in that bucket. Quantitative numerical
  53947. 40:17:30data uh falls into two classes, discrete
  53948. 40:17:34or continuous. And so data with a final
  53949. 40:17:37set of values which can be categorized
  53950. 40:17:39class strength questions answered
  53951. 40:17:42correctly and runs hit in cricket. A lot
  53952. 40:17:45of times when you see this you can think
  53953. 40:17:47integer uh and a very restricted integer
  53954. 40:17:50i.e. you can only have 100 questions um
  53955. 40:17:53on a test. So you can it's very
  53956. 40:17:55discreet. I only have a 100 different
  53957. 40:17:56values that it can attain. So think
  53958. 40:17:59usually you're talking about integers
  53959. 40:18:01but within a very small range. They
  53960. 40:18:03don't have an open end or anything like
  53961. 40:18:04that.
  53962. 40:18:05Uh so discrete is very solid, simple to
  53963. 40:18:08count, set number. Continuous on the
  53964. 40:18:12other hand uh continuous data can take
  53965. 40:18:14any numerical value within a range. So
  53966. 40:18:17water pressure, weight of a person etc.
  53967. 40:18:20Usually we start thinking about float
  53968. 40:18:21values where they can get phenomenally
  53969. 40:18:24small in their in what they're worth.
  53970. 40:18:26And there's a whole series of values
  53971. 40:18:28that falls right between discrete and
  53972. 40:18:30continuous. Um you can think of the
  53973. 40:18:32stock market. You have dollar amounts.
  53974. 40:18:34It's still discreet, but it starts to
  53975. 40:18:37get complicated enough when you have
  53976. 40:18:39like, you know, jump in the stock market
  53977. 40:18:40from $525.33
  53978. 40:18:44to $580.67.
  53979. 40:18:48There's a lot of point values in there.
  53980. 40:18:49It'd still be called discreet, but you
  53981. 40:18:52start looking at it as almost continuous
  53982. 40:18:54because it does have such a variance in
  53983. 40:18:56it. Now uh we talk about n we did we
  53984. 40:18:59went over nominal and ordinal uh almost
  53985. 40:19:01true false charts and we looked at
  53986. 40:19:04quantitative and numerical data which
  53987. 40:19:06we're starting to get into numbers.
  53988. 40:19:08Discrete you can usually a lot of times
  53989. 40:19:10discreet will be put into it could be
  53990. 40:19:12put into true false but usually it's
  53991. 40:19:14not. Uh so we want to address this stuff
  53992. 40:19:15and the first thing we want to look at
  53993. 40:19:17is the very basic which is your algebra.
  53994. 40:19:19So we're going to take a look at linear
  53995. 40:19:21algebra. You can remember back when your
  53996. 40:19:24uklidian geometry uh we have a line.
  53997. 40:19:27Well, let's go through this. We have
  53998. 40:19:28linear algebra is the domain of
  53999. 40:19:30mathematics concerning linear equations
  54000. 40:19:33and their representations in vector
  54001. 40:19:35spaces and through matrices. I told you
  54002. 40:19:37we're going to talk about matrices. Uh
  54003. 40:19:40so a linear equation is simply um uh 2x
  54004. 40:19:44+ 4 y - 3 z = 10. Very linear. 10 x +
  54005. 40:19:5012.4 4 y = z. And now you can actually
  54006. 40:19:53solve these two equations by combining
  54007. 40:19:55them. Uh, and that's we're talking about
  54008. 40:19:57a linear equation.
  54009. 40:19:59In the vectors, we have a + b= c. Now,
  54010. 40:20:03we're starting to look at a direction.
  54011. 40:20:05And these values usually think of an xyz
  54012. 40:20:08plot. Um, so each one is a direction.
  54013. 40:20:11And the actual distance of like a
  54014. 40:20:14triangle A is C. And then your matrix
  54015. 40:20:17can describe all kinds of things. Um, I
  54016. 40:20:20find matrixes uh confuse a lot of
  54017. 40:20:22people, not because they're particularly
  54018. 40:20:25difficult, but because of the magnitude
  54019. 40:20:28and the different things are used for.
  54020. 40:20:31And a matrix is a chart or a um, you
  54021. 40:20:35know, think of a spreadsheet, but you
  54022. 40:20:36have your rows and your columns. And
  54023. 40:20:39you'll see here we have a * b= c. Very
  54024. 40:20:43important to know your counts. Uh, so
  54025. 40:20:47depending on how the math is being done,
  54026. 40:20:48what you're using it for, making sure
  54027. 40:20:50you have the same rows and the number of
  54028. 40:20:52columns or a single number, there's all
  54029. 40:20:54kinds of things that play in that that
  54030. 40:20:56can make matrixes confusing. Uh, but
  54031. 40:20:58really it has a lot more to do with what
  54032. 40:21:00domain you're working in. Uh, are you
  54033. 40:21:02adding in multiple polomials where you
  54034. 40:21:05have like uh uh ax^2 plus b y plus, you
  54035. 40:21:10know, you start to see that can be very
  54036. 40:21:12confusing versus a very straightforward
  54037. 40:21:14matrix. And let's just go a little
  54038. 40:21:16deeper into these because these are such
  54039. 40:21:18primary this is what we're here to talk
  54040. 40:21:20about is these different math uh
  54041. 40:21:22mathematical computations that come up.
  54042. 40:21:25So we're looking at linear equations.
  54043. 40:21:26Let's dig deeper into that one. An
  54044. 40:21:28equation having a maximum order of one
  54045. 40:21:30is called a linear equation. Uh so it's
  54046. 40:21:33linear because when you look at this we
  54047. 40:21:35have uh ax plus b= c which is a one
  54048. 40:21:38variable. We have two variable ax plus b
  54049. 40:21:42y = c ax plus b y plus z c cz z= d and
  54050. 40:21:46so forth. But all of these are to the
  54051. 40:21:50power of one. You don't see x squar. You
  54052. 40:21:52don't see x cubed. So we're talking
  54053. 40:21:54about linear equations. That's what
  54054. 40:21:55we're talking about in their addition.
  54055. 40:21:57If you have already dived into say
  54056. 40:22:00neural networks, you should recognize
  54057. 40:22:02this ax plus b y plus cz um setup plus
  54058. 40:22:06the intercept. uh which is basically
  54059. 40:22:08your your neural network each node
  54060. 40:22:10adding up all the different inputs and
  54061. 40:22:13we can drill down into that most common
  54062. 40:22:15formula is your y = mx + c.
  54063. 40:22:20So you have your uh y equals the m which
  54064. 40:22:24is your slope, your x value plus c which
  54065. 40:22:28is your um y intercept. They kind of
  54066. 40:22:31labeled it wrong here
  54067. 40:22:33threw me for a loop but the the c would
  54068. 40:22:35be your y intercept. So when you set x
  54069. 40:22:37equal to zero, y equals c. And that's
  54070. 40:22:40that's your y intercept right there. Uh
  54071. 40:22:43and that's they they just had reversed
  54072. 40:22:45value of y. When x equals 0, it equals
  54073. 40:22:48the y intercept, which is c. And your
  54074. 40:22:50slope gradient line, which is your m. So
  54075. 40:22:52you get your y = 2x + 3. And there's
  54076. 40:22:56lots of easy ways to compute this. This
  54077. 40:22:58why this is why we always start with the
  54078. 40:23:00most basic one when we're solving one of
  54079. 40:23:01these problems. And then of course the
  54080. 40:23:03um one of the most important takeaways
  54081. 40:23:05is the slope gradient of the line. Uh so
  54082. 40:23:08the slope is very important that m
  54083. 40:23:10value. Uh in this case we went ahead and
  54084. 40:23:12solved this. If you have y = 2x + 3 you
  54085. 40:23:16can see how it has a nice line graph
  54086. 40:23:18here on the right.
  54087. 40:23:20So matrixes a matrix refers to a
  54088. 40:23:23rectangular representation of an array
  54089. 40:23:25of numbers arranged in columns and rows.
  54090. 40:23:29So we're talking m rows by n columns
  54091. 40:23:31here. A11 is denotes the element of the
  54092. 40:23:34first row in the first column. Similarly
  54093. 40:23:37a12 and it's really pronounced a11 in
  54094. 40:23:40this particular setup. So it's row one
  54095. 40:23:43column one. A12 is a of row one column 2
  54096. 40:23:48uh first row and second column and so
  54097. 40:23:50on.
  54098. 40:23:52And there's a lot of ways to denote
  54099. 40:23:53this. I've seen these as like a capital
  54100. 40:23:56letter A, smaller case A for the top row
  54101. 40:23:58or I mean you can see where they can go
  54102. 40:24:01all kinds of different directions as far
  54103. 40:24:02as the value. You just take a moment to
  54104. 40:24:05realize there's need to be some
  54105. 40:24:06designation as far as what row it's in
  54106. 40:24:09and what column it's in. And we have our
  54107. 40:24:12uh basic operations. We have addition.
  54108. 40:24:14So when you think about addition, you
  54109. 40:24:16have uh two matrices of 2x two and you
  54110. 40:24:20just add each individual number in that
  54111. 40:24:23matrix and then when you get to the
  54112. 40:24:25bottom you have uh in this case the
  54113. 40:24:27solution is 12, 10 + 2 is 12, 5 + 3 is 8
  54114. 40:24:30and so on. And the same thing with
  54115. 40:24:32subtraction.
  54116. 40:24:34Now again you're counting matrices you
  54117. 40:24:37want to check your um dimensions of the
  54118. 40:24:39matrix the shape you'll see shape come
  54119. 40:24:42up a lot in programming. So we're
  54120. 40:24:44talking about dimensions we're talking
  54121. 40:24:46about the shape. If the two shapes are
  54122. 40:24:48equal this is what happens when you add
  54123. 40:24:51them together or subtract them. And we
  54124. 40:24:54have multiplication. When you look at
  54125. 40:24:56the multiplication you end up with a
  54126. 40:24:57very slightly different setup going.
  54127. 40:25:00Now, if we look at our last one, we're
  54128. 40:25:03um uh we're like, why? This always gets
  54129. 40:25:06to me when we get to matrices. They
  54130. 40:25:08don't really say why you multiply
  54131. 40:25:10matrices. Um you know, my first thought
  54132. 40:25:12is 1 * 2, 4 * 3. But if you look at
  54133. 40:25:15this, we get 1 * 2 + 4 * 3, 1 * 3 + 4 *
  54134. 40:25:205,
  54135. 40:25:22uh 6 * 2 + 3 * 3, 6 * 3 + 3 * 5. If
  54136. 40:25:27you're looking at these matrices, uh,
  54137. 40:25:29think of this more as an equation. And
  54138. 40:25:32so we have, uh, if you remember when we
  54139. 40:25:33back up here for our multiple line
  54140. 40:25:35equations, let's just go back up a
  54141. 40:25:37couple slides where we were looking at,
  54142. 40:25:39uh, two variable. So this is a two
  54143. 40:25:41variable equation. ax plus b y= c.
  54144. 40:25:45Um, and this is a way to make it very
  54145. 40:25:48quick to solve these variables. And
  54146. 40:25:50that's why you have the matrix, and
  54147. 40:25:51that's why you do
  54148. 40:25:53the multiplication the way they do. And
  54149. 40:25:56this is the dotproduct of uh 1 * 2 + 4 *
  54150. 40:26:013
  54151. 40:26:031 * 3 + 4 * 5
  54152. 40:26:07uh 6 * 2 + 3 * 3 6 * 3 + 3 * 5. And it
  54153. 40:26:13gives us a nice little 14, 23, 21, and
  54154. 40:26:1633 over here, which then can be used and
  54155. 40:26:19reduced down to a simple um formula as
  54156. 40:26:23far as solving the variables as you have
  54157. 40:26:25enough inputs. Uh and then in matrix
  54158. 40:26:27operations, when you're dealing with a
  54159. 40:26:29lot of matrices, uh now keep in mind
  54160. 40:26:32multiplying matrices is different than
  54161. 40:26:34finding the product of two matrices.
  54162. 40:26:36Okay? So we're talking about
  54163. 40:26:37multiplication, we're talking about
  54164. 40:26:39solving uh for equations. When you're
  54165. 40:26:42finding the product, you are just
  54166. 40:26:43finding one time two. Keep that in mind
  54167. 40:26:45because that does come up. I've had that
  54168. 40:26:47come up a number of times where I am
  54169. 40:26:49altering data and I get confused as to
  54170. 40:26:51what I'm doing with it. Uh transpose
  54171. 40:26:54flipping the matrix over it's diagonal.
  54172. 40:26:56Comes up all the time where you have you
  54173. 40:26:58still have 12, but instead of it being
  54174. 40:27:00uh 128, it's now 1214 821. You're just
  54175. 40:27:05flipping the columns and the rows. Uh
  54176. 40:27:07and then of course you can do an inverse
  54177. 40:27:09um changing the signs of the values
  54178. 40:27:11across this main diagonal. And you can
  54179. 40:27:13see here we have the inverse a to the
  54180. 40:27:15minus1 and ends up with uh instead of 12
  54181. 40:27:188 14 12 it's now -22 -12 vectors uh
  54182. 40:27:23vector just means we have
  54183. 40:27:26a value and a direction and we have down
  54184. 40:27:30four numbers here on our vector.
  54185. 40:27:33uh in mathematics a one-dimensional
  54186. 40:27:35matrix is called a vector. Uh so if you
  54187. 40:27:39have your xplot and you have a single
  54188. 40:27:41value that values along the x- axis and
  54189. 40:27:44it's a single dimension. If you have two
  54190. 40:27:46dimensions you can think about putting
  54191. 40:27:48them on a graph. You might have x and
  54192. 40:27:50you might have y and each value denotes
  54193. 40:27:53a direction. And then of course the
  54194. 40:27:55actual distance is going to be the
  54195. 40:27:57hypothesis of that triangle. Uh and you
  54196. 40:27:59can do that with three dimensionals x y
  54197. 40:28:01and z. uh and you can do it all the way
  54198. 40:28:03to nth dimensions. So when they talk
  54199. 40:28:06about the k means uh for categorizing
  54200. 40:28:09and how close data is together, they
  54201. 40:28:12will compute that based on the
  54202. 40:28:14Pythagorean theorem. So you would take
  54203. 40:28:16uh the square of each value, add them
  54204. 40:28:18all together and find the square root
  54205. 40:28:20and that gives you a distance as far as
  54206. 40:28:22where that point is, where that vector
  54207. 40:28:24exists or an actual point value. And
  54208. 40:28:26then you can compare that point value to
  54209. 40:28:29another one and it makes a very easy
  54210. 40:28:31comparison versus comparing uh 50 or 60
  54211. 40:28:34different numbers. And that brings us up
  54212. 40:28:36to gene vectors and I gene values. Uh I
  54213. 40:28:41gene vectors the vectors that don't
  54214. 40:28:43change their span while transformation
  54215. 40:28:46and I gene values the scalar values that
  54216. 40:28:49are associated to the vectors.
  54217. 40:28:52Conceptually you can think of the vector
  54218. 40:28:54as your picture. you have a picture.
  54219. 40:28:56It's um uh two dimensions x and y. And
  54220. 40:29:00so when you do those two dimensions and
  54221. 40:29:02those two values or whatever that value
  54222. 40:29:04is um that is that point but the values
  54223. 40:29:09change when you skew it and so if we
  54224. 40:29:12take and we have a vector a and that's a
  54225. 40:29:16set value uh b is um your is your you
  54226. 40:29:20have a and b which is your hygiene
  54227. 40:29:21vector. Two is the i gene value. So,
  54228. 40:29:25we're altering all the values by two.
  54229. 40:29:28That means we're u maybe we're
  54230. 40:29:30stretching it out one direction, making
  54231. 40:29:31it tall if you're doing picture editing.
  54232. 40:29:34Um that that's one of the places this
  54233. 40:29:36comes in. But you can see when you're
  54234. 40:29:38transforming uh your different
  54235. 40:29:40information, how you transform it is
  54236. 40:29:43then your hygiene value. And you can see
  54237. 40:29:45here uh vector after line transition
  54238. 40:29:49uh we have 3 a is the hygiene vector.
  54239. 40:29:52Three is the hygiene value. So A doesn't
  54240. 40:29:55change. That's whatever we started with.
  54241. 40:29:57That's your original picture. And three
  54242. 40:29:59uh is skewing it one direction and maybe
  54243. 40:30:02uh B is being skewed another direction.
  54244. 40:30:05And so you have a nice tilted picture
  54245. 40:30:06because you've altered it by those by
  54246. 40:30:08the hygiene values.
  54247. 40:30:10So let's go ahead and pull up a demo on
  54248. 40:30:13linear algebra. And to do this, I'm
  54249. 40:30:16going to go through my trusted Anaconda
  54250. 40:30:19into my Jupiter notebook. and we'll
  54251. 40:30:22create a new uh notebook called linear
  54252. 40:30:25algebra. Since we are working in Python,
  54253. 40:30:28uh we're going to use our numpy. I
  54254. 40:30:30always import that as np or numpy array.
  54255. 40:30:33Probably the most popular um module for
  54256. 40:30:36doing matrixes and things in
  54257. 40:30:39given that this is part of a series. I'm
  54258. 40:30:41not going to go too much into numpy. Uh
  54259. 40:30:43we are going to go ahead and create two
  54260. 40:30:45different variables. A for a numpy array
  54261. 40:30:4710 15 and b 29.
  54262. 40:30:51We'll go ahead and run this. And you can
  54263. 40:30:52see there's our two arrays 105 29. And I
  54264. 40:30:55went ahead and added a space there in
  54265. 40:30:57between so it's easier to read. And
  54266. 40:31:00since it's the last line, we don't have
  54267. 40:31:02to put the print statement on it unless
  54268. 40:31:04you want. We can simp but we can simply
  54269. 40:31:06do a plus b. So when I run this, uh, we
  54270. 40:31:10have 10 15 29 and we get 30 24, which is
  54271. 40:31:15what you expect. 10 + 20 15 + 9. You
  54272. 40:31:19could almost look at this addition as
  54273. 40:31:21being um
  54274. 40:31:24just adding up the columns on here
  54275. 40:31:26coming down. And if we wanted to do it a
  54276. 40:31:28different way, we could also do a t plus
  54277. 40:31:32b dot t. Remember that t flips them. And
  54278. 40:31:35so if we do that, we now get them uh we
  54279. 40:31:39now have 304 going the other way. We
  54280. 40:31:42could also do something kind of fun.
  54281. 40:31:44There's a lot of different ways to do
  54282. 40:31:45this. Uh, as far as a plus b, I can also
  54283. 40:31:49do a plus b. T and you're going to see
  54284. 40:31:53that that will come out the same. The 30
  54285. 40:31:5524 whether I transpose a and b or
  54286. 40:31:57transpose them both at the end.
  54287. 40:32:01And likewise, we can very easily
  54288. 40:32:03subtract two vectors. I can go a minus
  54289. 40:32:06b. And we run that and we get - 106. Now
  54290. 40:32:11remember, this is the last line in this
  54291. 40:32:13particular section. That's why I don't
  54292. 40:32:14have to put the print around it. Um, and
  54293. 40:32:17just like we did before, we can
  54294. 40:32:20transpose either the individual or we
  54295. 40:32:22can transpose the main setup and then we
  54296. 40:32:25get a minus 106 going the other way.
  54297. 40:32:30Now, we didn't mention this in our
  54298. 40:32:32notes, but you can also do a scalar
  54299. 40:32:35multiplication.
  54300. 40:32:37Let me just put down scaler so you can
  54301. 40:32:38remember that. Uh what we're talking
  54302. 40:32:41about here is I have uh this array here
  54303. 40:32:45u and if I go a time u uh we'll take the
  54304. 40:32:51value two we'll multiply it by every
  54305. 40:32:52value in here. So 2 * 30 is 60 2 * 15
  54306. 40:32:58and just like we did before
  54307. 40:33:01um this happens a lot because when
  54308. 40:33:03you're doing matrices you do need to
  54309. 40:33:04flip them you get 6030 coming this way.
  54310. 40:33:08So in numpy uh we have what they call
  54311. 40:33:11dotproduct
  54312. 40:33:14and uh what this this in a
  54313. 40:33:16twodimensional vectors it is the
  54314. 40:33:18equivalent of two matrix multiplication
  54315. 40:33:21and remember we were talking about
  54316. 40:33:22matrix multiplication
  54317. 40:33:24uh where it is the well let's walk
  54318. 40:33:27through it
  54319. 40:33:30we'll go ahead and start by defining two
  54320. 40:33:32um numpy arrays we'll have uh 10 20 256
  54321. 40:33:37or our u and our E uh and then we're
  54322. 40:33:39going to go ahead and do if we take
  54323. 40:33:43the values uh and if you remember
  54324. 40:33:45correctly
  54325. 40:33:47an array like this would be 10 * 25 + 20
  54326. 40:33:52* 6. We'll go ahead and uh print that.
  54327. 40:34:00There we go.
  54328. 40:34:02And then we'll go ahead and do the uh np
  54329. 40:34:05dot of u comma
  54330. 40:34:10v.
  54331. 40:34:12And we'll find when we do this, we go
  54332. 40:34:14and run this uh we're going to get uh
  54333. 40:34:17370
  54334. 40:34:18370.
  54335. 40:34:20So this is a strain multiplication where
  54336. 40:34:22they use it to solve uh linear algebra
  54337. 40:34:26uh when you have multiple numbers going
  54338. 40:34:28across. And so this could be very
  54339. 40:34:30complicated. We could have a whole
  54340. 40:34:31string of different variables going in
  54341. 40:34:33here. But for this we get a nice uh
  54342. 40:34:35value for our dot multiplication
  54343. 40:34:39and we did um addition earlier which was
  54344. 40:34:42just your basic addition. Uh and of
  54345. 40:34:44course a matrix you can get very
  54346. 40:34:46complicated on these or in this case
  54347. 40:34:48we'll go ahead and do um let's create
  54348. 40:34:51two complex matrixes.
  54349. 40:34:55This one is a matrix of um you know 1210
  54350. 40:34:5946 431. We'll just print out A so you
  54351. 40:35:02can see what that looks like. Here's
  54352. 40:35:04print A.
  54353. 40:35:06We print A out. You can see that we have
  54354. 40:35:09a um 2x3
  54355. 40:35:13layer matrix for A. And we can also put
  54356. 40:35:16together always kind of fun when you're
  54357. 40:35:18playing with print values. Uh we could
  54358. 40:35:20do something like this. We could go in
  54359. 40:35:22here. There we go. Uh, we could print a.
  54360. 40:35:25We have it end with uh equals a run. And
  54361. 40:35:29this kind of gives it a nice look. Uh,
  54362. 40:35:31here's your matrix. That's all this is.
  54363. 40:35:33Comma, n means it just tags it on the
  54364. 40:35:35end. That's all all that is doing on
  54365. 40:35:37there. And then we can simply add in
  54366. 40:35:39what is a plus b. And you should already
  54367. 40:35:42guess because this is the same as what
  54368. 40:35:43we did before. There's no difference.
  54369. 40:35:45Uh, we do a simple vector addition. We
  54370. 40:35:47have 12 + 2 is 14, 10 + 8 is 18. And so
  54371. 40:35:51on. And just like we did the uh matrix
  54372. 40:35:54addition, we can also do a minus b and
  54373. 40:35:58do our matrix subtraction.
  54374. 40:36:01And we look at this uh we have what? 12
  54375. 40:36:03- 2 is 10. 10 - 8 um where are we?
  54376. 40:36:11Oh, there we go. 8 min
  54377. 40:36:15confusing what I'm looking at. I should
  54378. 40:36:16have reprinted out the original numbers.
  54379. 40:36:18Uh but we can see here 12 - 2 is of
  54380. 40:36:21course 10. 10 - 8 is 2. Uh 4 - 46 is -
  54381. 40:36:2542 and so forth. So same as a
  54382. 40:36:28subtraction as before, we just call it
  54383. 40:36:30matrix subtraction. It's identical.
  54384. 40:36:33Now if you remember up here, we had
  54385. 40:36:35scalar addition where we're adding just
  54386. 40:36:37one number to a matrix. You can also do
  54387. 40:36:41scalar multiplication. Uh and so simply
  54388. 40:36:44if you have a single value A and you
  54389. 40:36:46have B which is your array, we can also
  54390. 40:36:48do A * B. When we run that, uh, you can
  54391. 40:36:53see here we have 2 * 4 is 8. Uh, 5 * 4
  54392. 40:36:57is 20 and so forth. You're just
  54393. 40:36:59multiplying the four across each one of
  54394. 40:37:01these values. And this is an interesting
  54395. 40:37:03one that comes up. A little bit of a
  54396. 40:37:06brain teaser is matrix and vector
  54397. 40:37:09multiplication.
  54398. 40:37:11And so when we're looking at this,
  54399. 40:37:14uh, we are just do a regular arrays. It
  54400. 40:37:17doesn't necessarily have to be a numpy
  54401. 40:37:18array. We have a
  54402. 40:37:21which has our um array of arrays and b
  54403. 40:37:25which is a single array and so we can
  54404. 40:37:28from here
  54405. 40:37:30do the dot
  54406. 40:37:33a b and this is going to return two
  54407. 40:37:36values and the first value is that it's
  54408. 40:37:39you could say it's like uh um we're
  54409. 40:37:41doing the this array b array first with
  54410. 40:37:45a and then with a second one and so it
  54411. 40:37:47splits it up so you have a matrix of
  54412. 40:37:49vector multiplication and you mix and
  54413. 40:37:50match. When you get into really
  54414. 40:37:52complicated uh backend stuff, this
  54415. 40:37:54becomes more common because you're now
  54416. 40:37:56you got layers upon layers of data and
  54417. 40:37:59so you you'll end up with a matrix and a
  54418. 40:38:02set of uh vector matrices. Do you want
  54419. 40:38:04to multiply?
  54420. 40:38:06Now, keep in mind that if you're doing
  54421. 40:38:09data science, a lot of times you're not
  54422. 40:38:11looking at this. This is what's going on
  54423. 40:38:12behind the scenes. So if you're in um
  54424. 40:38:15the scikit looking at sklearn where
  54425. 40:38:17you're doing linear regression models,
  54426. 40:38:20this is some of the math that's hidden
  54427. 40:38:21behind the scenes that's going on. Other
  54428. 40:38:24times you might find yourself having to
  54429. 40:38:26do part of this and manipulate the data
  54430. 40:38:28around so it fits right and then you go
  54431. 40:38:30back in and you run it through the
  54432. 40:38:31scikit. And if we can do um up here
  54433. 40:38:36where we did a uh matrix and vector
  54434. 40:38:39multiplication, we can also do matrix to
  54435. 40:38:41matrix multiplication. And if we run
  54436. 40:38:43this where we have the two matrices, uh
  54437. 40:38:45you can see we have a very complicated
  54438. 40:38:47array that of course comes out on there
  54439. 40:38:48for our dot. And just to reiterate it,
  54440. 40:38:52we have our transpose a matrix which is
  54441. 40:38:54your T. And so if we create a matrix A
  54442. 40:38:57and then we do transpose it, you can see
  54443. 40:38:59how it flips it from 5 10 15 20 25 30 to
  54444. 40:39:045 15 25 10 20 30 uh rows and columns.
  54445. 40:39:10And certainly with the math, uh, this
  54446. 40:39:12comes up a lot. Um, it also comes up a
  54447. 40:39:15lot with XY plotting. When you put it
  54448. 40:39:17into piplot, you have one format where
  54449. 40:39:19they're looking at pairs of numbers and
  54450. 40:39:21then they want all of X's and all Y's.
  54451. 40:39:24So, you know, the transpose is an
  54452. 40:39:25important tool both for your math and
  54453. 40:39:27for plotting and all kinds of things.
  54454. 40:39:30Another tool that we didn't discuss uh
  54455. 40:39:32is your identity matrix. Uh and this one
  54456. 40:39:37is more definition.
  54457. 40:39:40Uh the identity matrix. Um we have here
  54458. 40:39:43one where we just did uh two. So it
  54459. 40:39:46comes down as one 0 0 1 uh 1 0 0 1 0. It
  54460. 40:39:51creates a diagonal of one. And what that
  54461. 40:39:53is is when you're doing your identities,
  54462. 40:39:55you could be comparing all your
  54463. 40:39:58different features to the different
  54464. 40:40:00features and how they correlate. And of
  54465. 40:40:02course when you have uh feature one
  54466. 40:40:04compared to feature one to itself it is
  54467. 40:40:06always one uh where usually it's between
  54468. 40:40:10zero one depending on how well
  54469. 40:40:12correlates. So when we're talking about
  54470. 40:40:14identity matrix that's what we're
  54471. 40:40:16talking about right here is that you
  54472. 40:40:18create this preset matrix and then you
  54473. 40:40:20might adjust these numbers depending on
  54474. 40:40:22what you're working with and what the
  54475. 40:40:23domain is. And then another thing we can
  54476. 40:40:26do uh to kind of wrap this up. We'll hit
  54477. 40:40:28you with the most complicated uh um
  54478. 40:40:30piece of this puzzle here is an inverse
  54479. 40:40:34um a matrix. And let's just go ahead and
  54480. 40:40:37put the um it's a lengthy description.
  54481. 40:40:41Let's go and put the description. This
  54482. 40:40:43is straight out of the uh the website
  54483. 40:40:46for um numpy. Uh so given a square
  54484. 40:40:50matrix A, here's our square matrix A,
  54485. 40:40:53which is 2 1 0 0 1 0 1 2 1. Keep in mind
  54486. 40:40:573x3, it's square. It's got to be equal.
  54487. 40:40:59It's going to return the matrix A
  54488. 40:41:02inverse satisfying dot A um A inverse.
  54489. 40:41:06So here's our matrix multiplication.
  54490. 40:41:10Um and then of course it equals the dot
  54491. 40:41:13uh yeah a inverse of a um with an
  54492. 40:41:17identity shape of uh a dotshaped zero.
  54493. 40:41:20This is just reshaping the identity.
  54494. 40:41:23That's a little complicated there. Uh so
  54495. 40:41:25we go and have our here's our array. Uh
  54496. 40:41:27we'll go ahead and run this. And you can
  54497. 40:41:30see what we end up with is we end up
  54498. 40:41:32with uh an array 0.5 minus 0.5 and so
  54499. 40:41:36forth with our 211 going down to 1 0 0 1
  54500. 40:41:400 1 2 1. Um getting into a little deep
  54501. 40:41:44on the math understanding when you need
  54502. 40:41:47this is probably really is is what's
  54503. 40:41:49really important when you're doing data
  54504. 40:41:50science versus uh handwriting this out
  54505. 40:41:54and looking up the math and handwriting
  54506. 40:41:55all the pieces out. you do need to know
  54507. 40:41:57about the linear algorithm inverse of a.
  54508. 40:42:00Uh so if it comes up, you can easily
  54509. 40:42:02pull it up or at least remember where to
  54510. 40:42:03look it up. You took a look at the
  54511. 40:42:06algebra side of it. Let's go ahead and
  54512. 40:42:07take a look at the calculus side of uh
  54513. 40:42:10what's going on here with the machine
  54514. 40:42:11learning. So calculus, oh my goodness,
  54515. 40:42:14and differential equations, you got to
  54516. 40:42:16throw that in there because that's all
  54517. 40:42:18part of the bag of tricks, especially
  54518. 40:42:21when you're doing large neural networks,
  54519. 40:42:23but also comes up in many other areas.
  54520. 40:42:25The good news is most of it's already
  54521. 40:42:27done for you in the back end. Uh so when
  54522. 40:42:29it comes up, you really do need to
  54523. 40:42:30understand from the data science, not
  54524. 40:42:32data analytics. Data analytics means
  54525. 40:42:34you're digging deep into actually
  54526. 40:42:36solving these math equations. U and a
  54527. 40:42:39neural network is just a giant
  54528. 40:42:41differential equation. Uh so we talk
  54529. 40:42:43about calculus uh we're going to go
  54530. 40:42:45ahead and understand it by talking about
  54531. 40:42:49cars versus time and speed. uh so helps
  54532. 40:42:53to calculate the spontaneous rate of
  54533. 40:42:56change.
  54534. 40:42:58Uh so suppose we plot a graph of the
  54535. 40:43:00speed of a car with respect to time. So
  54536. 40:43:02as you can see here going down the
  54537. 40:43:04highway probably merged into the highway
  54538. 40:43:06from an on-ramp. So I had to accelerate
  54539. 40:43:09so my speed went way up uh stuck in
  54540. 40:43:12traffic merged into the traffic. Traffic
  54541. 40:43:15opens up and I accelerate again up to
  54542. 40:43:16the speed limit and u maybe it peters
  54543. 40:43:19off up there. So you can look at this as
  54544. 40:43:22as um the speed versus time. I'm getting
  54545. 40:43:25faster and faster because I'm
  54546. 40:43:26continually accelerating. And if I hit
  54547. 40:43:29the brakes, it go the other way. So the
  54548. 40:43:31rate of change of speed with respect of
  54549. 40:43:33time is nothing but acceleration. How
  54550. 40:43:36fast are we accelerating? The
  54551. 40:43:38acceleration is the area between the
  54552. 40:43:40start point of x and the end point of
  54553. 40:43:42delta x. Uh so we can calculate a simple
  54554. 40:43:46if you had x and delta x we could put a
  54555. 40:43:48line there and that slope of the line is
  54556. 40:43:51our acceleration.
  54557. 40:43:53Now that's pretty easy when you're doing
  54558. 40:43:55linear algebra but I don't want to know
  54559. 40:43:58it just for that line and those two
  54560. 40:44:00points. I want to know it across the
  54561. 40:44:02whole of what I'm working with. That's
  54562. 40:44:04where we get into calculus. So when we
  54563. 40:44:06talk about the distance between x and
  54564. 40:44:08delta x it has to be the smallest
  54565. 40:44:10possible near to zero in order to
  54566. 40:44:12approximate the acceleration.
  54567. 40:44:15Uh so the idea is that instead of I mean
  54568. 40:44:17if you ever did took a basic calculus
  54569. 40:44:19class they would draw bars down here and
  54570. 40:44:22you would divide this area up um let's
  54571. 40:44:25go back up a screen. you divide this
  54572. 40:44:27area of this time period up into maybe
  54573. 40:44:3010 sections and you'd use that and you
  54574. 40:44:32could calculate the acceleration between
  54575. 40:44:34each one of those 10 sections kind of
  54576. 40:44:35thing. Uh and then we just keep making
  54577. 40:44:38that space smaller and smaller until
  54578. 40:44:40delta x is almost uh infantismally
  54579. 40:44:44small. And so we get a function of a uh
  54580. 40:44:48equals a limit as h goes to zero of a
  54581. 40:44:51function of a plus h minus a function of
  54582. 40:44:54a over h. And that is you're computing
  54583. 40:44:57the slope of the line.
  54584. 40:45:00We're just computing that slope under
  54585. 40:45:01smaller and smaller and smaller samples.
  54586. 40:45:04Uh and that's what calculus is. Calculus
  54587. 40:45:06is the integral. You can see down here
  54588. 40:45:08we have our nice uh integral sign. Looks
  54589. 40:45:11like a giant s. And that's what that
  54590. 40:45:13means is that we've taken this down to
  54591. 40:45:16as small as we can for that sampling. Uh
  54592. 40:45:20so we're talking about calculus. Finding
  54593. 40:45:22the area under the slope is the main
  54594. 40:45:24process in the integration. Similar
  54595. 40:45:27small intervals are made of the smallest
  54596. 40:45:29possible length of x plus delta x where
  54597. 40:45:32delta x approaches almost an infantismly
  54598. 40:45:35small space. And then it helps to find
  54599. 40:45:37the overall acceleration by summing up
  54600. 40:45:39all the lengths together. Uh so we're
  54601. 40:45:42summing up all the accelerations from
  54602. 40:45:44the beginning to the end. And so here's
  54603. 40:45:46our integral. we sum of a of x * d ofx =
  54604. 40:45:50a + c. Uh that is our basic calculus
  54605. 40:45:55here. So when we talk about
  54606. 40:45:57multivvariant calculus, uh multivariate
  54607. 40:46:01calculus deals with functions that have
  54608. 40:46:02multiple variables and you can see here
  54609. 40:46:05we start getting into some very
  54610. 40:46:06complicated equations. Um uh change in w
  54611. 40:46:10over change of time equals change of w
  54612. 40:46:13over change of z. the differential of z
  54613. 40:46:16to dx differential of x to dt. It gets
  54614. 40:46:19pretty complicated. Uh and it really
  54615. 40:46:21translates into the multivariate
  54616. 40:46:23integration using double integrals. And
  54617. 40:46:25so you have the the sum of the sum of f
  54618. 40:46:28ofxy of d of a equals the sum from c to
  54619. 40:46:31d and a to b of f ofxy dx dy equals uh
  54620. 40:46:36the sum of a to b sum of c to d of fxy
  54621. 40:46:40dy dx.
  54622. 40:46:42understanding the very specifics of
  54623. 40:46:44everything going on in here and actually
  54624. 40:46:46doing the math is usually calculus one,
  54625. 40:46:50calculus 2, and differential equations.
  54626. 40:46:52Uh so you're talking about three
  54627. 40:46:54fulllength courses to dig into and solve
  54628. 40:46:57these math equations. What we want to
  54629. 40:47:00take from here is we're talking about
  54630. 40:47:01calculus. Uh we're talking about summing
  54631. 40:47:04of all these different slopes. And so
  54632. 40:47:07we're still solving a linear uh
  54633. 40:47:09expression. We're still solving y = mx +
  54634. 40:47:13b, but we're doing this for
  54635. 40:47:15infantismally small x's. And then we
  54636. 40:47:17want to sum them up. That's what this
  54637. 40:47:19integral sign means. The the sum of a of
  54638. 40:47:21x d of x= a plus c.
  54639. 40:47:25And when you see these very complicated
  54640. 40:47:27uh multivariate differentiation using
  54641. 40:47:29the chain rule uh when we come in here
  54642. 40:47:31and we have the change of w to the
  54643. 40:47:34change of t equals the change of w dz uh
  54644. 40:47:37and so forth. That's what's going on
  54645. 40:47:40here. That's what these means. We're
  54646. 40:47:41basically looking for the area under the
  54647. 40:47:43curve which really comes to how is the
  54648. 40:47:46change changing and speed's going up.
  54649. 40:47:49How is that changing? And then you end
  54650. 40:47:51up with a multiple layer. So if I have
  54651. 40:47:53three layers of neural networks, how is
  54652. 40:47:55the third layer changing based on the
  54653. 40:47:57second layer changing which is based on
  54654. 40:47:59the first layer changing? And you get
  54655. 40:48:01the picture here that now we have a very
  54656. 40:48:03complicated uh multivariate integration
  54657. 40:48:06um with integrals.
  54658. 40:48:08The good news is we can solve this uh
  54659. 40:48:11mathematically and that's what we do
  54660. 40:48:13when you do neural networks and reverse
  54661. 40:48:14propagation. Uh so the nice thing is
  54662. 40:48:17that you don't have to solve this on
  54663. 40:48:18paper unless you're a data analysis and
  54664. 40:48:20you're working on the back end of
  54665. 40:48:22integrating these formulas and building
  54666. 40:48:24the script to actually build them. So we
  54667. 40:48:26talk about applications of calculus. Uh
  54668. 40:48:28it provides us the tools to build an
  54669. 40:48:30accurate predictive model. Um so it's
  54670. 40:48:32really behind the scenes we want to
  54671. 40:48:34guess at what the change of the change
  54672. 40:48:36of the change is.
  54673. 40:48:38That's a little goofy. I I know I just
  54674. 40:48:40threw that out there. It's kind of a
  54675. 40:48:41meta term. But if you can guess how
  54676. 40:48:44things are going to change, then you can
  54677. 40:48:46guess what the new numbers are.
  54678. 40:48:48Multivariate calculus explains the
  54679. 40:48:51change in our target variable in
  54680. 40:48:52relation to the rate of change in the
  54681. 40:48:54input variables. So there's our multiple
  54682. 40:48:57variables going in there. If uh one
  54683. 40:49:00variable is changing, how does it affect
  54684. 40:49:01the other variable? And then in gradient
  54685. 40:49:04descent, calculus is used to find the
  54686. 40:49:07local and global maxima. And this is
  54687. 40:49:10really big. Uh we're actually going to
  54688. 40:49:12have a whole section here on gradient
  54689. 40:49:14descent because it is really I mean I
  54690. 40:49:18talked about neural networks and how you
  54691. 40:49:19can see how the different layers go in
  54692. 40:49:21there, but gradient descent is one of
  54693. 40:49:23the most key things for trying to guess
  54694. 40:49:26the best answer to something. So let's
  54695. 40:49:30take a look at the code behind gradient
  54696. 40:49:33descent. And uh before we open up the
  54697. 40:49:36code, let's just do real quick uh
  54698. 40:49:39gradient descent.
  54699. 40:49:42Let's say we have a curve like this. And
  54700. 40:49:44most common is that this is going to
  54701. 40:49:47represent your error. Oops.
  54702. 40:49:51Error. There we go. Error. Ah, hard to
  54703. 40:49:55read there. And I want to make the error
  54704. 40:49:57as low as possible. And so what I'm
  54705. 40:50:00looking at it is I want to find this
  54706. 40:50:02line here which is the minimum value. So
  54707. 40:50:07we're looking for the minimum and it
  54708. 40:50:10does that by uh sampling there and then
  54709. 40:50:14it based on this it guesses it might be
  54710. 40:50:16someplace here and it goes hey this is
  54711. 40:50:18still going down. It goes here and then
  54712. 40:50:21goes back over here and then goes a
  54713. 40:50:23little bit closer and it's just playing
  54714. 40:50:25a high low until it gets to that spot,
  54715. 40:50:28that bottom spot. And so we want to
  54716. 40:50:30minimize the error in uh on the flip
  54717. 40:50:34note, you could also want to be
  54718. 40:50:37maximizing something. You want to get
  54719. 40:50:38the best output of it. Uh that's simply
  54720. 40:50:41uh minus the value. Uh so if you're
  54721. 40:50:44looking for where the peak is, this is
  54722. 40:50:46the same as a negative for where the
  54723. 40:50:49valley is and looking for that valley.
  54724. 40:50:52Uh that's all that is and this is a way
  54725. 40:50:53of finding it. So the cool thing is um
  54726. 40:50:57all the heavy lifting's done. Um I
  54727. 40:50:59actually ended up putting together one
  54728. 40:51:01of these a while back as uh when I
  54729. 40:51:03didn't know about sidekick and I was
  54730. 40:51:05just starting. Boy, it's a long while
  54731. 40:51:08back and uh is playing high low. How do
  54732. 40:51:11you play high low? not get stuck in the
  54733. 40:51:13valleys, uh, figure out these curves and
  54734. 40:51:16things like that. Well, you do that and
  54735. 40:51:18the back end is all the calculus and
  54736. 40:51:19differential equations to calculate this
  54737. 40:51:21out. The good news is you don't have to
  54738. 40:51:24do those. Uh, so instead, we're going to
  54739. 40:51:27put together the code and let's go ahead
  54740. 40:51:31and see what we can do with that.
  54741. 40:51:36So, uh, guys in the back put together a
  54742. 40:51:38nice little piece of code here, which is
  54743. 40:51:40kind of fun.
  54744. 40:51:41uh some things we're going to note and
  54745. 40:51:43this is this is really important stuff
  54746. 40:51:45because when you start doing your data
  54747. 40:51:47science and digging into your machine
  54748. 40:51:49learning models uh you're going to find
  54749. 40:51:52these things are stumbling blocks. Uh
  54750. 40:51:54the first one is current x. Where do we
  54751. 40:51:56start at? Uh keep in mind your model
  54752. 40:52:00that you're working with is very
  54753. 40:52:02generic. So whatever you use to minimize
  54754. 40:52:04it the first question is where do we
  54755. 40:52:06start? Um, and we started at this cuz
  54756. 40:52:09the algorithm starts at x= 3. So, we
  54757. 40:52:12arbitrarily picked five. Learning rate
  54758. 40:52:15is uh how many bars to skip going one
  54759. 40:52:17way or the other. Uh, I'm in fact, I'm
  54760. 40:52:19going to separate that a little bit
  54761. 40:52:20because these two are really important.
  54762. 40:52:22Um, if we're dealing with something like
  54763. 40:52:24this where we're talking about um uh
  54764. 40:52:26well, here's our here's the function
  54765. 40:52:28we're going to use our um gradient of
  54766. 40:52:31our function um 2 * x + 5. Keep it
  54767. 40:52:34simple. So that's a function we're going
  54768. 40:52:36to work with. So if I'm dealing with
  54769. 40:52:38increments of a th00and 0.1 is going to
  54770. 40:52:41be a very long time. And if I'm dealing
  54771. 40:52:44with increments of 0.001,
  54772. 40:52:47uh 0.1 is going to skip over my answer.
  54773. 40:52:50So I won't get a very good answer. Um
  54774. 40:52:52and then we look at precision. This
  54775. 40:52:54tells us when to stop the algorithm. So
  54776. 40:52:56again, very specific to what you're
  54777. 40:52:59working on. uh if you're working with
  54778. 40:53:01money and you don't convert it into a
  54779. 40:53:05float value uh you might be dealing with
  54780. 40:53:080.01 which is a penny that might be your
  54781. 40:53:11precision you're working with. Um and
  54782. 40:53:14then of course the previous step size
  54783. 40:53:16max iterations uh we want something to
  54784. 40:53:18cut out at a certain point. Usually
  54785. 40:53:20that's built into a lot of minimization
  54786. 40:53:22functions. And then here's our actual uh
  54787. 40:53:25formula we're going to be working with.
  54788. 40:53:27And then we come in, we go while
  54789. 40:53:29previous step size is greater than
  54790. 40:53:31precision and its is less than max its
  54791. 40:53:36say that 10 times fast. Um
  54792. 40:53:40we're just saying if it's uh if we're if
  54793. 40:53:42we're still greater than our precision
  54794. 40:53:43level, we still got to keep digging
  54795. 40:53:45deeper. Um and then we also don't want
  54796. 40:53:47to go past a thou or whatever this is, a
  54797. 40:53:50million or 10,000 uh running. That's
  54798. 40:53:52actually pretty high. um almost never do
  54799. 40:53:54max iterations more than like 100 or
  54800. 40:53:57200. Rare occasions you might go up to
  54801. 40:53:59four or 500 if it's depending on the
  54802. 40:54:01problem you're working with. Uh so we
  54803. 40:54:04have our previous equals our current.
  54804. 40:54:06That way we can track timewise.
  54805. 40:54:08Uh the current now equals the current
  54806. 40:54:10minus the rate times the formula of our
  54807. 40:54:13previous x. So now we've generated our
  54808. 40:54:16new version. Uh previous step size
  54809. 40:54:19equals the absolute current previous.
  54810. 40:54:22Uh, so we're looking for the change in x
  54811. 40:54:25itters equals iterations + one. That's
  54812. 40:54:27so we know to stop if we get too far.
  54813. 40:54:30And then we're just going to print the
  54814. 40:54:31local minimum occurs at x on here. And
  54815. 40:54:35if we go ahead and run this,
  54816. 40:54:37uh, you can see right here it gets down
  54817. 40:54:39to this point and it says, hey, um,
  54818. 40:54:43local minimum is minus 3.3222
  54819. 40:54:46for this particular series we created.
  54820. 40:54:49Uh, and this is created off of our
  54821. 40:54:50formula here. lambda x2 * x + 5. Now,
  54822. 40:54:55when I'm running this stuff, uh you'll
  54823. 40:54:57see this come up a lot
  54824. 40:55:00and uh with the sklearn kit and and one
  54825. 40:55:04of the nice reasons of breaking this
  54826. 40:55:05down the way we did is I could go over
  54827. 40:55:08those top pieces. Uh those top pieces
  54828. 40:55:10are everything when you start looking at
  54829. 40:55:12these minimization toolkits in built-in
  54830. 40:55:15code. And so from um we'll just do it's
  54831. 40:55:19actually docs.cipi.org
  54832. 40:55:26and we're looking at the scikit. There
  54833. 40:55:30we go. Um optimize minimize. You can
  54834. 40:55:34only minimize one value. You have the
  54835. 40:55:36function that's going in. This function
  54836. 40:55:38can be very complicated. Uh so we used a
  54837. 40:55:41very simple function up here. It could
  54838. 40:55:43be there's all kinds of things that
  54839. 40:55:46could be on there. And there's a number
  54840. 40:55:47of methods to solve this as far as how
  54841. 40:55:49they shrink down. Uh and your x knot.
  54842. 40:55:52There's your there's your start value.
  54843. 40:55:54So your function, your start value. Um
  54844. 40:55:57there's all kinds of things that come in
  54845. 40:55:58here that we can look at which we're not
  54846. 40:56:00going to. Um optimization automatically
  54847. 40:56:03creates constraints bounds. Some of this
  54848. 40:56:06it does automatically, but you really
  54849. 40:56:08the big thing I want to point out here
  54850. 40:56:10is you need to have a starting point.
  54851. 40:56:12You want to start with something that
  54852. 40:56:13you already know is mostly the answer.
  54853. 40:56:15Uh if you don't, then it's going to have
  54854. 40:56:17a heck of a time trying to calculate it
  54855. 40:56:18out.
  54856. 40:56:20Or you can write your own little script
  54857. 40:56:21that does this and and does a high low
  54858. 40:56:23guessing and tries to find the max
  54859. 40:56:25value. That brings us to statistics.
  54860. 40:56:29What this is kind of all about is
  54861. 40:56:30figuring things out. Lot of vocabulary
  54862. 40:56:33and statistics. Uh so statistics, well,
  54863. 40:56:36I guess it's all relative. It's
  54864. 40:56:38definitely not an ed class. Uh so a
  54865. 40:56:40bunch of stuff going on. Statistics.
  54866. 40:56:42Statistics concerns with the collection,
  54867. 40:56:45organization, analysis, interpretation,
  54868. 40:56:48and presentation of data. That is a
  54869. 40:56:52mouthful. Um so we have from end to end
  54870. 40:56:56we're
  54871. 40:56:58valid, what does it mean? How do we
  54872. 40:57:00organize it? Um how do we analyze it?
  54873. 40:57:03Then you got to take those analysis and
  54874. 40:57:04interpret it into something that uh
  54875. 40:57:06people can use. kind of reduce it to
  54876. 40:57:09understandable. Um, and nowadays you
  54877. 40:57:11have to be able to present it. If you
  54878. 40:57:13can't present it, then no one else is
  54879. 40:57:14going to understand what the heck you
  54880. 40:57:15did.
  54881. 40:57:17So, we look at the terminologies. Uh,
  54882. 40:57:20there is a lot of terminologies
  54883. 40:57:22depending on what domain you're working
  54884. 40:57:24in. So clearly if you're working in um a
  54885. 40:57:28domain that deals with
  54886. 40:57:31viruses and tea cells and and how does
  54887. 40:57:36you know where does that come from and
  54888. 40:57:37you're studying the different people
  54889. 40:57:38then you're going to have a population.
  54890. 40:57:40if you are working with um mechanical
  54891. 40:57:44gear um you know a little bit different
  54892. 40:57:46if you're looking for the wobbling
  54893. 40:57:48statistics uh to know when to replace a
  54894. 40:57:51rotor on a machine or something like
  54895. 40:57:52that uh that can be a big deal. You
  54896. 40:57:54know, we have these huge fans that turn
  54897. 40:57:57in our sewage processing systems. And so
  54898. 40:58:01those fans, they start to wobble and hum
  54899. 40:58:03and do different things that the sensors
  54900. 40:58:05pick up. At one point, do you replace
  54901. 40:58:07them? Instead of waiting for it to
  54902. 40:58:08break, in which case it cost a lot of
  54903. 40:58:10money. Instead of replacing a bushing,
  54904. 40:58:11you're replacing the whole fan unit. Uh
  54905. 40:58:14an interesting project that came up for
  54906. 40:58:16our city a while back. Uh so population,
  54907. 40:58:19all objects are measurements whose
  54908. 40:58:21properties are being observed. Uh so
  54909. 40:58:24that's your population all the objects.
  54910. 40:58:27It's easy to see it with people because
  54911. 40:58:28we have our population and large. Um but
  54912. 40:58:32in the case of the sewer fans we're
  54913. 40:58:34talking about how the fan units. That's
  54914. 40:58:36the population of fans that we're
  54915. 40:58:37working with.
  54916. 40:58:39You have a parameter a matrix uh that is
  54917. 40:58:42used to represent a population or
  54918. 40:58:44characteristic.
  54919. 40:58:45You have your sample a subset of the
  54920. 40:58:48population studied. You don't want to do
  54921. 40:58:50them all because then you don't have a
  54922. 40:58:51if you come up with a conclusion for
  54923. 40:58:53everyone, you don't have a way of
  54924. 40:58:55testing it. So you take a sample. Uh
  54925. 40:58:57sometimes you don't have a choice. You
  54926. 40:58:58can only take a sample of what's going
  54927. 40:58:59on. You can't u study the whole
  54928. 40:59:02population. And a variable, a metric of
  54929. 40:59:05interest for each person or object in a
  54930. 40:59:07population.
  54931. 40:59:09Types of sampling. We have probabilistic
  54932. 40:59:12approach. uh selecting samples from a
  54933. 40:59:14larger population using a method based
  54934. 40:59:17on the theory of probability
  54935. 40:59:20and we'll go into a little bit more
  54936. 40:59:21deeper on these. We have random
  54937. 40:59:23systematic stratified and then you have
  54938. 40:59:26nonprobabilistic approach selecting
  54939. 40:59:28samples based on the subjective judgment
  54940. 40:59:30of the researcher rather than random
  54941. 40:59:33selection. Uh it has to do with
  54942. 40:59:35convenience trying to reach a quota um
  54943. 40:59:37or snowball. Uh and they're very biased.
  54944. 40:59:41That's one of the reasons you'll see
  54945. 40:59:42this big stamp on it says biased. Uh so
  54946. 40:59:44you got to be very careful on that. So
  54947. 40:59:47probabilistic sampling uh when we talk
  54948. 40:59:50about a random sampling, we select
  54949. 40:59:52random size samples from each group or
  54950. 40:59:54category. So we it's as random as you
  54951. 40:59:56can get. Uh we talk about systematic
  54952. 40:59:59sampling. We're selecting randomsiz
  54953. 41:00:02samples from each group or category with
  54954. 41:00:04a fixed periodic interval. Uh so we kind
  54955. 41:00:08of split it up. This would be like a
  54956. 41:00:09time setup or different categories. And
  54957. 41:00:12you might ask your question, what is a
  54958. 41:00:13category or a group? Uh if you look at
  54959. 41:00:17I'm going to go back a window. Let's say
  54960. 41:00:19we're studying um economics of different
  54961. 41:00:21of an area. Um we know pretty much that
  54962. 41:00:25based on their culture, where they came
  54963. 41:00:28from, they might need to be separated.
  54964. 41:00:31And so uh and when I say separated, I
  54965. 41:00:33don't mean separated from their their uh
  54966. 41:00:35place where they live. I mean, as far as
  54967. 41:00:38the analysis, we want to look at the
  54968. 41:00:39different groups and make sure they're
  54969. 41:00:41all represented. So, if we had like an
  54970. 41:00:4480% uh of a group that is uh say
  54971. 41:00:47Hispanic and or Indian and also in that
  54972. 41:00:51same area, we have 20% 20% who are let's
  54973. 41:00:55call our expatriots. They left America
  54974. 41:00:57and they're nice and uh your Caucasian
  54975. 41:00:59group. We might want to sample a group
  54976. 41:01:02that is representative of both. Uh, so
  54977. 41:01:05we're talking about stratified sampling
  54978. 41:01:08and we're talking about groups. Those
  54979. 41:01:09are the groups we're talking about. And
  54980. 41:01:10it brings us to stratified sampling,
  54981. 41:01:12selecting approximately equalized
  54982. 41:01:14samples from each group or category. Uh,
  54983. 41:01:17this way we can actually separate the
  54984. 41:01:19categories and give us an insight into
  54985. 41:01:22the different cultures and how that
  54986. 41:01:24might affect them in that area. Uh so
  54987. 41:01:26you can see these are very very
  54988. 41:01:28different kind of depends on what you're
  54989. 41:01:30working with um as far as your data and
  54990. 41:01:33what you're studying. And so we can see
  54991. 41:01:35here just to go a little bit more we'd
  54992. 41:01:37have selecting 25 employees from a
  54993. 41:01:39company of 250 employees randomly. Don't
  54994. 41:01:41care anything about them. What groups
  54995. 41:01:42they're in, which office are in,
  54996. 41:01:44nothing. Um and we might be selecting
  54997. 41:01:47one employee from every 50 unique
  54998. 41:01:49employees in a company of 250 employees.
  54999. 41:01:52And then we have selecting one employee
  55000. 41:01:54from every branch in the company office.
  55001. 41:01:56So we have all the different branches.
  55002. 41:01:58There's our group or our categories by
  55003. 41:02:00the branch. And the category could
  55004. 41:02:01depend on what you're studying. So it
  55005. 41:02:03has a lot of variation on there. You see
  55006. 41:02:06this kind of grouping and categorizing
  55007. 41:02:08is also used to generate a lot of
  55008. 41:02:10misinformation.
  55009. 41:02:12Uh so if you only study one group and
  55010. 41:02:14you say this is what it is, then
  55011. 41:02:16everybody assumes that's what it is for
  55012. 41:02:18everybody. And so you got to be very
  55013. 41:02:20careful of that. and it's very unethical
  55014. 41:02:21thing to kind of do. So, types of
  55015. 41:02:24statistics. Uh we talk about statistics,
  55016. 41:02:27we're going to talk about descriptive
  55017. 41:02:29and inferential statistics. There are so
  55018. 41:02:33many different terms in statistics to
  55019. 41:02:35break it up. Uh so we so we're talking
  55020. 41:02:37about a particular setup. So we're
  55021. 41:02:41talking about descriptive and
  55022. 41:02:42inferential uh statistics. You the base
  55023. 41:02:45of the word describe is pretty solid.
  55024. 41:02:48you're describing the data. What does it
  55025. 41:02:51look like? With inferial statistics,
  55026. 41:02:54we're going to take that from the small
  55027. 41:02:55population to a large population. So, if
  55028. 41:02:58you're working with a drug company, uh
  55029. 41:02:59you might look at the data and say,
  55030. 41:03:01"These people were helped by this drug.
  55031. 41:03:04They did uh 80% better as far as their
  55032. 41:03:07health or 80% better survival rate than
  55033. 41:03:09the people um who did not have the drug.
  55034. 41:03:13So, we can infer that that drug will
  55035. 41:03:15work in the greater populace and will
  55036. 41:03:16help people. So that's where you get
  55037. 41:03:18your inferential. Uh so we are
  55038. 41:03:20predicting how it's going to affect the
  55039. 41:03:22greater population.
  55040. 41:03:24So descriptive statistics it is used to
  55041. 41:03:26describe the basic features of data and
  55042. 41:03:28form the basis of quantitative analysis
  55043. 41:03:31of data. So we have a measure of central
  55044. 41:03:34tendencies. We have your mean, median
  55045. 41:03:36and mode. And then we have a measure of
  55046. 41:03:39spread like your range, your
  55047. 41:03:40interquartile range, your variance and
  55048. 41:03:43your standard deviation. And we're going
  55049. 41:03:45to look at all these a little deeper
  55050. 41:03:46here in a second. Uh but one of them you
  55051. 41:03:49can think of is um how the data
  55052. 41:03:53difference differences you know what's
  55053. 41:03:55the max men range all that stuff is your
  55054. 41:03:58spread and anything that's just a single
  55055. 41:04:00number is usually your central uh
  55056. 41:04:02tendencies measure of central
  55057. 41:04:04tendencies. So we talk about the mean it
  55058. 41:04:07is the average of the set of values
  55059. 41:04:08considered. what is the average outcome
  55060. 41:04:11of whatever's going on? And then your
  55061. 41:04:13median separates the higher half and the
  55062. 41:04:16lower half of data.
  55063. 41:04:19Uh so where's the center point of all
  55064. 41:04:21your different data points? So your mean
  55065. 41:04:24might have some a couple really big
  55066. 41:04:26numbers that skew it uh so that the
  55067. 41:04:29average is much higher than if you took
  55068. 41:04:31those outliers out where the median
  55069. 41:04:34would by separating the high from the
  55070. 41:04:36low might give you a much lower number.
  55071. 41:04:39you might look at and say, "Oh, that's
  55072. 41:04:40that's odd. Why is the average so much
  55073. 41:04:42higher than the median?" Well, it's
  55074. 41:04:44because you have some outliers, or why
  55075. 41:04:45is it so much lower? And then the mode
  55076. 41:04:47is the most frequent appearing value.
  55077. 41:04:50Uh, this is really interesting. If
  55078. 41:04:51you're studying economics and how people
  55079. 41:04:53are doing, you might find that the most
  55080. 41:04:55common um income like in the US was at
  55081. 41:04:59one point 24,000 a year where the
  55082. 41:05:02average was closer to 80,000. And it's
  55083. 41:05:05like, wow, what a difference. Well,
  55084. 41:05:07there's some people have a lot of money
  55085. 41:05:09and so that skews that way up. So the
  55086. 41:05:11average person is not making that kind
  55087. 41:05:13of money. And then you look at the
  55088. 41:05:14median income and you're like, well, the
  55089. 41:05:16median income is a little bit closer to
  55090. 41:05:18the average. Uh so it does create a very
  55091. 41:05:20interesting way of looking at the data.
  55092. 41:05:22Again, these are all uh central
  55093. 41:05:25tendencies, single numbers you can look
  55094. 41:05:26at for the whole spread of the data.
  55095. 41:05:29And we look at the measure of central
  55096. 41:05:32tendencies. The mean is the average
  55097. 41:05:33marks of a students in a classroom. So
  55098. 41:05:35here we have the mean sum of the marks
  55099. 41:05:37of the students total number of students
  55100. 41:05:40and as we talked about the median uh if
  55101. 41:05:42we have 0 through 10 and we take half
  55102. 41:05:46the numbers and put them on one side of
  55103. 41:05:47the line half the numbers on the other
  55104. 41:05:49side of the line uh we end up with five
  55105. 41:05:51in the middle and then the mode what
  55106. 41:05:53mark was scored by most of the students
  55107. 41:05:55in a test in a simple case where most
  55108. 41:05:58people scored like an 82% and got
  55109. 41:06:01certain problems wrong easy to figure
  55110. 41:06:03out. uh not so easy when you have
  55111. 41:06:06different areas where like you have like
  55112. 41:06:08the um oh let's go back to economy a
  55113. 41:06:11little bit more difficult to calculate
  55114. 41:06:12if you have a large group that scores
  55115. 41:06:14that makes 30,000 and a slightly bigger
  55116. 41:06:17group that makes 26,000. So what do you
  55117. 41:06:19put down for the mode? Uh certainly
  55118. 41:06:21there's a number of ways to calculate
  55119. 41:06:22that and there's actually a different
  55120. 41:06:24variations depending on what you're
  55121. 41:06:25doing. So now we're looking at a measure
  55122. 41:06:28of spread uh range. What's the
  55123. 41:06:30difference between the highest and the
  55124. 41:06:31lowest value? First thing you want to
  55125. 41:06:33look at, you know, it's we had everybody
  55126. 41:06:35in the test scored between 60 and 100%,
  55127. 41:06:37somebody got 100% or maybe 60 to 90%. It
  55128. 41:06:41was so hard that a lot of people could
  55129. 41:06:42not get 100%. Um, and you have your
  55130. 41:06:46interquartile range. Quartortiles divide
  55131. 41:06:49a rankorder data set into four equal
  55132. 41:06:52parts.
  55133. 41:06:53very common thing to do as part of all
  55134. 41:06:55the basic packages whether you're
  55135. 41:06:57working in uh dataf frames with pandas
  55136. 41:07:01whether you're working in scala whether
  55137. 41:07:03you're working in R um you'll see this
  55138. 41:07:05come up where they have range your min
  55139. 41:07:07your max and then it'll have your
  55140. 41:07:09interquartile range how does it look
  55141. 41:07:11like in each quarter of data variance
  55142. 41:07:13measures how far each number in the set
  55143. 41:07:16is from the mean and therefore from
  55144. 41:07:18every other number in the set uh so you
  55145. 41:07:21have like a how much turbulence is going
  55146. 41:07:23on in this data. And then the standard
  55147. 41:07:26deviation, it is the measure of the
  55148. 41:07:28variance or the dispersion of a set of
  55149. 41:07:30values from the mean. And you'll usually
  55150. 41:07:33see uh if I'm doing a graph, I might
  55151. 41:07:35have the value graphed. Um and then
  55152. 41:07:37based on the the error, I might graph
  55153. 41:07:40graph the standard deviation and the
  55154. 41:07:42error on the graph as a background so
  55155. 41:07:44you can see how far off it is. Uh so
  55156. 41:07:47standard deviation is used a lot. So
  55157. 41:07:49measurement of spread uh marks of a
  55158. 41:07:51student out of a 100 uh we have here
  55159. 41:07:54from 50 to 63 or 50 to 90 uh so the
  55160. 41:07:58range maximum marks minimum marks we
  55161. 41:08:00have 90 to 45 and the spread of that is
  55162. 41:08:0245 90 - 45 and then we have the
  55163. 41:08:05interquartile range using the same marks
  55164. 41:08:08over there you can see here where the
  55165. 41:08:10median is and then there's the first
  55166. 41:08:13quarter the second quarter and the third
  55167. 41:08:15quarter based on splitting it apart by
  55168. 41:08:17those values
  55169. 41:08:19And to understand the variance and
  55170. 41:08:20standard deviation, we first need to
  55171. 41:08:22find out the mean. Uh so here's our our
  55172. 41:08:25you know calculating the average there.
  55173. 41:08:27We end up at approximately 66 for the
  55174. 41:08:29average. And then we look at that the
  55175. 41:08:31variance once we know the means we can
  55176. 41:08:33do equals the marks minus the mean
  55177. 41:08:35squared. Why is it squared? Uh because
  55178. 41:08:39one, you want to make sure it's you
  55179. 41:08:41don't have like if you if you're putting
  55180. 41:08:43all this stuff together, you end up with
  55181. 41:08:45an error as far as one's negative, one's
  55182. 41:08:47positive, one's a little higher, one's a
  55183. 41:08:49little lower. Uh so you always see the
  55184. 41:08:52squared value and over the total
  55185. 41:08:54observations. And so the standard
  55186. 41:08:56deviation equals the square root of the
  55187. 41:08:58variance, which is approximately 16. And
  55188. 41:09:02if you were looking at um a predictable
  55189. 41:09:04model, you would be looking at the
  55190. 41:09:06deviation based on the error. How much
  55191. 41:09:09error does it have? Uh that's again
  55192. 41:09:12really important to know if you're if
  55193. 41:09:14your prediction is predicting something,
  55194. 41:09:16what's a chance of it being way off or
  55195. 41:09:18just a little bit off.
  55196. 41:09:21Now that we've looked at the um tools as
  55197. 41:09:24far as some of the basics for doing your
  55198. 41:09:26statistics and what we're talking about,
  55199. 41:09:28let's go ahead and pull up a little demo
  55200. 41:09:30and show you what that looks like in
  55201. 41:09:31Python code. Uh so you can get some
  55202. 41:09:33little hands-on here. For that, let's go
  55203. 41:09:35back into our Jupyter notebook in
  55204. 41:09:37Python. Now, almost all of this you can
  55205. 41:09:40do in numpy. Last time we worked um in
  55206. 41:09:42numpy. This time we're going to go ahead
  55207. 41:09:44and use pandas. And if you remember from
  55208. 41:09:47pandas on here, uh this is basically a
  55209. 41:09:50data frame, rows, columns. Let's just go
  55210. 41:09:53ahead and do a print df. head
  55211. 41:09:58and run that.
  55212. 41:10:00And you can see we have uh the name
  55213. 41:10:02Jane, Michael, William, Rosie, Hannah,
  55214. 41:10:04and their salaries on here. And of
  55215. 41:10:06course, instead of having to do all
  55216. 41:10:08those hand calculations and add
  55217. 41:10:09everything together and divide by the
  55218. 41:10:11total, we can do something very simple
  55219. 41:10:13on this uh like use the command mean in
  55220. 41:10:17pandas. And so if I go ahead and do this
  55221. 41:10:19print df, pick our column salary because
  55222. 41:10:22we want to find the means of that
  55223. 41:10:24colery.
  55224. 41:10:25We want to find the means of that
  55225. 41:10:27column. Uh and we go and print this out.
  55226. 41:10:29And you can see that the uh average
  55227. 41:10:32income on here is 71,000.
  55228. 41:10:35Uh, and let's just go ahead and do this.
  55229. 41:10:37We'll go ahead and put in uh means.
  55230. 41:10:42And if we're going to do that, we also
  55231. 41:10:44might want to find the median.
  55232. 41:10:48And the median is uh very similar except
  55233. 41:10:52it actually is just median. Uh we're
  55234. 41:10:54used to means and average. It's kind of
  55235. 41:10:56interesting that those are they use the
  55236. 41:10:58two different words. Uh there can be in
  55237. 41:11:01some computations slight differences but
  55238. 41:11:03for the most part the means is the
  55239. 41:11:05average. Uh and then the median oops
  55240. 41:11:08let's put a
  55241. 41:11:12median here. DF salary that way it
  55242. 41:11:14displays a little better. We can see the
  55243. 41:11:16median is 54 um000. So the halfway mark
  55244. 41:11:20is significantly below the average. Why?
  55245. 41:11:23Because we have somebody in here who
  55246. 41:11:24makes 189,000.
  55247. 41:11:26Darn you Rosie for throwing off our
  55248. 41:11:28numbers. Uh but that's something you'd
  55249. 41:11:30want to notice. This is this is the
  55250. 41:11:32difference between these is huge and so
  55251. 41:11:34is what is the meaning behind that when
  55252. 41:11:36you're studying a populace and looking
  55253. 41:11:37at uh the different data coming in. And
  55254. 41:11:40of course we also want to find out hey
  55255. 41:11:42what's the most uh common income that
  55256. 41:11:46people make in this little tiny sample.
  55257. 41:11:48And so we'll go ahead and do the mode.
  55258. 41:11:51And you can see here with the mode uh
  55259. 41:11:53it's at 50,000.
  55260. 41:11:55So this is this is very telling that
  55261. 41:11:57most people are making 50,000. The
  55262. 41:12:00middle point is at 54,000. So half the
  55263. 41:12:03people are making more than that. What
  55264. 41:12:05that tells me is that if the most common
  55265. 41:12:08income is way is below the median, then
  55266. 41:12:13there's a few there's a SK, you know,
  55267. 41:12:14there's a a lot of high salaries going
  55268. 41:12:16up, but there's some really low salaries
  55269. 41:12:18in there. And so this trend which is
  55270. 41:12:21very common in statistic you when you're
  55271. 41:12:23analyzing the economy and different
  55272. 41:12:26people's income is pretty common and the
  55273. 41:12:28bigger difference between these is also
  55274. 41:12:31very important when we're studying
  55275. 41:12:32statistics. Uh and when you hear someone
  55276. 41:12:35just say hey the average income was you
  55277. 41:12:37might start asking questions at that
  55278. 41:12:39point. Why aren't you talking about the
  55279. 41:12:40median income? Why aren't you talking
  55280. 41:12:42about the mode the most common income?
  55281. 41:12:44What are you hiding? Uh and if you're
  55282. 41:12:47doing these analysis, you should be
  55283. 41:12:48looking at these saying, "Hey, why why
  55284. 41:12:49are this discrepancies? Why are these so
  55285. 41:12:51different?" And of course, with any uh
  55286. 41:12:53analysis, it's important to find out the
  55287. 41:12:55minimum
  55288. 41:12:57and the maximum. So, we'll go ahead.
  55289. 41:12:59It's just simply uh um min'll pull up
  55290. 41:13:04your minimum and then do max pulls up
  55291. 41:13:06the maximum. pretty straightforward on
  55292. 41:13:09as far as um translating it and knowing
  55293. 41:13:13what your you know what the your lowest
  55294. 41:13:15value and what your highest value is
  55295. 41:13:16here. Um which you'll use to generate
  55296. 41:13:19like a spread later on. And real quick
  55297. 41:13:22on no mode mode, uh note that it puts
  55298. 41:13:24mode zero. Like I said, there's a couple
  55299. 41:13:26different ways you can compute the mode.
  55300. 41:13:28Um although, you know, standard one's
  55301. 41:13:30pretty good. We can of course do the
  55302. 41:13:32range, which is your max minus your min.
  55303. 41:13:35So now we have a range of 149,000
  55304. 41:13:38between the upper end and the lower end.
  55305. 41:13:40And you might want to be looking up the
  55306. 41:13:42individual values on all of these. But
  55307. 41:13:45it turns out there is a describe
  55308. 41:13:49feature in pandas.
  55309. 41:13:51And so in pandas we can actually do df
  55310. 41:13:54salary describe. And if we do this you
  55311. 41:13:56can see we have that there's seven uh
  55312. 41:13:58setups. Here's our mean. Um, our
  55313. 41:14:01standard deviation, which we didn't
  55314. 41:14:03compute yet, which would just be a STD.
  55315. 41:14:06And you got to be a little careful
  55316. 41:14:07because when it computes it, it looks
  55317. 41:14:08for axes and things like that. Uh, we
  55318. 41:14:10have our minimum value, and here's our
  55319. 41:14:12cortiles,
  55320. 41:14:14uh, our maximum value, and then of
  55321. 41:14:16course the name salary. Uh, so these are
  55322. 41:14:18the these are the basic statistics. You
  55323. 41:14:20can pull them up and just describe. This
  55324. 41:14:22is a dictionary. So I could actually do
  55325. 41:14:25something like um in here I could
  55326. 41:14:28actually go uh count and run. And now it
  55327. 41:14:32just prints the count. Uh so because
  55328. 41:14:34this is a dictionary, you can pull any
  55329. 41:14:36one of these values out of here. It's
  55330. 41:14:38kind of a quick and dirty way to pull
  55331. 41:14:40all the different information and then
  55332. 41:14:42split it up and depending on what you
  55333. 41:14:43need. Now if I just walked in and gave
  55334. 41:14:46you this information um in a meeting, at
  55335. 41:14:49some point you would just kind of fall
  55336. 41:14:51asleep. That's what I would do anyway.
  55337. 41:14:55Um, so we want to go ahead and and see
  55338. 41:14:57about graphing it here. And we'll go
  55339. 41:14:59ahead and put it into a histogram and
  55340. 41:15:01plot that graph on it of the salaries.
  55341. 41:15:04And let's just go ahead and put that in
  55342. 41:15:06here. So we do our map plot inline.
  55343. 41:15:09Remember that's a Jupiter's notebook
  55344. 41:15:10thing. Uh, a lot of the new version of
  55345. 41:15:13the mapplot library does it
  55346. 41:15:14automatically, but just in case I always
  55347. 41:15:17put it in there. Uh, import mattplot
  55348. 41:15:18library piplot as plt. That's my
  55349. 41:15:21plotting.
  55350. 41:15:23And then we have our data frame. Uh I
  55351. 41:15:25don't I guess I really don't need to
  55352. 41:15:26respell the data frame. Maybe we could
  55353. 41:15:28just remind oursel what's in it. So
  55354. 41:15:30we'll go ahead and just uh print
  55355. 41:15:32DF. That way we still have it. And then
  55356. 41:15:35we have our salary. DF salary
  55357. 41:15:38salary.plot history title salary
  55358. 41:15:40distribution color gray. Uh plot AXV
  55359. 41:15:44line salary the mean value. So, we're
  55360. 41:15:47going to take the mean value um color
  55361. 41:15:50violet line style dash. This is just all
  55362. 41:15:53making it pretty. Uh what color dash
  55363. 41:15:56line width of two that kind of thing.
  55364. 41:15:59And the median. And let's go ahead and
  55365. 41:16:00run this just so you can see what we're
  55366. 41:16:02talking about.
  55367. 41:16:04And so up here we are taking on our
  55368. 41:16:07plot. Um so here's the data. Here's our
  55369. 41:16:10our data frame printed out so you can
  55370. 41:16:12see it with the salaries. We're looking
  55371. 41:16:14at the salary distribution and just look
  55372. 41:16:16at this the way they're the salary is
  55373. 41:16:18distributed. Um you have our in this
  55374. 41:16:22case we did let's see we had red for the
  55375. 41:16:25median we have violet
  55376. 41:16:28for our average or mean and you can just
  55377. 41:16:32see how it really here's our outlier.
  55378. 41:16:35Here's our person who makes a lot of
  55379. 41:16:36money. Here's the um average and here's
  55380. 41:16:40the median. Um, and so as you look at
  55381. 41:16:42this, you can say, "Wow." Um, based on
  55382. 41:16:44the average, it really doesn't tell you
  55383. 41:16:46much about what people are really taking
  55384. 41:16:47home. All it does is tell you how much
  55385. 41:16:50money is in this, you know, what the
  55386. 41:16:52average salary is. So, some of the
  55387. 41:16:55things you want to take away in addition
  55388. 41:16:57to this is that it's very easy to plot
  55389. 41:17:01um an AXV line. These are these up and
  55390. 41:17:04down lines for your markers. Um, and as
  55391. 41:17:08you display display the data, I mean,
  55392. 41:17:09you can add all kinds of things to this
  55393. 41:17:11and get really complicated. Keeping it
  55394. 41:17:13simple is pretty straightforward. I look
  55395. 41:17:14at this and I can see we have a major
  55396. 41:17:16outlier out here. We can definitely do a
  55397. 41:17:18histogram and stuff like that. Um, but
  55398. 41:17:20you know, picture's worth a thousand
  55399. 41:17:22words. What you really want to make sure
  55400. 41:17:24you take away is that we can do a basic
  55401. 41:17:26describe which pulls all this
  55402. 41:17:29information out and we can print any of
  55403. 41:17:31the individual information from the
  55404. 41:17:33describe uh because this is a
  55405. 41:17:35dictionary.
  55406. 41:17:38And so if we want to go ahead and look
  55407. 41:17:40up um the mean value, we can also do
  55408. 41:17:43describe mean. So if you're doing a lot
  55409. 41:17:45of statistics, uh being able to
  55410. 41:17:49doesn't have the print on there, so it's
  55411. 41:17:50only going to print um the last one,
  55412. 41:17:52which happens to be the mean. Uh you can
  55413. 41:17:54very easily reference any one of these.
  55414. 41:17:56And then you can also, if you're doing
  55415. 41:17:58something a little bit more complicated
  55416. 41:17:59and you don't need just the basics, you
  55417. 41:18:01can come through and pull any one of the
  55418. 41:18:04individual um
  55419. 41:18:07references from the from the pandas on
  55420. 41:18:10here. So now we've had a chance to
  55421. 41:18:13describe our data. Uh let's get into
  55422. 41:18:16inferential statistics. Inferial
  55423. 41:18:19statistics allows you to make
  55424. 41:18:20predictions or inferences from data. And
  55425. 41:18:24you can see here we have a nice little
  55426. 41:18:25picture movie ratings and um if we took
  55427. 41:18:29this group of people and said hey how
  55428. 41:18:31many people like the movie dislike it
  55429. 41:18:33can't say and then you ask just a random
  55430. 41:18:36person who comes out of the movie who
  55431. 41:18:37hasn't been in this study uh you can
  55432. 41:18:39infer that 55% chance of saying liked
  55433. 41:18:4335% chance of saying disliked or a 10 or
  55434. 41:18:4611% chance of can't say. So that that's
  55435. 41:18:49real basics of what we're talking about
  55436. 41:18:51is you're going to infer that the next
  55437. 41:18:53person is going to follow these
  55438. 41:18:54statistics.
  55439. 41:18:57Uh so let's look at point estimation. Uh
  55440. 41:19:00it is a process of finding an
  55441. 41:19:01approximate value for a population's
  55442. 41:19:03parameter like mean or average from
  55443. 41:19:06random samples of the population. Let's
  55444. 41:19:09take an example of testing vaccines for
  55445. 41:19:11COVID 19. Uh vaccines and flu bugs, all
  55446. 41:19:14that. It's a pretty big thing of how do
  55447. 41:19:16you test these out and make sure they're
  55448. 41:19:17going to work on the populace. A group
  55449. 41:19:20of people are chosen from the
  55450. 41:19:21population. Medical trials are
  55451. 41:19:23performed. Results are generalized for
  55452. 41:19:26the whole population. So here's a
  55453. 41:19:28protected here's our small group up here
  55454. 41:19:30where we've selected them. We run
  55455. 41:19:32medical trials on them and then the
  55456. 41:19:33results work for the population. You
  55457. 41:19:35nice diagram with the arrows going back
  55458. 41:19:37and forth and the very scary co virus in
  55459. 41:19:40the middle of one. And let's take a look
  55460. 41:19:42at the applications of inferial
  55461. 41:19:44statistics.
  55462. 41:19:46Very central is what they call
  55463. 41:19:47hypothesis testing uh and the confidence
  55464. 41:19:51interval which go with that. And then as
  55465. 41:19:54we get into
  55466. 41:19:56probability, we get into our binomial
  55467. 41:19:59theorem, our normal distribution and
  55468. 41:20:01central limit theorem. Hypothesis
  55469. 41:20:04testing. Hypothesis testing is used to
  55470. 41:20:06measure the plausibility of a hypothesis
  55471. 41:20:09assumption by using sample data. Now
  55472. 41:20:13when we talk about theorems, theory,
  55473. 41:20:17hypothesis,
  55474. 41:20:19uh keep in mind that if you are in a
  55475. 41:20:21philosophy class, theory is the same as
  55476. 41:20:25hypothesis where theorem is a scientific
  55477. 41:20:29uh statement that is something that has
  55478. 41:20:30been proven although it is always up for
  55479. 41:20:33debate because in science we always want
  55480. 41:20:35to make sure things are up to debate. So
  55481. 41:20:37a hypothesis is the same as a phil
  55482. 41:20:39philosophical class calling a theory
  55483. 41:20:41where theory in science is not the same.
  55484. 41:20:44Theory in science says this has been
  55485. 41:20:45well proven. Gravity is a theory. Uh so
  55486. 41:20:48if you want to debate the theory of
  55487. 41:20:50gravity try jumping up and down. If you
  55488. 41:20:52want to have a theory about why the
  55489. 41:20:54economy is collap collapsing in your
  55490. 41:20:56area that is a philosophical debate.
  55491. 41:21:00Very important. I've heard people mix
  55492. 41:21:01those up and it is a pet peeve of mine.
  55493. 41:21:04When we talk about hypothesis testing,
  55494. 41:21:06the steps involved in hypothesis testing
  55495. 41:21:08is first we formulate a hypothesis. We
  55496. 41:21:11figure out the right test to test our
  55497. 41:21:13hypothesis. We execute the test and we
  55498. 41:21:16make a decision. And so when you're
  55499. 41:21:18talking about hypothesis, you're usually
  55500. 41:21:20trying to disprove it. If you can't
  55501. 41:21:22disprove it and it works for all the
  55502. 41:21:25facts, then you might call that a
  55503. 41:21:27theorem at some point. So in a use case,
  55504. 41:21:30uh let's consider an example. We have
  55505. 41:21:32four students. were given a task to
  55506. 41:21:33clean a room every day. Sounds like
  55507. 41:21:35working with my kids. They decided to
  55508. 41:21:37distribute the job of cleaning the room
  55509. 41:21:39among themselves. They did so by making
  55510. 41:21:41four chits which has their names on it
  55511. 41:21:44and the name that gets picked up has to
  55512. 41:21:46do the cleaning for that day. Rob took
  55513. 41:21:48the opportunity to make chits and wrote
  55514. 41:21:50everyone's name on it. So here's our
  55515. 41:21:52four people, Nick, Rob, Imlia, Imlia,
  55516. 41:21:55and Summer.
  55517. 41:21:58Now Rick, Imlia and Summer are asking us
  55518. 41:22:00to decide whether Rob has done some
  55519. 41:22:03mischief in preparing the chits i.e
  55520. 41:22:05whether Rob has written his name on one
  55521. 41:22:07of the chit. For that we will find out
  55522. 41:22:09the probability of Rob getting the
  55523. 41:22:11cleaning job on first day, second day,
  55524. 41:22:13third day and so on till 12 days. The
  55525. 41:22:16probability of Rob getting the job
  55526. 41:22:18decreases every day. I.e. his turn never
  55527. 41:22:21comes up. Then definitely he has done
  55528. 41:22:23some mischief while making the chits. So
  55529. 41:22:26the probability of Rob not doing work on
  55530. 41:22:28day one is uh three out of four. There's
  55531. 41:22:30a 75 chance that he didn't do work. Uh
  55532. 41:22:33two days 34s * 34s equals.56.
  55533. 41:22:383 days you have 3/4 34 34 which
  55534. 41:22:41equals42.
  55535. 41:22:43Uh when you get to day 12 it's 0032
  55536. 41:22:46which is less than 0.05.
  55537. 41:22:48Remember this 0.05 uh that comes up a
  55538. 41:22:51lot when we're talking about um certain
  55539. 41:22:54values when we're looking at statistics.
  55540. 41:22:57Rob is cheating as he wasn't chosen for
  55541. 41:22:5912 consecutive days. That's a very high
  55542. 41:23:01probability when on day 12 he still
  55543. 41:23:05hasn't gotten the job cleaning the room.
  55544. 41:23:08So we come up to our important important
  55545. 41:23:10terminologies.
  55546. 41:23:12We have null hypothesis.
  55547. 41:23:15a general statement that states that
  55548. 41:23:16there is no relationship between two
  55549. 41:23:18measured phenomenon or no assoc
  55550. 41:23:20association among the groups.
  55551. 41:23:23Alternative hypothesis contrary to the
  55552. 41:23:26null hypothesis it states whenever
  55553. 41:23:28something is happening a new theory is
  55554. 41:23:30preferred instead of an old one. And so
  55555. 41:23:33the two hypothesis go hand in hand. Uh
  55556. 41:23:36so your null this is always interesting
  55557. 41:23:38in in we're talking about data science
  55558. 41:23:40and the math behind it. It's about
  55559. 41:23:42proving that the things have no
  55560. 41:23:44correlation. Null hypothesis says these
  55561. 41:23:47two have zero relation to each other.
  55562. 41:23:49Where the alternative hypothesis says,
  55563. 41:23:51hey, we found a relation. This is what
  55564. 41:23:53it is. We have p value. The p value is
  55565. 41:23:56the probability of finding the observed
  55566. 41:23:58or more extreme results when the null
  55567. 41:24:00hypothesis of a study question is true.
  55568. 41:24:04And the t value, it is simply the
  55569. 41:24:06calculated difference represented in
  55570. 41:24:08units of standard error. The greater the
  55571. 41:24:10magnitude of t, the greater the evidence
  55572. 41:24:12against the null hypothesis. And you can
  55573. 41:24:14look at the t value as being specific to
  55574. 41:24:17the test you're doing where the p value
  55575. 41:24:20is derived from your t value and you're
  55576. 41:24:22looking for what they call the 5% or the
  55577. 41:24:250.05
  55578. 41:24:26showing that it has a high correlation.
  55579. 41:24:28So digging in deeper, let's assume that
  55580. 41:24:31a new drug is developed with the goal of
  55581. 41:24:33lowering the blood pressure more than
  55582. 41:24:35the existing drug. And this is a good
  55583. 41:24:37one because uh the null value here isn't
  55584. 41:24:40that you don't have any drug. The null
  55585. 41:24:41value here is that it's better than the
  55586. 41:24:43existing drug. The new drug doesn't
  55587. 41:24:45lower the blood pressure more than the
  55588. 41:24:47existing drug. Now if we get that uh
  55589. 41:24:50that says our null hypothesis is
  55590. 41:24:52correct. There is no correlation and the
  55591. 41:24:54new drug is not doing its job. The
  55592. 41:24:57alternative hypothesis the new drug does
  55593. 41:25:00significantly lower the blood pressure
  55594. 41:25:02more than the existing drug. Uh, yay, we
  55595. 41:25:05got a new drug out there. And that's our
  55596. 41:25:06alternative hypothesis or the H1 or HA.
  55597. 41:25:10And we look at the p value results from
  55598. 41:25:13the evidence like medical trials showing
  55599. 41:25:15positive results which will reject the
  55600. 41:25:17null hypothesis. And again, they're
  55601. 41:25:19looking for um a 0.05 or 5%. And the t
  55602. 41:25:23value comparing all the positive test
  55603. 41:25:25results and finding means of different
  55604. 41:25:27samples in order to test hypothesis. So
  55605. 41:25:30this is specific to the test. how uh
  55606. 41:25:32what percentage of increase did they
  55607. 41:25:34have and this leads us to the confidence
  55608. 41:25:37intervals. Uh a confidence interval is a
  55609. 41:25:40range of values we are sure our true
  55610. 41:25:42values of observations lie in. Let's say
  55611. 41:25:45you asked a dog owner around you and
  55612. 41:25:47asked them how many cans of food do you
  55613. 41:25:50buy for your uh per year for your dog.
  55614. 41:25:53Through calculations you got to know
  55615. 41:25:55that the on an average around 95% of the
  55616. 41:25:58people bought around 200 to 300 cans of
  55617. 41:26:00food. Hence we can say that we have a
  55618. 41:26:03confidence interval of 230 where 95% of
  55619. 41:26:06our values lie in that spread data
  55620. 41:26:09spread. Uh and this the graph really
  55621. 41:26:12helps a lot. So you can start seeing
  55622. 41:26:13what you're looking at here where you
  55623. 41:26:15have the 95%. You have your peak in this
  55624. 41:26:18case it's a normal distribution. So you
  55625. 41:26:19have the nice bell curve equal on both
  55626. 41:26:21sides. It's not asymmetrical. And 95% of
  55627. 41:26:24all the values lie within a very small
  55628. 41:26:26range. And then you have your outliers
  55629. 41:26:28the 2.5% going each way.
  55630. 41:26:31So we touched upon hypothesis uh and
  55631. 41:26:34we're going to move into probability. Uh
  55632. 41:26:36so you have your hypothesis. Once you've
  55633. 41:26:38generated your hypothesis, we want to
  55634. 41:26:39know the probability of something
  55635. 41:26:41occurring. Probability is a measure of
  55636. 41:26:43the likelihood of an event to occur. Any
  55637. 41:26:46event can be predicted with total
  55638. 41:26:47certainty and can only be predicted as a
  55639. 41:26:50likelihood of its occurrence. So any
  55640. 41:26:53event cannot be predicted with total
  55641. 41:26:54certainty. It can only be predicted as a
  55642. 41:26:56likelihood of its occurrence. Uh score
  55643. 41:26:59prediction. how good you're going to do
  55644. 41:27:01in whatever sport you're in, weather
  55645. 41:27:04prediction, stock prediction, if you've
  55646. 41:27:07studied physics and chaos theory, even
  55647. 41:27:09the location of the chair you're sitting
  55648. 41:27:11on has a probability that it might move
  55649. 41:27:133 ft over. Granted, that probability is
  55650. 41:27:16one in like uh I think we calculated as
  55651. 41:27:18under one in trillions upon trillions.
  55652. 41:27:21So, it's the better the probability, the
  55653. 41:27:24more likely it's going to happen. There
  55654. 41:27:25are some things that have such a low
  55655. 41:27:26probability that we don't see them. So
  55656. 41:27:29we talk about a random variable. Uh
  55657. 41:27:31random variable is a variable whose
  55658. 41:27:33possible values are numerical outcomes
  55659. 41:27:35of a random phenomena. So uh we have the
  55660. 41:27:38coin toss. How many heads will occur in
  55661. 41:27:41the series of 20 coin flips? Probably
  55662. 41:27:44you know the on average there are 10,
  55663. 41:27:45but you really can't know because it's
  55664. 41:27:47very random. How many times a red ball
  55665. 41:27:49is picked from a bag of balls if there's
  55666. 41:27:51equal number of of red balls and blue
  55667. 41:27:54balls and green balls in there. How many
  55668. 41:27:56times the sum of digits on two dice uh
  55669. 41:27:59result or five each? Um so you know
  55670. 41:28:02there's how often you're going to roll
  55671. 41:28:04two fives on your pair of dice. So in a
  55672. 41:28:07use case uh let's consider the example
  55673. 41:28:09of rolling two dice. We have a random
  55674. 41:28:11variable outcome equals y. You can take
  55675. 41:28:12values 2 3 4 5 6 7 8 9 10 11 12. So we
  55676. 41:28:17have a random variable and a combination
  55677. 41:28:19of dice and instead of looking at how
  55678. 41:28:22many times um both dice were roll five
  55679. 41:28:25let's go ahead and look at a total sum
  55680. 41:28:27of five and you have in as far as your
  55681. 41:28:29random variables you can have a one four
  55682. 41:28:31equals 5 4 1 2 3 32 so four of those
  55683. 41:28:35roles can be four if you look at all the
  55684. 41:28:38different options you have four of those
  55685. 41:28:40random rolls can be a five and if we
  55686. 41:28:43look at the total number
  55687. 41:28:46which happens to be 36 different
  55688. 41:28:48options. Uh you can see that we have
  55689. 41:28:51four out of 36 chance every time you
  55690. 41:28:53roll the dice that you're going to roll
  55691. 41:28:54a total of five. You're going to have an
  55692. 41:28:56outcome of five. And uh we'll look a
  55693. 41:28:59little deeper as to what that means. Uh
  55694. 41:29:01but you could think of that at what
  55695. 41:29:03point if someone never rolls a five or
  55696. 41:29:05they always roll a five, can you say,
  55697. 41:29:07"Hey, that person's probably cheating."
  55698. 41:29:09uh we'll look a little closer at the
  55699. 41:29:11math behind that but let's just consider
  55700. 41:29:13this as one of the cases is rolling two
  55701. 41:29:15dice and gambling. There's also a
  55702. 41:29:17binomial distribution. It is the
  55703. 41:29:19probability of getting success or
  55704. 41:29:21failure as an outcome in an experiment
  55705. 41:29:23or trial that is repeated multiple
  55706. 41:29:25times. And the key is is by meaning two
  55707. 41:29:29binomial. Uh so passing or failing an
  55708. 41:29:31exam, winning or losing a game and
  55709. 41:29:34getting either head or tails. So if you
  55710. 41:29:36ever see binomial distribution, it's
  55711. 41:29:38based on a um true false kind of setup.
  55712. 41:29:41You win or lose. Let's consider a uh use
  55713. 41:29:45case and let's consider the game of
  55714. 41:29:47football between two clubs Barcelona and
  55715. 41:29:51Dortmund. The teams will have to play a
  55716. 41:29:53total of four matches and we have to
  55717. 41:29:55find out the chances of Barcelona
  55718. 41:29:58winning the series. So we look at the
  55719. 41:30:00total games and we're looking at five
  55720. 41:30:02different games or matches. Let's say
  55721. 41:30:04that the winning chance for Barcelona is
  55722. 41:30:0675% or 75. That means at each game they
  55723. 41:30:10have a 75% chance that they're going to
  55724. 41:30:12win that game and losing chances are 25%
  55725. 41:30:15or 0.25. Clearly 75 plus 0.25 equals 1.
  55726. 41:30:20So that accounts for 100% of the game.
  55727. 41:30:22Probability for getting K wins in n
  55728. 41:30:25matches is calculated.
  55729. 41:30:27And we we're talking like so if you have
  55730. 41:30:29five games uh and you want to know if I
  55731. 41:30:31play um how many wins in those five
  55732. 41:30:34games should I get? What's a percentage
  55733. 41:30:36on those? And the probability for
  55734. 41:30:38getting k wins in n matches is
  55735. 41:30:41calculated by px= k= n k p the k q to
  55736. 41:30:47the n minus k. Here p is the probability
  55737. 41:30:50of success and q is the probability of
  55738. 41:30:53failure. And so we can do total games of
  55739. 41:30:55n equals 5 where k equals 012345.
  55740. 41:31:00P which is the chance of winning is 75.
  55741. 41:31:04Q the chance of losing equals 1 minus p
  55742. 41:31:07which equals 1 - 0075 which equals 0.25.
  55743. 41:31:11The probability that Barcelona will lose
  55744. 41:31:13all of the matches can then just plug in
  55745. 41:31:15the numbers and we end up with a
  55746. 41:31:1809765625.
  55747. 41:31:23So very small chance they're going to
  55748. 41:31:25lose all their matches.
  55749. 41:31:27And we can plug in uh the value for two
  55750. 41:31:29matches. Probability that Barcelona will
  55751. 41:31:32win at least two matches is 00878. And
  55752. 41:31:36of course we can go on to probability
  55753. 41:31:38that Barcelona will win three matches
  55754. 41:31:40the 26 and of course four matches and so
  55755. 41:31:44on. And it's always nice to take this
  55756. 41:31:46information um and let let's find the
  55757. 41:31:48cumulative discrete probabilities for
  55758. 41:31:50each of the outcomes where Barcelona has
  55759. 41:31:53won three or more matches x= 3 x= 4 x= 5
  55760. 41:31:58and we end up with the p =264 plus 395 +
  55761. 41:32:02237 which equals89.
  55762. 41:32:05In reality the probability of Barcelona
  55763. 41:32:08winning the series is much higher than
  55764. 41:32:1075. And it's always nice to uh put out a
  55765. 41:32:15nice graph so you can actually see the
  55766. 41:32:17number of wins to the probability and
  55767. 41:32:19how that pans out with our binomial
  55768. 41:32:22case. Continuing in our important
  55769. 41:32:24terminology, location, the location of
  55770. 41:32:27the center of the graph depends on the
  55771. 41:32:29mean value. And uh this is some very
  55772. 41:32:32important things. So much of the data we
  55773. 41:32:35look at and when you start looking at
  55774. 41:32:36probabilities almost always has a
  55775. 41:32:38normalized look like the graph in the
  55776. 41:32:40middle.
  55777. 41:32:42uh but you do have left skewed where the
  55778. 41:32:44data is skewed off to the left and you
  55779. 41:32:45have more stuff happening off to the
  55780. 41:32:47left and you have right skewed data and
  55781. 41:32:49so when this comes up and these
  55782. 41:32:50probabilities come up where they're
  55783. 41:32:52skewed it's really important to take a
  55784. 41:32:54closer look at that uh mostly you end up
  55785. 41:32:56with a normalized set of data but you
  55786. 41:32:58got to also be aware that sometimes it's
  55787. 41:33:00a skewed data and then the height height
  55788. 41:33:03of the slope inversely depends upon the
  55789. 41:33:05standard deviation
  55790. 41:33:07so you can see down here the standard
  55791. 41:33:09deviation is really large it kind of
  55792. 41:33:10squishes it out. And if the standard
  55793. 41:33:12deviation is small, then most of your
  55794. 41:33:15data is going to hit right there in the
  55795. 41:33:16middle. You're going to have a nice
  55796. 41:33:17peak. Um, and so being aware of this
  55797. 41:33:19that you might have a probability that
  55798. 41:33:22fits certain data, but it has a lot of
  55799. 41:33:24outliers. So you're if you have a really
  55800. 41:33:26high standard deviation, um, if you're
  55801. 41:33:28doing stock market analysis,
  55802. 41:33:31this means your predictions are probably
  55803. 41:33:33not going to make you much money. uh
  55804. 41:33:35where if you have a very small
  55805. 41:33:36deviation, you might be right on target
  55806. 41:33:38and set to become a millionaire. Which
  55807. 41:33:40leads us to the zcore. Zcore tells you
  55808. 41:33:43how far from the mean a data point is.
  55809. 41:33:46It is measured in terms of standard
  55810. 41:33:48deviations from the mean. Around 68% of
  55811. 41:33:51the results are found between one
  55812. 41:33:53standard deviation. Around 95% of the
  55813. 41:33:56results are found between two standard
  55814. 41:33:58deviations.
  55815. 41:33:59And you read the symbols. Of course,
  55816. 41:34:01they love to throw some Greek letters in
  55817. 41:34:02there. we have mu minus 2 sigma. Mu is
  55818. 41:34:07just a quick way. It's that kind of
  55819. 41:34:09funky u. It just means the mean. Uh and
  55820. 41:34:12then the sigma is the standard
  55821. 41:34:14deviation. And that's the o with a
  55822. 41:34:16little arrow off to the right or the
  55823. 41:34:18little waggly tail going up. The o with
  55824. 41:34:20a with a line on it. Uh so mu minus 2
  55825. 41:34:23sigma is your uh 95% of the results are
  55826. 41:34:27found between two standard deviations.
  55827. 41:34:30The central limit theorem. This goes
  55828. 41:34:33back to the skew. If you remember, we
  55829. 41:34:35were looking at the skew values on this
  55830. 41:34:37previous slide. Have left skewed,
  55831. 41:34:40normalized, and right skewed. When we're
  55832. 41:34:42talking about it being skewed or not
  55833. 41:34:44skewed, the distribution of the sample
  55834. 41:34:46means will be approximately normally
  55835. 41:34:49distributed, evenly distributed, not
  55836. 41:34:51skewed. If you take large random samples
  55837. 41:34:54from the population with the mean mu and
  55838. 41:34:57the standard deviation sigma with
  55839. 41:35:00replacement
  55840. 41:35:02and you can see here um uh of course we
  55841. 41:35:04have our uh mu minus 2 sigma and the
  55842. 41:35:07spread down here the mean the median and
  55843. 41:35:09the mode and so when you're talking
  55844. 41:35:11about very large populations
  55845. 41:35:14these numbers should come together and
  55846. 41:35:15you shouldn't have a skewed value. If
  55847. 41:35:17you do that's a flag that something's
  55848. 41:35:19wrong. That's why this is so important
  55849. 41:35:22to be aware of what's going on with your
  55850. 41:35:24data, where your samples are coming
  55851. 41:35:26from, and the math behind it. And if
  55852. 41:35:29you're going to do all this, we got to
  55853. 41:35:30jump into conditional probability. The
  55854. 41:35:34conditional probability of an event A is
  55855. 41:35:36a probability that the event will occur
  55856. 41:35:39given the knowledge that an event B has
  55857. 41:35:41already occurred. And you'll see this as
  55858. 41:35:44Baze theorem. B A Y S bay. Uh, and this
  55859. 41:35:48is read. I mean, you have these funky
  55860. 41:35:50looking little P brackets. A B. This is
  55861. 41:35:54the probability of A being true while B
  55862. 41:35:58is already true. And you have the
  55863. 41:36:00probability of B being true when A is
  55864. 41:36:02already true. So, P B of A probability
  55865. 41:36:06of A being true divided by the
  55866. 41:36:08probability of B being true. And we talk
  55867. 41:36:11about BA's theorem which occurred back
  55868. 41:36:13in the 1800s when he discovered this.
  55869. 41:36:16This is such an important formula and
  55870. 41:36:18it's really it's not if you actually do
  55871. 41:36:20the math you could just kind of do um um
  55872. 41:36:23XY equals J K and then you divide them
  55873. 41:36:27out and you're going to see the same
  55874. 41:36:28math but it works with probabilities
  55875. 41:36:30which makes it really nice. And so if
  55876. 41:36:33you have a s you might have uh eight or
  55877. 41:36:35nine different studies going on in
  55878. 41:36:38different areas different people have
  55879. 41:36:39done the studies they brought them
  55880. 41:36:41together. Um if we look at today's co
  55881. 41:36:44virus the virus spread uh certainly the
  55882. 41:36:47studies done in China versus the studies
  55883. 41:36:50the way they're done in the US that data
  55884. 41:36:52is different in each of those studies
  55885. 41:36:54but if you can find a place where it
  55886. 41:36:56overlaps where they're studying the same
  55887. 41:36:58thing together you can then compute the
  55888. 41:37:01changes that you need to make in one
  55889. 41:37:02study to make them equal and this is
  55890. 41:37:05also true if you have a study of uh um
  55891. 41:37:09one group and you want to find out more
  55892. 41:37:10about it. So this formula is very
  55893. 41:37:13powerful. Uh it really has to do with
  55894. 41:37:15the data collection part of the math and
  55895. 41:37:17data science and understanding where
  55896. 41:37:19your data is coming from and how you're
  55897. 41:37:21going to combine different studies in
  55898. 41:37:23different groups. And we'll go ahead and
  55899. 41:37:25go into a use case. Uh let's find out
  55900. 41:37:27the chance of a person getting lung
  55901. 41:37:29disease due to smoking. Uh and this is
  55902. 41:37:32kind of interesting the way they word
  55903. 41:37:33this. Um let's say that according to
  55904. 41:37:35medical report provided by the hospital
  55905. 41:37:38states that around 10% of all patients
  55906. 41:37:41they treated suffered lung lung disease.
  55907. 41:37:44Uh so we have kind of a generic medical
  55908. 41:37:46report. They further found out uh by a
  55909. 41:37:50survey that 15% of the patients that
  55910. 41:37:52visit them smoke. So we have 10% that
  55911. 41:37:55are lung disease and um 15% of the
  55912. 41:37:58patients smoke. And finally, 5% of the
  55913. 41:38:02people continued smoke even when they
  55914. 41:38:04had lung disease. Uh not the brightest
  55915. 41:38:07choice um but you know it is an
  55916. 41:38:09addiction so it can be really difficult
  55917. 41:38:10to kick. And so we can look at the
  55918. 41:38:13probability of a uh prior probability of
  55919. 41:38:1510% people having lung disease. And then
  55920. 41:38:18probability b probability that a patient
  55921. 41:38:21smokes is 15%.
  55922. 41:38:24Uh and the probability of B um if B then
  55923. 41:38:28A. The probability of a patient smokes
  55924. 41:38:30even though they have lung disease is
  55925. 41:38:335%. And probability of A is B.
  55926. 41:38:36Probability that the patient will have
  55927. 41:38:38lung disease if they smoke. And then
  55928. 41:38:40when you put the formulas together, uh
  55929. 41:38:42you get a nice solution here. You get
  55930. 41:38:43the probability of A of B, probability
  55931. 41:38:45that the patient will have lung disease
  55932. 41:38:47if they smoke. And you can just plug the
  55933. 41:38:49numbers right in and we get a 3.33%
  55934. 41:38:53chance. Hence, there is a 3.33% chance
  55935. 41:38:56that a person who smokes will get a lung
  55936. 41:38:58disease. So, we're going to pull up a
  55937. 41:39:01little Python code, always my favorite,
  55938. 41:39:03roll up the sleeves. Keep in mind, we're
  55939. 41:39:06going to be doing this um kind of like
  55940. 41:39:08the backend way so that you can see
  55941. 41:39:11what's going on. And then later on we're
  55942. 41:39:14going to create um we'll get into
  55943. 41:39:17another demo which shows you some of the
  55944. 41:39:19tools that are already pre-built for
  55945. 41:39:20this. Let's start by creating a set. So
  55946. 41:39:25we're going to create a set with curly
  55947. 41:39:26braces. This means that our set has um
  55948. 41:39:30only unique values. So you have a list
  55949. 41:39:34uh you have your tupils which can never
  55950. 41:39:36change and then you have um in this case
  55951. 41:39:39the the set. So 47, you can't create a
  55952. 41:39:4347, 4. It'll delete the four out. So
  55953. 41:39:46it's only unique values. And if you use
  55954. 41:39:49dictionaries,
  55955. 41:39:51quick reminder, this should look
  55956. 41:39:53familiar because it is a dictionary uh
  55957. 41:39:56where you have a value and that value is
  55958. 41:39:58assigned to or that key is assigned to a
  55959. 41:40:01value. Uh so you could have a key value
  55960. 41:40:03set up as a dictionary. So it's like a
  55961. 41:40:05dictionary without the value. It's just
  55962. 41:40:07the keys and they all have to be unique.
  55963. 41:40:11And if we run this, we have a set of 47.
  55964. 41:40:17We can also take a list, a regular um
  55965. 41:40:20setup. And I'm going to go ahead and
  55966. 41:40:21just throw in another number in here,
  55967. 41:40:22four, and run it. Uh, and you can see
  55968. 41:40:25here if I take my list 1 2 3 4, and I
  55969. 41:40:29convert it to a set, and here it is. My
  55970. 41:40:32set from list equals set my list.
  55971. 41:40:36The result is 1 2 3 4. So, it just
  55972. 41:40:38deletes that last four right out of
  55973. 41:40:40there.
  55974. 41:40:42And with the sets, you can also go in
  55975. 41:40:44there and um print here is my set. My
  55976. 41:40:48set uh three is in the set. And then if
  55977. 41:40:51you do three in my set,
  55978. 41:40:54that's going to be a logic function. Uh
  55979. 41:40:57and one in my set, six is not in the
  55980. 41:41:00set, and so forth. If we run this,
  55981. 41:41:04we get three is in the set true one is
  55982. 41:41:06in the set false because 357 is another
  55983. 41:41:09one. Six is in the set uh six is not in
  55984. 41:41:12the set. So not in my set. You can also
  55985. 41:41:16use this with a list. We could have just
  55986. 41:41:18used 357 and it would have um the same
  55987. 41:41:22response on there is three and usually
  55988. 41:41:25you do if three is in but three in my
  55989. 41:41:27set is still works on a just a regular
  55990. 41:41:29list. And we'll go ahead and do a little
  55991. 41:41:32iteration. We're going to do kind of the
  55992. 41:41:34dice one. Remember um uh 1 2 3 4 5 6.
  55993. 41:41:38And so we're going to bring in an
  55994. 41:41:39iteration tool and import product as
  55995. 41:41:42product.
  55996. 41:41:44And uh I'll show you what that means in
  55997. 41:41:46just a second. So we have our two dice.
  55998. 41:41:48We have dice A and it's going to be a
  55999. 41:41:51set of values. Um they can only have one
  56000. 41:41:53value for each one. That's why they put
  56001. 41:41:55it in a set. And if you remember from
  56002. 41:41:57range, it is up to seven. So this is
  56003. 41:42:00going to be 1 2 3 4 5 6. It will not
  56004. 41:42:03include the seven. And the same thing
  56005. 41:42:05for our dice B.
  56006. 41:42:08And then we're going to do is we're
  56007. 41:42:09going to create a list which is the
  56008. 41:42:12product of A and B. So what's um a + b?
  56009. 41:42:17And if we go ahead and run this uh it'll
  56010. 41:42:19print that out. And you'll see um in
  56011. 41:42:22this case when they say product because
  56012. 41:42:23it's an iteration tool,
  56013. 41:42:27we're talking about creating a tupole of
  56014. 41:42:29the two. So we've now created a tupole
  56015. 41:42:31of all possible outcomes of the dice
  56016. 41:42:34where dice A is one to three one to six
  56017. 41:42:37and dice B is 1 to six. And you can see
  56018. 41:42:38one to one, one to two, one to three and
  56019. 41:42:40so forth. You remember we had a slide on
  56020. 41:42:42this earlier where we talked about um
  56021. 41:42:45the different all the different outcomes
  56022. 41:42:47of a dice. We can play around with this
  56023. 41:42:50a little bit. Uh we can do in dice
  56024. 41:42:52equals two divi dice faces 1 2 3 4 5 6.
  56025. 41:42:57Uh another way of doing what we did
  56026. 41:42:59before and then we can create an event
  56027. 41:43:00space where we have a set which is the
  56028. 41:43:03product of the dice faces repeat equals
  56029. 41:43:06end dice. And we'll go ahead and just
  56030. 41:43:07run this. And you can see here it just
  56031. 41:43:10again puts it through all the different
  56032. 41:43:12possible variables we can have. And then
  56033. 41:43:15if we wanted to take the same uh set on
  56034. 41:43:17here and print them all out like we had
  56035. 41:43:20before uh we can just go through for
  56036. 41:43:22outcome and event space. Outcome end
  56037. 41:43:25equals. So the event space is creating
  56038. 41:43:30a sequence and as you can see here when
  56039. 41:43:32we print it out it stacks them versus
  56040. 41:43:34going through and putting them in a nice
  56041. 41:43:36line.
  56042. 41:43:38and we'll go ahead and do something. Um,
  56043. 41:43:40let's go print. Since we have the end
  56044. 41:43:43printing with a comma, that just means
  56045. 41:43:45it's just going to it's not going to hit
  56046. 41:43:47the return going down to the next line.
  56047. 41:43:49Uh, and we'll go ahead and do the length
  56048. 41:43:54of our event space. Uh, that'll be an
  56049. 41:43:57important variable we're going to want
  56050. 41:43:58to know in a minute.
  56051. 41:44:01And of course, if I get carried away
  56052. 41:44:02with my typing of length, uh, we'll
  56053. 41:44:04print it twice and it'll give me an
  56054. 41:44:06error. Uh so we have 36 different
  56055. 41:44:09possible variations here
  56056. 41:44:12and we might want to calculate something
  56057. 41:44:14like um what about the multiple of
  56058. 41:44:16three? What if we want to have
  56059. 41:44:19uh the probability of the multiple of
  56060. 41:44:21three in our setup?
  56061. 41:44:25And so uh we can put together the code
  56062. 41:44:27for the outcome in event space of xy
  56063. 41:44:30equals outcome if x + y
  56064. 41:44:34remainder 3. So, we're going to divide
  56065. 41:44:36by three and look at the remainder and
  56066. 41:44:37it equals zero.
  56067. 41:44:40Then it's a favorable outcome and we're
  56068. 41:44:42going to pop that outcome on the end
  56069. 41:44:43there.
  56070. 41:44:45And we'll turn it into a set. So, the
  56071. 41:44:47favor outcome equals a set. Not
  56072. 41:44:50necessary uh because we know it's not
  56073. 41:44:52going to be repeating itself, but just
  56074. 41:44:54in case, we'll go ahead and do that.
  56075. 41:44:58And if we want to print out the outcome,
  56076. 41:45:01we can go ahead and see what that looks
  56077. 41:45:03like. And you can see here these are all
  56078. 41:45:05uh multiples of three. Uh 1 plus 2 is 3,
  56079. 41:45:085 + 4 is 9, which divided by 3 is 3, and
  56080. 41:45:11so forth.
  56081. 41:45:15And just like we looked up the length uh
  56082. 41:45:17of the one before, let's go ahead and
  56083. 41:45:18print the length of our f outcome so we
  56084. 41:45:24can see what that looks like.
  56085. 41:45:30There we go.
  56086. 41:45:32And of course, I did forget to add the
  56087. 41:45:34print in the middle because we're
  56088. 41:45:35looping through and putting an end on
  56089. 41:45:37the on the setup on there. So, we're
  56090. 41:45:38going to put the print in there. And if
  56091. 41:45:40I run this, you can see um
  56092. 41:45:46we end up with 12. So, we have 36 total
  56093. 41:45:49options. Uh we have 12 that are multiple
  56094. 41:45:53that um add up to a multiple of three.
  56095. 41:45:57And we can easily conver compute the
  56096. 41:45:59probability of this uh by simply taking
  56097. 41:46:02the length of our favorable outcome over
  56098. 41:46:05the length of the event space.
  56099. 41:46:09And if we print it out, let me put that
  56100. 41:46:10in there. Probability
  56101. 41:46:13last line. So we just type it in. We end
  56102. 41:46:15up with a 3333 chance. And it's roughly
  56103. 41:46:19a third.
  56104. 41:46:21And we might want to make this look
  56105. 41:46:23nice. So let's go ahead and put in
  56106. 41:46:24another line there. The probability of
  56107. 41:46:26getting the sum which is a multiple of
  56108. 41:46:27three is
  56109. 41:46:303333.
  56110. 41:46:34We can compute the same thing for five
  56111. 41:46:36dice.
  56112. 41:46:38And if we do this for five dice and go
  56113. 41:46:40ahead and run it, you can see we just
  56114. 41:46:42have a huge amount of choices. So it
  56115. 41:46:45just goes on and on down here. And we
  56116. 41:46:47can look at the uh length of the event
  56117. 41:46:51space.
  56118. 41:46:59And we have over 7,776
  56119. 41:47:02choices. That's a lot of choices.
  56120. 41:47:05And if we want to ask the question like
  56121. 41:47:07we did above, uh what is the sum where
  56122. 41:47:10the sum is a multiple of five but not a
  56123. 41:47:12multiple of three? We can go through all
  56124. 41:47:15of these different options. And then uh
  56125. 41:47:18you can see here uh d1 d2 d3 d4 d5
  56126. 41:47:22equals the outcome. And if uh you add
  56127. 41:47:24these all together and the
  56128. 41:47:27division by five does not have a
  56129. 41:47:29remainder of zero but the remainder is
  56130. 41:47:32also of a division by three is not equal
  56131. 41:47:35to zero. So the multiple of five is
  56132. 41:47:38equal to zero but the multiple of three
  56133. 41:47:39is not. We can just appin that on here
  56134. 41:47:42and then we can look at that uh
  56135. 41:47:44favorable outcome. We'll go ahead and
  56136. 41:47:47set that and we'll just take a look at
  56137. 41:47:48this. What's our length of our favorable
  56138. 41:47:52outcome?
  56139. 41:47:58It's always good to see what we're
  56140. 41:47:59working with. And so we have 94 out of
  56141. 41:48:02776.
  56142. 41:48:06And then of course we can just do a
  56143. 41:48:08simple division to get the probability
  56144. 41:48:10on here. What's the probability that
  56145. 41:48:11we're going to roll a multiple of five
  56146. 41:48:14when you add them together?
  56147. 41:48:16but not a multiple of three. And so
  56148. 41:48:19we're just going to divide those two
  56149. 41:48:20numbers. And you can see here we get
  56150. 41:48:22uh.16255
  56151. 41:48:24or 11.62%.
  56152. 41:48:30And so you can really have a nice visual
  56153. 41:48:32that this is not really complicated math
  56154. 41:48:35right here on probabilities. uh it's
  56155. 41:48:37just how many options do you have and
  56156. 41:48:39how many of those are you possibly going
  56157. 41:48:41to be able to um come up with with the
  56158. 41:48:44solution you're looking for. And this
  56159. 41:48:46leads us to a confusion matrix. A
  56160. 41:48:49confusion matrix is a table which is
  56161. 41:48:51used to describe the performance of a
  56162. 41:48:53classification model on a set of test
  56163. 41:48:55data for which the true values are
  56164. 41:48:57known. And so you'll see on the left we
  56165. 41:49:00have the predicted and the actual and we
  56166. 41:49:03have a negative uh false negative
  56167. 41:49:06positive true positive
  56168. 41:49:08um and then we have false positive and
  56169. 41:49:11true negative. And you can think of this
  56170. 41:49:14as your predicted model. What does that
  56171. 41:49:17mean? That means if you divided your
  56172. 41:49:19data and you use twothird of it to
  56173. 41:49:21create the model, you might then test it
  56174. 41:49:24against an actual case for the last
  56175. 41:49:25third to see how well it comes out. How
  56176. 41:49:27many times was it uh true positive
  56177. 41:49:30versus uh false positive? It gave a
  56178. 41:49:33false positive response. And you can
  56179. 41:49:35imagine in medical uh situations, this
  56180. 41:49:38is a pretty big deal. You don't want to
  56181. 41:49:40give a false positive. So you might
  56182. 41:49:42adjust your model accordingly so you
  56183. 41:49:44don't have a false positive. Say with a
  56184. 41:49:46co virus test, it'd be better to have a
  56185. 41:49:48false negative and then go back and get
  56186. 41:49:50retested than to have 30% false
  56187. 41:49:53positives where then the test is pretty
  56188. 41:49:55much invalid. So in a use case uh like
  56189. 41:49:58cancer prediction, let's consider an
  56190. 41:50:00example where a cancer prediction model
  56191. 41:50:02is put to the test for its accuracy and
  56192. 41:50:04precision. Actual result of a person's
  56193. 41:50:07medical report is compared with the
  56194. 41:50:09prediction made by the machine learning
  56195. 41:50:11model. And so you can see here here's
  56196. 41:50:13our actual predicted uh whether they
  56197. 41:50:15have cancer or not. You know cancer a
  56198. 41:50:17big one. You don't want to have a uh
  56199. 41:50:19false positive. I mean a false negative.
  56200. 41:50:22In other words, you don't want to have
  56201. 41:50:23it tell you that you don't have cancer
  56202. 41:50:25when you do. So that would be something
  56203. 41:50:27you'd really be looking for in this
  56204. 41:50:29particular domain. You don't want a
  56205. 41:50:31false negative. Uh and this is again,
  56206. 41:50:34you know, you've created a model, you
  56207. 41:50:35have hundreds of people or thousands of
  56208. 41:50:38pieces of data that come in. There's a
  56209. 41:50:40real famous case study where they have
  56210. 41:50:42the imagery and all the measurements
  56211. 41:50:43they take and there's about 36 different
  56212. 41:50:45measurements they take. And then if you
  56213. 41:50:48run the a basic model, you want to know
  56214. 41:50:50just how accurate it is. How many um
  56215. 41:50:52negative results do you have that are
  56216. 41:50:54either telling people they have cancer
  56217. 41:50:56that don't or telling people that don't
  56218. 41:50:57have cancer that they do? And then we
  56219. 41:50:59can take these numbers and we can feed
  56220. 41:51:02them into our accuracy, our precision,
  56221. 41:51:04and our recall. Uh so accuracy,
  56222. 41:51:07precision, and recall, accuracy metric
  56223. 41:51:09to measure how accurately the results
  56224. 41:51:11are predicted. And this is your um total
  56225. 41:51:15um true where you got the right results.
  56226. 41:51:17you add them together, the true
  56227. 41:51:18positive, the true negative over all the
  56228. 41:51:21results. So what percentage of them were
  56229. 41:51:23accurate versus what were wrong. We talk
  56230. 41:51:26about precision is a metric to measure
  56231. 41:51:28how many of the correctly predicted
  56232. 41:51:30cases are actually turned out to be
  56233. 41:51:32positive. Uh so we have a precision on
  56234. 41:51:36true positive. Again, if you're talking
  56235. 41:51:38about like uh COVID testing with the
  56236. 41:51:42viruses, uh you really want this to be a
  56237. 41:51:44a high number. you want this true um
  56238. 41:51:47that to be the center point where you
  56239. 41:51:49might have the opposite if you're
  56240. 41:51:51dealing with cancer where you want no
  56241. 41:51:53false negatives. Uh so this is your
  56242. 41:51:56metric on here. Precision is your test
  56243. 41:51:58positive uh true positive plus uh false
  56244. 41:52:02positive. And then your recall how many
  56245. 41:52:05of the actual positive cases we were
  56246. 41:52:07able to predict quickly with our model.
  56247. 41:52:09Uh so test positive is the test positive
  56248. 41:52:12plus the false negative on there. And
  56249. 41:52:15we'll want to go ahead and do a demo on
  56250. 41:52:17the naive bay classifier. Before I get
  56251. 41:52:21too far into uh naive baze classifier
  56252. 41:52:24because we're going to pull it from the
  56253. 41:52:25sklearn or the scikit. Um let's go ahead
  56254. 41:52:30kind of an interesting page here for
  56255. 41:52:31classifiers. When you go into the
  56256. 41:52:33sklearn kit, there's a lot of ways to do
  56257. 41:52:35classification. I'll just zoom up in
  56258. 41:52:37here so you can see some of the titles.
  56259. 41:52:40Uh there's everything from the nearest
  56260. 41:52:42neighbor linear
  56261. 41:52:44uh but we're going to be focusing on the
  56262. 41:52:46naive bays over here. And this is just
  56263. 41:52:50um a sample data set that they put
  56264. 41:52:52together. And you can see how some of
  56265. 41:52:54these have a very different output. The
  56266. 41:52:57naive bay remember is set up as probably
  56267. 41:53:00the most simplified uh calculator or um
  56268. 41:53:03set of predictions out there. And so
  56269. 41:53:05what we've been talking about with the
  56270. 41:53:07true false and stuff like that where
  56271. 41:53:08there's a uh
  56272. 41:53:11an belief that there is a independent
  56273. 41:53:13assumption between the features where
  56274. 41:53:14the features are very assumed to have
  56275. 41:53:17some kind of connection uh then we can
  56276. 41:53:19go ahead and use that for the
  56277. 41:53:21prediction. And so that's what we're
  56278. 41:53:23using as a naive bay classifier versus
  56279. 41:53:26many of the other classifiers that are
  56280. 41:53:27out there.
  56281. 41:53:30For this we're going to use uh the
  56282. 41:53:32social network ads. It's a little data
  56283. 41:53:35set on here and let me go and just open
  56284. 41:53:38that up the file. Uh here we go. It has
  56285. 41:53:42user ID, gender, age, estimated salary,
  56286. 41:53:45uh purchased. And so we have you can see
  56287. 41:53:48the user ID, male 19, uh estimated
  56288. 41:53:52salary 19,000 and purchased zero. Uh so
  56289. 41:53:56it's either going to make a purchase or
  56290. 41:53:57not. So look at that last one. 01. We
  56291. 41:54:01should be thinking of binomials. we
  56292. 41:54:03should be thinking of simple naive base
  56293. 41:54:05classifier kind of setup.
  56294. 41:54:09So if we close this out, we're going to
  56295. 41:54:10go ahead and import our numpy as np.
  56296. 41:54:14We're nice to have a a good visual of
  56297. 41:54:16our data. So we'll put in our mattplot
  56298. 41:54:18library. Here's our pandas, our data
  56299. 41:54:21frame.
  56300. 41:54:23Uh and then we're going to go ahead and
  56301. 41:54:24import the data set. And the data set's
  56302. 41:54:26going to be we're going to read it from
  56303. 41:54:28the social network ads.csv. Then we're
  56304. 41:54:31going to print the head just so you can
  56305. 41:54:32see it again uh even though I showed you
  56306. 41:54:34it in the file. And X equals the data
  56307. 41:54:37set I location uh two three values and Y
  56308. 41:54:40is going to be the four uh column 4. Let
  56309. 41:54:43me just run this so it's a little easier
  56310. 41:54:45to go over that. Um you can see right
  56311. 41:54:47here we're going to be looking at uh 012
  56312. 41:54:50is age and estimated salary. So 2 three
  56313. 41:54:54and that's what I location just means um
  56314. 41:54:57that we're looking at the number versus
  56315. 41:55:00a regular location. Uh regular location
  56316. 41:55:02you'd actually say age and estimated
  56317. 41:55:04salary.
  56318. 41:55:06And then column four is did they make a
  56319. 41:55:08purchase? They purchased something. Uh
  56320. 41:55:10so those are the three columns we're
  56321. 41:55:12going to be looking at when we do this.
  56322. 41:55:13And we've gone ahead and imported these
  56323. 41:55:15and imported the data. So now our data
  56324. 41:55:17set is all set with this information in
  56325. 41:55:19it.
  56326. 41:55:23And we'll need to go ahead and split the
  56327. 41:55:24data up. Uh so we need our from the
  56328. 41:55:26sklearn model selection we can import
  56329. 41:55:29train test split. Uh this does a nice
  56330. 41:55:32job. We can set the random state so it
  56331. 41:55:34randomly picks the data. And we're just
  56332. 41:55:36going to take uh 25% of it is going to
  56333. 41:55:39go into the test our x test and our y
  56334. 41:55:41test and the 75% will go to x train and
  56335. 41:55:44y train. That way once we create our
  56336. 41:55:48model, we can then have data to see just
  56337. 41:55:50how accurate or how well it has
  56338. 41:55:52performed with our um prediction.
  56339. 41:55:56The next step in pre-processing our data
  56340. 41:55:59is to go ahead and do feature scaling.
  56341. 41:56:02Now, a lot of this is start to look
  56342. 41:56:04familiar. If you've done a number of the
  56343. 41:56:06other modules and setup, you should
  56344. 41:56:08start noticing that we bring in our
  56345. 41:56:10data. We take a look at what we're
  56346. 41:56:12working with. uh we go ahead and split
  56347. 41:56:14it up into training and testing. Uh in
  56348. 41:56:17this case, we're going to go ahead and
  56349. 41:56:18scale it. Scale it means we're putting
  56350. 41:56:20it between a value of minus1 and one uh
  56351. 41:56:24or someplace in that middle ground
  56352. 41:56:26there. This way, if you have any huge
  56353. 41:56:28set, you don't have this huge um setup.
  56354. 41:56:31If we go back up to here where salary uh
  56355. 41:56:33salary is 20,000 versus age 35, well,
  56356. 41:56:39there's a good chance with a lot of the
  56357. 41:56:40back-end math that 20,000 will skew the
  56358. 41:56:43results and the estimated salary will
  56359. 41:56:45have a higher impact than the age
  56360. 41:56:47instead of balancing them out and
  56361. 41:56:48letting the calculations weigh them
  56362. 41:56:50properly.
  56363. 41:56:52And finally, we get to actually create
  56364. 41:56:55our naive bay model.
  56365. 41:56:58Um, and then we're going to go ahead and
  56366. 41:57:00import the Gazian naive bays.
  56367. 41:57:04And the Gazian is is uh the most basic
  56368. 41:57:07one. That's what we're looking at now.
  56369. 41:57:09It turns out though, if you go to the SK
  56370. 41:57:12um learn kit, uh they have a number of
  56371. 41:57:15different ones you can pull in there.
  56372. 41:57:17There's a um Bernoli. I I've never used
  56373. 41:57:20that one. Categorical
  56374. 41:57:22um compliment. And here's our Gazian. Uh
  56375. 41:57:25so there's a number of different options
  56376. 41:57:27you can look at. Gazian when you come to
  56377. 41:57:30the naive bays is the most commonly
  56378. 41:57:32used. Uh so we're talking about the
  56379. 41:57:34naive bays that's usually what people
  56380. 41:57:36are talking about when they when they're
  56381. 41:57:37pulling this in. And one of the nice
  56382. 41:57:39things about the gazian if you go to
  56383. 41:57:41their website um to sklearn the naive
  56384. 41:57:44bay gazian there's a lot of cool
  56385. 41:57:46features. One of them is you can do
  56386. 41:57:47partial fit on here. Um that means if
  56387. 41:57:50you have a huge amount of data, you
  56388. 41:57:51don't have to process it all at on you
  56389. 41:57:53once. You can batch it into the Gausian
  56390. 41:57:57uh NB model. And there's many other
  56391. 41:57:59different things you can do with it as
  56392. 41:58:01far as fitting the data and how you um
  56393. 41:58:03manipulate it. We're just doing the
  56394. 41:58:06basics. So we're going to go ahead and
  56395. 41:58:07create our classifier. We're going to
  56396. 41:58:09equal the Gausian NB.
  56397. 41:58:12And then we're going to do a fit. We're
  56398. 41:58:13going to fit our training data and our
  56399. 41:58:15training solution. So, X-Rain, Y train,
  56400. 41:58:20and we'll go ahead and run this. Uh,
  56401. 41:58:22it's going to tell us that it it ran the
  56402. 41:58:24code right there.
  56403. 41:58:26And now we have our trained classifier
  56404. 41:58:29model. So, the next step is we need to
  56405. 41:58:32go ahead and run a prediction. We're
  56406. 41:58:33going to do our Y predict equals the
  56407. 41:58:35classifier.predict
  56408. 41:58:37X test. So, here we fit the data and now
  56409. 41:58:40we're going to go ahead and predict.
  56410. 41:58:45And now we get to our confusion matrix.
  56411. 41:58:49Uh so from the sklearn matrix metrics
  56412. 41:58:52you can import your confusion matrix
  56413. 41:58:54just as saves you from doing all the
  56414. 41:58:56simple math. It does it all for you. And
  56415. 41:58:59then we'll go ahead and create our
  56416. 41:59:00confusion metrics with the y test and
  56417. 41:59:02the y predict. So we have our actual and
  56418. 41:59:05we have our predicted value.
  56419. 41:59:08And you can see from here this is the
  56420. 41:59:10chart we looked at. Here's predicted.
  56421. 41:59:11So, true positive, false positive, false
  56422. 41:59:14negative, true negative.
  56423. 41:59:18And if we go ahead and run this, there
  56424. 41:59:20we have it. 653725.
  56425. 41:59:24And in this particular uh prediction, we
  56426. 41:59:27had 65 uh or predicted the truth as far
  56427. 41:59:30as a a purchase. They're going to make a
  56428. 41:59:32purchase, and we guessed three wrong.
  56429. 41:59:35And then we had 25 we predicted would
  56430. 41:59:37not purchase, and seven of them did. So,
  56431. 41:59:40there's our our confusion matrix.
  56432. 41:59:44At this point, if you were uh with your
  56433. 41:59:46shareholders or a board meeting, um you
  56434. 41:59:49would start to hear some snoozing if
  56435. 41:59:50they were looking at the numbers and you
  56436. 41:59:52say, "Hey, here's my confusion mat uh
  56437. 41:59:54matrix." So, let's go ahead and
  56438. 41:59:56visualize the results.
  56439. 41:59:58We're going to pull from the map plot
  56440. 42:00:00library colors import listed color map.
  56441. 42:00:04Um, and this is actually my machine's
  56442. 42:00:06going to throw an error because this is
  56443. 42:00:09being um because of the way the setup
  56444. 42:00:12is. I have a newer version on here than
  56445. 42:00:14when they put together the demo. And we
  56446. 42:00:17need our um X set and our Y set, which
  56447. 42:00:19is our X train and Y train. And then
  56448. 42:00:22we'll create our X1, X2. And we'll put
  56449. 42:00:25that into a grid. Uh, and we set our X
  56450. 42:00:28set minimum stop and our X set max stop.
  56451. 42:00:31And if you come all the way over here,
  56452. 42:00:33we're going to step 0. 001. This is
  56453. 42:00:35going to give us a nice line, uh, is
  56454. 42:00:37what that's doing. And then we're going
  56455. 42:00:38to plot the contour, uh, plot the x
  56456. 42:00:41limit, plot the y limit, and put the
  56457. 42:00:44scatter plot in there. And let's go
  56458. 42:00:46ahead and run this. Uh, to be honest,
  56459. 42:00:48when I'm doing these graphs, there's so
  56460. 42:00:50many different ways to do that. There's
  56461. 42:00:52so many different ways to put this code
  56462. 42:00:53together to show you what we're doing.
  56463. 42:00:56it's uh a lot easier to pull up the
  56464. 42:00:59graph and then go back up and explain
  56465. 42:01:00it. So the first thing we want to note
  56466. 42:01:04here when we're looking at the data
  56467. 42:01:07is this is the training set.
  56468. 42:01:10And so we have those who didn't make a
  56469. 42:01:12purchase. We've drawn a nice area for
  56470. 42:01:14that that's defined by the naive bay
  56471. 42:01:17setup. And then we have those who did
  56472. 42:01:19make a purchase, the green. And you can
  56473. 42:01:21see that some of the green dots fall
  56474. 42:01:23into the red area and some of the red
  56475. 42:01:25dots fall into the green. So even our
  56476. 42:01:27training set isn't going to be 100%. Uh
  56477. 42:01:30we couldn't do that. And so we're
  56478. 42:01:32looking at our different data coming
  56479. 42:01:34down. Uh we can kind of arrange our x1
  56480. 42:01:37x2 so we have a nice plot going on. And
  56481. 42:01:40we're going to create the um contour.
  56482. 42:01:43That's that nice line that's drawn down
  56483. 42:01:45the middle on here with the red green.
  56484. 42:01:47Um that's what that's what this is doing
  56485. 42:01:49right here with the reshape and notice
  56486. 42:01:51that we had to uh do the t if you
  56487. 42:01:55remember from numpy um if you did the
  56488. 42:01:57numpy module um you end up with pairs
  56489. 42:02:00you know x uh x1 x2 x1 x2 next row and
  56490. 42:02:05so forth you have to flip it so it's all
  56491. 42:02:07one row you have all your x1's and all
  56492. 42:02:09your x2s. Um so this what we're kind of
  56493. 42:02:11looking for right here on this setup.
  56494. 42:02:15Uh, and then the scatter plot is of
  56495. 42:02:17course um your scattered data across
  56496. 42:02:19there. We're just going through all the
  56497. 42:02:20points that puts these nice little dots
  56498. 42:02:22onto our setup on here. And we have our
  56499. 42:02:25estimated salary and our H. And then of
  56500. 42:02:27course the dots are did they make a
  56501. 42:02:29purchase or not. And just a quick note,
  56502. 42:02:31this is kind of funny. You can see up
  56503. 42:02:33here where it says X set Y set equals uh
  56504. 42:02:36X train Y train, which seems kind of a
  56505. 42:02:39little weird to do. Um, this is because
  56506. 42:02:42this is probably originally a
  56507. 42:02:43definition. Uh, so it's its own module
  56508. 42:02:46that could be called over and over
  56509. 42:02:47again. And which is really a good way to
  56510. 42:02:50do it because the next thing we're going
  56511. 42:02:51to want to do is do the exact same
  56512. 42:02:52thing, but we're going to visualize the
  56513. 42:02:55test set results. Uh, that way we can
  56514. 42:02:57see what happened with our test group,
  56515. 42:02:59our 25%.
  56516. 42:03:02And you can see down here we have um the
  56517. 42:03:04test set. Uh, and it, if you look at the
  56518. 42:03:07two graphs next to each other, this one
  56519. 42:03:09obviously has um 75% of the data, so
  56520. 42:03:12it's going to show a lot more. This is
  56521. 42:03:14only 25% of the data. You can see that
  56522. 42:03:17there's a number that are kind of on the
  56523. 42:03:19edge as to whether they could guess by
  56524. 42:03:21age and income they're going to make a
  56525. 42:03:22purchase or not. U, but that said, it
  56526. 42:03:25still is pretty clear. It's pretty good
  56527. 42:03:27as far as how much the estimate is and
  56528. 42:03:28how good it does.
  56529. 42:03:31Now, graphs are really effective for
  56530. 42:03:34showing people what's going on, but you
  56531. 42:03:37also need to have the numbers. And so,
  56532. 42:03:39we're going to do from sklearn, we're
  56533. 42:03:41going to import metrics, and then we're
  56534. 42:03:43going to print our metrics
  56535. 42:03:44classification port from the Y test and
  56536. 42:03:46the Y predict.
  56537. 42:03:49And you can see here we have precision
  56538. 42:03:52uh precision of zeros is 90. There's our
  56539. 42:03:54recall
  56540. 42:03:5696. We have an F1 score and a support.
  56541. 42:04:00And we have our precision, the recall on
  56542. 42:04:03getting it right. Uh, and then we can do
  56543. 42:04:05our accuracy, the macro average, and the
  56544. 42:04:07weighted average. Uh, so you can see it
  56545. 42:04:10pulls in pretty good as far as um how
  56546. 42:04:13accurate it is. You could say it's going
  56547. 42:04:15to be about 90% is going to guess
  56548. 42:04:18correctly um that it that they're not
  56549. 42:04:21going to purchase. And we had an 89%
  56550. 42:04:23chance that they are going to purchase.
  56551. 42:04:25Um, and then the other numbers as you
  56552. 42:04:27get down have a little bit different
  56553. 42:04:29meaning, but it's pretty straightforward
  56554. 42:04:31on here. Here's our accuracy, and here's
  56555. 42:04:33our micro average, and the weighted
  56556. 42:04:35average, and everything else you might
  56557. 42:04:36need. And if you forgot the exact
  56558. 42:04:38definition of accuracy, it is the true
  56559. 42:04:42positive, true negative over all of the
  56560. 42:04:44different setups. Precision is your true
  56561. 42:04:47positive over all positives, true and
  56562. 42:04:50false. And recall is a true positive
  56563. 42:04:53over true positive plus false negative.
  56564. 42:04:56And we can just real quick flip back
  56565. 42:04:58there so you can see those numbers on
  56566. 42:05:00here. Uh here's our precision, here's
  56567. 42:05:03our recall, and here's our accuracy on
  56568. 42:05:06this.
  56569. 42:05:07>> Welcome to this exciting journey into
  56570. 42:05:09the world of statistics for data
  56571. 42:05:11science. Have you ever wondered how data
  56572. 42:05:13transforms from raw numbers into
  56573. 42:05:15powerful insights that drive decisions?
  56574. 42:05:18Well, statistics is the magic behind it
  56575. 42:05:21all. Today, we will uncover how
  56576. 42:05:23statistical methods help us summarize
  56577. 42:05:25data, model uncertaintity, test
  56578. 42:05:28hypothesis, and find relationships that
  56579. 42:05:31can predict the future. So, buckle up.
  56580. 42:05:33We are about to turn numbers into
  56581. 42:05:35knowledge. Without any further ado,
  56582. 42:05:37let's get started. Now, to start off,
  56583. 42:05:40here's a key question. What are
  56584. 42:05:42statistics in data science? Now,
  56585. 42:05:44statistics is the science of collecting,
  56586. 42:05:46analyzing and interpreting data. By
  56587. 42:05:49applying statistical methods, we can
  56588. 42:05:51uncover patterns in the data and make
  56589. 42:05:54informed decisions. Now, as we continue,
  56590. 42:05:57you notice how these essential concepts
  56591. 42:05:59will form the backbone of many data
  56592. 42:06:02science practices. Let's explore the key
  56593. 42:06:04functions of statistics in data science.
  56594. 42:06:07First, statistic helps summarize data
  56595. 42:06:10using measures like mean, median and
  56596. 42:06:13variance. Next, it models uncertaintity
  56597. 42:06:16with probability and distributions. So,
  56598. 42:06:19we can better understand risk and
  56599. 42:06:21variability in our data. It also tests
  56600. 42:06:25hypothesis such as when we use AB
  56601. 42:06:28testing to compare different outcomes.
  56602. 42:06:30Statistics finds relationship through
  56603. 42:06:33methods like regression and correlation
  56604. 42:06:36revealing how variables impact each
  56605. 42:06:38other. And finally, all these tools
  56606. 42:06:41enable datadriven decision-m turning raw
  56607. 42:06:44numbers into actionable insights. Now,
  56608. 42:06:47let's talk about why does statistics
  56609. 42:06:49matter in data science. Statistics form
  56610. 42:06:51the backbone of data science, providing
  56611. 42:06:54the mathematical framework needed to
  56612. 42:06:56make sense of data and draw reliable
  56613. 42:06:59conclusions. Without statistics, it
  56614. 42:07:01would be impossible to turn raw data
  56615. 42:07:03into meaningful insights or make
  56616. 42:07:05confident evidence-based decisions. Now
  56617. 42:07:09let's have a look at the main branches
  56618. 42:07:11of statistics. Descriptive statistics
  56619. 42:07:14and inferential statistic. So first
  56620. 42:07:16let's talk about the definition. Then we
  56621. 42:07:19have got methods and measures. Now in
  56622. 42:07:21the case of descriptive statistics, now
  56623. 42:07:24let's have a look at the main branches
  56624. 42:07:26of statistics. So basically there are
  56625. 42:07:28two core branches. Descriptive
  56626. 42:07:30statistics and inferial statistics.
  56627. 42:07:33Descriptive statistics summarizes and
  56628. 42:07:36describes data using measures like mean,
  56629. 42:07:39median, mode, range and standard
  56630. 42:07:41deviation. Its purpose is to organize
  56631. 42:07:43and present data typically with charts,
  56632. 42:07:46graph or summary tables. The scope of
  56633. 42:07:49descriptive statistics is limited to the
  56634. 42:07:52sample data itself. Now on the other
  56635. 42:07:54hand, inferential statistics makes
  56636. 42:07:56inferences about populations based on
  56637. 42:07:59samples. It uses methods like hypothesis
  56638. 42:08:03testing, confidence intervals and
  56639. 42:08:05regression. The purpose here is to draw
  56640. 42:08:08conclusion and make predictions with
  56641. 42:08:11common examples including AB testing and
  56642. 42:08:14survey analysis. The scope of inferial
  56643. 42:08:17statistics extends beyond the sample to
  56644. 42:08:20the larger population. Understanding
  56645. 42:08:22both these branches is essential for
  56646. 42:08:25analyzing and interpreting data in any
  56647. 42:08:28data science project. Now let's take a
  56648. 42:08:30closer look at the descriptive
  56649. 42:08:32statistics starting with measures of
  56650. 42:08:35central tendency. The first measure here
  56651. 42:08:37is mean which is the average of all the
  56652. 42:08:40values simply calculated as the sum of
  56653. 42:08:43all the data points divided by the total
  56654. 42:08:46count. Next is the median which
  56655. 42:08:48represents the middle value when the
  56656. 42:08:50data is sorted from lowest to highest.
  56657. 42:08:53This is especially useful when dealing
  56658. 42:08:55with skewed distributions. And finally,
  56659. 42:08:58the mode is the most frequently
  56660. 42:09:00occurring value in the data set, helping
  56661. 42:09:03us identify common patterns or repeated
  56662. 42:09:06outcomes. Let's continue our deep dive
  56663. 42:09:09into descriptive statistics by looking
  56664. 42:09:11at the measures of variability. First,
  56665. 42:09:14we've got the range. This is simply the
  56666. 42:09:16difference between the maximum and
  56667. 42:09:18minimum value in a data set showing us
  56668. 42:09:21the spread of our data which measures
  56669. 42:09:23the average of the squared differences
  56670. 42:09:26from the mean. This tells us how much
  56671. 42:09:29the values in our data set differ from
  56672. 42:09:31the average. Closely related is the
  56673. 42:09:34standard deviation which is the square
  56674. 42:09:36root of the variance. It gives us a more
  56675. 42:09:39intuitive sense of how much the values
  56676. 42:09:43typically deviate from the mean. And
  56677. 42:09:45finally, we've got the interquartile
  56678. 42:09:47range or we say IQR. This shows us the
  56679. 42:09:50range of the middle 50% of our data,
  56680. 42:09:54helping us understand how data is
  56681. 42:09:56distributed across the center and avoid
  56682. 42:09:59the effect of outliers. Understanding
  56683. 42:10:02these four measures allows us to
  56684. 42:10:04summarize not just the center of our
  56685. 42:10:06data, but how spread out and varied our
  56686. 42:10:10data set is. Now let's look at some
  56687. 42:10:12practical applications of descriptive
  56688. 42:10:15statistics. One major use is the data
  56689. 42:10:18exploration and summarization where we
  56690. 42:10:20quickly get an overview and basic
  56691. 42:10:23understanding of complex data sets
  56692. 42:10:25helping track performance and detect
  56693. 42:10:27problems early in fields like
  56694. 42:10:29manufacturing or operations. And
  56695. 42:10:32finally, they are central to business
  56696. 42:10:34reporting and dashboards where concise
  56697. 42:10:37summaries are essential for managers to
  56698. 42:10:40review trends and make datadriven
  56699. 42:10:42decisions. So in short, descriptive
  56700. 42:10:45statistics help transform raw data into
  56701. 42:10:48clear actionable information across many
  56702. 42:10:51business, scientific and operational
  56703. 42:10:53context. Let's explore the first type of
  56704. 42:10:56data in statistics, qualitative or
  56705. 42:10:59categorical data. This type of data
  56706. 42:11:01includes descriptive information that
  56707. 42:11:04cannot be measured numerically such as
  56708. 42:11:06categories or labels and yes or no
  56709. 42:11:09responses. These are all about qualities
  56710. 42:11:12or characteristics rather than
  56711. 42:11:14quantities. Qualitative data can be
  56712. 42:11:17further divided into two types. Nominal
  56713. 42:11:20where the categories have no specific
  56714. 42:11:22order and ordinal where the categories
  56715. 42:11:25do have an order or ranking. Recognizing
  56716. 42:11:28and classifying qualitative data is very
  56717. 42:11:31important as it affects how information
  56718. 42:11:33is analyzed and interpreted in
  56719. 42:11:36statistics. Now let's discuss the second
  56720. 42:11:38main type of data in statistics which is
  56721. 42:11:41quantitative or numerical data. This
  56722. 42:11:44type of data includes information that
  56723. 42:11:46can be measured and expressed with
  56724. 42:11:48numbers making it ideal for mathematical
  56725. 42:11:51analysis. Quantitative data is further
  56726. 42:11:54divided into two main categories. First,
  56727. 42:11:57there is discrete data. These are
  56728. 42:12:00countable values with specific fixed
  56729. 42:12:02points such as the number of students,
  56730. 42:12:05cars sold or website clicks. Second,
  56731. 42:12:08there is continuous data which includes
  56732. 42:12:10infinite possible values within a given
  56733. 42:12:13range. Examples of continuous data
  56734. 42:12:16include height, weight, temperature or
  56735. 42:12:18time. Understanding the distinction
  56736. 42:12:21between discrete and continuous data is
  56737. 42:12:23very important as it determines which
  56738. 42:12:26statistical methods and visualizations
  56739. 42:12:29will be most appropriate. Now let's dive
  56740. 42:12:31into the fundamentals of probability.
  56741. 42:12:34Probability measures the likelihood of
  56742. 42:12:36an event occurring and it's always
  56743. 42:12:38expressed as a value between 0 and 1.
  56744. 42:12:42Here are some key concepts. If the
  56745. 42:12:44probability or P equals to zero, that
  56746. 42:12:47means the event will never occur. If P
  56747. 42:12:50is equals to 1, the event will always
  56748. 42:12:53occur. And if P is equals to 0.5, the
  56749. 42:12:56event has an equal chance of occurring
  56750. 42:12:58or not occurring. It's truly a 50/50%
  56751. 42:13:02scenario. Now, understanding these basic
  56752. 42:13:04principle help us quantify uncertaintity
  56753. 42:13:07and make informed predictions about
  56754. 42:13:10future outcomes. Now let's look at the
  56755. 42:13:13different types of probability. First
  56756. 42:13:15there's classical probability. This is
  56757. 42:13:18based on equally likely outcomes such as
  56758. 42:13:20flipping a fair coin or rolling a
  56759. 42:13:23balanced die. Next empirical probability
  56760. 42:13:26which relies on observed frequency. It's
  56761. 42:13:30calculated from actual data such as the
  56762. 42:13:32proportion of rainy days over the past
  56763. 42:13:35month. Let's say for example the
  56764. 42:13:37probability of it's raining today given
  56765. 42:13:40that it's cloudy. Understanding these
  56766. 42:13:43three type help us choose the right
  56767. 42:13:45approach for different situations.
  56768. 42:13:47Whether we are predicting outcomes,
  56769. 42:13:50analyzing data or making decisions under
  56770. 42:13:53uncertaintity. Now probability has a
  56771. 42:13:55wide range of powerful applications in
  56772. 42:13:58data science. First of all, it is used
  56773. 42:14:00in predictive modeling and machine
  56774. 42:14:02learning where algorithms estimate
  56775. 42:14:05future outcomes based on existing data.
  56776. 42:14:08Probability also plays a key role in
  56777. 42:14:11risk management and decision making
  56778. 42:14:13helping businesses and researchers
  56779. 42:14:15evaluate the likelihood of different
  56780. 42:14:18scenarios and plan accordingly.
  56781. 42:14:21Conditional probability calculating the
  56782. 42:14:23chance of one event given that the
  56783. 42:14:26another has occurred. This is crucial in
  56784. 42:14:29fields like healthcare, fraud detection
  56785. 42:14:31and marketing analytics.
  56786. 42:14:35And finally, probability is foundational
  56787. 42:14:37in AB testing and experimental design,
  56788. 42:14:41allowing us to measure the effectiveness
  56789. 42:14:43of new strategies or products. These
  56790. 42:14:46applications show how probability
  56791. 42:14:48enables smarter evidence-driven progress
  56792. 42:14:51in modern data science. Let's look at
  56793. 42:14:54one of the most common probability
  56794. 42:14:56distributions, the normal distribution.
  56795. 42:14:59This distribution is famous for its
  56796. 42:15:01bell-shaped symmetric curve, which shows
  56797. 42:15:04that most values cluster around the mean
  56798. 42:15:07with fewer and fewer values appearing as
  56799. 42:15:10you move away from the center. Now many
  56800. 42:15:13natural phenomena like heights, test
  56801. 42:15:16scores and measurement errors tend to
  56802. 42:15:18follow this pattern making the normal
  56803. 42:15:21distribution a key concept in
  56804. 42:15:23statistics. It's characterized by two
  56805. 42:15:26main parameters. The mean which
  56806. 42:15:28determines the center of the curve and
  56807. 42:15:30the standard deviation which controls
  56808. 42:15:33its spread.
  56809. 42:15:34Recognizing the distribution help
  56810. 42:15:36analysts make predictions, calculate
  56811. 42:15:39probabilities, and apply statistical
  56812. 42:15:41techniques to real world data. Next,
  56813. 42:15:44let's explore the binomial distribution,
  56814. 42:15:47which is another common probability
  56815. 42:15:49distribution. The binomial distribution
  56816. 42:15:51is discrete and is used for situations
  56817. 42:15:54with binary outcomes like success or
  56818. 42:15:57failure. It is based on a fixed number
  56819. 42:16:00of trials where each trial has a
  56820. 42:16:03constant probability of success such as
  56821. 42:16:05flipping a coin a certain number of
  56822. 42:16:07times or tracking pass fail rates.
  56823. 42:16:10Examples of binomial experiment includes
  56824. 42:16:13coin flips and measuring how many
  56825. 42:16:15students pass or fail a test. This
  56826. 42:16:18distribution help us model and analyze
  56827. 42:16:20outcomes when only two possibilities
  56828. 42:16:23exist in each trial. Now let's focus on
  56829. 42:16:26the poison distribution. Another key
  56830. 42:16:29type of probability distribution. The
  56831. 42:16:31poion distribution is a discrete
  56832. 42:16:33distribution specifically used to model
  56833. 42:16:36rare events. It's particularly helpful
  56834. 42:16:39for modeling events that occur
  56835. 42:16:41independently over a fixed interval of
  56836. 42:16:44time or space. For example, it predicts
  56837. 42:16:47how many times an event like a customer
  56838. 42:16:50arriving on a website, receiving a visit
  56839. 42:16:53might happen in a certain period.
  56840. 42:16:55Typical examples include customer
  56841. 42:16:57arrivals at a store, defect rates, and
  56842. 42:17:00manufacturing or counts of website
  56843. 42:17:02visits over a set period. This
  56844. 42:17:05distribution is great tool for
  56845. 42:17:07understanding and predicting random
  56846. 42:17:09independent events that don't happen
  56847. 42:17:12very often, but are important to track.
  56848. 42:17:14It's a branch that allows us to make
  56849. 42:17:17inferences, predictions, or
  56850. 42:17:19generalizations about a larger
  56851. 42:17:21population using sample data. Here are
  56852. 42:17:24some key concepts. First is the
  56853. 42:17:26population versus sample. The population
  56854. 42:17:29represents the entire group we want to
  56855. 42:17:31know about while the sample is the
  56856. 42:17:33subset we actually collect data from.
  56857. 42:17:36Next concept is sampling distribution
  56858. 42:17:39which refers to the distribution of a
  56859. 42:17:41statistic across multiple samples from
  56860. 42:17:44the same population. Standard error
  56861. 42:17:47measures how much the sample statistic
  56862. 42:17:49is expected to vary due to random
  56863. 42:17:52sampling. And finally, margin of error
  56864. 42:17:55tells us how much we can expect our
  56865. 42:17:57estimates to differ from the true
  56866. 42:18:00population value. Together these concept
  56867. 42:18:03form the foundation for drawing reliable
  56868. 42:18:05insights from sample data in inferential
  56869. 42:18:08statistics.
  56870. 42:18:10Let's review some of the main techniques
  56871. 42:18:12used in inferial statistic. First, we've
  56872. 42:18:15got hypothesis testing. This method
  56873. 42:18:17allows us to test claims or ideas about
  56874. 42:18:21population parameters based on sample
  56875. 42:18:24data helping us determine if observed
  56876. 42:18:27results are statistically significant
  56877. 42:18:29and reliability of our estimate. Next
  56878. 42:18:32are confidence intervals. These provide
  56879. 42:18:34a range of likely values for a
  56880. 42:18:36population parameter giving us a sense
  56881. 42:18:39of possible variation reliability of our
  56882. 42:18:42estimate. And lastly, regression
  56883. 42:18:44analysis is used to model and analyze
  56884. 42:18:47the relationships between variables,
  56885. 42:18:49allowing us to make predictions, uncover
  56886. 42:18:52trends, and understand how changes in
  56887. 42:18:54one factor might affect another. Now,
  56888. 42:18:57let's walk through the hypothesis
  56889. 42:18:59testing process. The first step is to
  56890. 42:19:02formulate hypothesis. Start with null
  56891. 42:19:04hypothesis represented as Hnot, which
  56892. 42:19:08states that there is no effect or
  56893. 42:19:10difference. Then there's the alternative
  56894. 42:19:13hypothesis represented as H1 which
  56895. 42:19:16suggests that an effect or difference
  56896. 42:19:19does exist. Now after setting up the
  56897. 42:19:21hypothesis the next step is to choose a
  56898. 42:19:24significance level noted by alpha.
  56899. 42:19:27Common choices for significance levels
  56900. 42:19:29include 0.05 5% 01 1% 010 which is 10%.
  56901. 42:19:36The significance level is important
  56902. 42:19:38because it controls the likelihood of
  56903. 42:19:40making a type one error also known as
  56904. 42:19:42false positive. And now each step in the
  56905. 42:19:46process is critical for ensuring that
  56906. 42:19:48statistical results are both meaningful
  56907. 42:19:50and reliable. The next step is to
  56908. 42:19:53collect and analyze sample data. Ensure
  56909. 42:19:55the data is represented by choosing a
  56910. 42:19:58good sample. Then calculate the
  56911. 42:20:00appropriate and test statistic for your
  56912. 42:20:03hypothesis test. Once that's done, it's
  56913. 42:20:06time to make a decision. Compare the p
  56914. 42:20:08value to the chosen significance level
  56915. 42:20:11alpha. Now, if the p value is less than
  56916. 42:20:13or equals to alpha, you reject the null
  56917. 42:20:16hypothesis. If the p value is greater
  56918. 42:20:19than alpha, you fail to reject the null
  56919. 42:20:21hypothesis. And finally, interpret your
  56920. 42:20:24results. Draw your conclusions with
  56921. 42:20:26respect to the context and problem at
  56922. 42:20:29hand. Always keeping the bigger picture
  56923. 42:20:31in mind. Following these step helps
  56924. 42:20:34ensure your hypothesis test is robust,
  56925. 42:20:37clear and meaningful. Let's review the
  56926. 42:20:40common types of hypothesis test used in
  56927. 42:20:42statistic. The one sample test compares
  56928. 42:20:45the mean of a sample to a known value
  56929. 42:20:48and the one sample zed test is used when
  56930. 42:20:51the population standard deviation is
  56931. 42:20:53known. Next, we have two sample test.
  56932. 42:20:56The independent samples t test compares
  56933. 42:20:59to the means of two different groups
  56934. 42:21:02while paired samples t test compares
  56935. 42:21:04before and after measurements for the
  56936. 42:21:07same subject. For categorical data test,
  56937. 42:21:10the shear test checks for the
  56938. 42:21:12independence or goodness of fit. And
  56939. 42:21:16fiser's exact test is useful for small
  56940. 42:21:19sample sizes. And lastly, non-parametric
  56941. 42:21:22tests such as the man Whitney U test and
  56942. 42:21:25Wil Coxson signed rank test serve as
  56943. 42:21:28alternatives to the t test when data
  56944. 42:21:31doesn't meet certain parametric
  56945. 42:21:33assumptions. Choosing the right
  56946. 42:21:35hypothesis test depends on your data
  56947. 42:21:37type and the specific question you want
  56948. 42:21:40to answer. Let's explore the central
  56949. 42:21:42limit theorem or CLT of foundation for
  56950. 42:21:45inferial statistics. The central limit
  56951. 42:21:48theorem states that as the sample size
  56952. 42:21:51increases, the distribution of sample
  56953. 42:21:53means approaches a normal distribution
  56954. 42:21:56even if the original data is a normally
  56955. 42:21:59distributed. This sample works for any
  56956. 42:22:01population distribution making it
  56957. 42:22:04incredibly powerful. Now for good
  56958. 42:22:06results, the sample size should
  56959. 42:22:08typically be 30 or more. What's
  56960. 42:22:11interesting is the sample mean
  56961. 42:22:12distribution which will have the same
  56962. 42:22:15mean as the population. The standard
  56963. 42:22:18error which measures variability by the
  56964. 42:22:20sample mean equals sigma / the square
  56965. 42:22:23roo<unk> of n. And this gets smaller as
  56966. 42:22:26the sample size grows. And finally the
  56967. 42:22:29formula shown here which is z is equ= to
  56968. 42:22:32x -
  56969. 42:22:35sigma divided by sigma over the square
  56970. 42:22:38root of n. Let's standardize and compare
  56971. 42:22:41sample means. The CLT makes most
  56972. 42:22:44parametric statistics possible and is
  56973. 42:22:47the backbone for many statistical test.
  56974. 42:22:49Let's review the main types of
  56975. 42:22:51regression analysis which are used to
  56976. 42:22:54model and understand relationships
  56977. 42:22:56between variables. First up is linear
  56978. 42:22:59regression. This technique analyzes the
  56979. 42:23:02relationship between a continuous
  56980. 42:23:04variable and another variable resulting
  56981. 42:23:07in a straight line. It's commonly used
  56982. 42:23:09in predicting sales or prices. Next is
  56983. 42:23:13logistic regression. Unlike linear
  56984. 42:23:15regression, this method is used for
  56985. 42:23:18binary outcomes, helping estimate
  56986. 42:23:20probabilities such as whether an email
  56987. 42:23:23is spam or a patient has a disease.
  56988. 42:23:26Moving to multiple regression, this
  56989. 42:23:28allows us to account for the effect of
  56990. 42:23:31several variables at once, modeling more
  56991. 42:23:34complex relationships like determining
  56992. 42:23:36house prices. Lastly, polomial
  56993. 42:23:39regression which is used for nonlinear
  56994. 42:23:42relationships. The resulting curve
  56995. 42:23:44rather than a straight line lets us
  56996. 42:23:46capture growth trends and other patterns
  56997. 42:23:48that aren't linear. Understanding these
  56998. 42:23:51types of regression help analysts choose
  56999. 42:23:54the right model for the data and
  57000. 42:23:56business's problem at hand. Let's
  57001. 42:23:59clarify the important differences
  57002. 42:24:00between correlation and causation.
  57003. 42:24:03Correlation is statistical measure of
  57004. 42:24:06how two variables move together.
  57005. 42:24:08Correlation values range from minus 10
  57006. 42:24:10to + one. But remember correlation does
  57007. 42:24:14not imply causation. On the other hand,
  57008. 42:24:16causation means that one variable
  57009. 42:24:18actually causes changes in another.
  57010. 42:24:21Establishing causation requires
  57011. 42:24:24controlled experimentation and is much
  57012. 42:24:27stronger relationship than simple
  57013. 42:24:29correlation. Always be careful when
  57014. 42:24:32interpreting results. Just because two
  57015. 42:24:34variables move together doesn't mean one
  57016. 42:24:36cause the other. Let's look at some
  57017. 42:24:39common problems with interpreting
  57018. 42:24:41correlation and causation. First is the
  57019. 42:24:43third variable problem which happens
  57020. 42:24:46when a hidden variable affects both
  57021. 42:24:48variables in a question leading to a
  57022. 42:24:50misleading connection. There's also this
  57023. 42:24:52directionality problem where it's
  57024. 42:24:54unclear which variable is causing others
  57025. 42:24:57to change but both are actually caused
  57026. 42:25:00by a third variable which is the hot
  57027. 42:25:02weather. This demonstrate that
  57028. 42:25:03correlation does not mean one variable
  57029. 42:25:06causes the other. So always look for
  57030. 42:25:08hidden factors before assuming
  57031. 42:25:10causation. Let's talk about statistical
  57032. 42:25:13errors specifically type one and type
  57033. 42:25:15two errors. Type one error also called
  57034. 42:25:18as false positive occurs when we reject
  57035. 42:25:21the true null hypothesis. In other
  57036. 42:25:23words, we wrongly conclude there's an
  57037. 42:25:26effect when there's actually isn't. The
  57038. 42:25:28probability of making this error is
  57039. 42:25:31equal to the significance level alpha.
  57040. 42:25:34For example, concluding a drug works
  57041. 42:25:36when it actually doesn't. On the other
  57042. 42:25:38hand, type two error is false negative
  57043. 42:25:42means failing to reject a false null
  57044. 42:25:44hypothesis. That's when we miss a real
  57045. 42:25:47effect or difference. The probability of
  57046. 42:25:50type two error is beta. For example,
  57047. 42:25:53missing the real effect of a drug and
  57048. 42:25:55seeing it doesn't work when it actually
  57049. 42:25:57does. Understanding these errors is key
  57050. 42:26:00for designing good experiments and
  57051. 42:26:03interpreting statistical results
  57052. 42:26:05properly. Let's look at three common
  57053. 42:26:07sampling methods using statistics. The
  57054. 42:26:10first one is random sampling. Here every
  57055. 42:26:12individual in the population has a equal
  57056. 42:26:15chance of being selected where every nth
  57057. 42:26:18individual is chosen. Random sampling is
  57058. 42:26:21crucial for ensuring representative
  57059. 42:26:23sample. Next is stratified sampling. The
  57060. 42:26:27population is divided onto homogeneous
  57061. 42:26:29subgroups or strata and then a random
  57062. 42:26:33sample is drawn from each group. This
  57063. 42:26:35approach makes sure all the groups are
  57064. 42:26:38represented in the sample. And finally,
  57065. 42:26:40cluster sampling divides the population
  57066. 42:26:43into clusters, often based on geography.
  57067. 42:26:46Entire clusters are then randomly
  57068. 42:26:48selected. It's cost effective and useful
  57069. 42:26:51for large spread out populations.
  57070. 42:26:54Choosing the right method ensures the
  57071. 42:26:56data truly represents the whole
  57072. 42:26:59population and strengthens the study's
  57073. 42:27:01conclusions. So guys, let's wrap up with
  57074. 42:27:04some real world applications of
  57075. 42:27:06statistics. In business analytics,
  57076. 42:27:08statistics are used for AB testing to
  57077. 42:27:11optimize websites, customer segmentation
  57078. 42:27:14and targeting, sales forecasting and
  57079. 42:27:17demand planning and quality control and
  57080. 42:27:20also process improvement. These
  57081. 42:27:22techniques help businesses make smarter
  57082. 42:27:24datadriven decisions every day.
  57083. 42:27:27Statistics play a vital role in
  57084. 42:27:29healthcare and medicine as well. They
  57085. 42:27:32are key for analyzing clinical trial
  57086. 42:27:34results, conducting epidemological
  57087. 42:27:37studies, evaluating treatment
  57088. 42:27:39effectiveness, and identifying risk
  57089. 42:27:42factors. By using these approaches,
  57090. 42:27:44healthcare researchers and practitioners
  57091. 42:27:47can improve patient outcomes in public
  57092. 42:27:49health. From businesses to medicine,
  57093. 42:27:52statistics transform raw information
  57094. 42:27:54into actionable insights that create
  57095. 42:27:57real impact. Statistics has a huge
  57096. 42:27:59impact in technology, data science and
  57097. 42:28:02finance and power recommener systems
  57098. 42:28:04that personalize what users see. In
  57099. 42:28:07finance, statistic help with risk
  57100. 42:28:10assessment and management, optimizing
  57101. 42:28:12investment portfolios, determining
  57102. 42:28:15credit scores, and supporting market
  57103. 42:28:17research and analysis. Now, these
  57104. 42:28:19applications show how statistical
  57105. 42:28:22techniques help make smarter decisions
  57106. 42:28:24and solve complex challenges across high
  57107. 42:28:27impact industry. Are you one of the many
  57108. 42:28:30who dreams of becoming a data scientist?
  57109. 42:28:32Keep watching this video if you're
  57110. 42:28:34passionate about data science because we
  57111. 42:28:36will tell you how does it really work
  57112. 42:28:38under the hood. Emma is a data
  57113. 42:28:39scientist. Let's see how a day in her
  57114. 42:28:41life goes while she's working on a data
  57115. 42:28:44science project. Well, it is very
  57116. 42:28:46important to understand the business
  57117. 42:28:47problem first. In our meeting with the
  57118. 42:28:49clients, Emma asks relevant questions,
  57119. 42:28:52understands and defines objectives for
  57120. 42:28:55the problem that needs to be tackled.
  57121. 42:28:57She's a curious soul who asks a lot of
  57122. 42:29:00wise, one of the many traits of a good
  57123. 42:29:02data scientist. Now, she ges up for data
  57124. 42:29:05acquisition to gather and scrape data
  57125. 42:29:07from multiple sources like web servers,
  57126. 42:29:10logs, databases, APIs, and online
  57127. 42:29:12repositories. Oh, it seems like finding
  57128. 42:29:15the right data takes both time and
  57129. 42:29:17effort. After the data is gathered comes
  57130. 42:29:20data preparation. This step involves
  57131. 42:29:22data cleaning and data transformation.
  57132. 42:29:25Data cleaning is the most time consuming
  57133. 42:29:27process as it involves handling many
  57134. 42:29:29complex scenarios. Here Emma deals with
  57135. 42:29:32inconsistent data types, misspelled
  57136. 42:29:34attributes, missing values, duplicate
  57137. 42:29:37values and whatnot. Then in data
  57138. 42:29:39transformation she modifies the data
  57139. 42:29:42based on defined mapping rules. In a
  57140. 42:29:44project ETL tools like talent and
  57141. 42:29:46Informatica are used to perform complex
  57142. 42:29:49transformations that helps the team to
  57143. 42:29:51understand the data structure better.
  57144. 42:29:53Then understanding what you actually can
  57145. 42:29:55do with your data is very crucial. For
  57146. 42:29:57that Emma does exploratory data analysis
  57147. 42:30:00with the help of EDA. She defines and
  57148. 42:30:03refineses the selection of feature
  57149. 42:30:04variables that will be used in the model
  57150. 42:30:07development. But what if Emma skips this
  57151. 42:30:09step? She might end up choosing the
  57152. 42:30:11wrong variables which will produce an
  57153. 42:30:13inaccurate model. Thus, exploratory data
  57154. 42:30:15analysis becomes the most important
  57155. 42:30:18step. Now, she proceeds to the core
  57156. 42:30:20activity of a data science project which
  57157. 42:30:22is data modeling. She repetitively
  57158. 42:30:24applies diverse machine learning
  57159. 42:30:26techniques like KN&N, decision tree,
  57160. 42:30:29knives based to the data to identify the
  57161. 42:30:32model that best fits the business
  57162. 42:30:34requirements. She trains the models on
  57163. 42:30:36the training data set and tests them to
  57164. 42:30:38select the best performing model. Emma
  57165. 42:30:41prefers Python for modeling the data.
  57166. 42:30:43However, it can also be done using R and
  57167. 42:30:46SAS. Well, the trickiest part is not yet
  57168. 42:30:49over. Visualization and communication.
  57169. 42:30:51Emma meets the clients again to
  57170. 42:30:53communicate the business findings in a
  57171. 42:30:55simple and effective manner to convince
  57172. 42:30:58the stakeholders. She uses tools like
  57173. 42:31:00Tableau, PowerBI and ClickView that can
  57174. 42:31:03help her in creating powerful reports
  57175. 42:31:05and dashboards. And then finally, she
  57176. 42:31:07deploys and maintains the model. She
  57177. 42:31:10tests the selected model in a
  57178. 42:31:11pre-production environment before
  57179. 42:31:13deploying it in the production
  57180. 42:31:15environment which is the best practice.
  57181. 42:31:17Right? After successfully deploying it,
  57182. 42:31:19she uses reports and dashboards to get
  57183. 42:31:22realtime analytics. Further, she also
  57184. 42:31:25monitors and maintains the project's
  57185. 42:31:27performance. Well, that's how Emma
  57186. 42:31:29completes the data science project. We
  57187. 42:31:31have seen the daily routine of a data
  57188. 42:31:33scientist is a whole lot of fun, has a
  57189. 42:31:35lot of interesting aspects and comes
  57190. 42:31:37with its own share of challenges. Now,
  57191. 42:31:40let's see how data science is changing
  57192. 42:31:42the world. Data science techniques along
  57193. 42:31:44with genomic data provides a deeper
  57194. 42:31:47understanding of genetic issues and
  57195. 42:31:49reaction to particular drugs and
  57196. 42:31:50diseases. Logistic companies like DHL,
  57197. 42:31:54FedEx have discovered the best routes to
  57198. 42:31:56ship, the best suited time to deliver,
  57199. 42:31:58the best mode of transport to choose,
  57200. 42:32:00thus leading to cost efficiency. With
  57201. 42:32:03data science, it is possible to not only
  57202. 42:32:06predict employee attrition, but to also
  57203. 42:32:08understand the key variables that
  57204. 42:32:10influence employee turnover. Also, the
  57205. 42:32:13airline companies can now easily predict
  57206. 42:32:15flight delay and notify the passengers
  57207. 42:32:18beforehand to enhance their travel
  57208. 42:32:20experience. Well, if you're wondering,
  57209. 42:32:22there are various roles offered to a
  57210. 42:32:24data scientist like data analyst,
  57211. 42:32:27machine learning engineer, deep learning
  57212. 42:32:29engineer, data engineer, and of course,
  57213. 42:32:32data scientist. The median base salaries
  57214. 42:32:35of a data scientist can range from
  57215. 42:32:37$95,000 to $165,000.
  57216. 42:32:41So that was about the data science. Are
  57217. 42:32:43you ready to be a data scientist? If
  57218. 42:32:45yes, then start today. The world of
  57219. 42:32:48data. Picture this. You're shopping
  57220. 42:32:50online and suddenly you see a product
  57221. 42:32:52that feels like it was made just for
  57222. 42:32:55you. How did they know? It's not by
  57223. 42:32:58chance. It's data science. Data science
  57224. 42:33:00help businesses understand what you
  57225. 42:33:02like, predict what you'll need next, and
  57226. 42:33:05improve the way we shop and use
  57227. 42:33:07technology. And here's the best part.
  57228. 42:33:10Data science isn't just about watching
  57229. 42:33:12Netflix. It's one of the fastest growing
  57230. 42:33:15careers in the world right now. In fact,
  57231. 42:33:17the US Bureau of Labor Statistic says
  57232. 42:33:20that data science jobs are expected to
  57233. 42:33:23grow 36% by 2033, way faster than most
  57234. 42:33:28of the other jobs. Companies everywhere
  57235. 42:33:30are using data to make smarter
  57236. 42:33:32decisions. That means the demand for
  57237. 42:33:35data scientists is huge. And let's talk
  57238. 42:33:38about the salary. You're probably
  57239. 42:33:39wondering how much can I earn actually?
  57240. 42:33:42Well, for entry- level position, data
  57241. 42:33:44scientists in the US are earning around
  57242. 42:33:47$152,000
  57243. 42:33:49per year right now. And by 2025, some
  57244. 42:33:52can make as much as $230,000.
  57245. 42:33:55And in India, starting salaries range
  57246. 42:33:57from 50,000 rupees to 1 lakh per month.
  57247. 42:34:01And experienced professionals can earn
  57248. 42:34:03more than 5 lakh rupees per month.
  57249. 42:34:05That's impressive, right? But the best
  57250. 42:34:07part is as a data scientist, you won't
  57251. 42:34:09just stop here. The skills you develop
  57252. 42:34:12in this role like machine learning, data
  57253. 42:34:14visualization, and statistics are highly
  57254. 42:34:17transferable and crucial for moving into
  57255. 42:34:20AI roles. So whether it's becoming an AI
  57256. 42:34:22engineer or an AI specialist, the
  57257. 42:34:24foundation you build in data science
  57258. 42:34:27will help you level up and pursue
  57259. 42:34:29exciting hyping opportunities in the AI
  57260. 42:34:32field. Now if you're thinking this
  57261. 42:34:34sounds great but where do I even start?
  57262. 42:34:37Well that's exactly what the
  57263. 42:34:39professional certificate course in data
  57264. 42:34:41science from IIT Kpur and Simply Learn
  57265. 42:34:44is designed to do. It will get you
  57266. 42:34:46started and make sure you're ready for
  57267. 42:34:48this booming industry. In this 11 month
  57268. 42:34:50live online interactive program, you
  57269. 42:34:53will learn the skills you need to become
  57270. 42:34:54a data science professional. No more
  57271. 42:34:57theory, no more fluff. You get hands-on
  57272. 42:34:59projects, life classes and mentorship
  57273. 42:35:01from IT Kpur faculty plus real world
  57274. 42:35:04industry expert who will help you build
  57275. 42:35:06your skills. So this course comes with
  57276. 42:35:09exciting amazing features that makes
  57277. 42:35:11learning even more impactful. Eight
  57278. 42:35:13times high interaction in live online
  57279. 42:35:16classes with industry expert. Regular
  57280. 42:35:18live online classes conducted by
  57281. 42:35:20experienced professionals who bring real
  57282. 42:35:22world knowledge into every session.
  57283. 42:35:25You'll also get to master 13 plus key
  57284. 42:35:27skills including generative AI, prompt
  57285. 42:35:29engineering, charge, expendable AI,
  57286. 42:35:32conversional AI, NLP, and many more.
  57287. 42:35:34These are the skills that top companies
  57288. 42:35:36use every day. And you'll gain hands-on
  57289. 42:35:38experience with 14 plus industry tools
  57290. 42:35:41like Python, SQL, Tableau, Dali 2,
  57291. 42:35:43Midjourney, TensorFlow, and more. And by
  57292. 42:35:46the end of this course, you will be
  57293. 42:35:48ready to tackle real world challenges
  57294. 42:35:50using these powerful tools and
  57295. 42:35:52techniques. But we are not just talking
  57296. 42:35:54about textbook and theory. You'll work
  57297. 42:35:57on 25 plus real world projects giving
  57298. 42:36:00you hands-on experience with the tools
  57299. 42:36:02and skills you'll use in the industry.
  57300. 42:36:05For example, our first project would be
  57301. 42:36:07about sales analysis. You'll use Python
  57302. 42:36:10to analyze a clothing company fourth
  57303. 42:36:12quarter sales data across Australian
  57304. 42:36:15state, helping the company make informed
  57305. 42:36:17decisions. The next project would be
  57306. 42:36:19about employee performance analysis
  57307. 42:36:21where you learn how to build machine
  57308. 42:36:23learning models to understand the
  57309. 42:36:24factors influencing employees turnover.
  57310. 42:36:27Our third project would be about
  57311. 42:36:29e-commerce which will help Amazon
  57312. 42:36:31improve its recommendation engine to
  57313. 42:36:33offer better recommendation to
  57314. 42:36:35customers. Upon successful completion,
  57315. 42:36:37you'll receive a program certificate
  57316. 42:36:39directly issued by the ENI city academy
  57317. 42:36:41IT Kpool within 45 days of completing
  57318. 42:36:44your cohort. This prestigious
  57319. 42:36:46certificate will help you boost your
  57320. 42:36:48resume and show potential employees that
  57321. 42:36:50you have mastered the skills needed to
  57322. 42:36:53succeed. You'll also benefit from master
  57323. 42:36:55classes delivered by distinguished IT
  57324. 42:36:56Kur faculty who bring their deep
  57325. 42:36:58expertise in this course. Plus, you will
  57326. 42:37:01be exposed to trending tools like
  57327. 42:37:02chargeb2 geni and prompt engineering.
  57328. 42:37:05Now, if you're wondering who will be
  57329. 42:37:06teaching all of this, then the IT
  57330. 42:37:08carpool faculty is here for you. These
  57331. 42:37:10experts have been in this field for
  57332. 42:37:12years and have worked with the top
  57333. 42:37:14companies. You'll also get master
  57334. 42:37:16classes from them and they will guide
  57335. 42:37:17you through the learning process. You're
  57336. 42:37:19not just learning from a textbook. You
  57337. 42:37:21are learning from the people who have
  57338. 42:37:23been there and done that. And the best
  57339. 42:37:25part is once you have learned the
  57340. 42:37:26skills, simply learns career assistance
  57341. 42:37:28team will help you take to the next step
  57342. 42:37:31which will help you to build a killer
  57343. 42:37:32resume, show you how to stand out top
  57344. 42:37:35recruiters and even give you access to
  57345. 42:37:37mock interviews. Plus, you will also
  57346. 42:37:39gain access to exclusive networking
  57347. 42:37:41events and hackathons to connect with
  57348. 42:37:43industry professionals. Along with
  57349. 42:37:45Simply Learn's job assistant, you'll
  57350. 42:37:47also get access to ID Kpur's career
  57351. 42:37:49services helping you connect with top
  57352. 42:37:51recruiters and land interviews with
  57353. 42:37:53leading tech companies. So, upon
  57354. 42:37:55finishing this course, you'll be ready
  57355. 42:37:56for top data science roles like data
  57356. 42:37:58scientists, machine learning engineer,
  57357. 42:38:00AI specialist, and business analyst. The
  57358. 42:38:02good news is top companies like Amazon,
  57359. 42:38:05EY, Fidelity Investment, Johnson and
  57360. 42:38:07Johnson, Borafhone, Accenture, Infosys
  57361. 42:38:10and Nvidia is looking for professionals
  57362. 42:38:13just like you. And with salaries in the
  57363. 42:38:15US hitting around $230,000 plus and in
  57364. 42:38:18India reaching up to five lakh per
  57365. 42:38:20month, your career outlook looks great.
  57366. 42:38:23So what will you be actually learning in
  57367. 42:38:25this course? So here's a sneak peek of
  57368. 42:38:27the syllabus which you'll be learning in
  57369. 42:38:28this course which is the foundation in
  57370. 42:38:30Python, SQL and mathematics, core data
  57371. 42:38:33science like machine learning, data
  57372. 42:38:34visualization, NLP, special topics like
  57373. 42:38:37GNI, charge GBT, prompt engineering and
  57374. 42:38:40you'll also work on industry projects
  57375. 42:38:41which we have already mentioned before.
  57376. 42:38:43You'll also have the option to choose
  57377. 42:38:45electives like data storytelling with
  57378. 42:38:47PowerBI and business analytics with
  57379. 42:38:49Excel to tailor your learning experience
  57380. 42:38:51and focus on the areas that interest you
  57381. 42:38:53the most. So don't wait, hurry up and
  57382. 42:38:56enroll now and find the course link in
  57383. 42:38:57the description box below and in the pin
  57384. 42:38:59comments. Have you ever wondered how
  57385. 42:39:01your favorite online store seems to know
  57386. 42:39:03exactly what you are looking for? Every
  57387. 42:39:06time you browse, add to cart or wish
  57388. 42:39:08list an item, you are leaving clues
  57389. 42:39:10about your style, favorite colors,
  57390. 42:39:12brands, and even shopping times. Data
  57391. 42:39:15scientists jump in, analyze these
  57392. 42:39:16patterns, and create a super
  57393. 42:39:18personalized shopping experience.
  57394. 42:39:20Suddenly, the store is showing you just
  57395. 42:39:22the right pieces at just the right time.
  57396. 42:39:24Almost like it's reading your mind.
  57397. 42:39:27That's data science. Turning your clicks
  57398. 42:39:29into a shopping spree crafted just for
  57399. 42:39:31you. Hello everyone. Welcome back to
  57400. 42:39:33Simply Learn's YouTube channel. If
  57401. 42:39:35you're already a data science enthusiast
  57402. 42:39:37or just got curious about this exciting
  57403. 42:39:39field, you're in the right place. Today
  57404. 42:39:41in this video, I'm diving into 10
  57405. 42:39:44essential steps to help you become the
  57406. 42:39:46next in- demand data scientist and land
  57407. 42:39:48that dream job. No more waiting. Let's
  57408. 42:39:51dive right in and get you on the path to
  57409. 42:39:53your future in data science. So, let's
  57410. 42:39:55see the 10 essential steps to become the
  57411. 42:39:58next data scientist in demand. Step
  57412. 42:40:00number one is programming languages.
  57413. 42:40:02Starting with Python is a beginner is a
  57414. 42:40:05great move because it's simple,
  57415. 42:40:06versatile, and widely used in data
  57416. 42:40:08science. Python straightforward syntax
  57417. 42:40:10makes it beginner friendly, helping you
  57418. 42:40:12grasp programming basics quickly and
  57419. 42:40:15dive into data science libraries like
  57420. 42:40:17pandas, numpy and mattplotive with ease.
  57421. 42:40:20Adding R to your skill set is valuable
  57422. 42:40:22because it excels at statistical
  57423. 42:40:24analysis and data visualization two
  57424. 42:40:26essential parts of data science. You can
  57425. 42:40:29be comfortable with Python and R within
  57426. 42:40:31a month or two. So moving on to the next
  57427. 42:40:33step that is version control system.
  57428. 42:40:35Learning a version control system like
  57429. 42:40:38Git is essential because it allows you
  57430. 42:40:40to track, manage, and collaborate and
  57431. 42:40:42code effectively. With Git, you can save
  57432. 42:40:44different versions of your work, making
  57433. 42:40:46it easy to backtrack if something goes
  57434. 42:40:48wrong or to experiment without losing
  57435. 42:40:50progress. This is especially useful when
  57436. 42:40:53working with complex data science
  57437. 42:40:55projects where you might try out
  57438. 42:40:56different models of analysis techniques.
  57439. 42:40:59One or two weeks for practice along with
  57440. 42:41:01Python and R is good to get start. Now
  57441. 42:41:03moving on to the third step that is data
  57442. 42:41:06structures and algorithms. Learning data
  57443. 42:41:08structures and algorithms is crucial for
  57444. 42:41:10becoming a data scientist because they
  57445. 42:41:12provide the foundation for efficient
  57446. 42:41:13data handling and problem solving. Data
  57447. 42:41:16structures like arrays, stacks, cues and
  57448. 42:41:18trees help you store and organize data
  57449. 42:41:20in ways that make it easier and faster
  57450. 42:41:22to access, process and analyze.
  57451. 42:41:24Algorithms on the other hand give you
  57452. 42:41:26strategies to perform tasks like
  57453. 42:41:28searching, sorting and optimizing data
  57454. 42:41:30operations which are essential for
  57455. 42:41:32handling large data sets. While many
  57456. 42:41:34candidates struggle with the essay,
  57457. 42:41:36mastering it gives you an age helping
  57458. 42:41:38you stand out in the interviews and
  57459. 42:41:40shine as a skilled data scientist
  57460. 42:41:42capable of tackling the toughest data
  57461. 42:41:44problems. Spend about two months in
  57462. 42:41:46this, you will get in the shape for
  57463. 42:41:48sure. Now moving on to the step number
  57464. 42:41:50four that is SQL. Learning SQL is
  57465. 42:41:53essential for data scientists because it
  57466. 42:41:55enables you to access, manage, and
  57467. 42:41:57manipulate data directly within
  57468. 42:41:58databases where most real world data
  57469. 42:42:02resides. With SQL, you can create new
  57470. 42:42:04tables, alter existing ones, delete
  57471. 42:42:07unnecessary records, and run queries to
  57472. 42:42:10filter, sort, and aggregate data. These
  57473. 42:42:13abilities allow you to retrieve, clean,
  57474. 42:42:15and organize data effectively. Core
  57475. 42:42:18skills needed for any data science role.
  57476. 42:42:20It's easy and you don't have to spend
  57477. 42:42:22more than a month to have a deep
  57478. 42:42:24understanding of it. Now moving on to
  57479. 42:42:26the fifth step that is mathematics and
  57480. 42:42:28statistics. Mathematics and statistics
  57481. 42:42:30are essential for data science because
  57482. 42:42:32they form the backbone of data analysis,
  57483. 42:42:34model building and interpretation.
  57484. 42:42:36Topics like linear algebra, calculus,
  57485. 42:42:39probability and statistics gives data
  57486. 42:42:42scientists the tools to understand data
  57487. 42:42:44patterns, perform accurate analysis and
  57488. 42:42:46make datadriven decisions. Mastering
  57489. 42:42:48these areas enables you to build robust
  57490. 42:42:50models, validate results and tackle
  57491. 42:42:53complex problems confidently making you
  57492. 42:42:56a well-rounded and skilled data
  57493. 42:42:58scientist. Make sure you spend two
  57494. 42:43:00months to grasp this topics. Now moving
  57495. 42:43:02on to the step number six that is data
  57496. 42:43:04prep-processing and visualization.
  57497. 42:43:06Learning data prep-processing and
  57498. 42:43:08visualization is essential for a data
  57499. 42:43:10scientist because these skills make you
  57500. 42:43:13data accurate, insightful and easy to
  57501. 42:43:16understand. Python libraries like NumPy
  57502. 42:43:18and Panders are crucial for manipulating
  57503. 42:43:21and creating data, enabling you to
  57504. 42:43:23handle missing values, filter out noise,
  57505. 42:43:26and prepare data for analysis. Once the
  57506. 42:43:28data is ready, visualization lets you
  57507. 42:43:31uncover patterns and communicate results
  57508. 42:43:33effectively. Libraries like Mattplot tip
  57509. 42:43:36and Seaborn help create clear, impactful
  57510. 42:43:39visuals, allowing you to interpret
  57511. 42:43:41trends and convey insights in a way
  57512. 42:43:43that's easily understood by others.
  57513. 42:43:45Together with these tools make data
  57514. 42:43:47prep-processing and visualization
  57515. 42:43:49fundamentals for effective data science.
  57516. 42:43:51If you have a solid foundation on Python
  57517. 42:43:53and mathematics, you will get a good
  57518. 42:43:55understanding of data prep-processing
  57519. 42:43:56and visualization in a month or two. Now
  57520. 42:43:59moving on to the seventh step that is
  57521. 42:44:01machine learning fundamentals. Machine
  57522. 42:44:03learning fundamentals involve
  57523. 42:44:04understanding how algorithms enable
  57524. 42:44:06computers to learn from data and make
  57525. 42:44:08predictions on decisions without
  57526. 42:44:10explicit programming. The two main
  57527. 42:44:12categories are supervised learning and
  57528. 42:44:14unsupervised learning. In supervised
  57529. 42:44:16learning, models are trained on labelled
  57530. 42:44:18data to make predictions while in
  57531. 42:44:19unsupervised learning models find
  57532. 42:44:22patterns in unlabelled data. Popular
  57533. 42:44:25tools like TensorFlow, PyTorch help
  57534. 42:44:28build and train complex models
  57535. 42:44:30especially for deep learning. While
  57536. 42:44:32skyit learn is essential used for
  57537. 42:44:35simpler machine learning algorithms and
  57538. 42:44:37data prep-processing. These tools make
  57539. 42:44:39it easier to implement machine learning
  57540. 42:44:41fundamentals effectively and build
  57541. 42:44:44intelligent datadriven decisions.
  57542. 42:44:46Dedicate about three months to
  57543. 42:44:48understand the core of machine learning.
  57544. 42:44:50Now coming to the next step that is deep
  57545. 42:44:52learning. Deep learning is a subset of
  57546. 42:44:54machine learning that focuses on
  57547. 42:44:56algorithms inspired by the structures of
  57548. 42:44:58the human brain called neural networks.
  57549. 42:45:00Deep learning uses neural networks with
  57550. 42:45:02multiple layers often dozens or hundreds
  57551. 42:45:05to learn complex patterns from large
  57552. 42:45:07data sets. Specialized types like
  57553. 42:45:10convolutional neural networks that is
  57554. 42:45:12CNN's are great for image processing
  57555. 42:45:15while recurrent neural networks RNNs are
  57556. 42:45:18used for sequence data like text or time
  57557. 42:45:20series. Essential tools like TensorFlow,
  57558. 42:45:23PyTorch make building, training and
  57559. 42:45:25deploying deep learning models more
  57560. 42:45:27accessible, allowing you to create
  57561. 42:45:29powerful AI solutions across various
  57562. 42:45:32domains. I think it will take about 2
  57563. 42:45:34months to have a good hold on deep
  57564. 42:45:36learning concepts and how to implement
  57565. 42:45:38them. Now moving on to the ninth step
  57566. 42:45:40that is specializations. Once you have
  57567. 42:45:42grasped the deep learning, it's like
  57568. 42:45:44reaching a new level as a data
  57569. 42:45:46scientist. Just as doctors specialize in
  57570. 42:45:48areas in nephrology and cardiology, data
  57571. 42:45:52scientists often choose to specialize in
  57572. 42:45:54fields like natural language processing
  57573. 42:45:56or computer vision. Natural language
  57574. 42:45:59processing focuses on teaching machines
  57575. 42:46:02to understand and generate human
  57576. 42:46:03language enabling applications like
  57577. 42:46:06chatbot, sentiment analysis, and
  57578. 42:46:08language transition. It's about making
  57579. 42:46:11computers read, write, and even
  57580. 42:46:12interpret human emotions through text or
  57581. 42:46:15speech. Computer vision on the other
  57582. 42:46:17hand is all about enabling machines to
  57583. 42:46:19see and interpret images or videos. This
  57584. 42:46:22field powers innovations like facial
  57585. 42:46:24recognition, object detection and
  57586. 42:46:26autonomous driving. Now you don't need
  57587. 42:46:28to learn both. You can choose what
  57588. 42:46:30interests you the most. Now spend one to
  57589. 42:46:33two months diving deep into one of these
  57590. 42:46:35areas. Now moving on to the last but not
  57591. 42:46:38the least step that is big data. Big
  57592. 42:46:40data refers to extremely large volumes
  57593. 42:46:42of data generated rapidly from sources
  57594. 42:46:44like social media and sensors. For data
  57595. 42:46:47scientists, learning to handle big data
  57596. 42:46:49is crucial as it requires specialized
  57597. 42:46:52tools like Hadoop and Spark to analyze
  57598. 42:46:54and extract insights effectively. With
  57599. 42:46:57companies relying on datadriven
  57600. 42:46:58decisions, big data skills make you a
  57601. 42:47:00highly in- demand professional in the
  57602. 42:47:02field. Focus for about 2 months and you
  57603. 42:47:05will be able to spot trends and patterns
  57604. 42:47:06from data sets very easily. Once you're
  57605. 42:47:09ready, it's time to build a killer
  57606. 42:47:10resume packed with projects that
  57607. 42:47:12showcase your new skills. Start applying
  57608. 42:47:14to jobs on platforms like Noy and Indate
  57609. 42:47:17and supercharge your LinkedIn. Connect
  57610. 42:47:19with data scientists. See what skills
  57611. 42:47:21they are mastering and learn from their
  57612. 42:47:23journeys as well. Keep sharpening your
  57613. 42:47:25own skills and when the time comes, you
  57614. 42:47:27will be ready to crush those interviews
  57615. 42:47:29and land your dream data scientist role
  57616. 42:47:32in 2025.
  57617. 42:47:33>> Yeah. step by step we will go through
  57618. 42:47:35all of this and uh we'll make sure that
  57619. 42:47:38we learn everything and we bring
  57620. 42:47:39everything together towards the end
  57621. 42:47:41right without further ado let me just
  57622. 42:47:44straight away deep dive to business
  57623. 42:47:47right to learn data science
  57624. 42:47:54right
  57625. 42:47:56and with this data science there is also
  57626. 42:47:58something which is prefixed which is
  57627. 42:48:01applied data science
  57628. 42:48:06and suffix for this is with Python
  57629. 42:48:13right apply data science with Python
  57630. 42:48:15right so there are there are two key
  57631. 42:48:17concepts which are going to be a part of
  57632. 42:48:19this course the first one is the
  57633. 42:48:21knowledge about data science that what
  57634. 42:48:23data science is and then because we are
  57635. 42:48:26doing an applied course right we are
  57636. 42:48:27doing an applied course I will try to
  57637. 42:48:29tie up these concepts which we will
  57638. 42:48:32understand in data science with a tool
  57639. 42:48:35right which is Python for you right we
  57640. 42:48:37already know about uh 60 65% of Python
  57641. 42:48:42right which is the fundamental Python
  57642. 42:48:44and now we will be moving to the next
  57643. 42:48:46step to advanced Python
  57644. 42:48:52right and using Python right leveraging
  57645. 42:48:55Python we will be solving a lot of
  57646. 42:48:59problems of data science right using
  57647. 42:49:01this.
  57648. 42:49:03Okay. So, the first few sessions, right?
  57649. 42:49:06The first few sessions will be about
  57650. 42:49:09making you a breast with Python. What
  57651. 42:49:11Python is, right? What how and what
  57652. 42:49:14packages do we have? How do they work in
  57653. 42:49:16reality, right? And all those things.
  57654. 42:49:18And then we will be coupling it up with
  57655. 42:49:20data science concepts. And then finally
  57656. 42:49:22towards the end of the session in the
  57657. 42:49:24last few classes, we will be doing uh we
  57658. 42:49:27will be taking a real data set. And on
  57659. 42:49:29that data set we will be applying all
  57660. 42:49:31these concepts right to understand the
  57661. 42:49:33data better and we will be drawing
  57662. 42:49:35inferences from that to convert that
  57663. 42:49:38into information to take actionable
  57664. 42:49:40insights or using that actionable
  57665. 42:49:42insights taking a better decision.
  57666. 42:49:44Right? We'll do all of that in in the
  57667. 42:49:47actual way. Okay. So now guys if you
  57668. 42:49:51understand this then the next point of
  57669. 42:49:54contention is data sets
  57670. 42:49:58right one of the most famous keywords on
  57671. 42:50:02the planet right now right one of the
  57672. 42:50:04most famous keywords in the planet right
  57673. 42:50:05now do you think that these two things
  57674. 42:50:10okay let me put it different way what do
  57675. 42:50:12you think that can be the possible
  57676. 42:50:15explanation about this term data science
  57677. 42:50:18you know data You know science what do
  57678. 42:50:21you think is going to follow in these
  57679. 42:50:23sessions? What is data science to you as
  57680. 42:50:26per these two words? Okay. So this is
  57681. 42:50:29people made up of two words right data
  57682. 42:50:31and science right. So what we are trying
  57683. 42:50:35to do is we are
  57684. 42:50:40trying to understand data
  57685. 42:50:46right? We are trying to understand data
  57686. 42:50:48right and then do something to it right
  57687. 42:50:52understanding its science understanding
  57688. 42:50:54the uh nature the behavior of this data
  57689. 42:50:58and converting it in something called as
  57690. 42:51:01information
  57691. 42:51:04right do we know difference between data
  57692. 42:51:07and information data is something which
  57693. 42:51:10is completely raw okay it is completely
  57694. 42:51:14raw it has no meaning
  57695. 42:51:19right it has no meaning isn't it for
  57696. 42:51:21example
  57697. 42:51:23I give you these stats of some player
  57698. 42:51:27like suppose Sid Dhoni I give you stats
  57699. 42:51:29of Mahindra Singh Dhoni right that what
  57700. 42:51:31what what were his scores uh what is his
  57701. 42:51:34name what is his age and you know all
  57702. 42:51:37those things now everything is there
  57703. 42:51:40right but we don't know what to do about
  57704. 42:51:42it right do you think the score of Dhoni
  57705. 42:51:44has any context people it has any
  57706. 42:51:47context no right but when I deep down
  57707. 42:51:50but I when I go and deep dive about it
  57708. 42:51:52right what is the first thing you find
  57709. 42:51:54out of scores what is the first thing
  57710. 42:51:57you find out of score scores you try to
  57711. 42:52:00find the average of score isn't it that
  57712. 42:52:03in last 10 innings
  57713. 42:52:05right in last 10 innings before this
  57714. 42:52:07also you need something which is called
  57715. 42:52:09as a problem statement isn't it now for
  57716. 42:52:12example the problem statement is select
  57717. 42:52:15Selectors want to understand selectors
  57718. 42:52:17wants to understand that whether
  57719. 42:52:19Mahindra Singh Dhoni should be picked
  57720. 42:52:21up. So there's a problem now right
  57721. 42:52:24selectors want to see that whether Dhoni
  57722. 42:52:27is fit for the next tournament or not.
  57723. 42:52:31So what we will do we will now try to
  57724. 42:52:34take the mean of the scores right for
  57725. 42:52:40last 10 innings. And if this score is
  57726. 42:52:43suppose X, we will try to compare this
  57727. 42:52:45with Y. What is Y? Y is a reference,
  57728. 42:52:49right? Y is a reference that we want to
  57729. 42:52:52compare it against. Now when you are
  57730. 42:52:54doing this comparisons, when you are
  57731. 42:52:56applying these techniques to this score,
  57732. 42:52:59this is now slowly becoming information,
  57733. 42:53:03right? And at the end of the day once
  57734. 42:53:06you have the strike rate once you have
  57735. 42:53:09the mean score of Dhoni once you have
  57736. 42:53:12his age once you have his fitness score
  57737. 42:53:15all those things will now help you to
  57738. 42:53:18take this particular decision because
  57739. 42:53:21now what you have is called as
  57740. 42:53:23information right because this has
  57741. 42:53:25context
  57742. 42:53:28right this has meaning
  57743. 42:53:31and this is usually processed
  57744. 42:53:35Right? This is usually processed. Right?
  57745. 42:53:37This is usually processed. Now what did
  57746. 42:53:39we do? Now what did we do here? If you
  57747. 42:53:41will go and read about data science,
  57748. 42:53:43data science says,
  57749. 42:53:47data science says
  57750. 42:53:49it is the art of collecting,
  57751. 42:53:54right? cleaning,
  57752. 42:53:58analyzing,
  57753. 42:54:03modeling,
  57754. 42:54:08improving,
  57755. 42:54:12right? And visualizing,
  57756. 42:54:19right? Visualizing
  57757. 42:54:21the day, right? If a person is adept in
  57758. 42:54:25doing all these things, this person
  57759. 42:54:28people is cumulatively called a data
  57760. 42:54:30scientist. Right?
  57761. 42:54:33That person is called a data scientist.
  57762. 42:54:35Right? So before going to the definition
  57763. 42:54:37of data scientist, now I will give you
  57764. 42:54:39some more examples, right? I'll give you
  57765. 42:54:40some more examples. Data science people
  57766. 42:54:43as I said is a combination of these
  57767. 42:54:45things, right? You have to collect the
  57768. 42:54:47data,
  57769. 42:54:49right? You have to collect the data,
  57770. 42:54:52right? Right. And this has a lot of
  57771. 42:54:54things. Data can be connected from two
  57772. 42:54:56types in two types. One is primary
  57773. 42:55:01and the second one is secondary.
  57774. 42:55:05Right? What is the primary way of
  57775. 42:55:07collecting data? From your IoT devices,
  57776. 42:55:09right? From sensors,
  57777. 42:55:12from your inbuilt machines,
  57778. 42:55:16right? Then from your surveys
  57779. 42:55:20which you float, right? Questionnaires,
  57780. 42:55:25right? All these things are primary
  57781. 42:55:26ways. What is the secondary way of
  57782. 42:55:27collecting data?
  57783. 42:55:29Purchasing data,
  57784. 42:55:33right? Using internet data
  57785. 42:55:38because you have not generated it. You
  57786. 42:55:40are just using someone else's data.
  57787. 42:55:42Right? Something like uh transfer
  57788. 42:55:44learning.
  57789. 42:55:48What is transfer learning?
  57790. 42:55:50Transfer learning is a technique where
  57791. 42:55:52suppose I am bank A and you are bank B
  57792. 42:55:56right so bank A has created some model
  57793. 42:56:00right trained on their data now you are
  57794. 42:56:04going to use the exact same model right
  57795. 42:56:07you're going to use the exact same model
  57796. 42:56:09right maybe you're not seeing the data
  57797. 42:56:11but you are just using the property of
  57798. 42:56:13data like mean median mode and a lot of
  57799. 42:56:16modeling things which will come we will
  57800. 42:56:18learn about them and You use this model
  57801. 42:56:20on your particular data right so in a
  57802. 42:56:22way you did not have enough data to
  57803. 42:56:25create the model yourself but you are
  57804. 42:56:26now using someone else's model to run
  57805. 42:56:29your data on it right so this is called
  57806. 42:56:31as transfer learning so this kind of
  57807. 42:56:33collection is basically secondary data
  57808. 42:56:36collection so you can collect the data
  57809. 42:56:39right then you can perform data analysis
  57810. 42:56:46right you can perform data analysis
  57811. 42:56:48right How will you perform this data
  57812. 42:56:50analysis? Using complex
  57813. 42:56:53algorithms,
  57814. 42:56:57right? Using complex algorithms, right?
  57815. 42:56:59Some statistics,
  57816. 42:57:03right? You can use artificial
  57817. 42:57:06intelligence,
  57818. 42:57:12right? Artificial intelligence. You can
  57819. 42:57:14use machine learning.
  57820. 42:57:19Right.
  57821. 42:57:22Right. You can use all these things for
  57822. 42:57:24data analysis. Then you can transform
  57823. 42:57:31transform
  57824. 42:57:33the patterns
  57825. 42:57:39into predictions.
  57826. 42:57:43Right? You can transfer these patterns
  57827. 42:57:45into predictions, right? Which can be
  57828. 42:57:48used for business
  57829. 42:57:50decision making,
  57830. 42:57:55right? For business decision making.
  57831. 42:57:57Then you can validate the results,
  57832. 42:58:03right? And present the results,
  57833. 42:58:09right? So this is like a complete life
  57834. 42:58:11cycle of a data scientist, right? So
  57835. 42:58:14before going further, let me give you
  57836. 42:58:15what combinations do you need to have to
  57837. 42:58:18become a data scientist. The first one
  57838. 42:58:20is
  57839. 42:58:23domain knowledge,
  57840. 42:58:29right? So what is domain knowledge?
  57841. 42:58:32First of all, I told you right there
  57842. 42:58:33will be a problem, right? You'll be
  57843. 42:58:35solving a problem in any project of data
  57844. 42:58:38science. You'll be trying to solve a
  57845. 42:58:39problem right and the problem will be
  57846. 42:58:41belonging to a particular domain even if
  57847. 42:58:44you're working for yourself right even
  57848. 42:58:45if you're an entrepreneur then also
  57849. 42:58:47you'll be solving a problem. So this
  57850. 42:58:49domain knowledge part includes things
  57851. 42:58:51like understanding
  57852. 42:58:55it's a very important diagram
  57853. 42:58:58understanding
  57854. 42:59:00client requirement
  57855. 42:59:03right understanding the client
  57856. 42:59:05requirement right
  57857. 42:59:07important criterians
  57858. 42:59:12right important
  57859. 42:59:14criteria knowledge
  57860. 42:59:18Right? For example, to give an example,
  57861. 42:59:21suppose we have created a machine
  57862. 42:59:23learning model. Okay? Understanding the
  57863. 42:59:25data, we have created a machine learning
  57864. 42:59:27model whose accuracy is 90%. Right? Is
  57865. 42:59:3190% a good accuracy?
  57866. 42:59:34Yeah, fairly decent accuracy. Yes.
  57867. 42:59:37Suppose you have to predict sales,
  57868. 42:59:39right? You are selling something.
  57869. 42:59:40Suppose you are selling clothes and you
  57870. 42:59:42want to predict what will be the sales
  57871. 42:59:44for the next week. When you use this
  57872. 42:59:46model, whatever the output model gives
  57873. 42:59:48you, what is going to be the accuracy of
  57874. 42:59:50your output using this model?
  57875. 42:59:54How much accuracy?
  57876. 42:59:5790%.
  57877. 42:59:58But my question is model is 90%. But my
  57878. 43:00:02question to you is that if 90% accuracy
  57879. 43:00:05is on sales data, a person like me will
  57880. 43:00:08be very very happy. Okay? Very very
  57881. 43:00:10happy. I'll be probably dancing, right?
  57882. 43:00:13But if you try to apply the same model
  57883. 43:00:16right for a medical diagnosis case, will
  57884. 43:00:20you be interested in getting operated in
  57885. 43:00:23such a hospital or an institution where
  57886. 43:00:26the accuracy is coming as 90%. Domain
  57887. 43:00:29knowledge, right? Domain knowledge. We
  57888. 43:00:31need to understand what are the exact
  57889. 43:00:33requirements. We need to understand what
  57890. 43:00:36are the exact expectations,
  57891. 43:00:38right? And we need to know how much do
  57892. 43:00:41we need to pivot right? So the first
  57893. 43:00:43thing in data science is these
  57894. 43:00:45accuracies and everything are subjective
  57895. 43:00:47right they are subjective. So for that
  57896. 43:00:50you need domain knowledge. Domain
  57897. 43:00:52knowledge part very important guys very
  57898. 43:00:55important these three things which I'm
  57899. 43:00:56going to tell. Second part is people
  57900. 43:01:00the game changer right the second part
  57901. 43:01:02is
  57902. 43:01:04computer science.
  57903. 43:01:07Now if I take you back in history okay
  57904. 43:01:10if I take you back in history in 1980s
  57905. 43:01:13or somewhere then do you okay how many
  57906. 43:01:16of you think that data science is a new
  57907. 43:01:17concept
  57908. 43:01:19how many of you think that data science
  57909. 43:01:20is a new concept I hope you all it's not
  57910. 43:01:23a new concept everyone knows that yeah
  57911. 43:01:26it has been happening for ages just like
  57912. 43:01:30you guys will be shocked if you already
  57913. 43:01:32don't know AI was coined in the year
  57914. 43:01:341956
  57915. 43:01:361956 6 at the University of Dharma. AI
  57916. 43:01:40was coined by Paul McCarthy, right? And
  57917. 43:01:43we saw the boom of AI in the year 2010,
  57918. 43:01:48right? Such a long journey. Same case
  57919. 43:01:50with data science because people back in
  57920. 43:01:53the day data science was called as data
  57921. 43:01:56mining. Everyone heard about it data
  57922. 43:01:59mining, knowledge databases. Yeah, we
  57923. 43:02:02need we used to mine the data. Now, what
  57924. 43:02:05were the problems? What were the hiccups
  57925. 43:02:07of data mining? The hiccups for data
  57926. 43:02:09mining was that we were doing everything
  57927. 43:02:14everything manually
  57928. 43:02:18right now if I give you 100 points can
  57929. 43:02:21you calculate the mean
  57930. 43:02:23or let's say if I give you two points to
  57931. 43:02:26multiply 2 * 3 how much time will you
  57932. 43:02:30take?
  57933. 43:02:322 seconds.
  57934. 43:02:34Yep. How much time a computer will take?
  57935. 43:02:372 seconds. If I give you to multiply 2
  57936. 43:02:43489
  57937. 43:02:45multiplied by 200, how much time will
  57938. 43:02:47you take to calculate this? Say 5
  57939. 43:02:50seconds. How much time computer will
  57940. 43:02:52take? 2 seconds. Now if I give you to
  57941. 43:02:56multiply 2 48 9 into 15 1 95 4386
  57942. 43:03:02how much time will you take to calculate
  57943. 43:03:03this manually? Maybe say 1 minute
  57944. 43:03:0780 seconds 1 minute. How much time a
  57945. 43:03:09computer will take? Still 2 seconds
  57946. 43:03:12right? still 2 seconds right and if I
  57947. 43:03:15give you to calculate this over 200
  57948. 43:03:17times you will take 200 minutes right
  57949. 43:03:21using parallel computing computer will
  57950. 43:03:23still take about 3 to 5 seconds right so
  57951. 43:03:27are you understanding the power of
  57952. 43:03:28computer do you understand this concept
  57953. 43:03:31in this relationship what was happening
  57954. 43:03:33back then what computer science did
  57955. 43:03:36people was it revolutionized the way
  57956. 43:03:39data mining was happening and That thing
  57957. 43:03:42now is called as that again that thing
  57958. 43:03:46now is called as data science in which
  57959. 43:03:48computer science is one of the most
  57960. 43:03:50important contributors. So this is just
  57961. 43:03:52one reason. Now manually right manually
  57962. 43:03:55if I give you say 1 million rows of data
  57963. 43:03:59right 1 million rows of data right so
  57964. 43:04:02how many pages will you pages will you
  57965. 43:04:04need to store this data suppose your
  57966. 43:04:07notebook is like this
  57967. 43:04:09right these boxes and here you are
  57968. 43:04:11storing the data 1 million so maybe you
  57969. 43:04:13can buy n number of notebooks but now do
  57970. 43:04:17you think it's as easy in storing
  57971. 43:04:19something in computer because back in
  57972. 43:04:21the day people we had memory issues,
  57973. 43:04:25isn't it? Memory constraints.
  57974. 43:04:30So there is something called as Murray's
  57975. 43:04:32law, right? Which says as the
  57976. 43:04:35advancement in microprocessors will
  57977. 43:04:38increase, the price of microprocessor
  57978. 43:04:41will decrease, right? So this is what is
  57979. 43:04:42happening right now. Back in 1980s, if I
  57980. 43:04:46show you right guys, right? Yeah. 2.5
  57981. 43:04:50kg. Exactly. Right. It was size of a
  57982. 43:04:51fridge hard disk but now it fits in your
  57983. 43:04:54palm right. So this was enabled the
  57984. 43:04:57storage techniques right the processing
  57985. 43:04:59techniques the infrastructure right
  57986. 43:05:02things like big data what kind of data
  57987. 43:05:04do you think we will be dealing with
  57988. 43:05:06people in data science you all know the
  57989. 43:05:08term very famous term the kind of data
  57990. 43:05:12big data right everyone knows about big
  57991. 43:05:14data what is big data yes a data which
  57992. 43:05:17is fast right it has velocity veracity
  57993. 43:05:23variety right so This kind of data needs
  57994. 43:05:26to be stored. This kind of data needs to
  57995. 43:05:28be processed. So which thing brought all
  57996. 43:05:31these things into data science? It was
  57997. 43:05:33given to us by computer science, right?
  57998. 43:05:36So computer science people included
  57999. 43:05:38things like database management,
  58000. 43:05:44right? Data validation, right? Data
  58001. 43:05:48infrastructure,
  58002. 43:05:51right? Data infrastructure,
  58003. 43:05:53right? Then we had languages, computer
  58004. 43:05:57languages
  58005. 43:06:00which is Python right now for us. Right?
  58006. 43:06:04Again do you think when you do this
  58007. 43:06:07thing manually right suppose you do this
  58008. 43:06:09thing manually how easy do you think it
  58009. 43:06:11will become using something like Python
  58010. 43:06:14or any other computer language to create
  58011. 43:06:16complex models. How easy it will be to
  58012. 43:06:19do that to create the complexity in
  58013. 43:06:22models right where you can capture the
  58014. 43:06:24nonlinear nature isn't it people isn't
  58015. 43:06:28it
  58016. 43:06:30for example for example let me tell you
  58017. 43:06:33this
  58018. 43:06:342 4 6 8 10 dash what do you think is the
  58019. 43:06:40next number guys 12 if I tell you to
  58020. 43:06:43define this to me in okay let leave
  58021. 43:06:47leave
  58022. 43:06:48What do you think is going to be the
  58023. 43:06:49next number here?
  58024. 43:06:57What is the next number? 11.
  58025. 43:07:04Next number
  58026. 43:07:0825.
  58027. 43:07:09Perfect. Now guys, if I ask you to write
  58028. 43:07:12these numbers, right, the way you
  58029. 43:07:14predicted them, can you give me a
  58030. 43:07:15function f ofx is equal to what is the f
  58031. 43:07:19of x here?
  58032. 43:07:21It's 2x, right? It's 2x. f ofx is equal
  58033. 43:07:24to 2x. If I tell you to create a
  58034. 43:07:27function, it will be f of x is equal to
  58035. 43:07:292x. What will be the function here,
  58036. 43:07:32guys?
  58037. 43:07:34f ofx
  58038. 43:07:36will be equal to
  58039. 43:07:39x + 1. Yeah. x + 1. Yeah. And here f ofx
  58040. 43:07:46will be equal to
  58041. 43:07:48x². Yeah. Now the last example, right?
  58042. 43:07:52Last example.
  58043. 43:08:00What is the next number here? I don't
  58044. 43:08:02want the number. I want the function. I
  58045. 43:08:04want this so that I can generalize.
  58046. 43:08:07Isn't it? How did you reach this figure?
  58047. 43:08:11How many of you think it's not possible
  58048. 43:08:13to determine this? How many
  58049. 43:08:17of you
  58050. 43:08:20think
  58051. 43:08:23it is not
  58052. 43:08:26possible
  58053. 43:08:28to determine this?
  58054. 43:08:33Yeah. How many of you think what if I
  58055. 43:08:36just change this question and ask you
  58056. 43:08:39how many of you think it is not possible
  58057. 43:08:41to determine this
  58058. 43:08:44manually?
  58059. 43:08:47Same response. But now I say how many of
  58060. 43:08:50you think it is not possible to
  58061. 43:08:52determine this
  58062. 43:08:54with computers?
  58063. 43:08:56Will your answer still remain no? Do you
  58064. 43:08:59think I cannot approximate this function
  58065. 43:09:01using computers?
  58066. 43:09:03We have something called as deep
  58067. 43:09:07neural networks
  58068. 43:09:10and they are called as universal
  58069. 43:09:12function approximators.
  58070. 43:09:14Right? So this is the problem people.
  58071. 43:09:17This is the problem. Right? I will show
  58072. 43:09:18this to you when the time comes. Right?
  58073. 43:09:20I will remember this example and I will
  58074. 43:09:22show this to you. But now what I'm
  58075. 43:09:24trying to tell you is the things which
  58076. 43:09:26seemed impossible manually was solved by
  58077. 43:09:30what? It was solved by computers. The
  58078. 43:09:33distribution of this is like this. Can
  58079. 43:09:36you figure it out yourself? No. Right?
  58080. 43:09:40We cannot. Isn't it? We cannot do that.
  58081. 43:09:43So this kind of approximation will be
  58082. 43:09:45given by what? It will be only given by
  58083. 43:09:48machines. Right? And this is people what
  58084. 43:09:50data science is all about. Right? It is
  58085. 43:09:53what computer science did inside data
  58086. 43:09:56science. Right? I hope this is clear.
  58087. 43:09:59I'm assuming a lot of you will be going
  58088. 43:10:01for interviews and everything after
  58089. 43:10:03these course. Right? So this will be a
  58090. 43:10:05very very important thing for you to
  58091. 43:10:07know. Right? Often it is asked why data
  58092. 43:10:09science is having computer science in
  58093. 43:10:11it. Right? The reason is this. Okay? So
  58094. 43:10:14this is the role of computer science
  58095. 43:10:16inside data science. Now people the
  58096. 43:10:18third thing the third circle which is
  58097. 43:10:21one of the most
  58098. 43:10:23parts is
  58099. 43:10:26maths
  58100. 43:10:28and stats right mathematics statistics
  58101. 43:10:33which was optimization
  58102. 43:10:37right optimization of your models right
  58103. 43:10:42design
  58104. 43:10:45of model. Right? Now guys, if you look
  58105. 43:10:49carefully, if you look carefully, in
  58106. 43:10:52order to approximate this, what you what
  58107. 43:10:54will you be playing with? You will be
  58108. 43:10:56playing with a lot of data. You'll be
  58109. 43:11:00playing with a lot of mathematical to
  58110. 43:11:02mathematical concepts and statistical
  58111. 43:11:04concepts, isn't it? How did you what do
  58112. 43:11:07you call this? This is math, right? This
  58113. 43:11:09is statistics and mathematics. finding
  58114. 43:11:12mean, median, mode, standard deviations,
  58115. 43:11:15probability, statistics, all these will
  58116. 43:11:18lead to this kind of result, isn't it?
  58117. 43:11:20So, this becomes the third wheel of this
  58118. 43:11:23particular uh of of this particular
  58119. 43:11:25diagram. And this point of intersection,
  58120. 43:11:30right? The sweet point of intersection
  58121. 43:11:33is basically data science.
  58122. 43:11:39Yeah. is particularly data science.
  58123. 43:11:44Right? So now people this point okay
  58124. 43:11:47this point
  58125. 43:11:49is basically representing data
  58126. 43:11:53engineering right data engineering right
  58127. 43:11:56data engineering is people the part of
  58128. 43:11:59data science which enables us to capture
  58129. 43:12:02the correct data right how the data will
  58130. 43:12:05flow how the data will be stored how the
  58131. 43:12:08data will be cleaned right all this is
  58132. 43:12:10done by home it is done by data
  58133. 43:12:13engineering Just because so am I right
  58134. 43:12:15there you'll go for interview right
  58135. 43:12:16after this and try to fetch yourself
  58136. 43:12:18jobs in this domain data science AI ML
  58137. 43:12:23if my understanding is correct is that
  58138. 43:12:25the aim yes so now guys there will be
  58139. 43:12:27three types of companies
  58140. 43:12:29or let's say to simplify let's say two
  58141. 43:12:31types one is small and the other one is
  58142. 43:12:35big right so in a small organization if
  58143. 43:12:39you become a part of small or
  58144. 43:12:40organization and you are the data
  58145. 43:12:42scientist there You can be involved in
  58146. 43:12:45all of these things, right? All of these
  58147. 43:12:48things possible, right? Your bosses and
  58148. 43:12:51your management will expect you to
  58149. 43:12:53construct all these flows, right? Know
  58150. 43:12:56computer science, you should know maths
  58151. 43:12:58and stats and you should have the domain
  58152. 43:12:59knowledge and you will be asked to do
  58153. 43:13:01all of this. But if you are going to
  58154. 43:13:04become a part of a big organization,
  58155. 43:13:06usually all these roles are fragmented.
  58156. 43:13:09All of these roles are fragmented,
  58157. 43:13:12right? There's a separate data
  58158. 43:13:13infrastructure team. There's a separate
  58159. 43:13:15data governance team. Now guys, when you
  58160. 43:13:17go on to collect the data, can you
  58161. 43:13:19collect any sensitive data about is it
  58162. 43:13:21possible ethically it's not right? And
  58163. 43:13:24legally also it's not right. My question
  58164. 43:13:26to you is who will look after this
  58165. 43:13:28compliance? Whose responsibility indeed
  58166. 43:13:31it is to look after this compliance?
  58167. 43:13:33Data scientist. So this is about
  58168. 43:13:36fragmentation. If you are part of a big
  58169. 43:13:38organization, this thing will be taken
  58170. 43:13:40by someone else, right? But if you are a
  58171. 43:13:42part of a small organization, you will
  58172. 43:13:44be know you'll be expected to do this
  58173. 43:13:46all by yourself. In if if you are part
  58174. 43:13:48of a big organization then do you think
  58175. 43:13:50you need to have this domain knowledge?
  58176. 43:13:53The answer is no. Why? Because there
  58177. 43:13:56will be separate set of people who are
  58178. 43:13:57called as what? Who are called as
  58179. 43:14:00business analyst. Have you heard about
  58180. 43:14:01this position people? Business analyst.
  58181. 43:14:05What is a business analyst role? It is
  58182. 43:14:07basically a technical translator, right?
  58183. 43:14:10who knows technical, who knows domain
  58184. 43:14:12and that person goes and talks to the
  58185. 43:14:14client, talks to the client in a layman
  58186. 43:14:16language, convert it into technical
  58187. 43:14:18requirement coupled with the domain
  58188. 43:14:20knowledge and give you the document.
  58189. 43:14:22This is basically a medical engineering
  58190. 43:14:24problem. So we need this this this this
  58191. 43:14:27and you need to fulfill this this this
  58192. 43:14:29this criteria. But again, if you're part
  58193. 43:14:31of a small organization, who needs to
  58194. 43:14:33take care of that? Who needs to make
  58195. 43:14:35sure that you know everything about a
  58196. 43:14:36domain? You yourself, right? You
  58197. 43:14:40yourself, right? Then people, this area
  58198. 43:14:45usually represents whom? This area
  58199. 43:14:49represents research
  58200. 43:14:52and analysis.
  58201. 43:14:55Why?
  58202. 43:14:56Because they we have people who have
  58203. 43:14:58domain knowledge and we have people who
  58204. 43:15:00are knowledge of math, stats. Have you
  58205. 43:15:02heard about a position called as actury
  58206. 43:15:05in the world? Acturial science. Acturies
  58207. 43:15:08are people who are basically dealing
  58208. 43:15:11with uh domains which are very very
  58209. 43:15:14heavily data intrinsic. Right? For
  58210. 43:15:16example, finance domain, right? Finance
  58211. 43:15:18is all about numbers. So there we go
  58212. 43:15:20actal science and we have math, stats,
  58213. 43:15:23optimization, model development, all of
  58214. 43:15:25those things happening there. We are
  58215. 43:15:26also at at this point of time people
  58216. 43:15:28belonging to which section research and
  58217. 43:15:30data analysis right and then people
  58218. 43:15:34there's the third intersection right
  58219. 43:15:35there's a third the third intersection
  58220. 43:15:39which is this part and this people is
  58221. 43:15:42called as machine learning
  58222. 43:15:48right machine learning why machine
  58223. 43:15:50learning if you can combine the power of
  58224. 43:15:53maths stats right and you Combine the
  58225. 43:15:56power of computer science, you will find
  58226. 43:15:59yourself to be in a position where you
  58227. 43:16:01can call yourself a machine learning
  58228. 43:16:03engineer. Right? How does a machine
  58229. 43:16:05learning engineer becomes a data
  58230. 43:16:06scientist? When they couple it up with
  58231. 43:16:09the domain expertise, right? So in this
  58232. 43:16:12course people in this course we will
  58233. 43:16:14teach you computer science. We will
  58234. 43:16:16teach you a little bit about math stats.
  58235. 43:16:19But what we cannot teach you is domain
  58236. 43:16:21knowledge. Yeah.
  58237. 43:16:24and making the base of what we are about
  58238. 43:16:26to do. Very very important to
  58239. 43:16:28understand. Right? If you understand
  58240. 43:16:29this then half the battle is won. Right?
  58241. 43:16:33So now to answer the question which was
  58242. 43:16:36posted earlier uh answer to the question
  58243. 43:16:38which was posted earlier. There are
  58244. 43:16:40different things in the world of data
  58245. 43:16:41science. Right? So you can pick and
  58246. 43:16:43choose anything or you can do everything
  58247. 43:16:46by yourself. If you guys are engineers
  58248. 43:16:48then I think you can be at the sweet
  58249. 43:16:50spot going forward in life if you choose
  58250. 43:16:52a domain for yourself. For example, you
  58251. 43:16:54choose to be a part of automobile
  58252. 43:16:56industry, you choose to be a part of
  58253. 43:16:57medical industry, you choose to be a
  58254. 43:16:59part of say retail industry, you choose
  58255. 43:17:03to be a part of aeronautics industry,
  58256. 43:17:06right? You choose to be a part of
  58257. 43:17:08finance industry, right? So whatever you
  58258. 43:17:10will choose, this thing will get
  58259. 43:17:12developed over time, right? This is the
  58260. 43:17:14most difficult out of these three. I try
  58261. 43:17:16to give you the example right my domain
  58262. 43:17:19was agricultural industry right the agri
  58263. 43:17:22products I have worked extensively in
  58264. 43:17:25agricultural industry right so again for
  58265. 43:17:28now in my current role this is something
  58266. 43:17:30which I don't have I have this expertise
  58267. 43:17:32I have this expertise so same will be
  58268. 43:17:34with you and you guys will develop this
  58269. 43:17:36knowledge over the time now I will try
  58270. 43:17:39to give you an example right elections
  58271. 43:17:42to make you understand how data science
  58272. 43:17:45can be used in one particular use case.
  58273. 43:17:48Okay, maybe we can extend that to a lot
  58274. 43:17:50of other use cases and examples, right?
  58275. 43:17:53Talking about election season, right?
  58276. 43:17:55Talking about the election season,
  58277. 43:17:57right? We will we will try to understand
  58278. 43:18:00how do we use data science because it is
  58279. 43:18:03very extensively used in this data
  58280. 43:18:05science. Right? Now, let me talk about
  58281. 43:18:09the first phase. Let's say this is
  58282. 43:18:12pre-election
  58283. 43:18:13phase,
  58284. 43:18:15right?
  58285. 43:18:17Right. This is pre-election phase. In
  58286. 43:18:20this pre-election phase, what do you
  58287. 43:18:22think will be the tasks with which an
  58288. 43:18:24agency like Election Commission of India
  58289. 43:18:26will be doing? The first task can be
  58290. 43:18:30that they will be
  58291. 43:18:32doing the voter
  58292. 43:18:34registration,
  58293. 43:18:37right? Voter registration and data
  58294. 43:18:40management, isn't it?
  58295. 43:18:44Yeah, it will start with that.
  58296. 43:18:47And what will be the things inside this?
  58297. 43:18:49The first thing will be data collection,
  58298. 43:18:54right? First thing will be data
  58299. 43:18:55collection. So, you will collect the
  58300. 43:18:57data from all the registered voters,
  58301. 43:19:00right? Maybe it can be their demographic
  58302. 43:19:02information, where they live, what is
  58303. 43:19:04their age, what is their gender, right?
  58304. 43:19:07What is their past polling behavior?
  58305. 43:19:09Have they turned out previously or not?
  58306. 43:19:11Right? All these details we can collect.
  58307. 43:19:14Then can we also do data cleaning?
  58308. 43:19:19Because I've told you, right? That there
  58309. 43:19:21can be a lot of redundancies. Some
  58310. 43:19:23person's name can appear twice, right?
  58311. 43:19:26Some people can be a mismatch. Suppose
  58312. 43:19:28for example, we have learned this in
  58313. 43:19:30Python. Raghav.
  58314. 43:19:36Raghav.
  58315. 43:19:39Raghav.
  58316. 43:19:42Radha. Right.
  58317. 43:19:45Right. All these are what people?
  58318. 43:19:48This is belonging to the same name.
  58319. 43:19:50Right. This is me. But for a computer,
  58320. 43:19:53for a computer, how many ragavves are
  58321. 43:19:55there? Different ones. All are
  58322. 43:19:56different. Right? So example like these,
  58323. 43:19:59right? Some people might have died. They
  58324. 43:20:01might not be existing anymore. Right? So
  58325. 43:20:03all this part will be taken care where
  58326. 43:20:05people in the data cleaning process.
  58327. 43:20:08Right? Removing the duplicates, updating
  58328. 43:20:10the new addresses, right? Correcting the
  58329. 43:20:13information about every voter, all those
  58330. 43:20:15things, right? And now lastly, we can
  58331. 43:20:18also include a flavor of data analytics,
  58332. 43:20:24right? What will data analytics include
  58333. 43:20:25in this? We can analyze the demographic
  58334. 43:20:28data to identify eligible voters. Isn't
  58335. 43:20:32it? Yeah. People till now we haven't
  58336. 43:20:35understood this why it is not automated.
  58337. 43:20:38How will you pick up all the how how
  58338. 43:20:40will you pick up all the details and
  58339. 43:20:42nuances? Suppose you're filling a form
  58340. 43:20:44right by mistake you have. Suppose you
  58341. 43:20:46are 25 years of age. Suppose you have
  58342. 43:20:48written 250. So does that mean I remove
  58343. 43:20:51this? I remove this entry of yours
  58344. 43:20:54because by mistake you have written your
  58345. 43:20:56age as 250
  58346. 43:20:58is age 250 possible in our current world
  58347. 43:21:01never right so I will have to tell the
  58348. 43:21:04machine that because this is a mistake
  58349. 43:21:07please convert it to 25 isn't it this is
  58350. 43:21:11called as imputation so this is mostly a
  58351. 43:21:14manual task not manual but you have to
  58352. 43:21:16understand the problem manually and then
  58353. 43:21:20code it on
  58354. 43:21:21Suppose someone has written their state
  58355. 43:21:26as
  58356. 43:21:28E D L H I right and country
  58357. 43:21:35as I D I N right so what is this state
  58358. 43:21:41referring to what is this country
  58359. 43:21:42referring to the humans are very smart
  58360. 43:21:46India right we can say this is India and
  58361. 43:21:48if this is India then what is this
  58362. 43:21:50pointing out to
  58363. 43:21:52telly just because you said and it's a
  58364. 43:21:55it's a leading question I want you to
  58365. 43:21:56explain this tell me is it possible to
  58366. 43:21:59do this automatically no right we have
  58367. 43:22:02to employ manual rules right we have to
  58368. 43:22:05tell because we who is more intelligent
  58369. 43:22:08humans or machines humans right machines
  58370. 43:22:11are just more optimized right so we know
  58371. 43:22:14through human intelligence that this is
  58372. 43:22:16pointing to Delhi and this is not EDLHI
  58373. 43:22:19so this is data cleaning Right. Lastly,
  58374. 43:22:21we have data analytics. So, do you think
  58375. 43:22:23people based on these data points, we
  58376. 43:22:26can understand that who are the eligible
  58377. 43:22:30voters
  58378. 43:22:32and maybe who are not registered yet,
  58379. 43:22:34maybe who have not voted in the past.
  58380. 43:22:37Can we do all those analytics
  58381. 43:22:40and reach out to those peoples and
  58382. 43:22:42persons? That's the first part. This
  58383. 43:22:45just the first part pre-election phase.
  58384. 43:22:47Now moving on to the second part right
  58385. 43:22:51moving on to the second part let's say
  58386. 43:22:54uh we say public
  58387. 43:22:58opinion
  58388. 43:23:02regarding the polling right the polling
  58389. 43:23:04which is about to happen the first thing
  58390. 43:23:06will be you want to collect information
  58391. 43:23:08about people right so can can you go out
  58392. 43:23:12and reach all 1.8 8 billion people in
  58393. 43:23:16this country that what is their likable
  58394. 43:23:20vote for which party is it possible
  58395. 43:23:241.8 8 billion do you think it's possible
  58396. 43:23:26for 1 billion
  58397. 43:23:28do you think it's possible for 500
  58398. 43:23:30million do you think it's possible for
  58399. 43:23:32100 million no right so basically I'm
  58400. 43:23:36talking about what I have something
  58401. 43:23:38which is called as a population
  58402. 43:23:42right and if I have to study about this
  58403. 43:23:44population which is 1.8 8 billion people
  58404. 43:23:48which is impossible which you just said
  58405. 43:23:51what do I need to do should I stop my
  58406. 43:23:52process no right I will go and collect
  58407. 43:23:56something which is called a sample
  58408. 43:24:01right we always work in samples right
  58409. 43:24:05suppose someone says that a Coca-Cola
  58410. 43:24:08bottle does not contain 500 ml of liquid
  58411. 43:24:13which it claims right suppose someone
  58412. 43:24:15has put this allegation possible that
  58413. 43:24:17Coca-Cola bottles do not have 500 ml
  58414. 43:24:21liquid which they promise. Now there are
  58415. 43:24:23two ways to deal with this. Right? There
  58416. 43:24:25are two ways to deal with this. Either I
  58417. 43:24:27go and collect all the bottles of
  58418. 43:24:30Coca-Cola in the world. Possible
  58419. 43:24:33never right. So what will I do? I will
  58420. 43:24:36go and pick up handful of bottles.
  58421. 43:24:39Right? Handful of bottles. So what is
  58422. 43:24:41that handful of bottles? Those are
  58423. 43:24:43called as samples. One last thing.
  58424. 43:24:45Suppose someone says that because of an
  58425. 43:24:48industry
  58426. 43:24:49all the fishes of the lake are dying or
  58427. 43:24:53they are infected. Is it possible to go
  58428. 43:24:55and collect and check all the fishes in
  58429. 43:24:57the pond or a lake? No. Right. What will
  58430. 43:25:00we do? We will collect again handful of
  58431. 43:25:03fishes and we will test them. Right?
  58432. 43:25:06Again samples. Now how does the raw data
  58433. 43:25:10collected? Raw data as in I hope you
  58434. 43:25:13understand this. This is no more about
  58435. 43:25:16population
  58436. 43:25:17with this step being told. Now you're
  58437. 43:25:19dealing with samples. So now do you want
  58438. 43:25:22to ask me how is sample created? Yeah,
  58439. 43:25:24now I'm coming to that. Now guys, there
  58440. 43:25:26are a lot of second point is how to
  58441. 43:25:30sample right? How to sample
  58442. 43:25:34right? So we have sampling techniques
  58443. 43:25:36people. One is called as probabilistic
  58444. 43:25:38and one is called as nonrobabilistic.
  58445. 43:25:42Right? I will not go in detail right
  58446. 43:25:44now. I just want to tell you an overview
  58447. 43:25:46probabilistic is suppose uh you are
  58448. 43:25:48manufacturing t-shirts right you are
  58449. 43:25:50manufacturing t-shirts right and suppose
  58450. 43:25:53you created
  58451. 43:25:55100 lots
  58452. 43:25:58of
  58453. 43:26:00thousand t-shirts right so this is box
  58454. 43:26:03one box two box three box four up to up
  58455. 43:26:06to 100 right 100 boxes and in each box
  58456. 43:26:09how many t-shirts are there 1 th00and
  58457. 43:26:12right now suppose you are a Quality
  58458. 43:26:14inspector. You're a quality inspector.
  58459. 43:26:17Is it possible you for you to go over
  58460. 43:26:20all the 100 lots with all checking all
  58461. 43:26:23the thousand t-shirts one by one? No.
  58462. 43:26:25Right. What will you do? You will sample
  58463. 43:26:27again. You will sample. Now the most
  58464. 43:26:31common way of sampling these kind of
  58465. 43:26:34problems is probabilistic sampling. What
  58466. 43:26:37is probability? What is the probability
  58467. 43:26:38of getting heads or a tails when you
  58468. 43:26:40spin the when you flip the coin? equally
  58469. 43:26:42likely 1x2 and 1x2. What is the
  58470. 43:26:45probability of getting 1 2 3 4 5 6 on a
  58471. 43:26:48roll of a dice? 1x 6. Now what is the
  58472. 43:26:51probability of picking any t-shirt from
  58473. 43:26:54this first slot out of thousand
  58474. 43:26:56t-shirts?
  58475. 43:26:571 by,000.
  58476. 43:27:00Yes. So do you think all the t-shirts
  58477. 43:27:02have equally probable equal probability
  58478. 43:27:04of being picked up without any bias? If
  58479. 43:27:08you decide to draw five t-shirts, right,
  58480. 43:27:11from each of this box, right, and
  58481. 43:27:15suppose say two are defective and three
  58482. 43:27:19are not defective, what will you do?
  58483. 43:27:21Will you accept the lot or reject the
  58484. 43:27:23lot? We have majority of t-shirts of
  58485. 43:27:25non-deective
  58486. 43:27:27or let's say we will reject the lot. We
  58487. 43:27:29will reject the lot. We will reject the
  58488. 43:27:30lot. Let's say we will reject the lot.
  58489. 43:27:32Okay. Though this is basically
  58490. 43:27:34subjective as per the company policies
  58491. 43:27:37but let's say we rejected. Now people my
  58492. 43:27:41question to you is what if this entire
  58493. 43:27:44batch had only two defected t-shirts but
  58494. 43:27:47now what will happen? The entire batch
  58495. 43:27:49will be rejected.
  58496. 43:27:52Yes. Let's say you sampled one t-shirt.
  58497. 43:27:55Let's let's change the use case. I say
  58498. 43:27:58you sampled only one t-shirt and that
  58499. 43:28:00t-shirt was defected. Now you will
  58500. 43:28:02reject the batch and that is equally
  58501. 43:28:04likely case. So this is called as
  58502. 43:28:07probabilistic sampling people and there
  58503. 43:28:09is no way you can go back. There is no
  58504. 43:28:12way you cannot say that hey sir please
  58505. 43:28:15uh allow this batch to pass because
  58506. 43:28:17there is a chance that rest of the
  58507. 43:28:19t-shirts are not defective. No it is not
  58508. 43:28:21the way that happens. It happens
  58509. 43:28:24randomly. So right this is called as
  58510. 43:28:26random sampling.
  58511. 43:28:30Right? random sampling. Now suppose you
  58512. 43:28:34are doing a cancer research, right?
  58513. 43:28:37You're doing a cancer research, right?
  58514. 43:28:40So for your cancer research people, what
  58515. 43:28:42kind of people will you need? People who
  58516. 43:28:45had had cancer in the past, isn't it? To
  58517. 43:28:48know more about their problem, to know
  58518. 43:28:50more about their medical condition. So
  58519. 43:28:52is it possible people that in this use
  58520. 43:28:54case you can go and pick up any person
  58521. 43:28:57from the population and ask them
  58522. 43:28:59questions? No. Right? That is not
  58523. 43:29:02possible. So now is the probability
  58524. 43:29:05equally likely or it has changed when
  58525. 43:29:07you pick the sample? It has changed. Now
  58526. 43:29:10there is a bias which is introduced that
  58527. 43:29:13you only want people who had cancer.
  58528. 43:29:16Right? So that kind of sampling people
  58529. 43:29:19is called as nonprobabilistic sampling.
  58530. 43:29:22Right? Non-robabilistic sampling. Clear?
  58531. 43:29:26Now sampling technique. Right? Now third
  58532. 43:29:30thing in this same scheme can be people
  58533. 43:29:32what? It can be the data collection
  58534. 43:29:37right data collection mode right that
  58535. 43:29:40how do you collect the data? You can
  58536. 43:29:42float a survey
  58537. 43:29:45on say internet.
  58538. 43:29:49You can go and stand outside a mall
  58539. 43:29:53or office,
  58540. 43:29:56isn't it? How will you how will you can
  58541. 43:29:58probably interview someone,
  58542. 43:30:01right? Interview someone, right? You can
  58543. 43:30:03have a group discussion.
  58544. 43:30:06Yeah. All these techniques.
  58545. 43:30:09Yes. No, maybe. Right. In the same part,
  58546. 43:30:11public opinion polling, right? Now,
  58547. 43:30:14guys, uh so this was a brief
  58548. 43:30:16introduction, right? And this can be
  58549. 43:30:18extended to any industry. Right? As of
  58550. 43:30:22now, you can have example in the medical
  58551. 43:30:25science,
  58552. 43:30:29right? You can have an example in
  58553. 43:30:31automobile,
  58554. 43:30:34right? You can have an example in
  58555. 43:30:37retail,
  58556. 43:30:39right? Right. Then you can have example
  58557. 43:30:43in manufacturing.
  58558. 43:30:47Right? You can have example in
  58559. 43:30:49education,
  58560. 43:30:52right? You can have example in sports,
  58561. 43:30:56right? IPL analysis, cricket analysis,
  58562. 43:30:59all these sports analysis, right? These
  58563. 43:31:01are the most famous domains, right? They
  58564. 43:31:04are the most famous domains. They are
  58565. 43:31:06not topics, they are domains, right? In
  58566. 43:31:08which data science is used extensively.
  58567. 43:31:10So, I'm just going to check. Guys, in
  58568. 43:31:12automobiles, there's a biggest example,
  58569. 43:31:15self-driving cars.
  58570. 43:31:18Yeah, autonomous driving. How do you
  58571. 43:31:21think that's possible? Data science
  58572. 43:31:24again
  58573. 43:31:26like Tesla. Absolutely. Tesla is level
  58574. 43:31:29three. We have five levels.
  58575. 43:31:32Level three is narrow AI. Level four is
  58576. 43:31:36AGI and level five is super AI. We are
  58577. 43:31:39going to move first right with the
  58578. 43:31:43technical aspect right with the
  58579. 43:31:45technical aspect in our data science
  58580. 43:31:47course right
  58581. 43:31:57right which is based on Python
  58582. 43:32:00right because that is our base language
  58583. 43:32:02which we have learned so far right so in
  58584. 43:32:06Python people we will start and cover
  58585. 43:32:08four of the packages
  58586. 43:32:12right now and as we move on to machine
  58587. 43:32:15learning and other uh deep learning and
  58588. 43:32:17everything you will explore more and
  58589. 43:32:19more packages. The first package we have
  58590. 43:32:21to cover will be numpy.
  58591. 43:32:24Right? I'll explain you in detail what
  58592. 43:32:26numpy is. Then we will cover pandas.
  58593. 43:32:32Then we will cover mattplot lip
  58594. 43:32:36and then finally we will cover something
  58595. 43:32:38called as cbond.
  58596. 43:32:41Right? We will cover something called as
  58597. 43:32:43cbond. So these four packages inherently
  58598. 43:32:45we have to cover in Python to make sure
  58599. 43:32:51we are able to reduce
  58600. 43:32:55the time
  58601. 43:32:57in coding right we are able to reduce
  58602. 43:32:59the time in coding and using these
  58603. 43:33:02packages immediately help us in getting
  58604. 43:33:05the desired results right I hope you all
  58605. 43:33:07remember the concept of modules
  58606. 43:33:10we have covered in Python do we all
  58607. 43:33:12remember functions and modules. You can
  58608. 43:33:14use Jupyter notebook. If your Jupyter
  58609. 43:33:16notebook is not installed, you can use
  58610. 43:33:18something which is called as Google
  58611. 43:33:20Collab, right? Go to Google, type
  58612. 43:33:24Collab,
  58613. 43:33:26right? Let's say you write Collab,
  58614. 43:33:30right? And then you will see this
  58615. 43:33:32option. Click on Google Collab and it
  58616. 43:33:36will allow you to code in Python, right?
  58617. 43:33:40So we are now going to discuss about the
  58618. 43:33:43numpy package in python. Okay, numpy
  58619. 43:33:46package in python.
  58620. 43:33:49So numpy is a
  58621. 43:33:55fundamental
  58622. 43:34:00package
  58623. 43:34:02for data science in Python. Right? It is
  58624. 43:34:08one of the most fundamental packages for
  58625. 43:34:12practicing data science in Python.
  58626. 43:34:14Right? Why is that? Why is so why is
  58627. 43:34:17numpy so fundamental? What's is so
  58628. 43:34:19special? NumPy package
  58629. 43:34:25gives us a new data type
  58630. 43:34:29for handling
  58631. 43:34:32data in Python
  58632. 43:34:36called as
  58633. 43:34:39N D arrays, right? ND arrays which
  58634. 43:34:44stands for
  58635. 43:34:48this stands for
  58636. 43:34:51N dimensional
  58637. 43:34:54arrays right n dimensional arrays right
  58638. 43:34:57this stands for n dimensional arrays
  58639. 43:35:00so till now
  58640. 43:35:03till now we have studied
  58641. 43:35:07about
  58642. 43:35:09list integer
  58643. 43:35:12tpples,
  58644. 43:35:14strings,
  58645. 43:35:16right? Out of which
  58646. 43:35:20out of which
  58647. 43:35:25the data types
  58648. 43:35:29such as list
  58649. 43:35:32pupils have been used to store data,
  58650. 43:35:38right? store data, right?
  58651. 43:35:43And range
  58652. 43:35:46used for generating
  58653. 43:35:51new data which is primarily sequential.
  58654. 43:35:55Right?
  58655. 43:35:57Now there is one now there is one
  58656. 43:35:58problem right? There is one problem and
  58657. 43:36:01there should be a question that there
  58658. 43:36:04should be a question.
  58659. 43:36:09Why do we need a new data type
  58660. 43:36:15to work with data science,
  58661. 43:36:20right? Why do we need this? Yep. So the
  58662. 43:36:24answer to this question people the
  58663. 43:36:26answer to this question uh lies in a
  58664. 43:36:29small explanation right? Yeah lies in a
  58665. 43:36:33small explanation which is that
  58666. 43:36:36Python
  58667. 43:36:39is a
  58668. 43:36:42high
  58669. 43:36:44level
  58670. 43:36:47language right? Python is a highlevel
  58671. 43:36:51language, right? And
  58672. 43:36:56a highle language
  58673. 43:36:59is usually
  58674. 43:37:03very
  58675. 43:37:05distant
  58676. 43:37:08from
  58677. 43:37:10hardware.
  58678. 43:37:13A highle language is close to hardware
  58679. 43:37:15or distant from the hardware. Did you
  58680. 43:37:17not attend the Python programming
  58681. 43:37:19essentials?
  58682. 43:37:21What is the type of programming language
  58683. 43:37:24which is closest to hardware? A
  58684. 43:37:26low-level language.
  58685. 43:37:29If this is my hardware,
  58686. 43:37:33yeah, this is my OS.
  58687. 43:37:37This is my application layer. Right? So,
  58688. 43:37:40hard level langu language is here and
  58689. 43:37:43low-level languages here. Right? Which
  58690. 43:37:45is closest. So binary languages,
  58691. 43:37:48assembly languages,
  58692. 43:37:51right? All these are closest to the
  58693. 43:37:53hardware because where is the processing
  58694. 43:37:57happening? Where is the processing
  58695. 43:37:58happening of the data? At the hardware,
  58696. 43:38:02isn't it? Processing of data
  58697. 43:38:08is happening
  58698. 43:38:11at hardware.
  58699. 43:38:14No worry. Actually it is happening at
  58700. 43:38:16hardware right? What is processing?
  58701. 43:38:18Processing is signals of zeros and ones
  58702. 43:38:20right? What are zeros and ones? These
  58703. 43:38:23are electric signals.
  58704. 43:38:25These are the electric signals right?
  58705. 43:38:28Which is basically on and off. Right?
  58706. 43:38:31And it is communicated to the hardware
  58707. 43:38:33through the help of resistors
  58708. 43:38:36and microprocessors.
  58709. 43:38:40Right? Isn't it right? Why a computer
  58710. 43:38:43only knows zeros and ones? Because zero
  58711. 43:38:45is off and one is on which is the
  58712. 43:38:48electric current
  58713. 43:38:51right electric current to activate or
  58714. 43:38:54deactivate certain things right true and
  58715. 43:38:56false gates right so it is happening at
  58716. 43:38:59hardware so now people if you understand
  58717. 43:39:02this part then try to logically connect
  58718. 43:39:05it to what I'm going to say when you are
  58719. 43:39:09studying data science
  58720. 43:39:13what kind of data you'll be dealing with
  58721. 43:39:15big data,
  58722. 43:39:19right? And as the name suggests, it will
  58723. 43:39:21have a lot of volume,
  58724. 43:39:23right? It will have a lot of volume
  58725. 43:39:25other than a lot of other things, right?
  58726. 43:39:28It will be very very big. And on this
  58727. 43:39:31large volume of data, you'll be doing
  58728. 43:39:34processing.
  58729. 43:39:35You'll be doing processing. What does
  58730. 43:39:37processing means?
  58731. 43:39:40What does processing means? Operations.
  58732. 43:39:43So where is this operation happening?
  58733. 43:39:45This is happening in hardware
  58734. 43:39:48and for hardware which is the closest
  58735. 43:39:50language to hardware a low-level
  58736. 43:39:53language.
  58737. 43:39:54But now people but now we have a
  58738. 43:39:57situation in front of us. What is the
  58739. 43:39:59situation that what are we trying to do
  58740. 43:40:01data science with? What are we trying to
  58741. 43:40:03do data science with?
  58742. 43:40:06Python, right?
  58743. 43:40:09And Python is what?
  58744. 43:40:11A high level language,
  58745. 43:40:14isn't it? Yeah. So, there's a
  58746. 43:40:17discrepancy. Yeah. A big one
  58747. 43:40:20because hard level, high level language,
  58748. 43:40:24these
  58749. 43:40:26are slow
  58750. 43:40:29in processing,
  58751. 43:40:33right? These are very slow in
  58752. 43:40:35processing, right? So for these kind of
  58753. 43:40:38languages to handle this kind of data
  58754. 43:40:41and these kind of operations yeah is
  58755. 43:40:44very difficult right so let's let's keep
  58756. 43:40:48let's keep this part aside if you
  58757. 43:40:50understand this now let's go to the
  58758. 43:40:52second point right
  58759. 43:40:56when you learned Python
  58760. 43:41:00on a scale of 1 to 10 how easy was it
  58761. 43:41:05the ease of use of Python
  58762. 43:41:07It's relatively a higher number, right?
  58763. 43:41:10Relatively a higher number. So now guys,
  58764. 43:41:13when I talk about data science, right?
  58765. 43:41:16When I talk about
  58766. 43:41:18data science, okay?
  58767. 43:41:22Right? When I talk about data science,
  58768. 43:41:25my thing is that this will be used by
  58769. 43:41:30masses,
  58770. 43:41:32right? Will be used by masses. A lot of
  58771. 43:41:34people managers, programmers, business
  58772. 43:41:37analysts, data analysts, possible right
  58773. 43:41:41who are from nontechnical background who
  58774. 43:41:43don't know coding they also can do data
  58775. 43:41:45science because data science is a
  58776. 43:41:47general thing isn't it? Understanding
  58777. 43:41:49the data it should not be limited by
  58778. 43:41:52your capability to understand the uh
  58779. 43:41:55technicalities of a very complex
  58780. 43:41:57language. So for these people which
  58781. 43:42:00language is suitable which is Python
  58782. 43:42:03right? It is easiest to understand. It's
  58783. 43:42:05a high level language almost like
  58784. 43:42:06English. Yeah. So, Python is a simple
  58785. 43:42:09language. So, in this part people,
  58786. 43:42:11Python fits the bill, right? Which is a
  58787. 43:42:13bigger thing. In the second part, when
  58788. 43:42:16we talk about the operations,
  58789. 43:42:20we talk about the operations. In this
  58790. 43:42:23part, there is a problem, right? Python
  58791. 43:42:26fails,
  58792. 43:42:30right? Python fails, right? because it's
  58793. 43:42:33a highle language. We said that okay
  58794. 43:42:36there is a language called as C right
  58795. 43:42:39which is a middle level language
  58796. 43:42:45right and C language people is used to
  58797. 43:42:49create OS operating systems. It is used
  58798. 43:42:52to create networks
  58799. 43:42:55networking applications.
  58800. 43:43:00It is used to create games,
  58801. 43:43:03right? All the things which are close to
  58802. 43:43:07hardware,
  58803. 43:43:10C is used, right? So C fits this bill,
  58804. 43:43:15right? C fits this bill.
  58805. 43:43:18Python
  58806. 43:43:20said that okay, if C fits the bill and
  58807. 43:43:22Python is written
  58808. 43:43:26in C, written on C, right? It's written
  58809. 43:43:28on top of C language. Now what happened
  58810. 43:43:30was Python said okay if my intrinsic
  58811. 43:43:34data types my intrinsic processing is
  58812. 43:43:37not suitable for data science but my
  58813. 43:43:40highlevel nature is let's do one thing
  58814. 43:43:43let's take C language and use its power
  58815. 43:43:47right use its power that it is very
  58816. 43:43:51close to hardware and let's create a new
  58817. 43:43:55data type right let's create a new data
  58818. 43:43:59type which is written on top of C and
  58819. 43:44:04which can integrate with Python
  58820. 43:44:07seamlessly. And people this new data
  58821. 43:44:10type was called as array
  58822. 43:44:16and this was given to you by something
  58823. 43:44:18called as num py package. Right? Nump py
  58824. 43:44:23package. It defined a new data type
  58825. 43:44:26which was array. And along with defining
  58826. 43:44:28the array, it gave various operations.
  58827. 43:44:35Right? It gave various operations
  58828. 43:44:38one could
  58829. 43:44:40perform
  58830. 43:44:42on arrays,
  58831. 43:44:45right? One could perform on arrays.
  58832. 43:44:49It's simple, right? program the the
  58833. 43:44:51power of processing lied with C right it
  58834. 43:44:55was lying with C. So we developed a new
  58835. 43:44:59data type using C on top of Python and
  58836. 43:45:03that new data type was called as array
  58837. 43:45:06and this array was defined in a new
  58838. 43:45:09module which was called as numpy module
  58839. 43:45:13which told you how to create the arrays
  58840. 43:45:15and then how to manipulate those arrays
  58841. 43:45:19for doing data science. We had options
  58842. 43:45:22like Java, we had options like C, right?
  58843. 43:45:25We had options like forotron to be used
  58844. 43:45:28for data science but we chose Python
  58845. 43:45:30because of its simplicity and the simple
  58846. 43:45:33syntaxes that people from
  58847. 43:45:35non-programming background could also
  58848. 43:45:38use Python to do data science.
  58849. 43:45:42Right now the limitation was that
  58850. 43:45:44because it is slow because of being high
  58851. 43:45:46level we needed something which could
  58852. 43:45:49make it fast and that was using an
  58853. 43:45:52external data type which is not internal
  58854. 43:45:54to Python and that was array and this
  58855. 43:45:57array is defined inside a new uh module
  58856. 43:46:01or a library called as num py which is
  58857. 43:46:05numerical python right numerical python.
  58858. 43:46:09Back to the programming right where we
  58859. 43:46:12have understood right about this
  58860. 43:46:15question right.
  58861. 43:46:17So
  58862. 43:46:19numpy
  58863. 43:46:22essentially is
  58864. 43:46:25built on top of
  58865. 43:46:30C
  58866. 43:46:31language which is
  58867. 43:46:34compatible
  58868. 43:46:37with Python.
  58869. 43:46:40It leverages the power of closeness of C
  58870. 43:46:48with hardware,
  58871. 43:46:51right? C with hardware
  58872. 43:46:53which eventually
  58873. 43:46:56makes the processing
  58874. 43:47:00faster in
  58875. 43:47:05Python. Right?
  58876. 43:47:08So this is what the first part is right
  58877. 43:47:11this is what a first part is right now
  58878. 43:47:14the second thing is people so this is
  58879. 43:47:16about the performance bit right these
  58880. 43:47:18are the performance bit the above
  58881. 43:47:22is about the performance
  58882. 43:47:28of
  58883. 43:47:31right now coming to the next part which
  58884. 43:47:34is the memory efficiency Y right memory
  58885. 43:47:39efficiency right so
  58886. 43:47:44arrays created
  58887. 43:47:46by nump py in
  58888. 43:47:51python
  58889. 43:47:53are less memory
  58890. 43:47:57exhaustive
  58891. 43:48:00than lists in Python right and I will
  58892. 43:48:04prove these points later on to you
  58893. 43:48:06through code, right?
  58894. 43:48:08In list, right? In list,
  58895. 43:48:11each item is an object, right? Each item
  58896. 43:48:16is an object, right? I hope you remember
  58897. 43:48:18this, guys. Each item is an object,
  58898. 43:48:21right?
  58899. 43:48:23And it holds, right? It holds
  58900. 43:48:28meta information
  58901. 43:48:32like
  58902. 43:48:34references
  58903. 43:48:36and types,
  58904. 43:48:38right? Etc., right? A lot of information
  58905. 43:48:40it holds, right? This makes
  58906. 43:48:46list consume
  58907. 43:48:49more memory, right? But but
  58908. 43:48:54arrays in
  58909. 43:48:57numpy
  58910. 43:48:59are contiguous
  58911. 43:49:03which means
  58912. 43:49:06that they do not
  58913. 43:49:10create objects
  58914. 43:49:12but rather
  58915. 43:49:14directly store
  58916. 43:49:18the data
  58917. 43:49:20in
  58918. 43:49:22continuous
  58919. 43:49:24memory
  58920. 43:49:27blocks
  58921. 43:49:29one after another. Right? Also the
  58922. 43:49:33arrays are homogeneous in nature. Right?
  58923. 43:49:39You can only store the homogeneous data
  58924. 43:49:42in array unlike lists. In list you could
  58925. 43:49:45store different different data types.
  58926. 43:49:47Right? But in arrays you cannot right.
  58927. 43:49:49you have to store the same kind of data
  58928. 43:49:52in the array homogeneous. So it's a
  58929. 43:49:55contiguous memory block which is meaning
  58930. 43:49:58that you can store data in continuity
  58931. 43:50:02right one after another in the memory
  58932. 43:50:04block. So the access is faster the
  58933. 43:50:06memory location and the memory
  58934. 43:50:08efficiency is very higher right as
  58935. 43:50:11compared to the native data type like
  58936. 43:50:13lists or tpples in Python. Arrays
  58937. 43:50:17are basically
  58938. 43:50:20vectorzed operations
  58939. 43:50:23right I'll talk to about talk about to
  58940. 43:50:25you with vectors what are vectors right
  58941. 43:50:27they are the vectorzed operations and
  58942. 43:50:29they are way more convenient
  58943. 43:50:34to deal with as compared to list
  58944. 43:50:41in Python right in list you have to go
  58945. 43:50:44through a lot of loops, right? We saw
  58946. 43:50:47that we have to go through the list
  58947. 43:50:49comprehension. But you will see in
  58948. 43:50:52Python pandas, sorry, in Python numpy,
  58949. 43:50:55the vectorzed operations are very very
  58950. 43:50:58simple, right? They are very very
  58951. 43:51:00simple, right? Again, for all this, I
  58952. 43:51:02will give you examples, but uh it will
  58953. 43:51:05take some time because you have to
  58954. 43:51:07understand what arrays are first, right?
  58955. 43:51:13Okay. Then
  58956. 43:51:16the most important
  58957. 43:51:20other packages
  58958. 43:51:22which are pandas,
  58959. 43:51:26mattplot lib,
  58960. 43:51:28cb bond,
  58961. 43:51:30sklearn,
  58962. 43:51:32cypy
  58963. 43:51:34all are written
  58964. 43:51:36on top
  58965. 43:51:39of
  58966. 43:51:40numpy
  58967. 43:51:42package.
  58968. 43:51:44That's the reason for this reason
  58969. 43:51:47it's called as
  58970. 43:51:50fundamental
  58971. 43:51:52package.
  58972. 43:51:54All the other packages which make your
  58973. 43:51:56life easier as a data scientist where
  58974. 43:51:58you don't have to worry about code.
  58975. 43:52:02All these packages
  58976. 43:52:05make our data science
  58977. 43:52:09journey smooth
  58978. 43:52:12because we have to worry
  58979. 43:52:16less about code and more about
  58980. 43:52:22logic.
  58981. 43:52:23Yes. So for these for the understanding
  58982. 43:52:26of these packages it's very important
  58983. 43:52:29that we understand numpy first and then
  58984. 43:52:31we move forward right
  58985. 43:52:35now
  58986. 43:52:38then let's get started. So the first
  58987. 43:52:40step right the first step which you have
  58988. 43:52:42to uh see right the first step which you
  58989. 43:52:46have to see while using numpy package is
  58990. 43:52:49basically from where will you import the
  58991. 43:52:52package right from where will you import
  58992. 43:52:54the package.
  58993. 43:52:56So to import the package num py we write
  58994. 43:53:03import nump py as np right where np
  58995. 43:53:12is an alias right it's an alias
  58996. 43:53:17right so please do this import nump py
  58997. 43:53:20as np
  58998. 43:53:23right if you don't get any result for
  58999. 43:53:27this then you can write pip install
  59000. 43:53:30nump py right pip install nump py and
  59001. 43:53:34just execute this right when you will
  59002. 43:53:37execute this it will give you this
  59003. 43:53:39message or it will give you it will
  59004. 43:53:41download this package for you right pip
  59005. 43:53:45stands for
  59006. 43:53:47python
  59007. 43:53:49index package
  59008. 43:53:52right
  59009. 43:53:55it is basically like play store,
  59010. 43:53:59app store,
  59011. 43:54:01right? Or Windows store
  59012. 43:54:06for Python,
  59013. 43:54:08right?
  59014. 43:54:10So, you're just going to these Play
  59015. 43:54:12Store,
  59016. 43:54:16Windows Store, App Store of your Python
  59017. 43:54:19and asking them to download this for
  59018. 43:54:21you, right? It is also called as the
  59019. 43:54:25package manager
  59020. 43:54:28right pip
  59021. 43:54:30if it is done you can also check the
  59022. 43:54:32version you can say np dot
  59023. 43:54:35version
  59024. 43:54:38and it will give you the version of
  59025. 43:54:40numpy
  59026. 43:54:42right
  59027. 43:54:47numpy is an opensource
  59028. 43:54:52package,
  59029. 43:54:54right? Yeah, that's about it. Yeah, it's
  59030. 43:54:58an open-source package, people,
  59031. 43:55:01right? Open source package. And if you
  59032. 43:55:05want to see the code, you can go to
  59033. 43:55:10GitHub.
  59034. 43:55:12Go to Google, write the uh code for
  59035. 43:55:15numpy. It will show you it on GitHub.
  59036. 43:55:18Right. done. So let's say I okay so I
  59037. 43:55:23say
  59038. 43:55:25we have when we so okay before this let
  59039. 43:55:28me come to a little bit of theory before
  59040. 43:55:30I do this with you. So now guys I said
  59041. 43:55:33that num py
  59042. 43:55:36has arrays
  59043. 43:55:40as
  59044. 43:55:43data type
  59045. 43:55:45right and this is written on C which
  59046. 43:55:49runs directly
  59047. 43:55:52on hardware
  59048. 43:55:57also this numpy array
  59049. 43:56:01is basically basically a vector,
  59050. 43:56:05right? It's basically a vector. Now,
  59051. 43:56:07what is a vector, people? What is a
  59052. 43:56:09vector? A vector is a quantity which has
  59053. 43:56:14sign
  59054. 43:56:16plus magnitude,
  59055. 43:56:20right? It has a sign and magnitude. If I
  59056. 43:56:23say this is a cartition space, this is
  59057. 43:56:25I, this is J. And I say this this is 3 I
  59058. 43:56:32and 4 J right 3 I cap 4 Jcap. So this is
  59059. 43:56:37a vector right and this is the direction
  59060. 43:56:41guys. If you have studied elementary
  59061. 43:56:43maths you would know this.
  59062. 43:56:46Yes people this is a vector. If I draw
  59063. 43:56:49another like this
  59064. 43:56:52then this is another vector. So I will
  59065. 43:56:55call this say
  59066. 43:56:57uh 2 I and 5 J. Yeah, this is another
  59067. 43:57:01vector and this is the theta right. This
  59068. 43:57:05is the direction.
  59069. 43:57:06This is the direction. And what is the
  59070. 43:57:09magnitude?
  59071. 43:57:113 I 3² + 4² which is 9 + 16 which is 25
  59072. 43:57:18under root which is 5. So the magnitude
  59073. 43:57:22of vector is five and direction is equal
  59074. 43:57:25to theta. This is a vector quantity.
  59075. 43:57:29Now what is this? What is this? This is
  59076. 43:57:33a scalar.
  59077. 43:57:36This is a scalar. Yeah. Only magnitude
  59078. 43:57:43isn't it? This is scalar only magnitude.
  59079. 43:57:46And then I have 2a 3. Right? This is
  59080. 43:57:50what? This is a vector.
  59081. 43:57:55It has two dimensions
  59082. 43:57:58or one dimension. Only one dimension,
  59083. 43:58:02right?
  59084. 43:58:04This is one dimension vector.
  59085. 43:58:08Yep. One dimension vector. Now if I say
  59086. 43:58:11this 2 3 4 5, what is this called? This
  59087. 43:58:16is called a matrix,
  59088. 43:58:18right? Right? This is called a matrix
  59089. 43:58:20which is what collection of vectors
  59090. 43:58:27and this collection of m vectors matrix
  59091. 43:58:29is called as two-dimensional. Right? It
  59092. 43:58:32is called as two-dimensional.
  59093. 43:58:35Now if you have this
  59094. 43:58:43so these are stacked behind each other.
  59095. 43:58:47This is one. This is two. So this is
  59096. 43:58:48three right? So we have three layers
  59097. 43:58:55in matrix.
  59098. 43:58:58So how many dimensions will be this
  59099. 43:58:59people?
  59100. 43:59:01One dimension, two dimension and three
  59101. 43:59:03dimension. This is a threedimension
  59102. 43:59:06matrix,
  59103. 43:59:09right? Or threedimension vector
  59104. 43:59:13or three-dimension array.
  59105. 43:59:16I'll repeat once again. What is a single
  59106. 43:59:19value? A single value is called as a
  59107. 43:59:22scalar. Right? It only has magnitude.
  59108. 43:59:24Now when you have multiple values,
  59109. 43:59:27right? This is called as a vector. It
  59110. 43:59:29has a direction. It has a magnitude. And
  59111. 43:59:32this is single dimension. Now multiple
  59112. 43:59:36vectors right like this or maybe you can
  59113. 43:59:40say like this, right? Are you
  59114. 43:59:43understanding why I'm calling it one
  59115. 43:59:44dimension? It can be either this
  59116. 43:59:46dimension or it can be this dimension.
  59117. 43:59:48In any dimension you stack two vectors,
  59118. 43:59:51you will get yourself a matrix. Right?
  59119. 43:59:54You will get yourself a matrix which is
  59120. 43:59:56now two dimensions. It has rows and it
  59121. 44:00:00has columns.
  59122. 44:00:03Right?
  59123. 44:00:04Now if you stack multiple such matrix
  59124. 44:00:09one after another, right? This becomes a
  59125. 44:00:13threedimensional matrix. And this can go
  59126. 44:00:16up to how many dimensions people? How
  59127. 44:00:19many dimensions this can go up to? It
  59128. 44:00:21can go up to this
  59129. 44:00:25can
  59130. 44:00:27go up to n dimensions,
  59131. 44:00:32right? N dimensions,
  59132. 44:00:35right? So now do we get it? Why do we
  59133. 44:00:37call it ND
  59134. 44:00:40arrays?
  59135. 44:00:42Yeah, n dimension arrays,
  59136. 44:00:48right? That why are we calling something
  59137. 44:00:50what?
  59138. 44:00:52Right. N dimension arrays. We are only
  59139. 44:00:55capable of viewing three dimensions
  59140. 44:00:57people. It can go up to 100 dimensions,
  59141. 44:00:59500 dimensions, 1,000 dimensions, any
  59142. 44:01:03dimensions.
  59143. 44:01:04Scalas are least important.
  59144. 44:01:07vectors are more important. So now we
  59145. 44:01:10have
  59146. 44:01:14multiple
  59147. 44:01:15dimensions in
  59148. 44:01:19arrays
  59149. 44:01:23namely 0D,
  59150. 44:01:261D,
  59151. 44:01:292D, 3D and so on till the N D, right? NV
  59152. 44:01:36arrays. Right now let's create our first
  59153. 44:01:40array. Right? Let's create our first
  59154. 44:01:42array. And this will be a zero
  59155. 44:01:47dimension
  59156. 44:01:52array, right? Zero dimension array. How
  59157. 44:01:55will you create this? You will say a r0
  59158. 44:02:00is equal to np dot array and you will
  59159. 44:02:05mention what a scalar value is. What is
  59160. 44:02:07a scalar value? It is a simple value. I
  59161. 44:02:11say two, right? np dot array equal to
  59162. 44:02:13two. And when you will now check or
  59163. 44:02:17print the type of ar r0, it will tell
  59164. 44:02:22you class num py nd array. Right? It is
  59165. 44:02:27a zero dimension array. And if you want
  59166. 44:02:30to check
  59167. 44:02:33the dimension,
  59168. 44:02:35you just have to write a r0 dot end div,
  59169. 44:02:43right? And it shows you that there is
  59170. 44:02:46zero dimensions present. So what is this
  59171. 44:02:48in short? This is a scalar. I have
  59172. 44:02:54entered a single value people 2 200 500
  59173. 44:02:58whatever you want to enter. And the
  59174. 44:02:59syntax is np dot array np dot array.
  59175. 44:03:04You're instructing nump py package to
  59176. 44:03:06fetch the function array on method array
  59177. 44:03:10and convert this into that particular
  59178. 44:03:14data type right array zero
  59179. 44:03:18and when you check the type type is nd
  59180. 44:03:20array but what is the dimension of this
  59181. 44:03:22nd array this is zero which is nothing
  59182. 44:03:24but a scalar right we have created this
  59183. 44:03:29this is what we have created
  59184. 44:03:32Right
  59185. 44:03:35now guys,
  59186. 44:03:38if I created this list, right?
  59187. 44:03:45Say Lis is equal to
  59188. 44:03:50Yeah. So this was this was a list,
  59189. 44:03:53right? This was a list and this was this
  59190. 44:03:55is what people what is this that we have
  59191. 44:03:59just studied according to that? What
  59192. 44:04:01dimension is this list?
  59193. 44:04:03So what I'm trying to do is I will
  59194. 44:04:06create a
  59195. 44:04:09one deal
  59196. 44:04:14array with a or let's say from a list
  59197. 44:04:20right let's let me show you how do we do
  59198. 44:04:22that okay
  59199. 44:04:29so I say
  59200. 44:04:31hurry read the error name ar r not
  59201. 44:04:34defined. Why? Because you have created
  59202. 44:04:36array from a a r0. Come on hurry.
  59203. 44:04:41Right.
  59204. 44:04:43I will create a onedimensional array.
  59205. 44:04:44How will I do that people? I will say a
  59206. 44:04:46ar r1 is equal to np dot array. And can
  59207. 44:04:51I pass 1D list inside this? Or can I say
  59208. 44:04:55I can pass lis inside this? When I do
  59209. 44:04:58this people now see what will happen. A
  59210. 44:05:01R R1 will be equal to this right and if
  59211. 44:05:05you say let me say print a ar a ar a ar
  59212. 44:05:07a ar a ar a ar a ar a ar a ar a ar a ar
  59213. 44:05:08r r r r r r r r r r r r r r r r r r r r1
  59214. 44:05:10it will be like this okay this is your
  59215. 44:05:15a ar r r1 dot nim
  59216. 44:05:20you will see that it gives you one right
  59217. 44:05:22this is a onedimension array right on
  59218. 44:05:26dimension array
  59219. 44:05:43Yes. So for example
  59220. 44:05:45when I say scalar right when I say
  59221. 44:05:49scalar I say
  59222. 44:05:5335s right? when I say vector
  59223. 44:05:581D I say
  59224. 44:06:02uh
  59225. 44:06:03so these are suppose my marks
  59226. 44:06:06right now I say 35
  59227. 44:06:0940
  59228. 44:06:1150 right so now what are these my marks
  59229. 44:06:16in three subjects
  59230. 44:06:19yeah marks in three subjects this is my
  59231. 44:06:22say Hindi this is English and this is
  59232. 44:06:26maths right or let's say science because
  59233. 44:06:29not everyone has Hindi science English
  59234. 44:06:32and maths right guys now if I have to
  59235. 44:06:34create a matrix
  59236. 44:06:37of two dimension what will I write so
  59237. 44:06:40that means this is one student this is
  59238. 44:06:43one student people isn't it guys yes no
  59239. 44:06:47maybe so now in matrix we will have what
  59240. 44:06:50we will have multiple students
  59241. 44:06:55Yes. No. Maybe in a matrix people we
  59242. 44:06:58will have multiple students. Suppose
  59243. 44:07:00this was S1. Now you will have S_sub_1,
  59244. 44:07:03S_UB_2, S3, S4 like this.
  59245. 44:07:08And each student will have their own
  59246. 44:07:10individual list of marks. So can I say
  59247. 44:07:14that I'm making a nested list?
  59248. 44:07:19Can I say that people? I'm making a
  59249. 44:07:21nested list. So now let's make it okay
  59250. 44:07:28from a
  59251. 44:07:31nested list. Okay, a nested list.
  59252. 44:07:36So I'll say ar r2 is equal to np array
  59253. 44:07:39list. I will have to create a list
  59254. 44:07:41first. Lis2 is equal to. So this is my
  59255. 44:07:45first bracket. What is this bracket
  59256. 44:07:47representing? this bigger bracket. Now I
  59257. 44:07:49will put another bracket inside this and
  59258. 44:07:51I will write 1 1 22 33 3. I'll put a
  59259. 44:07:54comma again. Write a comma. Then I will
  59260. 44:07:58say
  59261. 44:07:594455 666 comma 778899
  59262. 44:08:05right I'll do this right now I will say
  59263. 44:08:09list to
  59264. 44:08:13two
  59265. 44:08:20when you will do this you will see that
  59266. 44:08:22an array like this has been created
  59267. 44:08:25Right?
  59268. 44:08:28like this AR R2
  59269. 44:08:30right this has been created right when
  59270. 44:08:34we check the dimension it is two
  59271. 44:08:36dimension array right
  59272. 44:08:40yeah marks of three different students
  59273. 44:08:43in three different subjects
  59274. 44:08:46and this is the same technique you can
  59275. 44:08:49create a three-dimension array how will
  59276. 44:08:51you create a three-dimension array
  59277. 44:08:52people
  59278. 44:08:54if I go here how will you create a
  59279. 44:08:56threedimension array
  59280. 44:08:59Now suppose I have data in this and this
  59281. 44:09:03is my master list. Okay, this is my
  59282. 44:09:04master list. In this I have data
  59283. 44:09:08and this can be represented like this.
  59284. 44:09:14This is my
  59285. 44:09:16first matrix isn't it? And this is the
  59286. 44:09:21vector inside this
  59287. 44:09:26V_sub_1, V_sub_2, V3.
  59288. 44:09:29Then this can be called as M1. And now
  59289. 44:09:34to create a three-dimension setup, how
  59290. 44:09:37many M1s do you need? You need multiple
  59291. 44:09:40M1s, isn't it? You need M1, M2, M3,
  59292. 44:09:44multiple matrix like this people.
  59293. 44:09:48like this matrix 1, matrix 2, matrix 3.
  59294. 44:09:54So what will you do? You will have you
  59295. 44:09:57will have what people?
  59296. 44:10:00You will have another yellow, right?
  59297. 44:10:04And you will have inside this yellow
  59298. 44:10:07multiple purples.
  59299. 44:10:12Isn't it
  59300. 44:10:13right? Again like this.
  59301. 44:10:27Yeah. Like this you will have it people.
  59302. 44:10:31So can I say people can I say that as I
  59303. 44:10:35am increasing the dimensions as I am
  59304. 44:10:41increasing
  59305. 44:10:46the dimensions
  59306. 44:10:49I am putting 1D sorry 0D right and okay
  59307. 44:10:54again in this V_sub_1 in this V_sub1 do
  59308. 44:10:58you think you will have multiple scalers
  59309. 44:11:00people can I say that
  59310. 44:11:05can I say multiple scalers create a
  59311. 44:11:08vector multiple vectors create a matrix
  59312. 44:11:13And multiple matrix create one
  59313. 44:11:15three-dimensional matrix.
  59314. 44:11:18Can I say that? Let me talk to you about
  59315. 44:11:21an image. Right?
  59316. 44:11:24Image, right? What is an image made up
  59317. 44:11:30of people?
  59318. 44:11:33What is an image made up of?
  59319. 44:11:38H
  59320. 44:11:41pixels.
  59321. 44:11:44Yes or no? No,
  59322. 44:11:47not frames. Frames is basically videos.
  59323. 44:11:52Pixels are creating an image, right? So,
  59324. 44:11:54we have how many pixels here? 1 2 1 2 3
  59325. 44:11:584 5 6 7 8 9 10 11 12 13 14 15 16. Right?
  59326. 44:12:04Suppose
  59327. 44:12:06this is one vector, right? This is one
  59328. 44:12:08vector and this is one scalar
  59329. 44:12:13right? Scalar 1, scalar 2, scalar 3,
  59330. 44:12:15scalar 4. And this will make vector
  59331. 44:12:18v_sub1.
  59332. 44:12:20This is v_sub_2, v_ub3, v4. And together
  59333. 44:12:24together can I call this m_sub_1 and
  59334. 44:12:28call this blue?
  59335. 44:12:30So for a colored image, how many
  59336. 44:12:33channels are there people? How many
  59337. 44:12:35channels are there? What do we call it?
  59338. 44:12:37We call it the image as
  59339. 44:12:41RGB
  59340. 44:12:43RGB image
  59341. 44:12:45that means red,
  59342. 44:12:48green,
  59343. 44:12:50blue.
  59344. 44:12:52So this is blue part. So similarly you
  59345. 44:12:55will have a red part in front of it.
  59346. 44:12:58Then you will have a green part and then
  59347. 44:13:01finally you will have a blue part. Yeah.
  59348. 44:13:04Are you understanding guys? Why do we
  59349. 44:13:07require three-dimensional arrays?
  59350. 44:13:09Yes. Suppose now you want to make a
  59351. 44:13:12change at this this pixel. So you will
  59352. 44:13:15go to the third layer which is the blue
  59353. 44:13:18layer. Then you will go to the third
  59354. 44:13:20column. You will go to the third column
  59355. 44:13:22and third row. And this is how you will
  59356. 44:13:24reach this pixel. Everyone? Yes. No.
  59357. 44:13:28Maybe.
  59358. 44:13:29Yes. So this is the reason why we need
  59359. 44:13:32to create a 3D array.
  59360. 44:13:36Right. To ingest information like this,
  59361. 44:13:39right? To ingest information like this.
  59362. 44:13:41So we can create a 3D array also. Right.
  59363. 44:13:50Right. And how did I tell you? How many
  59364. 44:13:52brackets will I have? Squared brackets.
  59365. 44:13:54I will have three squared brackets. So
  59366. 44:13:56now this is just one student.
  59367. 44:14:00Right. Now I'll put a comma here.
  59368. 44:14:04Right? Right, I'll put a comma here and
  59369. 44:14:06I will start.
  59370. 44:14:14Right, I'll do this. I'll say this list
  59371. 44:14:17three
  59372. 44:14:19a r3
  59373. 44:14:21list three ar r r3
  59374. 44:14:24a r3
  59375. 44:14:26r
  59376. 44:14:29and now you will see that there's a
  59377. 44:14:31threedimensional array
  59378. 44:14:33right guys
  59379. 44:14:42right this is a threedimensional
  59380. 44:14:44array
  59381. 44:14:51Yep. Moving on. There are multiple ways
  59382. 44:14:57create arrays, right? The first one we
  59383. 44:15:01have done.
  59384. 44:15:03So we have done
  59385. 44:15:05from lists.
  59386. 44:15:08From list we have done.
  59387. 44:15:11Then second will be
  59388. 44:15:13from uh we can create a zero array
  59389. 44:15:19right we can create
  59390. 44:15:22on's array
  59391. 44:15:24right then we can create custom array
  59392. 44:15:31right I will show this all to you
  59393. 44:15:36right I'll show this all to you so let's
  59394. 44:15:38start with the on's array sorry zero
  59395. 44:15:41those array
  59396. 44:15:44right what do you have to do you have to
  59397. 44:15:47write so let's say zero
  59398. 44:15:50ar r r0 okay a zero dimension zero array
  59399. 44:15:52so you say np dot zeros
  59400. 44:15:56right and you create
  59401. 44:15:58a two
  59402. 44:16:03right
  59403. 44:16:04a two right and if I say
  59404. 44:16:10uh 0
  59405. 44:16:12dot end
  59406. 44:16:14right you will see it is a onedimension
  59407. 44:16:17array
  59408. 44:16:20right it is a onedimension array
  59409. 44:16:24right so two is by default taken as so
  59410. 44:16:28let me just show this to you
  59411. 44:16:31it will look like this right it is taken
  59412. 44:16:33as horizontal what is the dimension of
  59413. 44:16:34this guys a vector what is the dimension
  59414. 44:16:37of this vector
  59415. 44:16:39no no it's 1. It is basically 2 + 1,
  59416. 44:16:43right? 2a 1.
  59417. 44:16:46Sorry, 1 comma 2. My bad.
  59418. 44:16:501 comma 2. Isn't it? Now, what if you
  59419. 44:16:53had to create a 2 + 1? What? What if you
  59420. 44:16:57had to create a 2 + 1? Right? So, let me
  59421. 44:17:00just show that to you.
  59422. 44:17:03So, I say wait
  59423. 44:17:08like this. Okay.
  59424. 44:17:13Now when I do this, I say a ar r r once
  59425. 44:17:17and I say here
  59426. 44:17:202, 1, right? I say 2a 1. Now you will
  59427. 44:17:26see people what will happen to this.
  59428. 44:17:31Now
  59429. 44:17:33because you have created right specified
  59430. 44:17:37two dimensions, right? What will this be
  59431. 44:17:39converted to now?
  59432. 44:17:41H what will this be converted to? This
  59433. 44:17:45will be converted to
  59434. 44:17:48a
  59435. 44:17:49two-dimension vector. By default, it was
  59436. 44:17:52this, right? Which was this vector,
  59437. 44:17:56right? By default, it was this vector,
  59438. 44:18:00right? What is the dimension of this
  59439. 44:18:01vector? 1 + how many values you put
  59440. 44:18:04here, right? 1 + 2 like this. So we call
  59441. 44:18:08this only one dimension. We call this
  59442. 44:18:10only one dimension. But now when I will
  59443. 44:18:13run this one, you will see yes 2 + 1.
  59444. 44:18:18And now you will see this is
  59445. 44:18:22two dimension, right? You see this is
  59446. 44:18:24two dimension. Now clear people just the
  59447. 44:18:27orientation has changed. But now you see
  59448. 44:18:29the brackets there are two brackets now
  59449. 44:18:31because what have you now instructed?
  59450. 44:18:33You have now instructed Python and
  59451. 44:18:36rather numpy to create the vector as 2 +
  59452. 44:18:401. So you have said give me this zero
  59453. 44:18:42and give me this zero here. So the
  59454. 44:18:45moment you do this you are now
  59455. 44:18:47specifying the rows and columns.
  59456. 44:18:52Now this 21 right let's say this is 21.
  59457. 44:18:58Now let me create a another two cross
  59458. 44:19:00two dimension matrix for you. Let me
  59459. 44:19:02call this 10
  59460. 44:19:05comma 10. Right? So how many rows and
  59461. 44:19:07columns will it have people?
  59462. 44:19:10How many rows and columns will it have?
  59463. 44:19:13I will say 21.
  59464. 44:19:19It will have 10 rows and 10 columns like
  59465. 44:19:22this. You saw this? Yeah. 10 rows and 10
  59466. 44:19:27columns. Right?
  59467. 44:19:31By default, it is 1 +2 like this. It is
  59468. 44:19:33a one-dimensional vector and this is a
  59469. 44:19:35two-dimensional vector. What about a
  59470. 44:19:37three-dimensional vector? 0 dot 0
  59471. 44:19:42uh a ar r3
  59472. 44:19:45is equal to np dot zeros, right? np
  59473. 44:19:48do.zer,
  59474. 44:19:50right?
  59475. 44:19:54H should I just write 3a 3a 3? And I
  59476. 44:19:59should check for this.
  59477. 44:20:02Yeah, you will have a 3 +3 vector 3 + 3
  59478. 44:20:06matrix with three matrix stacked behind
  59479. 44:20:08each other. Right? If I say four, this
  59480. 44:20:13will be four. So how do we read this?
  59481. 44:20:20number of layers,
  59482. 44:20:24number of rows and number of columns.
  59483. 44:20:30So, can you help me with a syntax? Can
  59484. 44:20:32you help me with a syntax which can
  59485. 44:20:35create me a 3D matrix of five layers,
  59486. 44:20:40three rows and three columns? What will
  59487. 44:20:42I write?
  59488. 44:20:44Five layers, three rows and three
  59489. 44:20:46columns. What will I write?
  59490. 44:20:535 33. Yeah, you'll get five layers. 1 2
  59491. 44:20:573 4 5 right like this. Suppose suppose
  59492. 44:21:02you have to represent
  59493. 44:21:06an image
  59494. 44:21:08with
  59495. 44:21:10with say
  59496. 44:21:13RGB channel
  59497. 44:21:15and
  59498. 44:21:19uh say 256 and 256 pixels. How will you
  59499. 44:21:24create this? So I'll say img right image
  59500. 44:21:29is equal to np dot
  59501. 44:21:32zeros and I will say inside this
  59502. 44:21:37three channel 256 cross 256
  59503. 44:21:42right and when you will just run this
  59504. 44:21:44image it will be like this right it will
  59505. 44:21:47be like this this is one channel this is
  59506. 44:21:51two channel and this is three channel
  59507. 44:21:55And uh we understood that numpy is a
  59508. 44:21:58fundamental package for data science in
  59509. 44:21:59Python. Right? This package gives us a
  59510. 44:22:02new data type to work with which is
  59511. 44:22:05called as n- dimensional arrays. Right?
  59512. 44:22:08Now what are arrays? What are arrays?
  59513. 44:22:11Arrays are nothing but vectors right
  59514. 44:22:15which are stored in a contiguous block
  59515. 44:22:17of memory which means they are stored
  59516. 44:22:20continuously one after another and they
  59517. 44:22:22do not get converted into the object
  59518. 44:22:24unlike the list and they are way faster
  59519. 44:22:27they are more memory efficient than list
  59520. 44:22:29I will prove this fact to you today with
  59521. 44:22:31the help of example through the help of
  59522. 44:22:32code right so why numpy arrays because
  59523. 44:22:36numpy is a package which is built on top
  59524. 44:22:38of C language which is compatible with
  59525. 44:22:40python and C being a middle level
  59526. 44:22:43language interacts directly with the
  59527. 44:22:45hardware. So whatever operation you run
  59528. 44:22:49in a fact that is getting directly
  59529. 44:22:50executed on the hardware itself. Right?
  59530. 44:22:53That is the reason why the uh
  59531. 44:22:56performance is way better when we try to
  59532. 44:22:58use the n dimensional arrays. Along with
  59533. 44:23:00this these syntaxes the type of syntaxes
  59534. 44:23:03we used to type uh in list I will show
  59535. 44:23:05that to you also today with the help of
  59536. 44:23:07example are way simpler when you try to
  59537. 44:23:10do them with nd arrays right so arrays
  59538. 44:23:15are basically vectorized operations
  59539. 44:23:17right and they're also very convenient
  59540. 44:23:20to deal with as compared to lists and
  59541. 44:23:21other native data types in python right
  59542. 44:23:24and the most important part is that in
  59543. 44:23:26our data science journey whatever other
  59544. 44:23:29packages packages we will use. Right?
  59545. 44:23:30Again, what are packages? They have
  59546. 44:23:32predefined things stored for you so that
  59547. 44:23:34you can leverage them and focus less on
  59548. 44:23:37code and more on logic. Right? You need
  59549. 44:23:39to be aware about the logic more than
  59550. 44:23:42the knowledge of the code. Right? So, we
  59551. 44:23:44have packages for that which contains
  59552. 44:23:46methods inside them which you can use
  59553. 44:23:49and uh without any further calculations
  59554. 44:23:52you can work with them directly. Right?
  59555. 44:23:56So that is how we started with numpy
  59556. 44:23:59right and the syntax to import numpy was
  59557. 44:24:02import numpy as np where np was an alias
  59558. 44:24:06right. Uh you could do pip install numpy
  59559. 44:24:08if someone did not have access to numpy
  59560. 44:24:10if numpy was not coming by default. You
  59561. 44:24:13can use pip install numpy which is
  59562. 44:24:15python index package right and uh this
  59563. 44:24:19is like play store app store window for
  59564. 44:24:21python. All the packages are stored in
  59565. 44:24:23pip and you can call pip you can ask pip
  59566. 44:24:26to download that package for you so that
  59567. 44:24:28you can use that right it's basically
  59568. 44:24:29the package manager so numpy is an open
  59569. 44:24:32source and if you want to see the code
  59570. 44:24:34you can go to github and check the code
  59571. 44:24:35out for yourself right in numpy we have
  59572. 44:24:38n dimensional arrays now the question
  59573. 44:24:40arises people that why do we need numpy
  59574. 44:24:43right so my answer to this particular
  59575. 44:24:45question is that when you will deal in
  59576. 44:24:48data science right when you will become
  59577. 44:24:49a data scientist you will be dealing
  59578. 44:24:51with data Right now my question is how
  59579. 44:24:53will you ingest how will you make the
  59580. 44:24:56machine ingest the data right there has
  59581. 44:24:59to be a way right for you to input the
  59582. 44:25:01data to the machine right to make
  59583. 44:25:03manipulations to the data to read the
  59584. 44:25:05data so all this is started as the base
  59585. 44:25:08package of numpy arrays right other than
  59586. 44:25:11that it becomes very difficult and
  59587. 44:25:13cumbersome for us to deal with that and
  59588. 44:25:15this provides us a lot of ease and
  59589. 44:25:17flexibility to deal with such massive
  59590. 44:25:20amounts of data which you will along
  59591. 44:25:21with me as we will move forward in this
  59592. 44:25:23particular course. Yeah, perfect. Now
  59593. 44:25:27people, let me just uh pull up the PBTs.
  59594. 44:25:31This is what we're discussing people. We
  59595. 44:25:34start with something which is called as
  59596. 44:25:36scalar, right? We start with something
  59597. 44:25:38which is called as scalar which is a
  59598. 44:25:41quantity which only has magnitude.
  59599. 44:25:43Right? In numpy language, this is also
  59600. 44:25:46called as 0D, right? It is called as
  59601. 44:25:48zero dimension. Then people we have
  59602. 44:25:50vector right and vector has a constant
  59603. 44:25:53dimension. It has only one dimension
  59604. 44:25:55right you can interpret it as a row or
  59605. 44:25:57you can interpret it as a column it
  59606. 44:25:59doesn't really matter because this is
  59607. 44:26:01only one single dimension right so
  59608. 44:26:03usually it will be written as five comma
  59609. 44:26:06blank right there will be nothing
  59610. 44:26:07written in front of it so this in numpy
  59611. 44:26:10terminology and nomenclature is called
  59612. 44:26:12as one dimension right when you move on
  59613. 44:26:16then you get combine couple of vectors
  59614. 44:26:19you get a shape and now that is called
  59615. 44:26:21as a matrix and In numpy terminology it
  59616. 44:26:24is called as two-dimension right and
  59617. 44:26:27when you try to stack multiple
  59618. 44:26:29two-dimension matrices one before with
  59619. 44:26:31one after each other or one before each
  59620. 44:26:32other then they become something called
  59621. 44:26:35as three-dimensional and now in my
  59622. 44:26:37capability I don't know what four
  59623. 44:26:38dimension looks like but there is a high
  59624. 44:26:40possibility that you have n dimensional
  59625. 44:26:43data right it has it has n we are
  59626. 44:26:45dealing with n dimension data right
  59627. 44:26:48suppose with this I also add time right
  59628. 44:26:52at t equal to 1 at t=2 that will serve
  59629. 44:26:55as the fourth dimension for this data
  59630. 44:26:57but how do how does it look like I don't
  59631. 44:26:59really know that right because humans
  59632. 44:27:00are only capable of visualizing 3D three
  59633. 44:27:04dimensions at max right so you can go to
  59634. 44:27:06n dimensions and hence the name n
  59635. 44:27:09dimensional arrays right nd arrays post
  59636. 44:27:12this right post this we moved on to
  59637. 44:27:15create certain things and I tried to
  59638. 44:27:17explain you the data right so the data
  59639. 44:27:20will look to you like this right You
  59640. 44:27:22might have a scalar quantity which is
  59641. 44:27:24marks right one marks right now if I go
  59642. 44:27:28on to vectors in one day it can be marks
  59643. 44:27:31of one student
  59644. 44:27:34right marks of one student in science
  59645. 44:27:38English
  59646. 44:27:40and maths right 35
  59647. 44:27:4340 and 50 out of say 50 right three
  59648. 44:27:47subjects so this will be characterized
  59649. 44:27:49as a vector right what will be the
  59650. 44:27:51dimension written for For this it will
  59651. 44:27:52be 3 comma nothing. This will be zero
  59652. 44:27:56right shape will be zero. For this
  59653. 44:27:58matrix suppose we have 1 2 3 four
  59654. 44:28:02students and each student will have
  59655. 44:28:05three marks.
  59656. 44:28:12Right? Each student will have three
  59657. 44:28:13marks. So what will be the shape of
  59658. 44:28:15this? We have four rows and three
  59659. 44:28:19columns. Right? So this will be the
  59660. 44:28:21shape right of this 2D matrix right and
  59661. 44:28:25now if you stack images right one behind
  59662. 44:28:28each other then it will be like this
  59663. 44:28:30right image I gave you an example so
  59664. 44:28:33this has 4 + 4 pixels so the shape will
  59665. 44:28:36be 3 + 4 + 4 right this will be the
  59666. 44:28:40shape for this particular 3D matrix
  59667. 44:28:43right guys so this is how you input the
  59668. 44:28:46data just to tell you a little bit more
  59669. 44:28:48since generative AI is very popular
  59670. 44:28:51these days. Right? So what if I tell you
  59671. 44:28:54the fact that the Chad GPT
  59672. 44:28:58which you use or you might have used
  59673. 44:29:02right has
  59674. 44:29:05never seen
  59675. 44:29:07a single
  59676. 44:29:10word
  59677. 44:29:13in its lifetime.
  59678. 44:29:17Right? All it sees
  59679. 44:29:21is numbers, right? Only numbers. How do
  59680. 44:29:25we see numbers? Suppose I say my
  59681. 44:29:29name
  59682. 44:29:30is
  59683. 44:29:32Raghav.
  59684. 44:29:35Raghav
  59685. 44:29:37is a
  59686. 44:29:40nice
  59687. 44:29:41name. Right? So these are two data,
  59688. 44:29:44right? These are two data points. Now we
  59689. 44:29:47all know that computers do not
  59690. 44:29:49understand these right computers do not
  59691. 44:29:51understand these right there is nothing
  59692. 44:29:53no understanding for computers to know
  59693. 44:29:55what text is right it only knows 0 and
  59694. 44:29:58one yes dhika right it only knows zeros
  59695. 44:30:00and ones so see how we will convert this
  59696. 44:30:02so there is something called as
  59697. 44:30:04vocabulary
  59698. 44:30:07right so vocabulary are nothing but the
  59699. 44:30:10unique words
  59700. 44:30:12right how many unique words do I have in
  59701. 44:30:14this my name is Raga four. This is not
  59702. 44:30:18unique. This is repeating. This is
  59703. 44:30:19repeating. Five, six. And this is
  59704. 44:30:22repeating. So I have six words. So now
  59705. 44:30:24guys, I will do something called as word
  59706. 44:30:27to
  59707. 44:30:29right where I will convert these words
  59708. 44:30:30into vectors. How will I convert them?
  59709. 44:30:32Look at this. So I will have suppose
  59710. 44:30:35this is S_sub_1,
  59711. 44:30:37this is S_sub_1 and this is S_sub_2,
  59712. 44:30:39right? So I will represent
  59713. 44:30:42S1 as
  59714. 44:30:47right. I will have a vector
  59715. 44:30:51of size six. How? I will say 1 0 0 0.
  59716. 44:30:59Right? 1 0 0. How many elements does it
  59717. 44:31:02have? Six elements. Right? Name will be
  59718. 44:31:050 1 0 0 0.
  59719. 44:31:08is will be 0 0 0 1 0 0 0
  59720. 44:31:13and ra will be 0 0 0 1 0 0 right this is
  59721. 44:31:18s1 my name is raghub now when it comes
  59722. 44:31:22to s_ub_2 right when it comes to s_ub_2
  59723. 44:31:26how will I enter this s2
  59724. 44:31:280 1 0 0 ragh what is is here
  59725. 44:31:3400 0 1 0 0 0
  59726. 44:31:37what is
  59727. 44:31:380 0 0 1 0 or nice is 0 0 0 1 and name
  59728. 44:31:46will be 0 1 0 0 0 0 right now this will
  59729. 44:31:51be the vector representation of these
  59730. 44:31:54two sentences just to tell you a fact
  59731. 44:31:58GBD3 right GPD3 model right GPD3 or
  59732. 44:32:03GPD3.5
  59733. 44:32:05they have vocabul vabulary
  59734. 44:32:09of 30,000 words, right? 30,000 words.
  59735. 44:32:13And each word, right? Each word
  59736. 44:32:19has a
  59737. 44:32:22dimension
  59738. 44:32:25of
  59739. 44:32:2612,500
  59740. 44:32:29numbers. Right? What do I mean? Suppose
  59741. 44:32:32I say Raghub.
  59742. 44:32:34So it will be one word and it will be
  59743. 44:32:37represented by five 11,500
  59744. 44:32:41different numbers
  59745. 44:32:45right and like raghub there will be
  59746. 44:32:4730,000 words in this GPD model right
  59747. 44:32:5230,000 words in this GPD model right and
  59748. 44:32:56this total number of parameters which
  59749. 44:32:58get trained in the neural network which
  59750. 44:33:00we will learn later on neural networks
  59751. 44:33:02they are almost close to 1
  59752. 44:33:0775
  59753. 44:33:09billion
  59754. 44:33:11parameters right 1.75 billion parameters
  59755. 44:33:15so why I'm telling you all this because
  59756. 44:33:18to showcase to you that what is the
  59757. 44:33:20importance of vectors right in the
  59758. 44:33:23entire machine learning and data science
  59759. 44:33:26yeah people so this is the reason people
  59760. 44:33:29now my question is how will you create
  59761. 44:33:30these vectors how will you read these
  59762. 44:33:32vectors Right? The answer is through
  59763. 44:33:35numpy package because it is the base
  59764. 44:33:38package. Clear people? Yeah. I hope
  59765. 44:33:41today's class will be uh interesting for
  59766. 44:33:43you because you will know the context.
  59767. 44:33:44Why are we doing it? Yeah. So I'll try
  59768. 44:33:46to show that to you how we convert
  59769. 44:33:48things to vectors.
  59770. 44:33:50Right? Okay. Let me go here now. Right.
  59771. 44:33:53Let me go here.
  59772. 44:33:55So people uh we started using numpy. So
  59773. 44:33:59I started with the zero dimension
  59774. 44:34:01arrays. Right? Zero dimension arrays. So
  59775. 44:34:03zero dimension is nothing but a scalar.
  59776. 44:34:05So I created a ar r0 which was np dot
  59777. 44:34:08array and I entered a single word single
  59778. 44:34:11uh element inside this which is nothing
  59779. 44:34:12but a scalar and then I checked the type
  59780. 44:34:15of uh ar0 also right and then I check
  59781. 44:34:19the dimension also. So the answer was
  59782. 44:34:21two class was numpy nd array and the
  59783. 44:34:24dimension was zero right exactly what I
  59784. 44:34:27had mentioned in my pb
  59785. 44:34:31right it will be having a zero dimension
  59786. 44:34:35like this right same thing has been
  59787. 44:34:38proven
  59788. 44:34:39right because it's a scalar now coming
  59789. 44:34:42on to one dimension right coming on to
  59790. 44:34:44one dimension I create a list which is
  59791. 44:34:46nothing but a one-dimension data type
  59792. 44:34:49right now I create an array which is a
  59793. 44:34:51ar r1 from array from this particular
  59794. 44:34:54list lis and then I check the type of
  59795. 44:34:58print ar1 check the type of ar1 and the
  59796. 44:35:01dimension right so when I execute
  59797. 44:35:07right then you will see that it was this
  59798. 44:35:10numpy array and the dimension was one
  59799. 44:35:13right exactly like this so if I show you
  59800. 44:35:15something else say print a ar r r1 one
  59801. 44:35:19dot shape
  59802. 44:35:26you will see it's 4 comma empty right
  59803. 44:35:28and I show it to you here
  59804. 44:35:35it will be empty right empty and this
  59805. 44:35:39will be comma 1 right so this means that
  59806. 44:35:42it has only four elements right if I
  59807. 44:35:45increase these elements to say 55 5 66
  59808. 44:35:4977
  59809. 44:35:51then it will become 7, blank, right?
  59810. 44:35:54Which means it has seven elements as a
  59811. 44:35:56vector. Now we create something with a
  59812. 44:35:59nested list right which is like this. So
  59813. 44:36:02with a nest I want one bracket which is
  59814. 44:36:05running outside right then inside this I
  59815. 44:36:08have one two and three lists inside one
  59816. 44:36:11list. Right? So this is a nested list.
  59817. 44:36:13This is marks of first student, second
  59818. 44:36:15student and the third student. Right? So
  59819. 44:36:19I do this and you see it is this right?
  59820. 44:36:22And I will show you the shape also
  59821. 44:36:30right. It will be 3 + 3 rows and three
  59822. 44:36:33columns. Three rows and three columns.
  59823. 44:36:36Right? Now similarly we can also create
  59824. 44:36:39a
  59825. 44:36:50the 3D matrix
  59826. 44:36:52right with
  59827. 44:36:56two levels right level one and level two
  59828. 44:37:00and 3 + 3 so the shape will be what
  59829. 44:37:03people can someone guess the shape
  59830. 44:37:06what will be the shape of this I've
  59831. 44:37:09shown you
  59832. 44:37:11Right? If this is 3 4 then what will be
  59833. 44:37:14this?
  59834. 44:37:16Three rows and three columns. Right? So
  59835. 44:37:19when you will execute you will get 2 3 3
  59836. 44:37:22right 2 3 3
  59837. 44:37:24right
  59838. 44:37:26right now people there are multiple ways
  59839. 44:37:28right there are multiple ways to create
  59840. 44:37:30arrays and we should know them because
  59841. 44:37:32all of these comes very very handy. Not
  59842. 44:37:35right now. I don't have enough context
  59843. 44:37:37to give you right now. But later on you
  59844. 44:37:39will see with me or with some other
  59845. 44:37:41trainer that how these will be used in
  59846. 44:37:43deep learning specifically, right? They
  59847. 44:37:45are the key of deep learning algorithms,
  59848. 44:37:49right? Where we initialize some weights,
  59849. 44:37:51we initialize some biases and those
  59850. 44:37:53initializations are nothing but
  59851. 44:37:55multi-dimensional numpy arrays, right?
  59852. 44:37:58Numpy arrays.
  59853. 44:38:01Okay.
  59854. 44:38:05Like for example, suppose I have this
  59855. 44:38:12I have to multiply this with some random
  59856. 44:38:15numbers, right? So how will you generate
  59857. 44:38:17these random numbers? You will generate
  59858. 44:38:19them through numpy. And you can generate
  59859. 44:38:21them in a specific kind of uh shape,
  59860. 44:38:25right? Which is 2 + 3. And then you can
  59861. 44:38:28multiply them. You can multiply the
  59862. 44:38:30matrices and you can get your output for
  59863. 44:38:33yourself. Right? So this is the way they
  59864. 44:38:36are used.
  59865. 44:38:38So we saw the first thing from list we
  59866. 44:38:41have already covered. Then now we are
  59867. 44:38:42moving on to creating zero arrays right.
  59868. 44:38:45So I create a zero array of one
  59869. 44:38:49dimension right of one dimension which
  59870. 44:38:51is 0 0. They are represented in floats
  59871. 44:38:54right. They are represented in floats
  59872. 44:38:560.0 zero. Right? Now you can create a
  59873. 44:39:00two-dimensional zero array. Right? You
  59874. 44:39:02can create a two-dimensional zero array
  59875. 44:39:04which is you have to mention just the
  59876. 44:39:05shape inside 2 + 1. So it will have two
  59877. 44:39:08rows and one columns, right? Two rows
  59878. 44:39:10and one columns. The difference here is
  59879. 44:39:12the difference here is that these are
  59880. 44:39:15one dimension and these are two
  59881. 44:39:17dimensions. Right? You have explicitly
  59882. 44:39:18mentioned the rows and columns. So you
  59883. 44:39:20can expand this to 10 + 10 also.
  59884. 44:39:27Right? You can expand in 10 + 10 or 10 +
  59885. 44:39:296 whatever you feel like yourself.
  59886. 44:39:31Right? It will have 10 rows and six
  59887. 44:39:35columns. Right?
  59888. 44:39:38Now you can also create
  59889. 44:39:41the 3D arrays 3D zero arrays.
  59890. 44:39:47Right?
  59891. 44:39:50Which is 5a 3a 3. What does five means?
  59892. 44:39:53What does five means? First element
  59893. 44:39:55represents the number of layers. So you
  59894. 44:39:57have five layers, right? It's five layer
  59895. 44:40:00deep. Then you have three rows and three
  59896. 44:40:03columns, right? So it will look
  59897. 44:40:05something like this.
  59898. 44:40:151
  59899. 44:40:182
  59900. 44:40:241 2 3 4 5 right like this something like
  59901. 44:40:27this right it will look like this tab 1
  59902. 44:40:302 3 1 2 3 1 2 3 right 5 33 okay yes the
  59903. 44:40:35number of matrices hurry what I
  59904. 44:40:37represented
  59905. 44:40:39right layers
  59906. 44:40:41rows
  59907. 44:40:43columns right layer rows and columns.
  59908. 44:40:46Clear?
  59909. 44:40:49So now when you execute this, you will
  59910. 44:40:50get an arrangement like this. Okay. I
  59911. 44:40:52try to show you this thing with another
  59912. 44:40:54example, right? Which is I created an
  59913. 44:40:57image of an RGB image of 256 cross 256
  59914. 44:41:01pixels, right? Which have all zeros
  59915. 44:41:03inside them. And this is how it was
  59916. 44:41:06created, right? Three layers RGB 256
  59917. 44:41:10256. So this is how the image will look
  59918. 44:41:12like
  59919. 44:41:14right.
  59920. 44:41:17This is how it will look like.
  59921. 44:41:20Now you can also create
  59922. 44:41:27arrays with ones. Right? Exactly the
  59923. 44:41:31same way you created it with zeros. I'll
  59924. 44:41:33give it to you. I'll give you 5 minutes
  59925. 44:41:34time to create them. I'll show you one.
  59926. 44:41:36So I say
  59927. 44:41:39a1 is equal to np dot
  59928. 44:41:44once
  59929. 44:41:46and inside I pass
  59930. 44:41:49two
  59931. 44:41:52I check a1. So this is a array like this
  59932. 44:41:55right? It is an array like this. Now you
  59933. 44:41:58create create two dimension
  59934. 44:42:02and
  59935. 44:42:04three dimension
  59936. 44:42:06arrays of one. It is basically
  59937. 44:42:11as a float. Hurry. It's represented as a
  59938. 44:42:13float. Right. It's represented as a
  59939. 44:42:15float. Okay.
  59940. 44:42:18Right.
  59941. 44:42:20Perfect. Right. Also guys uh with this
  59942. 44:42:24right also with this you can create the
  59943. 44:42:27custom arrays. Right. You can create the
  59944. 44:42:29custom arrays. Right. How do we create
  59945. 44:42:32custom arrays people? How do we create
  59946. 44:42:34the custom arrays?
  59947. 44:42:37You have created now zeros. You have
  59948. 44:42:39created now ones. Now what is left that
  59949. 44:42:43you create the custom arrays. Uh forget
  59950. 44:42:47about this. I will come to this later
  59951. 44:42:48on.
  59952. 44:42:50Right? Let's create
  59953. 44:42:53custom arrays. Right? So the syntax
  59954. 44:42:56remains the same. Right? I'll say cus
  59955. 44:42:59arr is equal to np.
  59956. 44:43:04Right? This is the syntax np.
  59957. 44:43:07Right? And you will say 6 + 6 and
  59958. 44:43:10suppose you want an array of all fours.
  59959. 44:43:13Right? This is the dimension 6 + 6. And
  59960. 44:43:16this value after comma is basically the
  59961. 44:43:18value which you want. You execute this
  59962. 44:43:21and you copy this paste this and you
  59963. 44:43:24will get the arrays of fours for
  59964. 44:43:27yourself. Right? If you want of 10, you
  59965. 44:43:30will get of 10. If you want 10.3,
  59966. 44:43:34you will get 10.3. Right? anything which
  59967. 44:43:37you want. If you want case, you will get
  59968. 44:43:40case, right? All the examples. So, let
  59969. 44:43:43me just show that to you.
  59970. 44:43:55Yep. Like this. Now guys, how did we
  59971. 44:43:58create
  59972. 44:44:01or how did we use
  59973. 44:44:04range in Python?
  59974. 44:44:07Can you use
  59975. 44:44:09range to generate
  59976. 44:44:13numbers between 20 to 50,
  59977. 44:44:19right? 20 to 50. Can you give me the
  59978. 44:44:22syntax quickly? How did you do that in
  59979. 44:44:24range?
  59980. 44:44:26How do we do that? We said R is equal to
  59981. 44:44:30range
  59982. 44:44:3220 to 51. Right? And then I said
  59983. 44:44:39I in R
  59984. 44:44:41print I,
  59985. 44:44:44right? And this is how I got the
  59986. 44:44:45numbers, right? So similar
  59987. 44:44:49to range in Python,
  59988. 44:44:52we have
  59989. 44:44:54a range in num py. Right? How do we use
  59990. 44:44:58a range? I say
  59991. 44:45:01uh a range
  59992. 44:45:05ar r is equal to np dot arange. Right?
  59993. 44:45:10And then same syntax I will say 20 to
  59994. 44:45:1351. Right? 20 to 51. And that's it. And
  59995. 44:45:17when I will check my AR range error, you
  59996. 44:45:21will see I have generated myself numbers
  59997. 44:45:23between 20 to 50 and a range in numpy.
  59998. 44:45:28Right? A range in numpy. Yes, if you
  59999. 44:45:32want a interval so you can use this say
  60000. 44:45:37a range one and after comma you pass the
  60001. 44:45:40third argument. Suppose it's three. So
  60002. 44:45:42now it will jump three times, right? 20
  60003. 44:45:4623 26 29 32 35 like this up till 50.
  60004. 44:45:52Now guys there is something which is
  60005. 44:45:54called as lind space
  60006. 44:46:00right. What is lindspace?
  60007. 44:46:02It stands for
  60008. 44:46:06linear spacing
  60009. 44:46:08which means
  60010. 44:46:10between two given numbers.
  60011. 44:46:15This function will fit the required
  60012. 44:46:21number of
  60013. 44:46:23numbers. Right? For example, suppose for
  60014. 44:46:27example,
  60015. 44:46:29we need to create an interval
  60016. 44:46:36from 0 to 1. People, there are infinite
  60017. 44:46:41numbers I can have between 0 to 1. Isn't
  60018. 44:46:44it?
  60019. 44:46:45Infinite numbers I can have between 0 to
  60020. 44:46:481. 0.0000000000001
  60021. 44:46:520 0 1 0 1 01 right I can go in the
  60022. 44:46:56infinite manner right now for example
  60023. 44:47:00you need to create numbers between 0 to
  60024. 44:47:0410 right and you want to create and want
  60025. 44:47:08to have
  60026. 44:47:1010 numbers in it right so how will you
  60027. 44:47:13do this it's not float it's about the
  60028. 44:47:16number theory right between 0 and one
  60029. 44:47:18you have infinite finite numbers, right?
  60030. 44:47:20So you say lindspace is equal to np dot
  60031. 44:47:25lindspace, right? np.tlind space. You
  60032. 44:47:28mention from 0 to 10, you want to have
  60033. 44:47:3110 numbers, right? And when you will
  60034. 44:47:34create lindspace,
  60035. 44:47:35you will see that these are the numbers
  60036. 44:47:38are there which have been created,
  60037. 44:47:40right? These are the numbers which have
  60038. 44:47:42been created,
  60039. 44:47:44right? Nine numbers. Now I say 100
  60040. 44:47:47numbers. I want evenly spaced 100
  60041. 44:47:50numbers, right? Evenly spaced 100
  60042. 44:47:53numbers. How are they even? You can
  60043. 44:47:55simply subtract one number from another
  60044. 44:47:57and the difference for all the numbers
  60045. 44:47:58will be exactly the same. 0.01 0 1 01.
  60046. 44:48:03Subtract any two numbers. It will be
  60047. 44:48:040.01 0 1 01. Right? Where do we need
  60048. 44:48:07this? We need this to plot the axises.
  60049. 44:48:10Right? When you will plot graphs, you
  60050. 44:48:12will need access between this interval.
  60051. 44:48:14You need five values that works like
  60052. 44:48:16this. Okay? Suppose you want from 0 to
  60053. 44:48:1910 five different values, right? You
  60054. 44:48:21will have five different values like
  60055. 44:48:23this, right? Between the gap of 0.5,
  60056. 44:48:27right? If I say 1 to 10, you will have
  60057. 44:48:30values like this,
  60058. 44:48:33right? Like this. Clear? 100 values like
  60059. 44:48:36this. Yeah. Clear guys. How do we use
  60060. 44:48:38lin space? Suppose you want from 0 to
  60061. 44:48:4210. Interval from 0 to 10 and 10 will be
  60062. 44:48:44included. Zero will not be included.
  60063. 44:48:47Right? We'll start from one. So I go to
  60064. 44:48:50one it will be from
  60065. 44:48:55one. Why is 0 not included then? Yeah. 0
  60066. 44:48:59is included. Right? 0 is also included
  60067. 44:49:01and 10 is also included. 100 numbers
  60068. 44:49:04between them. Right?
  60069. 44:49:06Clear? This is what lin space is. Now
  60070. 44:49:09guys, now
  60071. 44:49:12suppose you
  60072. 44:49:15lohan l space is basically used if you
  60073. 44:49:20want to create n numbers between the
  60074. 44:49:23range of numbers right between 0 to 1
  60075. 44:49:27right suppose between 0 to 1 you are
  60076. 44:49:30trying to plot a graph okay and your
  60077. 44:49:32values are 0.2 0.3 0.6 six right and you
  60078. 44:49:37want to draw a graph so you will have to
  60079. 44:49:39mark the axis right the x axis and the
  60080. 44:49:41y- axis so you can use lindspace there
  60081. 44:49:44and what will it do it will take the
  60082. 44:49:46range it will take the interval in
  60083. 44:49:48between you want to add the equal space
  60084. 44:49:51numbers and then the third argument here
  60085. 44:49:54will be that how many numbers do you
  60086. 44:49:57want between them so this syntax tells
  60087. 44:50:00you that
  60088. 44:50:02from
  60089. 44:50:040 to 1
  60090. 44:50:06give me 100 numbers. How are these
  60091. 44:50:09numbers? Equally
  60092. 44:50:12spaced
  60093. 44:50:13numbers. Equally spaced numbers, right?
  60094. 44:50:16So when you will execute this, you will
  60095. 44:50:19see that all there are 100 numbers which
  60096. 44:50:21have been generated, right? 100 numbers.
  60097. 44:50:23And all the numbers are equidistant from
  60098. 44:50:26each other because difference of every
  60099. 44:50:28single number from the next number is
  60100. 44:50:300.01
  60101. 44:50:3201.
  60102. 44:50:38Yep, that's what it does. Right
  60103. 44:50:43now guys, now suppose we want to
  60104. 44:50:48generate
  60105. 44:50:50random numbers, right? We want to
  60106. 44:50:52generate random numbers, right? Now
  60107. 44:50:54we're interested in generating random
  60108. 44:50:55numbers. So we have something called as
  60109. 44:51:00random
  60110. 44:51:02dot random right. What will it do?
  60111. 44:51:05Random.random
  60112. 44:51:06will generate
  60113. 44:51:10random
  60114. 44:51:11float numbers
  60115. 44:51:15between 0 to 1. Random float numbers
  60116. 44:51:18between 0 to 1. How will this happen?
  60117. 44:51:21You will say
  60118. 44:51:23rand rand is equal to np do. random dot
  60119. 44:51:29random and inside you will mention what
  60120. 44:51:32is the dimension that you seek. Suppose
  60121. 44:51:34I want 6 + 6. So when you will check
  60122. 44:51:38this you will have all numbers for 6 + 6
  60123. 44:51:42dimension right 6 + 6 matrix right now
  60124. 44:51:47every time you rerun this the numbers
  60125. 44:51:48will change because all of these are
  60126. 44:51:50random numbers
  60127. 44:51:52right all of these are random numbers
  60128. 44:51:56now I say 100 multiplied by rand rand
  60129. 44:52:02you will see all of them all these
  60130. 44:52:04numbers will be multiplied by 00 right
  60131. 44:52:07all of them
  60132. 44:52:09in one shopping.
  60133. 44:52:12Now just like this we can also create
  60134. 44:52:17random integers right how will we create
  60135. 44:52:20random integers guys
  60136. 44:52:23I say rand intore
  60137. 44:52:26rand is equal to np dot random dot rand
  60138. 44:52:33right and here you will specify that
  60139. 44:52:37what is the range of numbers you want
  60140. 44:52:39from so I say between 20 to 25 I need
  60141. 44:52:44random numbers and then I want it from
  60142. 44:52:48in a 3 + 3 format. Right? And now when
  60143. 44:52:51you will check your random you will get
  60144. 44:52:54random numbers generated like this.
  60145. 44:52:56Okay? Random numbers generated like
  60146. 44:52:58this.
  60147. 44:53:04Right?
  60148. 44:53:06If you say 3 + 3 + 3 you will get a 3 +
  60149. 44:53:103 + 3 matrix. Even if you will only say
  60150. 44:53:123, you will get a 1D.
  60151. 44:53:17Right guys? You can change the
  60152. 44:53:19dimension. So this is the range from
  60153. 44:53:22which you want to choose the random
  60154. 44:53:24numbers and this is the dimension you
  60155. 44:53:26want this matrix or vector to be in.
  60156. 44:53:30Now guys, we'll move on to the next part
  60157. 44:53:33which is basically properties and again
  60158. 44:53:37there are a lot of operations you'll
  60159. 44:53:38have to see it yourself right
  60160. 44:53:42properties and
  60161. 44:53:46attributes
  60162. 44:53:49of numpy
  60163. 44:53:53arrays right property and attributes of
  60164. 44:53:56numpy arrays.
  60165. 44:53:59Okay. Now guys, the first one in this
  60166. 44:54:02scheme of things is shape of array,
  60167. 44:54:06right? Shape of array. What is shape of
  60168. 44:54:08array? It tells you
  60169. 44:54:12the
  60170. 44:54:14dimensions of the array
  60171. 44:54:19stored in a
  60172. 44:54:23tle. Right?
  60173. 44:54:25For example, I say a ar r r r r r r r r
  60174. 44:54:29r r r r r r r r r r r r0
  60175. 44:54:30right a r r r r r r r r r r r r r r r r
  60176. 44:54:32r r r r r 1 a ar a ar a ar a ar a ar a
  60177. 44:54:34ar a ar a ar a ar a ar a r r r r r r r r
  60178. 44:54:34r r r r r r r r r r r r r 2 a ar r r3
  60179. 44:54:38right and then I say
  60180. 44:54:44print this
  60181. 44:54:47dot shape
  60182. 44:54:50right like this and you will see that it
  60183. 44:54:52will give you the shape of each array
  60184. 44:54:55Right? 0D, 1D, 2D and 3D. Right people?
  60185. 44:55:00Shape of the array.
  60186. 44:55:04Please try it out. We have used it one
  60187. 44:55:07or two times. But this is what shape of
  60188. 44:55:09array actually means.
  60189. 44:55:11Second is people
  60190. 44:55:14end
  60191. 44:55:16right is end
  60192. 44:55:20right end of array right it tells you
  60193. 44:55:26the rank of the array whether it's one
  60194. 44:55:30dimensional two dimensional zero
  60195. 44:55:31dimensional threedimensional four
  60196. 44:55:32dimensional
  60197. 44:55:34so again I will do the same and
  60198. 44:55:46right I'll say end and you will see it
  60199. 44:55:48will give you 0 1 2 3 zero dimension
  60200. 44:55:51zero rank one rank two rank and three
  60201. 44:55:53rank and it can go all the way up to end
  60202. 44:55:56rank
  60203. 44:55:58before this I should have also
  60204. 44:56:02printed these arrays
  60205. 44:56:09Right. These are the arrays.
  60206. 44:56:22Yep.
  60207. 44:56:24These are the arrays which we have and
  60208. 44:56:26these are the subsequent things, right?
  60209. 44:56:30Rank copy array. Then guys, the third
  60210. 44:56:32thing is the size of
  60211. 44:56:37array, right? It tells you
  60212. 44:56:42the number of elements inside. How many
  60213. 44:56:47elements do we have inside this array in
  60214. 44:56:49a ar r1? How many elements do we have? 1
  60215. 44:56:522 3 4 5 6 7. How many elements do we
  60216. 44:56:56have in this 2 + 2 ar2? 1 2 3 4 5 6 7 8
  60217. 44:57:009 which is rows multiplied by columns 3
  60218. 44:57:02* 3 right and how many elements do we
  60219. 44:57:05have in this 3D which is 2 * 3 * 3 which
  60220. 44:57:10is 18 right so now when you copy this
  60221. 44:57:14right you can use this
  60222. 44:57:17and say
  60223. 44:57:20size right it will say 17 918
  60224. 44:57:48Yep.
  60225. 44:57:57Fourth is people.
  60226. 44:58:00The D type
  60227. 44:58:03of array tells you the data type of the
  60228. 44:58:10array. Right? And I've told you we place
  60229. 44:58:14only
  60230. 44:58:16homogeneous
  60231. 44:58:18data in array right what will happen if
  60232. 44:58:22we don't do this I will show that to you
  60233. 44:58:24also right so we do this right and we
  60234. 44:58:28say
  60235. 44:58:32retype
  60236. 44:58:34and you will see in 64 all of them are
  60237. 44:58:37integers right all of them are integers
  60238. 44:58:40that's the reason we are getting in 64
  60239. 44:58:42right suppose I create a new array a ar
  60240. 44:58:45r new right let me say head
  60241. 44:58:50right hetro
  60242. 44:58:53heterogenous
  60243. 44:58:54and I say it is like n dot array
  60244. 44:59:01let's say like this
  60245. 44:59:07right like this now people when you will
  60246. 44:59:10Check
  60247. 44:59:12ar r
  60248. 44:59:16dot d type you will see it will give you
  60249. 44:59:18float just because of one floating point
  60250. 44:59:21number inside this entire array it gets
  60251. 44:59:24converted to float right between all the
  60252. 44:59:27integers if you put one float then it
  60253. 44:59:29will be taking float directly right now
  60254. 44:59:34let me just copy this and let's say
  60255. 44:59:36heterogenous one and let me add another
  60256. 44:59:38value which is string and say rather
  60257. 44:59:42right and when you will execute this it
  60258. 44:59:44will give you U32 U32 here is
  60259. 44:59:46representing strings right it is all
  60260. 44:59:48called as objects right these are all
  60261. 44:59:50string values right so precedences
  60262. 44:59:53strings greatest then float and then
  60263. 44:59:56your uh integers right if you place the
  60264. 45:00:00heterogenous data inside the numpy array
  60265. 45:00:04right you only need to put homogeneous
  60266. 45:00:06data in the array
  60267. 45:00:08now fifth is the item size
  60268. 45:00:12of array. Right? What is item size?
  60269. 45:00:17It gives you
  60270. 45:00:20the bite
  60271. 45:00:22occupied
  60272. 45:00:24by each element of an array. Right?
  60273. 45:00:29Because we assume that elements will be
  60274. 45:00:31homogeneous. It will give you the bite
  60275. 45:00:34occupied by each element of the array.
  60276. 45:00:37Only one element. Okay? So how will it
  60277. 45:00:39happen? So let's say a ar r r0
  60278. 45:00:43or let's say ar r r1 dot item size
  60279. 45:00:48right and you will get eight right. So
  60280. 45:00:51why eight? Because
  60281. 45:00:55because each data point right each data
  60282. 45:00:58point is occupying
  60283. 45:01:01the result is 8 bytes. Let me put this
  60284. 45:01:07here.
  60285. 45:01:13Right. Eight bytes
  60286. 45:01:15because each element in ARR1 is
  60287. 45:01:22occupying
  60288. 45:01:2464 bits which are
  60289. 45:01:30equivalent to which are equivalent to 8
  60290. 45:01:34bytes. Right? one bite is equal to 8
  60291. 45:01:37bits. So 64 bits will be equal to 8
  60292. 45:01:41bytes. Right? That is how it is giving
  60293. 45:01:43you the result. Now if you are
  60294. 45:01:45interested in knowing the entire bytes
  60295. 45:01:49right entire bytes then you say n bytes
  60296. 45:01:56will give you the
  60297. 45:02:00total bytes
  60298. 45:02:02occupied
  60299. 45:02:04by the elements of the array. Right? You
  60300. 45:02:08say print
  60301. 45:02:11a ar r1 dot
  60302. 45:02:13n bytes right and write
  60303. 45:02:18bytes it will be 56 bytes
  60304. 45:02:24right why because how many elements do
  60305. 45:02:26we have in our ar r1 1 2 3 4 5 6 7 right
  60306. 45:02:327 8 are 56 right 56 total bytes are
  60307. 45:02:36being occupied with by ar r1 one. Now
  60308. 45:02:39guys, the seventh one
  60309. 45:02:42is
  60310. 45:02:44as type right
  60311. 45:02:48in array.
  60312. 45:02:51This will help you change the data type
  60313. 45:02:57of the array. Right? Change the data
  60314. 45:03:00type of the array. Suppose I have a arr
  60315. 45:03:04type which is uh so I'll say print
  60316. 45:03:10d type right this is end 64 right and
  60317. 45:03:13now what I do is I say print
  60318. 45:03:20uh wait let me give you a structured way
  60319. 45:03:22print a r1 right let me say
  60320. 45:03:28array
  60321. 45:03:33R1
  60322. 45:03:39right D type of array.
  60323. 45:03:43Now
  60324. 45:03:46a ar r2 sorry a ar a ar a ar a ar a ar a
  60325. 45:03:47ar a ar a ar a ar a ar a r r r r r r r r
  60326. 45:03:48r r r r r r r r r r r r r1 is equal to a
  60327. 45:03:50ar r1
  60328. 45:03:52dot as type right dot as type and let's
  60329. 45:03:55say I want to convert this in np dot
  60330. 45:04:00right np dot
  60331. 45:04:02uh
  60332. 45:04:04int 32 right I want to create convert
  60333. 45:04:07this in uh ar r r int 32 right when I do
  60334. 45:04:11this and now when I will copy these same
  60335. 45:04:15things you will see for yourself. Right?
  60336. 45:04:18Now the D type was int 64 and now the DT
  60337. 45:04:20type is int 32. Right?
  60338. 45:04:24If I want I can do this conversion in
  60339. 45:04:27float also
  60340. 45:04:30float 64. Right? And then I will just
  60341. 45:04:33copy this
  60342. 45:04:35and I will paste it here.
  60343. 45:04:39Right? And now you will see now the
  60344. 45:04:42floating point has been activated.
  60345. 45:04:45Right? It has been now activated. We can
  60346. 45:04:48go till int. We can go till int 8.
  60347. 45:04:54Right?
  60348. 45:04:56I can go to 16.
  60349. 45:04:59Right? And I can do this.
  60350. 45:05:04I can go to int
  60351. 45:05:078 also.
  60352. 45:05:12Yeah, like this in date also 64 32.
  60353. 45:05:18So first one has to be
  60354. 45:05:22Yeah.
  60355. 45:05:2832
  60356. 45:05:38whatever right like this okay you can
  60357. 45:05:42convert this
  60358. 45:05:45also people this is later on conversion
  60359. 45:05:48you can define the data type of the
  60360. 45:05:54array while creation time also. How
  60361. 45:05:59would you do that? Suppose you are
  60362. 45:06:00creating a ar r11 and you say np dot
  60363. 45:06:04array right and suppose you take it from
  60364. 45:06:06a list and then you just put a comma and
  60365. 45:06:09say d type. So what will be the default
  60366. 45:06:12data type here people? If I just do this
  60367. 45:06:14if I just execute this what will be the
  60368. 45:06:17default data type? Int 64 is the default
  60369. 45:06:21isn't it? But now suppose I want to
  60370. 45:06:24change it right here. I say data type is
  60371. 45:06:26equal to float 32 right sorry float 64
  60372. 45:06:36np dot
  60373. 45:06:38sorry my bad
  60374. 45:06:42float 64 right and I say a ar r r11 and
  60375. 45:06:46this will be float 64 if you want float
  60376. 45:06:4932 it will also become float 32 right
  60377. 45:06:52right here while you define
  60378. 45:06:55Instead of using as type, you can do it
  60379. 45:06:57right here. Right? These things will
  60380. 45:06:58come in very handy people because you
  60381. 45:07:00will have to save memory because when
  60382. 45:07:02your data becomes very very big, you
  60383. 45:07:03will be always in a crunch for memory
  60384. 45:07:06like this. You want integers, then
  60385. 45:07:08integers will be like this
  60386. 45:07:11random.randent
  60387. 45:07:12like this. Suppose you want to generate
  60388. 45:07:160 to six, right? And suppose you want to
  60389. 45:07:19generate
  60390. 45:07:21say 100 numbers like this 0 to 6 the
  60391. 45:07:27scores
  60392. 45:07:281 to six like this randomly
  60393. 45:07:36right suppose you want to generate 100
  60394. 45:07:39scores for five different batsmen
  60395. 45:07:42randomly it will be like this batsman
  60396. 45:07:45number one batsman number to bat number
  60397. 45:07:48three, fourth and fifth. Right.
  60398. 45:07:54Yep.
  60399. 45:07:59Understand the data guys. Now it's the
  60400. 45:08:01time to understand the data.
  60401. 45:08:03Yep. Now guys, we have methods in numpy
  60402. 45:08:10arrays, right? Methods in numpy arrays.
  60403. 45:08:13So what are these methods? Now the first
  60404. 45:08:16method we have to learn is called as
  60405. 45:08:19reshape right. Reshape.
  60406. 45:08:29Yeah. So reshape is you can use this to
  60407. 45:08:31create
  60408. 45:08:34you can use this to create
  60409. 45:08:37a new shape of the array. Very very
  60410. 45:08:41powerful guys. Very powerful. One of the
  60411. 45:08:43most powerful methods in numpy is uh the
  60412. 45:08:47reshape right and how do we use reshape
  60413. 45:08:50suppose
  60414. 45:08:54we have a 1D array of 20 elements right
  60415. 45:09:01now to reshape this
  60416. 45:09:05reshape it we need to find the factors
  60417. 45:09:12Right. Factors of 20. They are what? 1
  60418. 45:09:1720 4 5
  60419. 45:09:212 10.
  60420. 45:09:23Right. The other factors.
  60421. 45:09:27The other factors. Now see what will I
  60422. 45:09:29do. Right. Now see what will I do. Let
  60423. 45:09:31me create a say random array. Right?
  60424. 45:09:34Random array. I say random
  60425. 45:09:39arr is equal to entprandom
  60426. 45:09:43dot rand right and let me say I want to
  60427. 45:09:47create it from 1 to 50 right and I want
  60428. 45:09:52it to be having 20 elements right so I
  60429. 45:09:57say random arrand
  60430. 45:10:00values inside this right random 20
  60431. 45:10:02values now see Now reshape
  60432. 45:10:08first
  60433. 45:10:11I will reshape in 1 + 20 right 1A 20 how
  60434. 45:10:18will I do that you just have to write
  60435. 45:10:22uh
  60436. 45:10:26print
  60437. 45:10:28a ar r r sorry sorry random dot ar r
  60438. 45:10:31random ar
  60439. 45:10:33dot reshape dot reshape and you just
  60440. 45:10:37pass in the dimension I say 1 20
  60441. 45:10:41right and when you will do this you will
  60442. 45:10:43see it is coming now in 1 20 format
  60443. 45:10:47right so let me just also write print
  60444. 45:11:02right 2D
  60445. 45:11:041A 20
  60446. 45:11:07right and now what I'll do is I'll copy
  60447. 45:11:10this and I will paste this and say 20
  60448. 45:11:15comma 1 right you will see it will be
  60449. 45:11:18like this 20 comma 1 immediately with
  60450. 45:11:21reshape right let me copy this
  60451. 45:11:25let me say 2D I'm still at 2D let me say
  60452. 45:11:292 10 right and you to see this is 2A 10.
  60453. 45:11:35Now I can reshape it in
  60454. 45:11:3910 2
  60455. 45:11:43right 10 2 right then I can reshape the
  60456. 45:11:48same thing
  60457. 45:11:52in
  60458. 45:11:544A 5 and I can reshape this in 5A 4
  60459. 45:11:58right like this guys are you able to see
  60460. 45:12:01the power one dimension I'm able to
  60461. 45:12:04create two dimensions
  60462. 45:12:05And now I will take it a step further
  60463. 45:12:08and I will write it in three dimensions.
  60464. 45:12:11Right? How will I write it in three
  60465. 45:12:12dimension? Let me say this 1 comma
  60466. 45:12:172a 10. Right? This is will also be three
  60467. 45:12:20dimensions. Let's sorry 2a 2a 5.
  60468. 45:12:24Let me say this. And now you will see I
  60469. 45:12:26can have this in three dimensions.
  60470. 45:12:29Right?
  60471. 45:12:32Yep. I can also say in three dimensions
  60472. 45:12:36like this. I want to have five layers
  60473. 45:12:40with two rows and two columns. Right? So
  60474. 45:12:43you will have it like this also. Right?
  60475. 45:12:46So this is how people we can reshape the
  60476. 45:12:49array. Very powerful. Very very
  60477. 45:12:51powerful.
  60478. 45:12:53Right? Very very powerful.
  60479. 45:12:57Right? And you can take this
  60480. 45:13:04and save print
  60481. 45:13:09random dot this and you can say print
  60482. 45:13:14dot shape.
  60483. 45:13:16Right? So this was the first one.
  60484. 45:13:27Yep. like this.
  60485. 45:13:29Also people if you want to visualize we
  60486. 45:13:32can also go this route.
  60487. 45:13:35We can have 10 comma 2 comma 1, right?
  60488. 45:13:40It will look like this, right? 10
  60489. 45:13:43layers. 10 layers you can have,
  60490. 45:13:47right? 10 layers you can have.
  60491. 45:13:52Great. Now, second method which we have
  60492. 45:13:54to learn is called as
  60493. 45:13:58transpose,
  60494. 45:14:00right? Transpose method. Right? What
  60495. 45:14:03does that do? It interchanges the
  60496. 45:14:07dimensions
  60497. 45:14:10like rows and columns, right? Yes.
  60498. 45:14:15Absolutely. Right. Suppose you have a
  60499. 45:14:18matrix, right? Which is
  60500. 45:14:2222, 33, 44, 55, 66, 77, right? This is
  60501. 45:14:29a. So now when you will a transpose it,
  60502. 45:14:32right? The dimension right now the shape
  60503. 45:14:35right now is 3A 2. Now this will become
  60504. 45:14:392a 3. And how this will happen? Rows
  60505. 45:14:42will now become columns and columns will
  60506. 45:14:44now become rows. Right? So let's make
  60507. 45:14:47column the rows. Right? Sorry columns
  60508. 45:14:49the rows. It will be 22 44 66
  60509. 45:14:5533 55 77. Right? So people in transpose
  60510. 45:15:00no information is lost. It is just a
  60511. 45:15:03change in the view right which is
  60512. 45:15:07happening right?
  60513. 45:15:1022 44 66 33 55 77.
  60514. 45:15:19Why do we need transposition? Suppose we
  60515. 45:15:23have two matrix.
  60516. 45:15:25This is 11th class mathematics. Right?
  60517. 45:15:28one has
  60518. 45:15:30a dimension of n cross m and the second
  60519. 45:15:33has dimension of a cross b. If you
  60520. 45:15:40want to multiply
  60521. 45:15:46these two matrix say
  60522. 45:15:50M_sub_1 and M_sub_2,
  60523. 45:15:54right? There needs to be a satisfaction
  60524. 45:15:56of condition. M should be equal to A.
  60525. 45:16:00Right? M should be equal to A. Right? M
  60526. 45:16:04should be equal to A. This should be
  60527. 45:16:05equal to this and the resultant vector
  60528. 45:16:08the resultant matrix which you will get
  60529. 45:16:10will be of n crossb dimension right. So
  60530. 45:16:14often times suppose this is n cross m
  60531. 45:16:17this is n cross m right m cross n and
  60532. 45:16:21you know that m is equal to a. So what
  60533. 45:16:24will you do? You will transpose this
  60534. 45:16:26matrix right? You will transpose this
  60535. 45:16:28matrix then it will become n cross m and
  60536. 45:16:31then m can be equivalent to a. Right?
  60537. 45:16:34For example, what I'm saying, we have
  60538. 45:16:36one matrix which is 2 + 3 and this
  60539. 45:16:38matrix is 2 + 5, right? So, can you
  60540. 45:16:42multiply these matrix people? Is 3 equal
  60541. 45:16:45to 2? The answer is no. Right? The
  60542. 45:16:48answer is no. So, what will you do? You
  60543. 45:16:50will just transpose this and this will
  60544. 45:16:52become 3 + 2 and this is 2 + 5. And now
  60545. 45:16:57you can multiply this and the resultant
  60546. 45:16:59will become 3 + 5 matrix. Right? So for
  60547. 45:17:04operations like these we need
  60548. 45:17:06transposition. Right? So how do we
  60549. 45:17:08transpose it?
  60550. 45:17:11How do we transpose it? So let's say
  60551. 45:17:14again
  60552. 45:17:15uh a ar r2 right this is a ar r2 and I
  60553. 45:17:20want to transpose it. So I say a r2 t is
  60554. 45:17:24equal to uh np.transpose transpose
  60555. 45:17:29sorry a r2 dot
  60556. 45:17:33transpose
  60557. 45:17:35right
  60558. 45:17:37and now when you will see ar r2 ts you
  60559. 45:17:41will see rows and columns have
  60560. 45:17:43interchanged right rows and columns have
  60561. 45:17:46interchanged with each other
  60562. 45:17:52or let me give you one more example
  60563. 45:17:56Uh if this is not clear, let me pick up
  60564. 45:18:02this again
  60565. 45:18:14right now. I say
  60566. 45:18:17dot reshape
  60567. 45:18:20into say
  60568. 45:18:232 + 10. Right? So this is 2 + 10. And
  60569. 45:18:26now when you want to transpose this so
  60570. 45:18:28I'll say this t is equal to this dot
  60571. 45:18:34transpose
  60572. 45:18:38t or a this and we can check this now
  60573. 45:18:44and it will be this. Sorry guys. So this
  60574. 45:18:48is going to be it will be like this
  60575. 45:18:51right transposed.
  60576. 45:18:53And now if you want to see this,
  60577. 45:19:03this was the original shape, right? Rows
  60578. 45:19:05and columns have now been interchanged,
  60579. 45:19:08transposed with each other.
  60580. 45:19:12Now guys, the third method,
  60581. 45:19:16the third method which is there with us
  60582. 45:19:18is called as flatten, right? It is
  60583. 45:19:21called as flatten.
  60584. 45:19:25Right? What does flatten do? It reduces
  60585. 45:19:29the dimension to one dimension. Right?
  60586. 45:19:33Any dimension you have, it reduces it to
  60587. 45:19:36one dimension. For example, I have this,
  60588. 45:19:40right? And now when I say
  60589. 45:19:44this
  60590. 45:19:46dot_f,
  60591. 45:19:48this will become this dot platin.
  60592. 45:19:53And when you will check this up, you
  60593. 45:19:56will see that it has now become one
  60594. 45:19:57dimension. No matter how many dimensions
  60595. 45:19:59you have, it will become one dimension.
  60596. 45:20:03Right? Let me take this again to show
  60597. 45:20:05you one more example.
  60598. 45:20:11Right?
  60599. 45:20:13And here I say dot reshape into
  60600. 45:20:18uh
  60601. 45:20:21uh 5 + 2 + 2 right I do this my random
  60602. 45:20:27ar r is this right it has five layers
  60603. 45:20:30two rows and two columns right so now I
  60604. 45:20:34say this
  60605. 45:20:36flatten is equal to this dot flatten
  60606. 45:20:41and If you will check it now again, you
  60607. 45:20:44will see it has now flattened it out.
  60608. 45:20:46Yep.
  60609. 45:20:49Now why do we need this? We need this
  60610. 45:20:51for a lot of statistical operations. We
  60611. 45:20:53need this to feed the data into the
  60612. 45:20:56algorithms. Right? As we will move
  60613. 45:20:59forward, you will understand the use of
  60614. 45:21:00flattening.
  60615. 45:21:02Right? Now guys, moving on and uh as
  60616. 45:21:06discussed, let me now show you
  60617. 45:21:10the power of
  60618. 45:21:13numpy,
  60619. 45:21:16right? Numpy
  60620. 45:21:18over
  60621. 45:21:20lists and other data types, right? I
  60622. 45:21:25will not take a lot of examples. Just a
  60623. 45:21:27second, guys.
  60624. 45:21:29Yeah. Okay.
  60625. 45:21:32Now guys, I told you that
  60626. 45:21:38less
  60627. 45:21:41take up
  60628. 45:21:44much more memory
  60629. 45:21:48as compared
  60630. 45:21:50to numpy arrays. Right? And I'm going to
  60631. 45:21:53prove this to you. Right? Now let me use
  60632. 45:21:58let me create a random sequence of
  60633. 45:22:00random numbers using range in Python
  60634. 45:22:04right and say I create range of 10,000
  60635. 45:22:08numbers right range of 10,000 numbers so
  60636. 45:22:11what will this give me this will give me
  60637. 45:22:13numbers from 0 to 99999 right continuous
  60638. 45:22:16numbers right so this is range I will
  60639. 45:22:20use
  60640. 45:22:22a range in
  60641. 45:22:24numpy Y to create
  60642. 45:22:28a similar
  60643. 45:22:32series of numbers
  60644. 45:22:35right so let's say array is equal to np
  60645. 45:22:39dot arange
  60646. 45:22:42right same thing same done by both right
  60647. 45:22:45I've shown you above also now let me
  60648. 45:22:48import sis
  60649. 45:22:51library right sis package and I will use
  60650. 45:22:55something called as get size of right
  60651. 45:22:58get size of. What does this do? Get size
  60652. 45:23:01of it's a method
  60653. 45:23:05which calculates
  60654. 45:23:07the bytes
  60655. 45:23:10occupied
  60656. 45:23:12by a single
  60657. 45:23:15element in
  60658. 45:23:18vanilla Python. What is vanilla Python?
  60659. 45:23:20It is the traditional Python,
  60660. 45:23:22right? vanilla Python.
  60661. 45:23:25So let me just show that to you. I'll
  60662. 45:23:26use this and I will say print. Now guys,
  60663. 45:23:30if I get
  60664. 45:23:33size of any random number from this
  60665. 45:23:36range, right? Any random number. Say I
  60666. 45:23:39get size of five, right? And I then
  60667. 45:23:43multiply that byes with the length of
  60668. 45:23:46random, right? With the length of random
  60669. 45:23:50this rand, right?
  60670. 45:23:53Right? With the length of random, do you
  60671. 45:23:55think I will get the bytes for the
  60672. 45:23:58entire
  60673. 45:23:59data structure? What am I saying is
  60674. 45:24:02suppose
  60675. 45:24:04uh I used range
  60676. 45:24:08five. So what will this give me? 0 1 2 3
  60677. 45:24:124. Right? This will be the output. So
  60678. 45:24:15now I say get
  60679. 45:24:18size of say I say two. Right? So suppose
  60680. 45:24:212 is x and then I multiply this with the
  60681. 45:24:25length of this series which is five. So
  60682. 45:24:28do you think I will get 5x which will
  60683. 45:24:31represent the number of bytes occupied
  60684. 45:24:33by the entire data type. Anything
  60685. 45:24:35randomly any random number this can be
  60686. 45:24:38three right? Why not hurry?
  60687. 45:24:42Why not?
  60688. 45:24:47All of these are integers. So integers
  60689. 45:24:50all of 64 bits assuming. So if you
  60690. 45:24:54calculate the side of size of this and
  60691. 45:24:56if you multiply with the total number of
  60692. 45:24:57numbers you will get the total size
  60693. 45:25:00isn't it?
  60694. 45:25:04Huh? Index is in
  60695. 45:25:10no no it's not about that it's about the
  60696. 45:25:12element right homogeneous elements
  60697. 45:25:13inside this.
  60698. 45:25:16I am saying when you use range five what
  60699. 45:25:19is going to be the output? 0 1 2 3 4
  60700. 45:25:22right now all these are elements
  60701. 45:25:27elements of range.
  60702. 45:25:31Right? All of them are elements of
  60703. 45:25:32range. Right? Now I'm saying if I fetch
  60704. 45:25:36the size of one element and multiply it
  60705. 45:25:41with the length of the entire range,
  60706. 45:25:45will I get the bytes occupied by the
  60707. 45:25:48entire range? For example, if I do this,
  60708. 45:25:52right? If I do this,
  60709. 45:25:56this is 28, right? 28 bytes people. 28
  60710. 45:26:00bytes
  60711. 45:26:02bytes are occupied
  60712. 45:26:06by one element of range right
  60713. 45:26:13one element of
  60714. 45:26:16range right now if I just I'm saying I'm
  60715. 45:26:20just saying if I multiply to find how
  60716. 45:26:24many numbers range has
  60717. 45:26:27how many numbers
  60718. 45:26:29range has
  60719. 45:26:31equal to 10,000
  60720. 45:26:35right so total
  60721. 45:26:39memory occupied
  60722. 45:26:41will be will be how much it will be 28
  60723. 45:26:48ult*lied by 10,000 which will be equal
  60724. 45:26:51to 28,000
  60725. 45:26:53yeah and how will you find this you will
  60726. 45:26:55say Print
  60727. 45:27:00this multiplied by length of RAM,
  60728. 45:27:07right? 28,000 bytes. Clear? Now, yes.
  60729. 45:27:11Now, this is for the range. Now, let me
  60730. 45:27:13use another thing. So, how will you
  60731. 45:27:15calculate the length of this array? What
  60732. 45:27:18what property and attribute will you use
  60733. 45:27:21people?
  60734. 45:27:22N bytes, right? N bytes will give you
  60735. 45:27:26total bytes occupied by the elements of
  60736. 45:27:27the array. Right? We will use n bytes
  60737. 45:27:29here. So I come back down and I say
  60738. 45:27:33using n bytes for arrays. Right? And you
  60739. 45:27:38will see what the result comes. Print
  60740. 45:27:43uh array dot n bytes. Right? And I say
  60741. 45:27:50bytes. Are you ready to see the result?
  60742. 45:27:52Do you see what has happened?
  60743. 45:27:55How many bytes this was taking? It was
  60744. 45:27:57taking 28,000 bytes. How many bytes this
  60745. 45:28:00is taking? This is taking 80,000 bytes.
  60746. 45:28:05This was taking 2 lakh 80,000. This is
  60747. 45:28:07taking 80,000. Two lakh extra bytes of
  60748. 45:28:10memory is taken by range.
  60749. 45:28:15Ran is range.
  60750. 45:28:24Yep. And if I just go to million
  60751. 45:28:28numbers,
  60752. 45:28:29right? Million numbers in both.
  60753. 45:28:34See the difference it becomes,
  60754. 45:28:37right? This is now 3 three 28 million
  60755. 45:28:42bytes it is taking and it is taking 8
  60756. 45:28:46million bytes. 20 million extra bytes
  60757. 45:28:49are occupied right now people do you
  60758. 45:28:53believe me? Yeah, that numpy wy are way
  60759. 45:28:56more efficient in memory management as
  60760. 45:28:59compared to the traditional data types
  60761. 45:29:00of Python. Yes. Okay, that's the first
  60762. 45:29:03part. Now second is people performance,
  60763. 45:29:07right? Performance. So what I'm going to
  60764. 45:29:09do is what I'm going to do is I am going
  60765. 45:29:12to
  60766. 45:29:14import
  60767. 45:29:16time, right? It's a it's a module in
  60768. 45:29:19Python, right? Suppose I say x is equal
  60769. 45:29:22to range
  60770. 45:29:26this much right. Okay. And then I have y
  60771. 45:29:31is equal to range say
  60772. 45:29:37this
  60773. 45:29:38to
  60774. 45:29:40this. Right? Both of them will have
  60775. 45:29:43equal amount of numbers. Same numbers
  60776. 45:29:45both of them will have. Right?
  60777. 45:29:47This will have say
  60778. 45:29:50uh 1 2 3 1 2 3 10 million values. 10
  60779. 45:29:56million values. This will also have 10
  60780. 45:29:59million values,
  60781. 45:30:03right? Both of them will have 10 million
  60782. 45:30:04values. Now what I'm trying to do is I
  60783. 45:30:06want to add them up right by bit by bit.
  60784. 45:30:09I want to add them up right. I want to
  60785. 45:30:11add first element of this to first
  60786. 45:30:13element of this. second of this to
  60787. 45:30:15second of this, third of this to third
  60788. 45:30:16of this like this. Okay, I want to do
  60789. 45:30:18this. Now what I'll do is I will run a
  60790. 45:30:21counter, right? I will run a counter
  60791. 45:30:24which is the start time,
  60792. 45:30:28right? And this is given by time dot
  60793. 45:30:31time, right? Which will give you the
  60794. 45:30:35this will give you the
  60795. 45:30:40current time, right? After this I will
  60796. 45:30:43run the operation. I will say C is equal
  60797. 45:30:45to X + Y
  60798. 45:30:49for X Y
  60799. 45:30:51in zip
  60800. 45:30:54X Y right in zip X Y right add X + Y bit
  60801. 45:31:02by bit element by element for X and Y in
  60802. 45:31:05zip zip is a function right which allows
  60803. 45:31:07you to do this operation sequentially
  60804. 45:31:10right sequentially right add the
  60805. 45:31:15elements of X and Y
  60806. 45:31:20element by element right element by
  60807. 45:31:24element
  60808. 45:31:28right element by element right and then
  60809. 45:31:31I'm going to print right so start time
  60810. 45:31:34will start and now I will say time dot
  60811. 45:31:37time which is now the end time minus
  60812. 45:31:40start time so this will give me the Time
  60813. 45:31:43taken for execution isn't it guys
  60814. 45:31:47will give me
  60815. 45:31:49delta of time which is equal to time
  60816. 45:31:53taken for operation
  60817. 45:31:56seconds
  60818. 45:31:58right these many seconds will be taken
  60819. 45:32:01right so let me run this and it takes
  60820. 45:32:04around say
  60821. 45:32:074.3 seconds right to do this right 4.3
  60822. 45:32:11seconds now Guys, see what happens. You
  60823. 45:32:14had to write this complex syntax in the
  60824. 45:32:18traditional Python. Now let me show this
  60825. 45:32:20on arrays. Right? What will happen on
  60826. 45:32:23arrays? I will say a is equal to np dot
  60827. 45:32:27a range.
  60828. 45:32:30Right? And inside a range I will pass
  60829. 45:32:33the same values what I have taken above.
  60830. 45:32:36Right? And I will say b is equal to np
  60831. 45:32:39dot
  60832. 45:32:43a range and I will pass the same values
  60833. 45:32:46inside
  60834. 45:32:48right exactly the same now what I'll do
  60835. 45:32:51is I will say same thing
  60836. 45:32:58right just I will change the execution
  60837. 45:33:01of C will now simply become people A + B
  60838. 45:33:06what is simple this or this
  60839. 45:33:09this or this
  60840. 45:33:13two right do you see the power if not I
  60841. 45:33:15will show this to you again right later
  60842. 45:33:17on and let me run this and you see the
  60843. 45:33:21difference now let me just increase a
  60844. 45:33:23couple of zeros right a couple of zeros
  60845. 45:33:28two zeros I'm increasing in both the use
  60846. 45:33:31cases
  60847. 45:33:36it is going on and on right let's see
  60848. 45:33:39See how much time it will take
  60849. 45:33:42to add say 2 million 1 billion numbers.
  60850. 45:33:481 billion numbers I have asked my system
  60851. 45:33:50to add and I want to see how much time
  60852. 45:33:53it takes.
  60853. 45:33:55Running running running.
  60854. 45:34:00Yep. Colonel has died. Kernel has died.
  60855. 45:34:03People,
  60856. 45:34:06I'll have to restart.
  60857. 45:34:09Right. I will have to import
  60858. 45:34:12numpy
  60859. 45:34:14as np. So let me just remove one zero
  60860. 45:34:19from both.
  60861. 45:34:26It is taking 4 seconds. Removing one
  60862. 45:34:29zero from here also. Right?
  60863. 45:34:33And when I do this it takes 1 second. Do
  60864. 45:34:37you see guys what is the difference in
  60865. 45:34:40performance also right for both of
  60866. 45:34:44these? Yeah.
  60867. 45:34:46And if you didn't understand this, let
  60868. 45:34:49me give you an example.
  60869. 45:34:52Range five. This is 5 to 10, right?
  60870. 45:35:00And this is basically
  60871. 45:35:02adding elements,
  60872. 45:35:06right? 5 7 9 11 13 Right. So this will
  60873. 45:35:13be what will the output of this? This
  60874. 45:35:15will be 0 1 2 3 4 and this will be
  60875. 45:35:19output what 5
  60876. 45:35:226 7 8 9 right so 0 + 5 5 6 + 1 7 7 + 2 9
  60877. 45:35:318 + 3 11 9 + 4 13 and the same thing if
  60878. 45:35:35I do here
  60879. 45:35:37then what will happen I say 5 I say 5
  60880. 45:35:42and 10 right and I say C is equal to a +
  60881. 45:35:46b and I say c. Same thing you get here.
  60882. 45:35:50Right? We can move on. Right? The next
  60883. 45:35:53bit guys which we have to understand the
  60884. 45:35:56next bit which we have to understand is
  60885. 45:35:58called as the indexing in numpy arrays.
  60886. 45:36:02Right? Indexing in numpy arrays. Right?
  60887. 45:36:05How do we index the elements? Right? How
  60888. 45:36:08do we index the elements?
  60889. 45:36:11indexing in
  60890. 45:36:13nump py
  60891. 45:36:15arrays right indexing in numpy arrays
  60892. 45:36:20so again you know indexing from basic
  60893. 45:36:23python so let's start with 1d for 1
  60894. 45:36:28arrays right I will use ar r r1
  60895. 45:36:35yeah this is a ar r1 now right this is a
  60896. 45:36:38ar r1 Okay.
  60897. 45:36:42Now people what I want to do is what I
  60898. 45:36:45want to do is I want to fetch right I
  60899. 45:36:50want to fetch right you can slice and
  60900. 45:36:53dice let's say dice
  60901. 45:36:5533 right so again as per our normal
  60902. 45:36:58indexing of list what is 33
  60903. 45:37:03what is the index of 33 people
  60904. 45:37:08two so you will say the same thing print
  60905. 45:37:11right A R R1 squared bracket 2 and you
  60906. 45:37:15will get 33 for yourself. Right? If you
  60907. 45:37:18wish to slice same things, right? 33 to
  60908. 45:37:23say 66. What is the index?
  60909. 45:37:2933 is 2. 2. Which one? 3 4 5 and 6.
  60910. 45:37:35Right? We will write 3 to six. Not five.
  60911. 45:37:37Hurry. Five is not included. Remember?
  60912. 45:37:41We print
  60913. 45:37:43a ar r r1
  60914. 45:37:452 is to 6 and you will get 33 44 55 66.
  60915. 45:37:52Right? Simple indexing. Please try it
  60916. 45:37:55out. Please try it out. And if you want
  60917. 45:37:58you can have this code also. You can
  60918. 45:38:00write this code. You will always have
  60919. 45:38:02clarity that why do we use it.
  60920. 45:38:06Moving on people. Moving on. Let's see
  60921. 45:38:09indexing.
  60922. 45:38:11in a 2D array. Right? And before I
  60923. 45:38:14explain this to you, let me take you
  60924. 45:38:16here. Right? So a 2D array will be what?
  60925. 45:38:22Right? This is a 2D array. It has three
  60926. 45:38:25rows and three columns. Right? Rows
  60927. 45:38:28columns. So now for rows indexing will
  60928. 45:38:32start from zero. So if you have to pitch
  60929. 45:38:36this particular row, right? So what will
  60930. 45:38:40be the index? It will be row 0. If you
  60931. 45:38:43have to fetch this particular row, the
  60932. 45:38:46index will be one. And if you have to
  60933. 45:38:48fetch this particular row, this the
  60934. 45:38:50index will be two. Similarly, for
  60935. 45:38:53column, if you have to fetch this
  60936. 45:38:55particular column, right, you will have
  60937. 45:38:58column is equal to zero. This particular
  60938. 45:39:01column, column equal to 1. This
  60939. 45:39:04particular column, column equal to two.
  60940. 45:39:07Right? indexing will start from minus
  60941. 45:39:09one again right 012
  60942. 45:39:12so let me just show that to you right
  60943. 45:39:14let's say I call ar r r2 this is my a r2
  60944. 45:39:18now
  60945. 45:39:20indexing
  60946. 45:39:23first
  60947. 45:39:25row right how will I do that I will say
  60948. 45:39:28print a ar r r2 and I will write how how
  60949. 45:39:33will I write this
  60950. 45:39:36h I will Write row 0, right? Row 0. So
  60951. 45:39:41what will this give me? This will give
  60952. 45:39:42me this, right? The syntax is
  60953. 45:39:49row space column. Right? So now if you
  60954. 45:39:52just pass one, it will give you only
  60955. 45:39:55rows, right?
  60956. 45:40:11row one, row two, row three. Right?
  60957. 45:40:13Similarly,
  60958. 45:40:16if you want to create it for columns,
  60959. 45:40:19right? What will you say?
  60960. 45:40:24Sorry,
  60961. 45:40:27uh
  60962. 45:40:28columns.
  60963. 45:40:30Uh
  60964. 45:40:32uh it was
  60965. 45:40:40zero
  60966. 45:40:45column 1 column 2
  60967. 45:40:49column 3. Right?
  60968. 45:40:54Right.
  60969. 45:40:58So what is my column 1? 76 89 98 76 89
  60970. 45:41:0398 right and if you want column 2
  60971. 45:41:10uh sorry
  60972. 45:41:13if you want column two this is the
  60973. 45:41:15column two and this is column 3 right
  60974. 45:41:18guys 90 999 99 independently
  60975. 45:41:23now if I ask you to fetch me a
  60976. 45:41:25particular element
  60977. 45:41:29element, right? Say I want you to fetch
  60978. 45:41:32me 78, right? How will you fetch 78? You
  60979. 45:41:36will say print a ar r2. What is the row
  60980. 45:41:39for 78? Which row does it belong to?
  60981. 45:41:42012. To which row 78 belongs? One row.
  60982. 45:41:46Which column it belongs to? 012.
  60983. 45:41:501. Hurry. Check again whether it belongs
  60984. 45:41:54to zero column, first column or second
  60985. 45:41:56column.
  60986. 45:42:00Right? And you will get uh sorry
  60987. 45:42:0501. So this is two. Right? 78. Right?
  60988. 45:42:10Second row first column. Right? Second
  60989. 45:42:13row first column.
  60990. 45:42:16How to define it as a
  60991. 45:42:27right? This is your matrix. Right? Now,
  60992. 45:42:30what are the index positions for this?
  60993. 45:42:33This row is zero. Row, first row, second
  60994. 45:42:37row. This is your zeroth column, first
  60995. 45:42:41column, second column. Yeah.
  60996. 45:42:47Yes. No. Maybe. Are we understanding
  60997. 45:42:49this? This much is clear. The indexing
  60998. 45:42:52of rows and columns. Now, now if you
  60999. 45:42:55have two, so the syntax is
  61000. 45:43:01syntax is say this is matrix A. So you
  61001. 45:43:05will say A in this A matrix you will
  61002. 45:43:07write row,
  61003. 45:43:10column. Right? Suppose I want to access
  61004. 45:43:13only zeroth row. So there will be no
  61005. 45:43:16column. So what will you get? Row number
  61006. 45:43:18zero.
  61007. 45:43:20Right? Row number zero. And you will
  61008. 45:43:22just put it like this. Or you can put a
  61009. 45:43:25comma and put colon. Colon means what?
  61010. 45:43:28Take everything. So I want all three
  61011. 45:43:30columns together. Zero row and all three
  61012. 45:43:32columns. So this will be your
  61013. 45:43:35show. Let me show that to you.
  61014. 45:43:41Right? See this zero row and all the
  61015. 45:43:44columns. So what will be your answer? 76
  61016. 45:43:4788 90. Right? Then if you want to access
  61017. 45:43:51the second row 89 90 99 89 90 99 one and
  61018. 45:43:57like this. Clear? Now similarly for
  61019. 45:44:00columns what will happen? You will take
  61020. 45:44:02all the rows
  61021. 45:44:04comma which column do you want? If you
  61022. 45:44:06say two what will be the result for
  61023. 45:44:08this? What numbers will you get for
  61024. 45:44:11this? Colon, 2, you will get 33 66 99.
  61025. 45:44:20Now coming to the element, right?
  61026. 45:44:22Suppose now you want to fetch 55s,
  61027. 45:44:26right? So what will you write? You will
  61028. 45:44:29write which row does it belong? Follow
  61029. 45:44:32the syntax. It belongs to the first row.
  61030. 45:44:34Which column does it belong to?
  61031. 45:44:37First column. So what will you get? 55.
  61032. 45:44:40Come to come come to this example. Now
  61033. 45:44:42you want 78 in this particular array or
  61034. 45:44:46matrix. Right? Where is 78? Which row
  61035. 45:44:49does it belong to?
  61036. 45:44:51Is it first row? Check carefully.
  61037. 45:44:55Zero row. First row. Second row.
  61038. 45:45:00Yeah. So I put two here. Now which
  61039. 45:45:03column does it belong to? First column.
  61040. 45:45:06Second column. Sorry. zero column, first
  61041. 45:45:08column, second column belongs to the
  61042. 45:45:10first column. So I put one here. So when
  61043. 45:45:13you put this syntax, you will get 78.
  61044. 45:45:17Clear? Now the third thing is slicing
  61045. 45:45:22through the array. Right? Now for this I
  61046. 45:45:26want I want 90 99 78 99. That means
  61047. 45:45:34what? I want this, this, this, and this.
  61048. 45:45:38Let me take you to the PPT first. Now,
  61049. 45:45:40what I'm asking you to fetch me? I'm
  61050. 45:45:41asking you to fetch me these four
  61051. 45:45:43numbers. Right? These four numbers. So,
  61052. 45:45:47here what will you write? Which rows are
  61053. 45:45:50included in this people? Which rows are
  61054. 45:45:52included in this?
  61055. 45:45:54Row one to all.
  61056. 45:45:58Right. One to all.
  61057. 45:46:00Yeah. Not two. One to all.
  61058. 45:46:04Right? And you leave everything like
  61059. 45:46:06this. If you have the last row, you
  61060. 45:46:08leave it empty after the colon, comma.
  61061. 45:46:12Which columns do you want for this? One
  61062. 45:46:15and two. Right? And when these will
  61063. 45:46:18intersect, when these will intersect,
  61064. 45:46:20you will get this area. You will get
  61065. 45:46:22this shaded area, green shaded area. So
  61066. 45:46:24I will say I need from column one to all
  61067. 45:46:29the columns. Right? So what will this
  61068. 45:46:31fetch you? What will this fetch you?
  61069. 45:46:33This will fetch you 44
  61070. 45:46:3655 66
  61071. 45:46:40and 77 88 99. Right? What will this
  61072. 45:46:46fetch you? This will fetch you 22 55 88
  61073. 45:46:5233 66 99. What are the commonalities
  61074. 45:46:56between both of these
  61075. 45:46:58access
  61076. 45:47:00both of these slices? What are the
  61077. 45:47:01commonalities? It is only this much
  61078. 45:47:05right
  61079. 45:47:0955 66 88 99 right so do you think you
  61080. 45:47:12will get your result yeah let's check it
  61081. 45:47:14here right I say print a ar r2 right and
  61082. 45:47:20this I say
  61083. 45:47:22I need row zero sorry row one to empty
  61084. 45:47:28and then column one empty and you will
  61085. 45:47:30get 90 99 78 90 indexing.
  61086. 45:47:35Yeah.
  61087. 45:47:37Now guys, let's check it for
  61088. 45:47:42three dimensions, right? 3D.
  61089. 45:47:45I say AR R3, right? This is my AR r3,
  61090. 45:47:48right? So, let me just put an example
  61091. 45:47:52here. Suppose my arrays are 1 1 22 2 3 4
  61092. 45:47:594 5 6 7 88 91. Right? This is my first.
  61093. 45:48:07Then behind this I have another matrix
  61094. 45:48:10which is uh
  61095. 45:48:13111 222 333
  61096. 45:48:17444 555 666
  61097. 45:48:21777 888 9999.
  61098. 45:48:25Right. And in my
  61099. 45:48:30third one, I have 11 1 1 222 33 33 3 4
  61100. 45:48:3744 44 44 44 44 44 44 44 44 44 44 44 44 4
  61101. 45:48:3855 555 66 66 66 66 66 6 7 8
  61102. 45:48:459
  61103. 45:48:46is that is it qualifying for a 3D array?
  61104. 45:48:50Can you access anything on this
  61105. 45:48:52particular array if given a chance
  61106. 45:48:55separately?
  61107. 45:48:56like this one. Now people, if you want
  61108. 45:48:59to move between layers, right? If you
  61109. 45:49:02want to move between layers, do you
  61110. 45:49:04think I have told you one particular
  61111. 45:49:07access which is what? Which is the
  61112. 45:49:10layers.
  61113. 45:49:12So what number will be given to this
  61114. 45:49:15layer? Layer number zero,
  61115. 45:49:18layer number one
  61116. 45:49:21and layer number two. So now if you have
  61117. 45:49:24to access this 555
  61118. 45:49:28which layer you will go to first you
  61119. 45:49:31will go to first layer. Then which row
  61120. 45:49:33will you go to? You will go to first row
  61121. 45:49:38and first column. So if you pass this
  61122. 45:49:42syntax what will you get? You will get 5
  61123. 45:49:45five5.
  61124. 45:49:47Right? Problem solved. The only
  61125. 45:49:49bottleneck was the layers part and you
  61126. 45:49:52have additional parameter or argument
  61127. 45:49:54for this layer.
  61128. 45:49:57So if I come back to my example and if
  61129. 45:50:00you have to access this 555 right how
  61130. 45:50:03will you do that? I say print a ar r2 a
  61131. 45:50:08r3 right and this I say and in this I
  61132. 45:50:11say what I have to access this 555 which
  61133. 45:50:14which layer number is this? This is
  61134. 45:50:17layer number zero. Right? This is layer
  61135. 45:50:20number zero. And this is layer number
  61136. 45:50:21one. So I say layer number one. Which
  61137. 45:50:25row is this
  61138. 45:50:27in this particular layer? It is layer
  61139. 45:50:29number. Sorry, it is row number one and
  61140. 45:50:32column number one. And when you do this,
  61141. 45:50:34you will get 5x5.
  61142. 45:50:36Right? If you want 999, what you will
  61143. 45:50:39do? You will change this to 2, 2, right?
  61144. 45:50:46If you want 777 or let's say if you want
  61145. 45:50:4998 what will you do for 98
  61146. 45:50:53layer number zero row number two column
  61147. 45:50:56number zero then you will get this 98
  61148. 45:51:02clear guys? Yep. Layer, row, column. In
  61149. 45:51:06the same way, you will slice it. Right.
  61150. 45:51:10You will slice it.
  61151. 45:51:12Yep.
  61152. 45:51:14Right. I'll write the syntax
  61153. 45:51:17so that you don't get confused. It is
  61154. 45:51:20array square bracket. Layer row column.
  61155. 45:51:27Right? Layer
  61156. 45:51:30row
  61157. 45:51:32column.
  61158. 45:51:34Right. This is the syntax.
  61159. 45:51:37Perfect.
  61160. 45:51:39Right. Great. We can also perform some
  61161. 45:51:44operations, right? Plus, minus,
  61162. 45:51:47multiplication, division between two
  61163. 45:51:50arrays very very easily, right? It
  61164. 45:51:53should not pose any problem to us,
  61165. 45:51:55right? We can do all the operations
  61166. 45:51:57which we want to, right? Between two
  61167. 45:51:59arrays, right? For example, right? You
  61168. 45:52:03had you have
  61169. 45:52:07right operations on arrays.
  61170. 45:52:10So let's say uh
  61171. 45:52:14let's see as a list right list. So we
  61172. 45:52:18have list is equal to 1 2 3 4 5 right
  61173. 45:52:23now suppose you want to square
  61174. 45:52:27each element of the list. Right? What
  61175. 45:52:31will you do? You will say s sq l is
  61176. 45:52:36equal to x to the power of 2 for x in
  61177. 45:52:44l right and then you will say print xq l
  61178. 45:52:50and you will get the squared of list
  61179. 45:52:52Right?
  61180. 45:53:01Original list squared list. Now
  61181. 45:53:05in array what will happen? Suppose I say
  61182. 45:53:10a 1 is equal to np dot array and l I
  61183. 45:53:15create the same array out of this list.
  61184. 45:53:18I say print
  61185. 45:53:21original
  61186. 45:53:23array
  61187. 45:53:24right I say A1 right this is my original
  61188. 45:53:27array same as this now I want to square
  61189. 45:53:30it
  61190. 45:53:33square the elements of arrays very very
  61191. 45:53:36simple nothing you require you just say
  61192. 45:53:40you just say
  61193. 45:53:43sq
  61194. 45:53:44a1 is equal to sq a1 1 is equal to a1 to
  61195. 45:53:51the power of 2, right? A1 to the power
  61196. 45:53:54of 2, right? And then you print
  61197. 45:53:59the
  61198. 45:54:05right you get the squared r. Suppose
  61199. 45:54:10you want to find the mean of
  61200. 45:54:21the mean of numbers
  61201. 45:54:24using list. Right? So what will you do?
  61202. 45:54:28You will say
  61203. 45:54:31mean is equal to sum of
  61204. 45:54:36l right list divided by length of list
  61205. 45:54:42right and you will get the means
  61206. 45:54:45right which is three for this one right
  61207. 45:54:48this original list. Now let me show you
  61208. 45:54:51in arrays
  61209. 45:54:54right. How will you do this? You will
  61210. 45:54:56say mean is equal to np dot mean and you
  61211. 45:55:01will just pass a1 right and when you
  61212. 45:55:04will check mean you will get 3.2
  61213. 45:55:08Right? Nothing like this direct right
  61214. 45:55:11direct like this right you can you have
  61215. 45:55:13I've already showed you add I've already
  61216. 45:55:15showed you uh square and then I believe
  61217. 45:55:19you can understand that what all
  61218. 45:55:21operations are possible using the arrays
  61219. 45:55:25right leveraging the power of arrays
  61220. 45:55:28right also guys in the arrays right in
  61221. 45:55:32the arrays what you can do is you can
  61222. 45:55:34perform
  61223. 45:55:36you can perform some string operations
  61224. 45:55:42right very powerful string operations
  61225. 45:55:44right so for example let's say I have a
  61226. 45:55:49I I have a array right I have an array
  61227. 45:55:55of names of people right so I say names
  61228. 45:56:01is equal to np dot array right and I say
  61229. 45:56:08Radha
  61230. 45:56:12uh then I say D
  61231. 45:56:15right and then I say
  61232. 45:56:19Maduk right these three names I have
  61233. 45:56:22right so you can check the names they
  61234. 45:56:24will be like in the array right and the
  61235. 45:56:26data type will be U6 which is a
  61236. 45:56:28representation of strings right now guys
  61237. 45:56:31suppose you want to capitalize you wish
  61238. 45:56:35to capitalize the names right of people
  61239. 45:56:40what will you do you will say print me
  61240. 45:56:44np docare right npcare dot capitalize
  61241. 45:56:50right capitalize and inside this you
  61242. 45:56:53will pass names
  61243. 45:56:56and you will see all the names have been
  61244. 45:56:58capitalized
  61245. 45:57:00right you see this R has been
  61246. 45:57:03capitalized D has been capitalized. M
  61247. 45:57:05has been capitalized.
  61248. 45:57:10Right?
  61249. 45:57:13You can convert them into upper if you
  61250. 45:57:15want. Print
  61251. 45:57:17np.care dot upper
  61252. 45:57:21names and you will have all of them in
  61253. 45:57:23caps lock. You can say print np.care
  61254. 45:57:28dot lower.
  61255. 45:57:31You will have them in lower.
  61256. 45:57:34Right? You can put the title. Right?
  61257. 45:57:36Suppose I say uh
  61258. 45:57:41title e title is equal to np dot array
  61259. 45:57:46and I say
  61260. 45:57:48rahov
  61261. 45:57:53go
  61262. 45:57:55right then I say
  61263. 45:57:59dhapati
  61264. 45:58:04right and I say mad
  61265. 45:58:10warm
  61266. 45:58:12right I say these three things now if I
  61267. 45:58:15say print
  61268. 45:58:18npcare
  61269. 45:58:20dot
  61270. 45:58:21title right and I say
  61271. 45:58:25titles you will see that all the words
  61272. 45:58:29will be in capitalized mode rael d and s
  61273. 45:58:33of dh sinapati m and V of MaduMa are now
  61274. 45:58:37capitalized. I can also
  61275. 45:58:41replace something if I wish to suppose I
  61276. 45:58:45want to replace Madhu with say suri
  61277. 45:58:50right I will say print
  61278. 45:58:53right np do.care care dot replace
  61279. 45:59:02right and you will say where you want to
  61280. 45:59:05replace I say title
  61281. 45:59:09and in this I want to replace
  61282. 45:59:12madu
  61283. 45:59:14with
  61284. 45:59:19suri
  61285. 45:59:20right and you will say it will be suri_1
  61286. 45:59:24suri one Right. Madu has been replaced
  61287. 45:59:26with
  61288. 45:59:29Right.
  61289. 45:59:33Right. If you want to calculate
  61290. 45:59:36the characters of strings,
  61291. 45:59:40right, you can do that. Print np.car
  61292. 45:59:46dot str length of
  61293. 45:59:50titles and you will get 11 characters
  61294. 45:59:52are there in here.
  61295. 45:59:5515 are there here and 11 are here in
  61296. 45:59:59this particular thing.
  61297. 46:00:01Right? You can do much more powerful
  61298. 46:00:04things also. Let me show you one complex
  61299. 46:00:06function. Right? Suppose I have f name
  61300. 46:00:12is equal to np dot array.
  61301. 46:00:16Right? And we have Ra
  61302. 46:00:25Madu
  61303. 46:00:27right and we have L name
  61304. 46:00:33arapati
  61305. 46:00:42worma. Right, we have these two things.
  61306. 46:00:45Now I can create a new array full name
  61307. 46:00:50by simply right by simply saying np.car
  61308. 46:00:54car dot add right and I say uh f name
  61309. 46:01:02right
  61310. 46:01:06comma
  61311. 46:01:11l
  61312. 46:01:13name right
  61313. 46:01:16Yes.
  61314. 46:01:26Yep.
  61315. 46:01:32Full name.
  61316. 46:01:34Yeah. Ra. I just was trying to add a
  61317. 46:01:38space
  61318. 46:01:40in between.
  61319. 46:01:43Anyway,
  61320. 46:01:54right. We can do that,
  61321. 46:01:58right? You can also people search in
  61322. 46:02:02arrays, right? Very powerful. Again,
  61323. 46:02:04search in arrays using where,
  61324. 46:02:08right?
  61325. 46:02:10Right. You can say suppose a 2 is equal
  61326. 46:02:14to
  61327. 46:02:16np dot array right and I'll say 1 1 22
  61328. 46:02:2033 3 4 4 5 66 right and now you have to
  61329. 46:02:24say a is equal to np dot where right and
  61330. 46:02:30in this you say a2 greater than 20
  61331. 46:02:36right and when you will check A
  61332. 46:02:43uh it's giving me the index. Why is it
  61333. 46:02:45giving me the index
  61334. 46:02:48or is it giving me the index?
  61335. 46:02:52Does it always return index?
  61336. 46:02:56One more thing is you can find a you can
  61337. 46:03:01find an
  61338. 46:03:03a letter
  61339. 46:03:06right through a letter. You can find an
  61340. 46:03:08element
  61341. 46:03:10through
  61342. 46:03:12a letter. Guys, these are all some
  61343. 46:03:14tricks which you should know because you
  61344. 46:03:16will be dealing with data and you need
  61345. 46:03:18to pull data, right? You need to pull
  61346. 46:03:19data a lot, right? Based on conditions
  61347. 46:03:21and based on things. Suppose you want to
  61348. 46:03:24find out the names which have G in them
  61349. 46:03:29right or R A in them. So how will you do
  61350. 46:03:31this? I will say print right and I will
  61351. 46:03:34say uh np do.care care dot find right
  61352. 46:03:41and I will say find this inside full
  61353. 46:03:46names right and find me ra a right ra a
  61354. 46:04:03two p
  61355. 46:04:07it uh right it returns true because it
  61356. 46:04:10has found it here right so I don't want
  61357. 46:04:14to tell you indexing through this but
  61358. 46:04:16anyway you should know this just just
  61359. 46:04:18assume this that I'm telling you to
  61360. 46:04:19write this okay because this is much
  61361. 46:04:21easier when we will go to pandas right
  61362. 46:04:26just uh write it as a syntax okay
  61363. 46:04:29greater than equal to zero I hope this
  61364. 46:04:32is clear
  61365. 46:04:33>> so let's start with the data science
  61366. 46:04:34interview questions and answers and The
  61367. 46:04:36number one problem we would be facing is
  61368. 46:04:38real world problem solving. And the
  61369. 46:04:40question one is handling missing data in
  61370. 46:04:43predictive modeling. So imagine you have
  61371. 46:04:45given a data set where 30% of the data
  61372. 46:04:48for key predictive variable is missing.
  61373. 46:04:50This variable is crucial for a
  61374. 46:04:52predictive model. How would you handle
  61375. 46:04:54this situation to ensure the integrity
  61376. 46:04:56and performance of your model? And
  61377. 46:04:58please describe your approach step by
  61378. 46:05:00step. So starting with the answer you
  61379. 46:05:02can start with handling missing data set
  61380. 46:05:04is a common challenge in data science
  61381. 46:05:06and it's important to address it
  61382. 46:05:08carefully to maintain the accuracy of
  61383. 46:05:10your model and here's how you could
  61384. 46:05:13approach this situation. The number one
  61385. 46:05:14point could be identify the missing
  61386. 46:05:16data. So first you need to understand
  61387. 46:05:18where the missing values are in your
  61388. 46:05:20data set. You can do this by using a
  61389. 46:05:23simple code in Python with libraries
  61390. 46:05:24like mandas. For example, you can use
  61391. 46:05:27the data dot isnull dot sum function
  61392. 46:05:31that will show you the count of missing
  61393. 46:05:33values in each column. Then you can
  61394. 46:05:35analyze the pattern. Determine if
  61395. 46:05:37there's a pattern to the missing data.
  61396. 46:05:39Is it random or is it missing for a
  61397. 46:05:41reason? This can affect your approach.
  61398. 46:05:43If the data is missing at random, the
  61399. 46:05:45methods you use might be different than
  61400. 46:05:47if the data is missing systematically.
  61401. 46:05:49So choosing a method for imputation.
  61402. 46:05:51Let's see the next method that is
  61403. 46:05:54choosing a method for imputation. So if
  61404. 46:05:56the missing data is numeric, you might
  61405. 46:05:58replace missing values with the mean or
  61406. 46:06:00median of that column. This is simple
  61407. 46:06:02and effective but can be used primarily
  61408. 46:06:05when the data is missing completely at
  61409. 46:06:06random. Then comes model based
  61410. 46:06:09imputation. Sometimes you can use other
  61411. 46:06:11variables in the data to predict missing
  61412. 46:06:13values using a regression model. This
  61413. 46:06:15can be more accurate but is also more
  61414. 46:06:17complex. Then we'll use the k nearest
  61415. 46:06:20neighbors can algorithm. But before that
  61416. 46:06:23we have a code snippet here that could
  61417. 46:06:26be used for the implementation of
  61418. 46:06:27imputation. You could use Python or R.
  61419. 46:06:30And now moving on we'll see the K
  61420. 46:06:32nearest neighbors algorithm. So this
  61421. 46:06:34method predicts the missing values based
  61422. 46:06:36on how closely related the data points
  61423. 46:06:38are to each other. So after imputation
  61424. 46:06:41it's crucial to check how your changes
  61425. 46:06:43have affected the overall data set and
  61426. 46:06:44model performance. Sometimes filling in
  61427. 46:06:47too many missing values can introduce
  61428. 46:06:49bias. And then we have visualization. To
  61429. 46:06:52help understand before and after the
  61430. 46:06:53imputation, you could visualize the
  61431. 46:06:55distribution of the variable using
  61432. 46:06:57histograms or box plots. This helps in
  61433. 46:07:00seeing how the imputation has changed
  61434. 46:07:02the statistical properties of the data.
  61435. 46:07:04And by following these steps, you can
  61436. 46:07:05handle missing data thoughtfully and
  61437. 46:07:07maintain the integrity of your
  61438. 46:07:09predictive model. Now moving to the
  61439. 46:07:11question number two that is based on
  61440. 46:07:13evaluating model overfitting. So the
  61441. 46:07:16question is you have developed a
  61442. 46:07:18predictive model but you suspect it
  61443. 46:07:20might be overfitting the training data.
  61444. 46:07:22How would you test and address the
  61445. 46:07:24issue? Please explain your steps and the
  61446. 46:07:26techniques you would use. So you could
  61447. 46:07:28start the answer by explaining what is
  61448. 46:07:31overfitting. So overfitting is a common
  61449. 46:07:33problem where model performs well on
  61450. 46:07:35training data but poorly on unseen data
  61451. 46:07:37indicating it's too closely fitted to
  61452. 46:07:39the training data specific details and
  61453. 46:07:41noise. So now we'll see a step-by-step
  61454. 46:07:44guide on how to address this. The number
  61455. 46:07:46one step is cross validation. So one
  61456. 46:07:48effective way to test for overfitting is
  61457. 46:07:51by using cross validation technique.
  61458. 46:07:53Cross validation involves splitting your
  61459. 46:07:55training data into multiple smaller sets
  61460. 46:07:57that is false and then training a model
  61461. 46:07:59on some of these set and validating it
  61462. 46:08:02on the others. So this helps you
  61463. 46:08:04understand if the model's good
  61464. 46:08:05performance is consistent across
  61465. 46:08:07different subsets of data. For example,
  61466. 46:08:09in Python you can use the cross value
  61467. 46:08:12score function from skarn.mmodel
  61468. 46:08:15selection. So this is the code and this
  61469. 46:08:19is the code snippet of Python that you
  61470. 46:08:21can use for the cross validation and
  61471. 46:08:23here we are importing from skarn that is
  61472. 46:08:26the module and we're importing
  61473. 46:08:29cross_well
  61474. 46:08:30score and here we have used the cross
  61475. 46:08:33val score function and then we have
  61476. 46:08:36printed the average cross validation
  61477. 46:08:38score and the next step we will do is
  61478. 46:08:40running cross validation model. So this
  61479. 46:08:42is your predictive model that you have
  61480. 46:08:44already built using scikit learn and
  61481. 46:08:46here's the x train these are the x input
  61482. 46:08:49features of your training data and y
  61483. 46:08:51train these are the output labels of
  61484. 46:08:53training data. So we are running gross
  61485. 46:08:55validation model here this is your
  61486. 46:08:57predictive model that you have already
  61487. 46:08:59built using scikitlearn. So x train here
  61488. 46:09:02that means these are the input features
  61489. 46:09:03of your training data and y train here
  61490. 46:09:06means these are the output labels of
  61491. 46:09:07training data and cv equal to 5. This
  61492. 46:09:10parameter tests the function to split
  61493. 46:09:11the data into five parts that is false.
  61494. 46:09:14And the model is trained on four of
  61495. 46:09:16these parts and the remaining part is
  61496. 46:09:18used for testing. So this process
  61497. 46:09:20rotates until each part has been used
  61498. 46:09:22for testing once and the printing
  61499. 46:09:24results that is score dot mean. So this
  61500. 46:09:27calculates the average of the scores
  61501. 46:09:29obtained from each gross validation for
  61502. 46:09:31this average score gives you an idea of
  61503. 46:09:34how well your model is likely to perform
  61504. 46:09:36on unseen data. A consistent score
  61505. 46:09:38across different polls suggests your
  61506. 46:09:40model is generalizing well rather than
  61507. 46:09:42overfitting to the training data. So now
  61508. 46:09:44moving to the next point that is
  61509. 46:09:46training versus validation error. So
  61510. 46:09:48plot the training and validation errors
  61511. 46:09:50as a function of training epochs or
  61512. 46:09:52complexity of the model. A model that
  61513. 46:09:54overfits will show a low error on
  61514. 46:09:56training data and a high error on
  61515. 46:09:58validation data as it trains further.
  61516. 46:10:00Then we have pruning the model. If you
  61517. 46:10:03confirm that the model is overfitting,
  61518. 46:10:05consider simplifying it. This might mean
  61519. 46:10:07reducing the number of parameters by
  61520. 46:10:09selecting fewer features using
  61521. 46:10:12regularization techniques like lasso or
  61522. 46:10:14ridge or choosing a less complex model.
  61523. 46:10:17After this step, we will move to
  61524. 46:10:18regularization technique step. So these
  61525. 46:10:20techniques add a penalty to the loss
  61526. 46:10:22function used to train the model which
  61527. 46:10:25can discourage complex models that
  61528. 46:10:27overfeit. Then we have common methods
  61529. 46:10:28that include L1 that is lasso and L2
  61530. 46:10:31ridge regularization. And here's how you
  61531. 46:10:34can add L2 regularization in Python. So
  61532. 46:10:37this is the code snippet here. And what
  61533. 46:10:38we have done here is we are creating the
  61534. 46:10:40ridge model and we have applied alpha
  61535. 46:10:43equal to 1.0. So this parameter controls
  61536. 46:10:45the strength of the regularization. A
  61537. 46:10:47higher alpha value increases the
  61538. 46:10:50regularization effect which helps reduce
  61539. 46:10:52model complexity and combat overfitting.
  61540. 46:10:55The alpha value can be tuned to find the
  61541. 46:10:58optimal balance between bias and
  61542. 46:11:00variance. And now coming for the fitting
  61543. 46:11:02the model. So model do fit and in that
  61544. 46:11:05we have X train and Y train that trains
  61545. 46:11:08the ridge model on the training data. It
  61546. 46:11:10adjusts the weight of the feature in X
  61547. 46:11:12train to predict the Y train while also
  61548. 46:11:15considering the regularization term.
  61549. 46:11:17This helps prevent the model from
  61550. 46:11:19fitting too closely to the noisy aspects
  61551. 46:11:21of the training data. And then we are
  61552. 46:11:23re-evaluating the model. After making
  61553. 46:11:25adjustments, it's important to
  61554. 46:11:27re-evaluate the model again using the
  61555. 46:11:29same cross validation technique to see
  61556. 46:11:32if the issue of overfitting has
  61557. 46:11:33improved. And then we have
  61558. 46:11:34visualization. To help illustrate or
  61559. 46:11:36ffitting, you could create a plot
  61560. 46:11:38showing the training and validation
  61561. 46:11:40errors or the number of epochs or model
  61562. 46:11:42complexity. So by using these
  61563. 46:11:44techniques, you can identify if your
  61564. 46:11:45model is all fitting and take steps to
  61565. 46:11:47correct it ensuring it performs well not
  61566. 46:11:50only on the training data but also on
  61567. 46:11:51new unseen data. So now moving to the
  61568. 46:11:53next question that is question number
  61569. 46:11:55three and it is based on realtime data
  61570. 46:11:57stream processing and the question is
  61571. 46:11:59you are tasked with building a model to
  61572. 46:12:01predict stock prices in real time. The
  61573. 46:12:04data comes in every second and you need
  61574. 46:12:06to update your predictions accordingly.
  61575. 46:12:08Describe how you would set up your
  61576. 46:12:09system to handle this type of data
  61577. 46:12:11effectively and what tools and
  61578. 46:12:13techniques would you use and why. So you
  61579. 46:12:15could start answering this question with
  61580. 46:12:17handling real-time data. So handling
  61581. 46:12:19real-time data especially for something
  61582. 46:12:21as volatile and fastpaced as stock
  61583. 46:12:23prices requires a robust system that can
  61584. 46:12:26process and analyze data quickly and
  61585. 46:12:28accurately. So here's how you could
  61586. 46:12:30approach this. We will set up such a
  61587. 46:12:32system and we'll have some steps. So
  61588. 46:12:35starting with the steps. So the first
  61589. 46:12:36step is choosing the right tools. The
  61590. 46:12:38right tool would be Apache Kafka. So
  61591. 46:12:40this is a popular tool for handling
  61592. 46:12:42real-time data that streams because it
  61593. 46:12:45allows you to publish and subscribe to
  61594. 46:12:46streams of records that is data and it
  61595. 46:12:49can handle high throughput with low
  61596. 46:12:50latency. Kafka acts as a buffer and
  61597. 46:12:53manages the flow of data ensuring that
  61598. 46:12:54your system doesn't get overwhelmed and
  61599. 46:12:57you can also use Apache Spark especially
  61600. 46:12:59Spark streaming is excellent for
  61601. 46:13:01processing the data. It can process data
  61602. 46:13:03in real time and perform complex
  61603. 46:13:05operations like windowing, grouping data
  61604. 46:13:07into chunks of a specified time period
  61605. 46:13:10and aggregating that is summarizing
  61606. 46:13:12data. So you can modify it and perform
  61607. 46:13:14the predicting of stock prices. And then
  61608. 46:13:17the step is data processing pipeline.
  61609. 46:13:19And the first step comes here is
  61610. 46:13:21injection. Data first enters the system
  61611. 46:13:23typically through Kafka which collects
  61612. 46:13:25data sent from the stock market and then
  61613. 46:13:27we do the processing. So spark streaming
  61614. 46:13:30takes over here. Here you can apply
  61615. 46:13:32transformations and run your predictive
  61616. 46:13:34models on the data. For example, you
  61617. 46:13:36might calculate moving averages or other
  61618. 46:13:38indicators that feed into your stock
  61619. 46:13:40price prediction model. And then comes
  61620. 46:13:42the output. Finally, the predictions are
  61621. 46:13:44outputed. This could be to a dashboard
  61622. 46:13:47for traders, an automated trading system
  61623. 46:13:49or even stored for further analysis. And
  61624. 46:13:52then we develop the model. Now comes the
  61625. 46:13:54model development. You would likely use
  61626. 46:13:56a machine learning model that can update
  61627. 46:13:57quickly and incorporate new data as it
  61628. 46:14:00arrives. models such as aim for time
  61629. 46:14:03series forecasting or more complex
  61630. 46:14:05machine learning models like rect neural
  61631. 46:14:07networks RNNs can be suitable. The model
  61632. 46:14:10should be retrained or fine-tuned
  61633. 46:14:12periodically with new data to ensure it
  61634. 46:14:14stays accurate. Now we'll come to
  61635. 46:14:16scalability and reliability. So ensure
  61636. 46:14:19your system can scale as data volume
  61637. 46:14:21increases. This might mean adding more
  61638. 46:14:23servers or optimizing your data
  61639. 46:14:24processing code. Implement monitoring to
  61640. 46:14:27catch any issues early like delays in
  61641. 46:14:29data processing or model performance
  61642. 46:14:31drops. And now we'll see the step that
  61643. 46:14:33is visualization and monitoring.
  61644. 46:14:35Consider setting up a real-time
  61645. 46:14:36dashboard that shows key metrics like
  61646. 46:14:39prediction accuracy and processing time.
  61647. 46:14:41This helps in quickly spotting when
  61648. 46:14:43something goes wrong. By setting up your
  61649. 46:14:45system with these tools and strategies,
  61650. 46:14:47you can effectively handle the challenge
  61651. 46:14:49of predicting stock prices in real time.
  61652. 46:14:51So now we'll move to the next question
  61653. 46:14:52that is question number four and this
  61654. 46:14:55will based on scalable data analytics.
  61655. 46:14:57So we have covered two questions that
  61656. 46:14:59were a bit code based questions and now
  61657. 46:15:02we'll see other questions that would be
  61658. 46:15:04based on scalable data analytics or they
  61659. 46:15:07might be on different areas and with the
  61660. 46:15:1013th question we'll start again with the
  61661. 46:15:12coding ones. So moving with the question
  61662. 46:15:14four that is based on scalable data
  61663. 46:15:16analytics and the question is given a
  61664. 46:15:18scenario where your organization
  61665. 46:15:20suddenly needs to scale its data
  61666. 46:15:22analysis capabilities due to an influx
  61667. 46:15:24of data that would be 10 times the
  61668. 46:15:27normal volume. How would you handle this
  61669. 46:15:29situation to ensure your data analytics
  61670. 46:15:31processes remain efficient and accurate?
  61671. 46:15:33What technologies would you consider and
  61672. 46:15:35what steps would you take? So you can
  61673. 46:15:37start answering this question with
  61674. 46:15:39handling a sudden increase in data
  61675. 46:15:41volume requires a strategic approach to
  61676. 46:15:43scaling your analytics infrastructure
  61677. 46:15:45without compromising on efficiency or
  61678. 46:15:47accuracy. So we'll see some steps from
  61679. 46:15:50that you could effectively manage this
  61680. 46:15:52scenario that you would start answering
  61681. 46:15:54the interviewer that we can start by
  61682. 46:15:56evaluating the current infrastructure's
  61683. 46:15:58ability to handle increased loads. This
  61684. 46:16:01includes assessing your databases,
  61685. 46:16:02servers and analytical tools to identify
  61686. 46:16:05potential bottlenecks or limitations.
  61687. 46:16:08Then you could move to next step that
  61688. 46:16:10would be choosing scalable technologies
  61689. 46:16:12to manage the increased data volume.
  61690. 46:16:14Consider leveraging cloud-based
  61691. 46:16:15solutions such as Amazon web services,
  61692. 46:16:17Google cloud platform or Microsoft
  61693. 46:16:19Azure. These platforms offer scalable
  61694. 46:16:21resources which can be adjusted
  61695. 46:16:23accordingly to the data load ensuring
  61696. 46:16:25you only pay for what you use. integrate
  61697. 46:16:27big data technologies like Apache Hadoop
  61698. 46:16:29for distributed storage and Apache Spark
  61699. 46:16:31for fast data processing. These tools
  61700. 46:16:33are designed to handle massive volumes
  61701. 46:16:35of data efficiently and can scale up to
  61702. 46:16:38meet standard increased demands. Now we
  61703. 46:16:41move to the next step that would be
  61704. 46:16:42optimizing data processing. So implement
  61705. 46:16:44data partitioning and indexing
  61706. 46:16:47strategies to improve the efficiency of
  61707. 46:16:49data queries. This will help in managing
  61708. 46:16:51large data sets by breaking them into
  61709. 46:16:53smaller manageable chunks and speeding
  61710. 46:16:55up search operations and use real-time
  61711. 46:16:58data processing frameworks like Apache
  61712. 46:17:00Kafka or Apache Flink which can handle
  61713. 46:17:02high throughput and provide timely
  61714. 46:17:04insights from large data streams. And
  61715. 46:17:06the next step would be automation and
  61716. 46:17:08monitoring. Automate routine data
  61717. 46:17:09processing task to reduce the manual
  61718. 46:17:11effort and speed up the analysis. This
  61719. 46:17:13can be done through scripting or using
  61720. 46:17:15workflow automation tools. Set up
  61721. 46:17:17comprehensive monitoring systems to
  61722. 46:17:19track the performance of your data
  61723. 46:17:21processes. Tools like Prometheus for
  61724. 46:17:23system monitoring and Graphana for
  61725. 46:17:25analytics and monitoring dashboards are
  61726. 46:17:27useful here. They help ensure that the
  61727. 46:17:29system is running smoothly and alert you
  61728. 46:17:32to potential issues before they become
  61729. 46:17:34critical. And the next step will be
  61730. 46:17:36regular evaluation and scaling.
  61731. 46:17:39Continuously evaluate the performance of
  61732. 46:17:40analytics infrastructure. As your data
  61733. 46:17:43grows, keep adjusting and scaling your
  61734. 46:17:45resources to maintain optimal
  61735. 46:17:46performance. Plan for periodic reviews
  61736. 46:17:48of your technology stack and
  61737. 46:17:50infrastructure to ensure they remain
  61738. 46:17:52aligned with your data needs and
  61739. 46:17:54organizational goals. By following these
  61740. 46:17:56steps, you can ensure that your data
  61741. 46:17:57analytics processes are prepared to
  61742. 46:17:59handle sudden surges in data volume
  61743. 46:18:01effectively maintaining the integrity
  61744. 46:18:03and speed of insights. So this was all
  61745. 46:18:06for the question four. Now moving to the
  61746. 46:18:08question five and this is based on
  61747. 46:18:10integrating machine learning models into
  61748. 46:18:12production and the question is you have
  61749. 46:18:14developed a machine learning model that
  61750. 46:18:15performs well in testing environment.
  61751. 46:18:17Now you need to integrate it into your
  61752. 46:18:19production environment where it will be
  61753. 46:18:21used in realtime applications. What
  61754. 46:18:23steps would you take to ensure the
  61755. 46:18:25successful deployment and operations of
  61756. 46:18:27the model in production? So we'll start
  61757. 46:18:29answering this by successfully deploying
  61758. 46:18:31a machine learning model into production
  61759. 46:18:33involves several critical steps to
  61760. 46:18:35ensure it performs as well in real time
  61761. 46:18:38operations as it does in testing. So you
  61762. 46:18:40would have a clear pathway to make the
  61763. 46:18:43interviewer understand. We will start
  61764. 46:18:45with the pathway with the first step
  61765. 46:18:46that would be model validation. So
  61766. 46:18:49before moving anything into production
  61767. 46:18:51revalidate your model's performance
  61768. 46:18:52using a separate validation data set.
  61769. 46:18:55This helps confirm that the model
  61770. 46:18:57generalizes well to new unseen data. The
  61771. 46:19:00next step will be preparing the
  61772. 46:19:02production environment. Ensure that the
  61773. 46:19:03production environment is ready to
  61774. 46:19:05handle the model. This includes setting
  61775. 46:19:07up the necessary hardware and software
  61776. 46:19:09ensuring that it can handle the expected
  61777. 46:19:11load and that all dependencies are
  61778. 46:19:14correctly installed and configured. Then
  61779. 46:19:16the next step comes that is model
  61780. 46:19:18wrapping. Wrap your model in an API that
  61781. 46:19:20is application programming interface
  61782. 46:19:22making it accessible to other parts of
  61783. 46:19:24your software infrastructure. Frameworks
  61784. 46:19:26like flask for Python can be used to
  61785. 46:19:28create a simple web server that listens
  61786. 46:19:30for data inputs and provides model
  61787. 46:19:32outputs. Then comes the next step that
  61788. 46:19:34is deployment strategies. Consider using
  61789. 46:19:37containerization tools like doer which
  61790. 46:19:39can help encapsulate your model and its
  61791. 46:19:41environment ensuring that it works
  61792. 46:19:43uniformly across different development
  61793. 46:19:46and production settings. And then we'll
  61794. 46:19:48use deployment strategies like blue
  61795. 46:19:50green deployment or canary releases to
  61796. 46:19:53minimize downtime and reduce the risk of
  61797. 46:19:55introducing a faulty model into
  61798. 46:19:57production. And then comes the next step
  61799. 46:19:59that is monitoring and logging.
  61800. 46:20:01Implement logging and monitoring to
  61801. 46:20:03track the model's performance and health
  61802. 46:20:05in real time. Tools like prompts for
  61803. 46:20:07monitoring and ELK elastic search log
  61804. 46:20:11statch kibbana for logging help in
  61805. 46:20:13quickly identifying and diagnosing
  61806. 46:20:15issues in production. And then comes the
  61807. 46:20:17next step that is performance tuning.
  61808. 46:20:19Monitor the model's performance over
  61809. 46:20:21time. If the model's performance
  61810. 46:20:23degrades or if new data shows different
  61811. 46:20:25patterns, you may need to retrain or
  61812. 46:20:28fine-tune the model to maintain
  61813. 46:20:29accuracy. And after this step, there's a
  61814. 46:20:32step for feedback loop. Set a feedback
  61815. 46:20:34loop where predictions and outcomes can
  61816. 46:20:36be compared. This feedback is crucial
  61817. 46:20:39for continuously improving the model and
  61818. 46:20:41catching any drift in data or changes in
  61819. 46:20:44external conditions that affect the
  61820. 46:20:45model. And after this comes a last step
  61821. 46:20:48that is legal and compliance checks.
  61822. 46:20:50Ensure all the data used by the model in
  61823. 46:20:52production complies with privacy laws
  61824. 46:20:54and regulations. This is crucial for
  61825. 46:20:56maintaining trust and legality
  61826. 46:20:58especially when handling sensitive
  61827. 46:21:00information. So by carefully planning
  61828. 46:21:02and executing these steps you can
  61829. 46:21:04smoothly transition your machine
  61830. 46:21:05learning model from a testing
  61831. 46:21:07environment to a fully functional
  61832. 46:21:09component of a production system. So
  61833. 46:21:11this was all about the question number
  61834. 46:21:12five. Now moving to the question number
  61835. 46:21:14six that would be based on datadriven
  61836. 46:21:16decision making. And the question is
  61837. 46:21:18your company wants to shift towards more
  61838. 46:21:20datadriven decision making. You have
  61839. 46:21:23been tasked with developing a strategy
  61840. 46:21:25to implement this. What steps would you
  61841. 46:21:27take to ensure that the data at all
  61842. 46:21:29levels of the organization is utilized
  61843. 46:21:31effectively to make informed decisions
  61844. 46:21:34and what challenges might you face and
  61845. 46:21:36how would you address them? So you can
  61846. 46:21:37start answering this by implementing a
  61847. 46:21:39datadriven decision-m strategy that will
  61848. 46:21:42require a comprehensive approach to
  61849. 46:21:44ensure that reliable data is accessible
  61850. 46:21:47and effectively used across all levels
  61851. 46:21:49of the organization. And now we can
  61852. 46:21:51develop and deploy this strategy. And
  61853. 46:21:53similarly you could tell this strategy
  61854. 46:21:55to the interviewer. So the number one
  61855. 46:21:57step will be assessing current data
  61856. 46:21:59infrastructure. Start by evaluating the
  61857. 46:22:01existing data infrastructure to
  61858. 46:22:03understand what data is available, how
  61859. 46:22:05it is stored and how it is currently
  61860. 46:22:06used. This assessment will help identify
  61861. 46:22:09gaps in data collection, storage and
  61862. 46:22:11access that need to be addressed. Now we
  61863. 46:22:14move to the next step that is developing
  61864. 46:22:16a data governance framework. Implement a
  61865. 46:22:19data governance framework that defines
  61866. 46:22:21who can access data, how it can be used
  61867. 46:22:23and who is responsible for maintaining
  61868. 46:22:25its quality. This framework ensures data
  61869. 46:22:28integrity and security which are
  61870. 46:22:29critical for making reliable decisions.
  61871. 46:22:32Now we'll move to the next step that is
  61872. 46:22:33training and empowerment. So train
  61873. 46:22:35employees at all levels on the
  61874. 46:22:37importance of datadriven decision making
  61875. 46:22:40and provide them with the tools and
  61876. 46:22:42knowledge necessary to analyze and
  61877. 46:22:44interpret data. This might include
  61878. 46:22:46training sessions, workshops and ongoing
  61879. 46:22:48support to ensure everyone can use data
  61880. 46:22:50effectively. Now move to the next step
  61881. 46:22:52that is implementing analytical tools.
  61882. 46:22:54So deploy user-friendly analytical tools
  61883. 46:22:56that can integrate seamlessly into the
  61884. 46:22:58daily workflows of employees. Tools like
  61885. 46:23:01Tableau, Microsoft PowerBI or even
  61886. 46:23:03advanced Excel techniques can provide
  61887. 46:23:05powerful data analysis capabilities
  61888. 46:23:07without requiring extensive technical
  61889. 46:23:09knowledge. After this we'll move to the
  61890. 46:23:12step that would be creating a
  61891. 46:23:13centralized data platform. Developer
  61892. 46:23:16centralized data platform where all
  61893. 46:23:18organizational data can be accessed and
  61894. 46:23:20analyzed. This platform should be
  61895. 46:23:22scalable and secure providing a single
  61896. 46:23:24source of truth for the organization.
  61897. 46:23:27And then we have the promoting a
  61898. 46:23:28datadriven culture. So foster culture
  61899. 46:23:31that values datadriven decision-m
  61900. 46:23:33encourage experimentation and learning
  61901. 46:23:35from datadriven initiatives. celebrate
  61902. 46:23:37successes and learn from failures to
  61903. 46:23:39continually improve the use of
  61904. 46:23:40datadriven in decision making and there
  61905. 46:23:43would be some challenges and solutions
  61906. 46:23:45for that. So one major challenge we know
  61907. 46:23:47here is resistance to change as some
  61908. 46:23:49employees may prefer traditional
  61909. 46:23:50decision-m methods. So address this by
  61910. 46:23:53demonstrating the tangible benefits of
  61911. 46:23:55datadriven decisions through pilot
  61912. 46:23:57projects and success stories. So data
  61913. 46:24:00silos can also hinder effective data use
  61914. 46:24:02promote cross department collaboration
  61915. 46:24:04and integrate disparate data sources to
  61916. 46:24:07overcome this challenge. After that you
  61917. 46:24:09can monitor and do continuous
  61918. 46:24:10improvement. So by systematically
  61919. 46:24:13implementing these steps you can
  61920. 46:24:14transform your organization into one
  61921. 46:24:16that leverages data at all levels to
  61922. 46:24:19make informed and effective decisions.
  61923. 46:24:21And after answering in these steps you
  61924. 46:24:23could make the interviewer have a truth
  61925. 46:24:26and a faith in you that you could make
  61926. 46:24:28these models. Now move to the next
  61927. 46:24:30question that is question number seven
  61928. 46:24:32and that is based on handling large data
  61929. 46:24:34set and the question is your project
  61930. 46:24:36involves analyzing extremely large data
  61931. 46:24:39sets potentially exceeding terabytes in
  61932. 46:24:41size. What strategies would you use to
  61933. 46:24:43manage and analyze such large data sets
  61934. 46:24:46effectively? Describe the tools and
  61935. 46:24:48techniques you might employ and you
  61936. 46:24:50could start this with answering that
  61937. 46:24:52working with large data sets especially
  61938. 46:24:54those in terabyte range presents unique
  61939. 46:24:56challenges in terms of storage
  61940. 46:24:58processing and analysis. So we'll have a
  61941. 46:25:00structured approach to handle these
  61942. 46:25:02challenges effectively. We'll start with
  61943. 46:25:04the data storage that would be use
  61944. 46:25:06distributed file systems. Consider using
  61945. 46:25:09a distributed file systems like Hadoop
  61946. 46:25:11distributed file system HDFS or Amazon
  61947. 46:25:14S3. These systems are designed to store
  61948. 46:25:16vast amounts of data across many servers
  61949. 46:25:18offering high availability and port
  61950. 46:25:21tolerance. And then comes the next step
  61951. 46:25:23that is data processing. Leverage big
  61952. 46:25:25data processing frameworks. Tools like
  61953. 46:25:27Apache Spark are ideal for processing
  61954. 46:25:29large data sets because they handle
  61955. 46:25:31distributed computing effectively. Spark
  61956. 46:25:34can perform data processing task much
  61957. 46:25:36faster than traditional disk based
  61958. 46:25:38processing due to its in-memory
  61959. 46:25:40computing capabilities. And next we
  61960. 46:25:42could start with efficient data
  61961. 46:25:44sampling. So there are many sampling
  61962. 46:25:46techniques that we can use. So when the
  61963. 46:25:48data set is too large to handle even
  61964. 46:25:50with powerful tools consider using data
  61965. 46:25:53sampling techniques to reduce the size
  61966. 46:25:55to a manageable level without losing
  61967. 46:25:57significant insights. Ensure that the
  61968. 46:25:59sample represents the whole data set
  61969. 46:26:00accurately. And then comes optimization
  61970. 46:26:03of data queries. Indexing and
  61971. 46:26:05partitioning. Optimize your data queries
  61972. 46:26:07by implementing indexing and
  61973. 46:26:08partitioning. This can drastically
  61974. 46:26:10reduce the time it takes to perform
  61975. 46:26:12queries by limiting the amounts of data
  61976. 46:26:14scan. And then we can do scalable
  61977. 46:26:16analytics. And then we'll move to the
  61978. 46:26:18next step that is scalable analytics.
  61979. 46:26:20And in that we could start with the
  61980. 46:26:21parallel computing. Use parallel
  61981. 46:26:23computing capabilities of frameworks
  61982. 46:26:25like spark or dask to analyze data
  61983. 46:26:28across multiple nodes. This helps in
  61984. 46:26:30scaling up your analytics operations to
  61985. 46:26:32handle large data sets effectively. And
  61986. 46:26:34now we'll move to the cloud-based
  61987. 46:26:36analytical tools. So consider using
  61988. 46:26:38cloud services like Google BigQuery or
  61989. 46:26:40AWS Red Shift which are designed to
  61990. 46:26:42handle massive data sets and complex
  61991. 46:26:45analytics with ease. And after this step
  61992. 46:26:47we'll move to data cleaning and
  61993. 46:26:48pre-processing. Here we will automate
  61994. 46:26:50pre-processing task. We'll use automated
  61995. 46:26:52tools to clean and pre-process data.
  61996. 46:26:55This includes handling missing values,
  61997. 46:26:57normalizing data and removing duplicates
  61998. 46:27:00which can be particularly challenging
  61999. 46:27:01with large data set. And after this
  62000. 46:27:03step, we'll move to the step that will
  62001. 46:27:05visualize large data set. So we'll use
  62002. 46:27:08specialized tools. That tools could be
  62003. 46:27:10Tableau or PowerBI that can handle large
  62004. 46:27:12data set by aggregating data and using
  62005. 46:27:15efficient backend technologies. For more
  62006. 46:27:18detailed exploration, tools like plotly
  62007. 46:27:20or bouquet can be used as they offer
  62008. 46:27:22capabilities to interactively visualize
  62009. 46:27:24large volumes of data. And after that,
  62010. 46:27:26there would be step for regular
  62011. 46:27:28maintenance and updates. That could be
  62012. 46:27:30continuously monitoring the data
  62013. 46:27:32quality. As new data comes in, you can
  62014. 46:27:35continuously monitor its quality. And
  62015. 46:27:37after this step, you could integrate all
  62016. 46:27:39these strategies and tools into your
  62017. 46:27:41workflow. And you can effectively manage
  62018. 46:27:43and extract valuable insights from
  62019. 46:27:45extremely large data sets thereby
  62020. 46:27:48supporting robust datadriven decision
  62021. 46:27:50making. And you could answer the whole
  62022. 46:27:52strategy to the interviewer. Now moving
  62023. 46:27:55to the question number eight that is
  62024. 46:27:56based on optimizing machine learning
  62025. 46:27:58models and the question is during model
  62026. 46:28:00development you have noticed that your
  62027. 46:28:01machine learning model is
  62028. 46:28:02underperforming. What steps would you
  62029. 46:28:04take to diagnose the problem and
  62030. 46:28:06optimize the model's performance? What
  62031. 46:28:08techniques and tools would you use? So
  62032. 46:28:10you can answer this by starting with the
  62033. 46:28:12optimizing and optimizing a machine
  62034. 46:28:14learning model that is underperforming
  62035. 46:28:17involves several steps to diagnose and
  62036. 46:28:19improve its accuracy and efficiency. And
  62037. 46:28:21here we will have structured approach to
  62038. 46:28:23tackle this issue and you could start
  62039. 46:28:25this with the number one step that is
  62040. 46:28:27diagnosing the problem. Evaluate model
  62041. 46:28:29metrics. Start by thoroughly evaluating
  62042. 46:28:31the performance metrics of your model.
  62043. 46:28:33For classification task, for
  62044. 46:28:35classification task, look at accuracy,
  62045. 46:28:37precision, recall and the F1 score. For
  62046. 46:28:40regression task, consider R squ, mean
  62047. 46:28:42squared error that is MSE and mean
  62048. 46:28:45absolute error that is MA. And then you
  62049. 46:28:48can move to the next step that is use
  62050. 46:28:50plots like ROC curves for classification
  62051. 46:28:52models and residual plots for regression
  62052. 46:28:54to visually assess with the model is
  62053. 46:28:57going wrong. After that, we'll move to
  62054. 46:28:58the next step that is data quality and
  62055. 46:29:00quantity check. Inspect the data that is
  62056. 46:29:03sometimes the quality and quantity of
  62057. 46:29:05data can be the root cause of poor model
  62058. 46:29:08performance. Ensure the data is clean,
  62059. 46:29:10well pre-processed and sufficient. Look
  62060. 46:29:12for issues like missing values, outliers
  62061. 46:29:14or imbalanced classes. And after this
  62062. 46:29:17we'll move to the feature engineering
  62063. 46:29:19step that would be experiment with
  62064. 46:29:21creating new features or transforming
  62065. 46:29:23existing ones to provide better
  62066. 46:29:25predictive power. And then we have the
  62067. 46:29:27next step that is model tuning and
  62068. 46:29:28configuration. After feature
  62069. 46:29:30engineering, we'll move to the next step
  62070. 46:29:32that is model tuning and configuration.
  62071. 46:29:34So, hyperparameter tuning. Use
  62072. 46:29:36techniques like grid search or random
  62073. 46:29:38search to find the optimal settings for
  62074. 46:29:40your model's parameters. Tools like
  62075. 46:29:41scikit learns, grid search CV or
  62076. 46:29:44randomized search CV can automate this
  62077. 46:29:46process. And there's a cross validation
  62078. 46:29:49that would implement cross validation to
  62079. 46:29:51ensure that the model's performance is
  62080. 46:29:53consistent across different subsets of
  62081. 46:29:55the data set. And then we have the next
  62082. 46:29:57step that is trying different models. So
  62083. 46:29:59experiment with algorithms here. If
  62084. 46:30:01initial models are underperforming, try
  62085. 46:30:03different algorithms that might be
  62086. 46:30:05better suited for the problem. For
  62087. 46:30:07instance, if you started with linear
  62088. 46:30:09regression and it's not performing well,
  62089. 46:30:11consider more complex models like random
  62090. 46:30:13forest or gradient boosting machines.
  62091. 46:30:16And after this, we have nseml methods
  62092. 46:30:18that we can use techniques like bagging,
  62093. 46:30:21boosting or stacking to combine the
  62094. 46:30:23predictions of multiple models to
  62095. 46:30:25improve overall performance. After this
  62096. 46:30:27step, we have feature selection that
  62097. 46:30:29includes reduce dimensionality. Use
  62098. 46:30:32techniques like principal component
  62099. 46:30:34analysis that is PCA to reduce the
  62100. 46:30:36number of features which might help in
  62101. 46:30:38improving model performance by removing
  62102. 46:30:39noise and redundancy. And then we have
  62103. 46:30:42select important features. So use model
  62104. 46:30:44based technique to identify and keep
  62105. 46:30:46only the most important features that
  62106. 46:30:48impact the outcome. And then comes the
  62107. 46:30:50last step that is regular updates and
  62108. 46:30:52retraining. So here you can monitor and
  62109. 46:30:54update that could be continuously
  62110. 46:30:56monitoring the model's performance over
  62111. 46:30:58time as new data becomes available
  62112. 46:31:00update and retrain the model to adapt to
  62113. 46:31:03any changes in underlying patterns and
  62114. 46:31:06after that you could have a consultation
  62115. 46:31:07and collaboration work with the other
  62116. 46:31:09teams and by methodically addressing
  62117. 46:31:11each of these areas you can diagnose why
  62118. 46:31:13your machine learning model is
  62119. 46:31:15underperforming and can take steps to
  62120. 46:31:17optimize its accuracy and efficiency. So
  62121. 46:31:19this was all about question number
  62122. 46:31:21eight. So let's start with the question
  62123. 46:31:22number nine and this is based on
  62124. 46:31:24handling unstructured data. So the
  62125. 46:31:26question is you are given a large amount
  62126. 46:31:28of unstructured data including text,
  62127. 46:31:30images and videos. What strategies would
  62128. 46:31:33you use to manage and analyze this type
  62129. 46:31:35of data effectively? Describe the tools
  62130. 46:31:38and techniques you might employ. So you
  62131. 46:31:40can start answering this question by
  62132. 46:31:42describing that dealing with
  62133. 46:31:44unstructured data can be challenging due
  62134. 46:31:46to its lack of predefined format or
  62135. 46:31:48structure. However, with the right
  62136. 46:31:49strategies and tools, you can
  62137. 46:31:51effectively manage and analyze it to
  62138. 46:31:54extract valuable insights and there will
  62139. 46:31:56be a approach how you can do that. So,
  62140. 46:31:58we will discuss the approach here and
  62141. 46:32:00starting with the steps. So, the number
  62142. 46:32:02one step will be data categorization and
  62143. 46:32:04organization. So, the number one step in
  62144. 46:32:08this step will be sorting and tagging.
  62145. 46:32:11We will begin by categorizing the data
  62146. 46:32:13into types that will be text, images or
  62147. 46:32:16videos. Use tagging to add metadata
  62148. 46:32:19which helps in organizing the data and
  62149. 46:32:21makes it easier to access and analyze
  62150. 46:32:23later. Then and after that particularly
  62151. 46:32:25for text data we'll use natural language
  62152. 46:32:28processing NLP. We will employ NLP
  62153. 46:32:30techniques to extract useful information
  62154. 46:32:32from text. Tools like NLTK, spacy or
  62155. 46:32:36even more advanced models like BERT can
  62156. 46:32:38help you perform tasks such as sentiment
  62157. 46:32:40analysis, entity recognition and topic
  62158. 46:32:43modeling. After that we will do text
  62159. 46:32:46indexing. We can use elastic search or
  62160. 46:32:48Apache sle to index large volumes of
  62161. 46:32:51text. These tools provide powerful
  62162. 46:32:53search capabilities and can handle
  62163. 46:32:55complex queries efficiently. And after
  62164. 46:32:57that we'll move to image data. And to
  62165. 46:32:59structure image data we'll use image
  62166. 46:33:01processing. We'll use libraries like
  62167. 46:33:03OpenCV for basic image processing tasks
  62168. 46:33:05such as filtering and transformations.
  62169. 46:33:07For more advanced image analysis,
  62170. 46:33:09consider deep learning models using
  62171. 46:33:11frameworks like TensorFlow or PyTorch.
  62172. 46:33:14And then we'll feature extraction. Apply
  62173. 46:33:16techniques to extract features from
  62174. 46:33:18images such as edges, textures or key
  62175. 46:33:20points which can be used for further
  62176. 46:33:22analysis or machine learning. And then
  62177. 46:33:24we'll come to video data. And here we'll
  62178. 46:33:26do video processing. and we'll use the
  62179. 46:33:28tools like fmpg that can be used for
  62180. 46:33:31basic video processing tasks such as
  62181. 46:33:33format conversion or extracting frames
  62182. 46:33:35for analyzing video content look at
  62183. 46:33:37machine learning models that can
  62184. 46:33:39classify or recognize activities in the
  62185. 46:33:41video and after this we'll move to
  62186. 46:33:43temporal analysis for videos temporal
  62187. 46:33:46components are important techniques like
  62188. 46:33:48sequence modeling or recurrent neural
  62189. 46:33:50networks RNNs can be useful to analyze
  62190. 46:33:53sequences of frames for activities or
  62191. 46:33:56events and then we'll move to data
  62192. 46:33:58storage and management. Here we'll use
  62193. 46:34:00the given volume and complexity of
  62194. 46:34:02unstructured data and use big data
  62195. 46:34:04platforms like Hadoop or cloud services
  62196. 46:34:06like AWS S3 for storage. These platforms
  62197. 46:34:09can scale up to handle large data sizes
  62198. 46:34:11and provide the necessary infrastructure
  62199. 46:34:13to store and retrieve unstructured data
  62200. 46:34:14efficiently. And then we have
  62201. 46:34:16visualization and reporting custom
  62202. 46:34:18dashboards that we'll create here. We
  62203. 46:34:20will develop custom dashboards using
  62204. 46:34:22tools like Tableau or PowerBI which can
  62205. 46:34:24integrate different data types and
  62206. 46:34:25provide a unified view of the analyzed
  62207. 46:34:27data. And after that we will do data
  62208. 46:34:29summarization. Tools that provide
  62209. 46:34:31summarization capabilities can help in
  62210. 46:34:33considering large volumes of
  62211. 46:34:34unstructured data into more manageable
  62212. 46:34:36and interpretable forms. And after that
  62213. 46:34:39we'll leverage these strategies and
  62214. 46:34:41tools and can effectively manage,
  62215. 46:34:43analyze and derive insights from
  62216. 46:34:45unstructured data which can be crucial
  62217. 46:34:47for making informed decisions in various
  62218. 46:34:49applications. And this is the path that
  62219. 46:34:51you can explore and explain to the
  62220. 46:34:53interviewer if this question has been
  62221. 46:34:55asked. Now moving to the question number
  62222. 46:34:5710 and that will be based on scaling AI
  62223. 46:34:59solutions in enterprise and the question
  62224. 46:35:01is your company wants to scale its AI
  62225. 46:35:04operations from a few initial pilot
  62226. 46:35:06projects to enterprisewide
  62227. 46:35:07implementation. What are the key
  62228. 46:35:09considerations and steps you would take
  62229. 46:35:11to ensure the successful scaling of AI
  62230. 46:35:13solutions across the organization and
  62231. 46:35:15what challenges might you face and how
  62232. 46:35:17would you address them? So you can start
  62233. 46:35:19answering this question with the scaling
  62234. 46:35:21AI solutions. You could answer him that
  62235. 46:35:23scaling AI solutions across an
  62236. 46:35:25enterprise requires careful planning and
  62237. 46:35:27strategic implementation to ensure
  62238. 46:35:29success and alignment with business
  62239. 46:35:32objectives and there should be a
  62240. 46:35:33strategic approach to implement this. So
  62241. 46:35:36starting with the approach and the
  62242. 46:35:38number one step will be that will be
  62243. 46:35:40strategic alignment. So identify
  62244. 46:35:43business objectives. Start by
  62245. 46:35:45identifying the business objectives that
  62246. 46:35:46the AI solutions are intended to
  62247. 46:35:48support. This ensures that the AI
  62248. 46:35:50initiatives are aligned with the company
  62249. 46:35:52strategic goals and can demonstrate
  62250. 46:35:54clear business value. And then comes the
  62251. 46:35:57stakeholder engagement. So engage
  62252. 46:35:59stakeholders from various departments
  62253. 46:36:01early in the process to gather input and
  62254. 46:36:03build support. This helps in
  62255. 46:36:05understanding diverse needs and ensures
  62256. 46:36:08broader acceptance of the AI solutions.
  62257. 46:36:10And after that comes the infrastructure
  62258. 46:36:12and technology. So there's an option
  62259. 46:36:14that is assess and upgrade
  62260. 46:36:16infrastructure. Evaluate whether your
  62261. 46:36:18current IT infrastructure can support
  62262. 46:36:20the expanded use of AI. You might need
  62263. 46:36:22to upgrade hardware, invest in cloud
  62264. 46:36:24solutions or adopt technologies that
  62265. 46:36:26facilitate AI processing and data
  62266. 46:36:28handling. And after that we have
  62267. 46:36:30standardization of tools. Standardize
  62268. 46:36:32the tools and platforms used for AI
  62269. 46:36:34development to ensure compatibility and
  62270. 46:36:37ease of maintenance across the
  62271. 46:36:38organizations. And after that we'll move
  62272. 46:36:40to data management. So robust data
  62273. 46:36:42governance that is to implement a strong
  62274. 46:36:45data governance framework to manage
  62275. 46:36:46enterprise data effectively. This
  62276. 46:36:48includes policies for data quality,
  62277. 46:36:50security and compliance especially
  62278. 46:36:52important when scaling AI solutions that
  62279. 46:36:54rely on vast amounts of data. And after
  62280. 46:36:56that we will come to data accessibility.
  62281. 46:36:58So ensure that data is accessible across
  62282. 46:37:00the organization but also secure against
  62283. 46:37:02unauthorized access. This involves
  62284. 46:37:04setting up secure data leaks or
  62285. 46:37:06warehouses that centralize data while
  62286. 46:37:08allowing controlled access. And then we
  62287. 46:37:11come to the next step that is talent and
  62288. 46:37:13training. So build AI competency that is
  62289. 46:37:16develop in-house AI expertise through
  62290. 46:37:18training programs and hiring. So this
  62291. 46:37:20build the necessary skills within the
  62292. 46:37:22organization to develop, manage and
  62293. 46:37:24scale AI solutions and after that you
  62294. 46:37:26can also perform cross functional AI
  62295. 46:37:28teams that could be forming cross
  62296. 46:37:30functional teams that include data
  62297. 46:37:32scientists, IT professionals and domain
  62298. 46:37:34experts. So this fosters collaboration
  62299. 46:37:36and ensure that AI solutions are
  62300. 46:37:38developed with a comprehensive
  62301. 46:37:39understanding. And after forming these
  62302. 46:37:41collaborative teams, we move to scalable
  62303. 46:37:43deployment models. So pilot test and
  62304. 46:37:46phase roll out. Before a full-scale
  62305. 46:37:48rollout, conduct pilot test to go the AI
  62306. 46:37:51solution effectiveness and integration
  62307. 46:37:53capabilities based on feedback, adjust
  62308. 46:37:56and then gradually deploy the solutions
  62309. 46:37:58across the organization. And then we
  62310. 46:38:00have modular and flexible design. So
  62311. 46:38:02design AI systems to be modular and
  62312. 46:38:05scalable allowing for adjustments and
  62313. 46:38:07expansions as needs and then we'll
  62314. 46:38:09monitor and do the continuous
  62315. 46:38:11improvement. So there will be
  62316. 46:38:13performance metrics that would establish
  62317. 46:38:14metrics to regularly assess the
  62318. 46:38:16performance of AI systems. We will
  62319. 46:38:18monitor these systems to ensure they met
  62320. 46:38:21expected outcomes and adapt as
  62321. 46:38:23necessary. And after that we have next
  62322. 46:38:25step that is addressing challenges. So
  62323. 46:38:28there could be cultural resistance that
  62324. 46:38:30there could be employees that would be
  62325. 46:38:32resisting to the changes but we have to
  62326. 46:38:34address this through continuous
  62327. 46:38:36education and by showcasing successful
  62328. 46:38:38AI use cases within the organizations
  62329. 46:38:41and by carefully considering these
  62330. 46:38:42aspects and methodically implementing
  62331. 46:38:45steps you can successfully scale AI
  62332. 46:38:47solutions across your enterprise driving
  62333. 46:38:49significant business value and
  62334. 46:38:50innovation. And that's all for question
  62335. 46:38:53number 10. Now we'll move to question
  62336. 46:38:55number 11 and that is based on ethical
  62337. 46:38:57considerations in data science. So the
  62338. 46:38:59question is in your data science
  62339. 46:39:01projects how do you ensure that ethical
  62340. 46:39:03considerations are addressed? Describe
  62341. 46:39:05the steps you take to identify and
  62342. 46:39:07mitigate ethical risk in your projects.
  62343. 46:39:09What frameworks or guidelines do you
  62344. 46:39:11follow? So you could start answering
  62345. 46:39:13this question with ethical
  62346. 46:39:14considerations that they're crucial in
  62347. 46:39:16data science to ensure that the
  62348. 46:39:17solutions and analyzes do not
  62349. 46:39:20advertently cause harm or bias. Here's
  62350. 46:39:23how you can ensure that. So there are
  62351. 46:39:25some steps and we will discuss those
  62352. 46:39:27steps. Starting with the number one that
  62353. 46:39:30is educate on ethical standards. So stay
  62354. 46:39:32informed about the ethical standards in
  62355. 46:39:34data science such as fairness,
  62356. 46:39:36accountability, transparency and
  62357. 46:39:38privacy. Organizations like the data
  62358. 46:39:40science association and the ACM have
  62359. 46:39:42codes of ethics that we refer to as
  62360. 46:39:44guidelines. And then we have ethical
  62361. 46:39:46risk assessment. Identify potential
  62362. 46:39:48ethical issues. That would be at the
  62363. 46:39:50beginning of each project. Conduct a
  62364. 46:39:52thorough assessment to identify any
  62365. 46:39:54potential ethical risk such as biases in
  62366. 46:39:57data or impact on vulnerable groups.
  62367. 46:40:00This involve reviewing the source of
  62368. 46:40:02data, the methodologies used for data
  62369. 46:40:04collection and the intended use of the
  62370. 46:40:06data analytics results. And then we have
  62371. 46:40:09stakeholder analysis. Engage with
  62372. 46:40:11stakeholders to understand the diverse
  62373. 46:40:13perspectives and potential impact of the
  62374. 46:40:16project. This helps in identifying
  62375. 46:40:18ethical issues that may not be apparent
  62376. 46:40:20from a purely technical standpoint. And
  62377. 46:40:23then we'll move to mitigation
  62378. 46:40:24strategies. Implementing bias mitigation
  62379. 46:40:27techniques. We will use statistical and
  62380. 46:40:29machine learning techniques to detect
  62381. 46:40:31and mitigate biases in data. This might
  62382. 46:40:34involve techniques like resampling,
  62383. 46:40:35reeing or using algorithms designed to
  62384. 46:40:38be fair. And then we have privacy
  62385. 46:40:40preserving methods. Employ methods such
  62386. 46:40:42as data anonymization, encryption or
  62387. 46:40:44differential privacy to protect
  62388. 46:40:46individual privacy when analyzing
  62389. 46:40:48sensitive data. Then we have other
  62390. 46:40:50methods that is transparency and
  62391. 46:40:52explanability. There we have model
  62392. 46:40:54explanability and after that coming to
  62393. 46:40:56documentation and reporting. So we have
  62394. 46:40:59to maintain thorough documentation of
  62395. 46:41:00data sources, model decisions and
  62396. 46:41:03methodologies. And then we have
  62397. 46:41:04continuous monitoring and feedback.
  62398. 46:41:06There you have to monitor outcomes and
  62399. 46:41:08the feedback mechanisms should be
  62400. 46:41:10applied. And then we have the panels
  62401. 46:41:12that is collaboration and advisory
  62402. 46:41:14panels. Then we have ethical review
  62403. 46:41:16boards. So for complex projects setting
  62404. 46:41:19up or consulting with an ethical review
  62405. 46:41:21board can provide oversight and diverse
  62406. 46:41:24perspectives on the ethical implications
  62407. 46:41:26of project methodologies. So by
  62408. 46:41:28proactively addressing ethical
  62409. 46:41:29considerations through these steps you
  62410. 46:41:32can ensure that your data science
  62411. 46:41:33projects uphold high ethical standards
  62412. 46:41:35and positively contribute to society
  62413. 46:41:38while minimizing harm. So this was all
  62414. 46:41:40about question 11. Now moving to
  62415. 46:41:42question number 12 that is based on time
  62416. 46:41:44series forecasting for business
  62417. 46:41:46decisions. So the question number 12 is
  62418. 46:41:48you are tasked with forecasting monthly
  62419. 46:41:50sales for a retail company using time
  62420. 46:41:53series data from the past 5 years. What
  62421. 46:41:56steps would you take to prepare and
  62422. 46:41:58analyze this data to make accurate
  62423. 46:42:00forecast? What specific tools or
  62424. 46:42:02techniques would you use and why? So we
  62425. 46:42:04can start answering this by time series
  62426. 46:42:06forecasting and we could address them
  62427. 46:42:08that it's a powerful tool for predicting
  62428. 46:42:10future events based on past data
  62429. 46:42:12especially in business context like
  62430. 46:42:14retail sales. So we will have a
  62431. 46:42:16structured approach here and we'll start
  62432. 46:42:18with data collection and cleaning.
  62433. 46:42:20First, you will gather data and ensure
  62434. 46:42:22that you have collected all relevant
  62435. 46:42:23data including monthly sale figures from
  62436. 46:42:25the past five years and also considering
  62437. 46:42:27including external factors that might
  62438. 46:42:29affect sales such as economic
  62439. 46:42:31indicators, holidays and promotional
  62440. 46:42:33activities. And then we'll proceed to
  62441. 46:42:35clean data. We will check for and handle
  62442. 46:42:37any inconsistencies or missing values.
  62443. 46:42:39And then we have data visualization.
  62444. 46:42:41Here we will plot the data. We'll use
  62445. 46:42:43plotting libraries like Matt Lib or
  62446. 46:42:45Seabbone in Python to visualize the
  62447. 46:42:47data. This will help in identifying
  62448. 46:42:49patterns, trends and seasonality. And
  62449. 46:42:51then we have decomposition of data. So
  62450. 46:42:53there's a seasonal decomposition and
  62451. 46:42:55we'll use statistical techniques to
  62452. 46:42:57decompose the data into trend seasonally
  62453. 46:43:00and residuals. So this can be
  62454. 46:43:02accomplished with tools like the
  62455. 46:43:04seasonal decompose function from the
  62456. 46:43:06stats models library in Python. And
  62457. 46:43:08we'll understand these components
  62458. 46:43:09separately and can improve the accuracy
  62459. 46:43:12of our forecast. And then the next step
  62460. 46:43:14is model selection and forecasting. So
  62461. 46:43:17there are two models that is a ara and s
  62462. 46:43:20IMA models. So we have to choose
  62463. 46:43:22appropriate forecasting models based on
  62464. 46:43:24data's characteristics. For instance,
  62465. 46:43:26AMA that is auto reggressive integrated
  62466. 46:43:29moving average. It is effective for
  62467. 46:43:31non-season data while SMA that is
  62468. 46:43:34seasonal AMA that is suitable for data
  62469. 46:43:37with seasonal patterns. And after
  62470. 46:43:39choosing the model we'll move to cross
  62471. 46:43:41validation. We will implement time
  62472. 46:43:43series specific cross validation
  62473. 46:43:45techniques like timebased splitting to
  62474. 46:43:47evaluate model performance and this will
  62475. 46:43:49ensure your model generalizes well on
  62476. 46:43:51unseen data and then we have model
  62477. 46:43:53fitting and diagnostics. We will fit the
  62478. 46:43:55model that is by using the cinemax class
  62479. 46:43:58from stat models that will fit your
  62480. 46:44:00model to the data and then we will
  62481. 46:44:02carefully select parameters based on AIC
  62482. 46:44:04that is a cake information criterion
  62483. 46:44:07that scores or thorough grid research
  62484. 46:44:09technique and then we can do the
  62485. 46:44:11diagnostics and forecast and validation
  62486. 46:44:14and after forecast validation we'll move
  62487. 46:44:16to iterative improvement. So there's a
  62488. 46:44:19feedback loop that should be mandatory
  62489. 46:44:21and there should be a regular update for
  62490. 46:44:24the model with new sales data and
  62491. 46:44:25refining the model as needed. So this
  62492. 46:44:28continuous improvement cycle helps adapt
  62493. 46:44:30to changing patterns in sales data. And
  62494. 46:44:33by following these steps and using these
  62495. 46:44:35tools, you can create robust forecast
  62496. 46:44:37that help the retail company plan better
  62497. 46:44:39and make informed decisions. So this was
  62498. 46:44:42all about question number 12. Now move
  62499. 46:44:43to question number 13 that is based on
  62500. 46:44:46customer segmentation using machine
  62501. 46:44:48learning. So the question is you are
  62502. 46:44:50given a data set containing demographic
  62503. 46:44:52and purchasing behavior data for a group
  62504. 46:44:54of customers. Your task is to segment
  62505. 46:44:56these customers into distinct groups
  62506. 46:44:58based on similarities in the purchasing
  62507. 46:45:00behavior and demographics. So what steps
  62508. 46:45:02would you take to perform this
  62509. 46:45:04segmentation and can you provide a
  62510. 46:45:05sample Python code snippet to illustrate
  62511. 46:45:08the initial stages of data handling and
  62512. 46:45:10model application. So we can start this
  62513. 46:45:12by explaining customer segmentation that
  62514. 46:45:14it's a powerful approach to tailor
  62515. 46:45:16marketing strategies and improve
  62516. 46:45:18customer service by identifying distinct
  62517. 46:45:20groups based on their behavior and
  62518. 46:45:22characteristics. And here also we have a
  62519. 46:45:25detailed approach for this task. So
  62520. 46:45:27we'll start with number one step that
  62521. 46:45:29would be data exploration and
  62522. 46:45:31pre-processing. So there will be initial
  62523. 46:45:33exploration that is beginning by
  62524. 46:45:35examining the data set to understand the
  62525. 46:45:37features available such as age, income,
  62526. 46:45:40purchase frequency etc. Then we'll look
  62527. 46:45:42for missing values or anomalies and
  62528. 46:45:45decide how to handle them. That could be
  62529. 46:45:47using imputation and then we'll move to
  62530. 46:45:49feature engineering. We'll create new
  62531. 46:45:51features that might be useful for
  62532. 46:45:52segmentation such as customer lifetime
  62533. 46:45:55value or average transaction amount.
  62534. 46:45:57We'll also use normalization that is
  62535. 46:45:59normalize the data to ensure that one
  62536. 46:46:01feature doesn't disproportionately
  62537. 46:46:03influence the model due to its scale.
  62538. 46:46:05We'll use standard scaling or minmax
  62539. 46:46:07scaling as appropriate. So then we'll
  62540. 46:46:10come to the next step that is choosing
  62541. 46:46:11the segmentation technique. And here we
  62542. 46:46:13have k means clustering. So this is a
  62543. 46:46:15popular method for customer
  62544. 46:46:16segmentation. Here we will decide on the
  62545. 46:46:18number of clusters by using techniques
  62546. 46:46:20like the elbow method or analysis to
  62547. 46:46:23determine the optimal cluster count. And
  62548. 46:46:26then we have model implementation and in
  62549. 46:46:29that we will use data preparation and
  62550. 46:46:31we'll prepare the data by selecting the
  62551. 46:46:33relevant features and applying any final
  62552. 46:46:35transformations and then we have model
  62553. 46:46:37fitting. We fit the C means clustering
  62554. 46:46:39model to the data and evaluate and
  62555. 46:46:42interpret analyzing clusters and after
  62556. 46:46:45analyzing clusters we'll move to the
  62557. 46:46:47next step that is strategic insights. We
  62558. 46:46:50will provide actionable insights based
  62559. 46:46:52on cluster characteristics such as
  62560. 46:46:54targeted marketing strategies for each
  62561. 46:46:56segment. And then we have iterative
  62562. 46:46:59refinement that is feedback
  62563. 46:47:01incorporation and we'll use business
  62564. 46:47:03feedback to refine the segmentation. If
  62565. 46:47:06additional data becomes available
  62566. 46:47:07incorporated to enhance the model and
  62567. 46:47:09now we'll see the sample Python code. So
  62568. 46:47:12for this first we'll import the
  62569. 46:47:14libraries and modules. As you can see on
  62570. 46:47:16the screen we have imported pandas
  62571. 46:47:19random forest classifier train test
  62572. 46:47:21split standard scaler classification
  62573. 46:47:24report and after that we will load the
  62574. 46:47:26data and and for that we have used the
  62575. 46:47:28pandas to read the data that is read ssv
  62576. 46:47:33and after that we are processing the
  62577. 46:47:34data that is data prep-processing we are
  62578. 46:47:37handling missing values and using the
  62579. 46:47:39forward fill or fil to fill missing
  62580. 46:47:42values in the data set and then we are
  62581. 46:47:44featuring scaling that is normalizing
  62582. 46:47:46the selected features that is feature
  62583. 46:47:48one, feature two and feature three using
  62584. 46:47:50standard scaler and then we'll move to
  62585. 46:47:52the next step that is data splitting.
  62586. 46:47:54We'll split the data set into training
  62587. 46:47:56and testing sets. So that test size
  62588. 46:47:58equal to 0.2 parameters specifies that
  62589. 46:48:0120% of the data will be used for testing
  62590. 46:48:03and then we'll train the model. We'll
  62591. 46:48:05initialize and train a random forest
  62592. 46:48:07classifier with 100 trees and a random
  62593. 46:48:10state for reproductibility and then
  62594. 46:48:12we'll evaluate the model. will make
  62595. 46:48:14predictions on the test set that is X
  62596. 46:48:18test using the train model and print a
  62597. 46:48:20classification report showing precision
  62598. 46:48:23recall F1 score and support for each
  62599. 46:48:26class. So this code demonstrates the
  62600. 46:48:28process of loading, pre-processing,
  62601. 46:48:31training and evaluating a machine
  62602. 46:48:33learning model that is random forest
  62603. 46:48:34classifier for predicting equipment
  62604. 46:48:37failures in a manufacturing plant. The
  62605. 46:48:39use of techniques such as data
  62606. 46:48:40prep-processing and splitting along with
  62607. 46:48:42the random forest classifier highlights
  62608. 46:48:44a standard flow for building predictive
  62609. 46:48:46maintenance models. So this was all
  62610. 46:48:48about the question number 13. So now
  62611. 46:48:50moving to the question number 14 that is
  62612. 46:48:52based on predictive customer churn and
  62613. 46:48:54the question is you are tasked with
  62614. 46:48:56developing a model to predict which
  62615. 46:48:58customers are likely to churn from a
  62616. 46:49:00subscription service. So what steps
  62617. 46:49:02would you take to build this model and
  62618. 46:49:04can you provide a sample Python code to
  62619. 46:49:06illustrate the data preparation and
  62620. 46:49:08model training process? So we'll start
  62621. 46:49:09answering this question about depicting
  62622. 46:49:12what is predicting customer churn. So
  62623. 46:49:14predicting customer churn is crucial for
  62624. 46:49:16businesses to implement detention
  62625. 46:49:18strategies proactively and we'll have a
  62626. 46:49:21detailed approach for building a
  62627. 46:49:22predictive model for this purpose.
  62628. 46:49:25Starting with data collection and
  62629. 46:49:26exploration and in this we will collect
  62630. 46:49:29data and after that we'll perform the
  62631. 46:49:31exploratory data analysis that is EDA.
  62632. 46:49:34We'll perform an initial analysis to
  62633. 46:49:36understand patterns and trends and then
  62634. 46:49:38we have feature engineering. We will
  62635. 46:49:40create new features and derive new
  62636. 46:49:42feature that might influence churn such
  62637. 46:49:44as change in usage pattern or service
  62638. 46:49:47upgrades. And then we'll handle the
  62639. 46:49:48missing values if we found any. And then
  62640. 46:49:50we'll encode categoral variables. We'll
  62641. 46:49:53use techniques like one hot encoding or
  62642. 46:49:55label encoding for categorial variables.
  62643. 46:49:58And then we have scale features to
  62644. 46:50:00normalize or standardize numerical
  62645. 46:50:02features to ensure they contribute
  62646. 46:50:04equally to the model's performance. And
  62647. 46:50:06then we'll select the model that is
  62648. 46:50:09we'll choose the appropriate model and
  62649. 46:50:11start with for the knowing handling
  62650. 46:50:13binary classification task that could be
  62651. 46:50:16with logistic regression, random forest
  62652. 46:50:18or gradient boosting machines. And after
  62653. 46:50:21selecting the model, we'll train the
  62654. 46:50:22model and evaluate it. So fit your model
  62655. 46:50:25on the training data and after that
  62656. 46:50:27evaluate the model using appropriate
  62657. 46:50:29metrics like accuracy, precision,
  62658. 46:50:31recall, F1 score and ROC to go its
  62659. 46:50:35performance. And then we'll optimize the
  62660. 46:50:37model using hyperparameter tuning. We'll
  62661. 46:50:40optimize the model parameter using grid
  62662. 46:50:42search or random search to improve
  62663. 46:50:44performance. And then we have feature
  62664. 46:50:46importance that is analyze and rank
  62665. 46:50:48features by their importance in
  62666. 46:50:50predicting churn to refine the model
  62667. 46:50:52further. And then and then the last step
  62668. 46:50:54is deployment and monitoring. We'll
  62669. 46:50:56deploy the model once validated deploy
  62670. 46:50:58the model into a production environment
  62671. 46:51:00where it can predict real-time churn. So
  62672. 46:51:02after deploying the model regularly
  62673. 46:51:04monitor the model to ensure it remains
  62674. 46:51:06effective over time as new data comes
  62675. 46:51:08in. So now we'll see the sample Python
  62676. 46:51:11code for this example. So starting with
  62677. 46:51:13the importing of libraries we will
  62678. 46:51:16import pandas numpy scikitlearn skarn
  62679. 46:51:19tensorflow and the tensorflow kas and
  62680. 46:51:23callbacks and after importing the
  62681. 46:51:25modules we'll start with data loading
  62682. 46:51:27we'll load the data set from a CSV file
  62683. 46:51:30named equipment data dot csv and that
  62684. 46:51:33with the pandas data frame and after
  62685. 46:51:36that we'll do the data prep-processing
  62686. 46:51:37we'll handle missing values and for that
  62687. 46:51:40we'll use forward fill to fill missing
  62688. 46:51:42values in the data set and then we have
  62689. 46:51:44feature scaling that will normalize the
  62690. 46:51:46selected features that is feature one,
  62691. 46:51:48feature two, feature three using
  62692. 46:51:49standard scaler and after that we'll use
  62693. 46:51:52the data splitting. We'll split the data
  62694. 46:51:54set into training and testing sets and
  62695. 46:51:56the test size will be equal to 0.2 and
  62696. 46:51:59this parameter specifies that 20% of the
  62697. 46:52:01data will be used for testing and after
  62698. 46:52:03that we'll start with building the
  62699. 46:52:05model. First we'll see sequential model
  62700. 46:52:07that initializes a sequential model
  62701. 46:52:10technique. And then we have dense layers
  62702. 46:52:12that adds two dense layers with 64 units
  62703. 46:52:15and value activation function. Then we
  62704. 46:52:17have dropout layers that adds two
  62705. 46:52:19dropout layers with a dropout rate of
  62706. 46:52:210.5 to reduce overfitting. After that
  62707. 46:52:24we'll do the model compilation. We'll
  62708. 46:52:26compile the model using the atom
  62709. 46:52:27optimizer and binary cross entropy loss
  62710. 46:52:30function for binary classification. And
  62711. 46:52:33there will be an early stopping that
  62712. 46:52:34will define an early stopping call back
  62713. 46:52:36to stop training when the validation
  62714. 46:52:38loss metric has stopped improving after
  62715. 46:52:40three blocks. And after training the
  62716. 46:52:43model, we will evaluate the model. And
  62717. 46:52:45evaluating the model on the test data
  62718. 46:52:47and print the loss and accuracy metrics.
  62719. 46:52:50So this code demonstrates the process of
  62720. 46:52:52loading, pre-processing, building,
  62721. 46:52:53compiling, training and evaluating a
  62722. 46:52:55deep learning model using TensorFlow and
  62723. 46:52:58KAS for predicting equipment failures in
  62724. 46:53:00a manufacturing plant. So the use of
  62725. 46:53:02techniques such as data prep-processing,
  62726. 46:53:04dropout regularization and early
  62727. 46:53:07stopping helps in building a robust deep
  62728. 46:53:09learning model for predictive
  62729. 46:53:10maintenance. So that's all with question
  62730. 46:53:12number 14. Now we'll start with question
  62731. 46:53:14number 15 that is based on deep learning
  62732. 46:53:16and NLP. And your question is you are
  62733. 46:53:18tasked with developing a sentiment
  62734. 46:53:20analysis model using deep learning to
  62735. 46:53:22understand customer opinions from
  62736. 46:53:24reviews. So what steps would you take to
  62737. 46:53:26build this model and can you provide a
  62738. 46:53:28sample Python code snippet to illustrate
  62739. 46:53:30how you would pre-process data and train
  62740. 46:53:32a simple deep learning model? So we'll
  62741. 46:53:34start answering this with sentiment
  62742. 46:53:36analysis that sentiment analysis using
  62743. 46:53:38deep learning allows businesses to coach
  62744. 46:53:41customer sentiment from text data like
  62745. 46:53:43reviews or comments effectively. And
  62746. 46:53:45we'll have a detailed approach for
  62747. 46:53:47building a sentiment analysis model.
  62748. 46:53:49We'll start with data collection and
  62749. 46:53:51cleaning. We will collect the data,
  62750. 46:53:52gather a substantial data set of text
  62751. 46:53:54reviews and their associated sentiments
  62752. 46:53:57typically labeled as positive, negative,
  62753. 46:53:59or neural. And then we'll clean the
  62754. 46:54:01data, pre-process the data by removing
  62755. 46:54:04noise such as HTML tags, special
  62756. 46:54:06characters, and so words. And we'll
  62757. 46:54:08normalize the text by converting it to
  62758. 46:54:10lower case. And then we have text
  62759. 46:54:12prep-processing. We'll convert text into
  62760. 46:54:14tokens, words, or phrases. And then we
  62761. 46:54:17have vectorization that transforms
  62762. 46:54:19tokens into numerical format using
  62763. 46:54:21techniques like word embeddings or TF
  62764. 46:54:23that is term frequency in document
  62765. 46:54:26frequency and then we'll use the padding
  62766. 46:54:29and then we have the option of model
  62767. 46:54:30selection. will choose a model
  62768. 46:54:32architecture based on a basic approach
  62769. 46:54:34and use a RNN or more advanced
  62770. 46:54:37architecture like LSTM that is long
  62771. 46:54:40short-term memory or GRU that is gated
  62772. 46:54:43recurrent units which are effective for
  62773. 46:54:45sequence data like text and then we have
  62774. 46:54:48model training we'll compile the model
  62775. 46:54:51define the model architecture and
  62776. 46:54:52compile it with a loss function suited
  62777. 46:54:55for classification like categoral cross
  62778. 46:54:57entropy and an optimizer like Adam and
  62779. 46:55:00then we'll train the model. We'll fit
  62780. 46:55:02the model on our pre-processed data.
  62781. 46:55:04We'll evaluate and optimize it.
  62782. 46:55:06Evaluating model performance. Here use
  62783. 46:55:08the metrics such as accuracy, precision,
  62784. 46:55:11recall, and F1 score to assess the
  62785. 46:55:13model. And then we have hyperparameter
  62786. 46:55:15tuning. We'll optimize the model by
  62787. 46:55:17adjusting parameters like learning rate,
  62788. 46:55:19number of layers and units per layer.
  62789. 46:55:22And then coming to deployment. We'll
  62790. 46:55:24deploy the model and integrate the model
  62791. 46:55:26into the existing review processing
  62792. 46:55:28pipeline. So it can automatically
  62793. 46:55:30classify new reviews. So let's see the
  62794. 46:55:32sample Python code and we'll have a
  62795. 46:55:34basic approach for that. Here we'll
  62796. 46:55:36import numpy tensorflow sequential
  62797. 46:55:39embedding LSTM dense stroke out. So
  62798. 46:55:42embedding converts positive integers
  62799. 46:55:44that is indexes into dense vectors of
  62800. 46:55:46fixed size and LSTM that is long
  62801. 46:55:48short-term memory layer that is used for
  62802. 46:55:51learning dependencies in sequence data.
  62803. 46:55:54And then we have dense that is a
  62804. 46:55:55regularly densed connected NN layer. And
  62805. 46:55:59then we will import pad sequences. And
  62806. 46:56:02after that we have the data set and the
  62807. 46:56:04sample text data representing customer
  62808. 46:56:06reviews that will store in variable
  62809. 46:56:09text. And then we have labels that has
  62810. 46:56:11binary labels indicating sentiment one
  62811. 46:56:13for positive zero for negative. And now
  62812. 46:56:16we'll start with the pre-processing of
  62813. 46:56:17data. Here we have declared that
  62814. 46:56:19tokenizer. We will initialize a
  62815. 46:56:21tokenizer that will help only the top
  62816. 46:56:24thousand most frequent words. And then
  62817. 46:56:26we have fit_on
  62818. 46:56:29text that is update the internal
  62819. 46:56:31vocabulary based on the list of text. It
  62820. 46:56:34essentially creates a dictionary of word
  62821. 46:56:36to index pairs. And then we have text to
  62822. 46:56:38sequences that will transform each text
  62823. 46:56:41in text to a sequence of integers. And
  62824. 46:56:44then we have pad sequences that will
  62825. 46:56:46ensure all sequences have the same
  62826. 46:56:48length by padding shorter sequences with
  62827. 46:56:50zeros up to the maximum length. And then
  62828. 46:56:53we'll start building the model. Here we
  62829. 46:56:55have sequential model that will set up a
  62830. 46:56:57linear stack of layers. And then we have
  62831. 46:56:59embedding layer that will map each word
  62832. 46:57:01index to an embedding vector of size 64.
  62833. 46:57:04So the input length is set to 10 that is
  62834. 46:57:06the length of the input sequences. Then
  62835. 46:57:09we'll start with LSTM layers. So two
  62836. 46:57:11LSTM layers are added. The first one
  62837. 46:57:13returns sequences to allow the next LSTM
  62838. 46:57:16layer to process these sequences. And
  62839. 46:57:18after that we have the dropout layer
  62840. 46:57:20that applies dropout with a rate of 0.5
  62841. 46:57:23of the first LSTM layer to reduce
  62842. 46:57:25overfitting. And after that we'll come
  62843. 46:57:27to dense layer that has output of a
  62844. 46:57:30single scalar that represents the
  62845. 46:57:32predicted setment and using sigmoid
  62846. 46:57:35activation to output a probability. And
  62847. 46:57:38now we'll start with model compilation
  62848. 46:57:40and training. So we'll configure the
  62849. 46:57:42model for training and we'll use binary
  62850. 46:57:44cross entropy as the loss function that
  62851. 46:57:46is suitable for binary classification
  62852. 46:57:48and the atom optimizer and tracks
  62853. 46:57:51constantly accuracy as a metric and then
  62854. 46:57:54we have the fit that trains the model
  62855. 46:57:56for a specified number of epochs that is
  62856. 46:57:58iterations or the entire data set and
  62857. 46:58:01then we'll predict the model that is
  62858. 46:58:03after training the model can predict the
  62859. 46:58:05sentiment of the reviews in the data
  62860. 46:58:06set. This is useful for checking how the
  62861. 46:58:08model performs on the training data
  62862. 46:58:10itself. So this breakdown explains each
  62863. 46:58:12step of the coding process detailing how
  62864. 46:58:14the data is prepared and how the model
  62865. 46:58:16is configured and then we'll compile it
  62866. 46:58:19and use for training and prediction. So
  62867. 46:58:21it's detailed explanation should help in
  62868. 46:58:24understanding how to implement a simple
  62869. 46:58:26LSTM model for sentiment analysis in
  62870. 46:58:28TensorFlow. Now moving to the question
  62871. 46:58:30number 16. So let's start with question
  62872. 46:58:33number 16 that is based on anomly
  62873. 46:58:35detection in transaction data. So the
  62874. 46:58:37question is you are tasked with
  62875. 46:58:39identifying unusual transactions in a
  62876. 46:58:41company's financial data that might
  62877. 46:58:43suggest fraudulent activity. So what
  62878. 46:58:45steps would you take to develop an
  62879. 46:58:46anomaly detection model and can you
  62880. 46:58:48provide a sample Python code snippet to
  62881. 46:58:51illustrate how you would pre-process the
  62882. 46:58:52data and apply an anomaly detection
  62883. 46:58:54technique. So we'll start answering this
  62884. 46:58:56with anomaly detection technique that is
  62885. 46:58:59anomaly detection is essential for
  62886. 46:59:00preventing fraud by identifying
  62887. 46:59:02transactions that deviate significantly
  62888. 46:59:05from typical patterns. And now we'll see
  62889. 46:59:07the structured approach to building an
  62890. 46:59:09anomaly detection model for transaction
  62891. 46:59:11data. We'll start with data collection
  62892. 46:59:14and cleaning and we'll collect all the
  62893. 46:59:16compiling transaction data which should
  62894. 46:59:18include details like transaction amount,
  62895. 46:59:20time, user ID and transaction type. Then
  62896. 46:59:22we'll move to feature engineering and
  62897. 46:59:24develop features that capture the
  62898. 46:59:25essence of transaction such as time of
  62899. 46:59:27day and the day of the week. And then we
  62900. 46:59:30have data normalization. We'll use
  62901. 46:59:31scaling techniques such as minmax
  62902. 46:59:33scaling or standardization to ensure
  62903. 46:59:35that the model is perfectly normalized.
  62904. 46:59:38And then we have choosing the anomaly
  62905. 46:59:40detection technique. So here we have to
  62906. 46:59:43choose the technique which is effective
  62907. 46:59:45for highdimensional data sets and works
  62908. 46:59:47for isolating anomalies instead of
  62909. 46:59:50profiling normal data points. After
  62910. 46:59:52choosing the anomaly technique will
  62911. 46:59:54train anomaly identification. We'll fit
  62912. 46:59:57the chosen model to the data and the
  62913. 46:59:59anomalies that would have been chosen
  62914. 47:00:01will be those transactions that the
  62915. 47:00:03model identifies and after this we come
  62916. 47:00:05to the last step that is review and
  62917. 47:00:07action. Here we have manual review that
  62918. 47:00:09is transactions flagged as potential
  62919. 47:00:11anomalies should be reviewed manually to
  62920. 47:00:13confirm fraudent activity and then we
  62921. 47:00:16have continuous improvement that is we
  62922. 47:00:18can regularly update the model with the
  62923. 47:00:19new data and feedback from the review
  62924. 47:00:21process to improve accuracy. And now
  62925. 47:00:23moving to the prediction that is after
  62926. 47:00:25training the model we can predict the
  62927. 47:00:27sentiment of the reviews in the data set
  62928. 47:00:29and this is useful for checking how the
  62929. 47:00:31model performs on the training data
  62930. 47:00:33itself. Now we'll see the Python code to
  62931. 47:00:36see how you can set up this model for
  62932. 47:00:39anomaly detection. We'll start by
  62933. 47:00:41importing the libraries and module and
  62934. 47:00:43after that we'll load and prepare data.
  62935. 47:00:45That is we'll load transaction data from
  62936. 47:00:47a CSV file into the pandas data frame.
  62937. 47:00:50And after that we'll convert the
  62938. 47:00:51transaction time column to date time
  62939. 47:00:53format which allows the extraction of
  62940. 47:00:55additional time based features. And
  62941. 47:00:58after that we'll perform feature
  62942. 47:00:59engineering that will extract the hour
  62943. 47:01:01of the day from the transaction time
  62944. 47:01:03column. This feature can be important as
  62945. 47:01:05transactions occurring at unusual hours
  62946. 47:01:07may be indicative of fraud. And then
  62947. 47:01:09we'll move to the normalization of data.
  62948. 47:01:12This will apply standard scaling to the
  62949. 47:01:14amount n of the day feature. This
  62950. 47:01:16normalization process involves
  62951. 47:01:17subtracting the mean and dividing by the
  62952. 47:01:19standard deviation for each feature
  62953. 47:01:22ensuring that the feature contribute
  62954. 47:01:23equally to the analysis and improving
  62955. 47:01:25the performance of many machine learning
  62956. 47:01:27algorithms. And after that we'll start
  62957. 47:01:29with anomaly detection with isolation
  62958. 47:01:31forest. That's a technique. We'll
  62959. 47:01:34initialize an isolation forest model
  62960. 47:01:36with 100 trees that is n estimators
  62961. 47:01:39equal to 100. Setting the proportions of
  62962. 47:01:41outliers that is contamination to 1% of
  62963. 47:01:44the data. So this parameter is crucial
  62964. 47:01:46as it influences the threshold of
  62965. 47:01:48marking an observation as an anomaly.
  62966. 47:01:51Then we'll fit the model to the scaled
  62967. 47:01:52amount n of the data and predict the
  62968. 47:01:55anomaly status for each transaction. And
  62969. 47:01:58then we'll start with filter and display
  62970. 47:02:00anomalies. We'll filter out transactions
  62971. 47:02:02identified as anomalies that is anomaly
  62972. 47:02:05equal equal to minus one. We'll display
  62973. 47:02:07these transactions which can be reviewed
  62974. 47:02:09manually to determine if they represent
  62975. 47:02:12actual fraud net activity. So this code
  62976. 47:02:14snippet provides a systematic approach
  62977. 47:02:16to detecting anomalies in transaction
  62978. 47:02:18data leveraging the isolation forest
  62979. 47:02:20algorithms ability to handle complex and
  62980. 47:02:23highdimensional data set effectively. So
  62981. 47:02:25the pre-processing steps ensured that
  62982. 47:02:27the data is appropriately formatted and
  62983. 47:02:29normalized for optional model
  62984. 47:02:30performance. So this was all about
  62985. 47:02:32question number 16. Now moving to
  62986. 47:02:34question number 17 and that is based on
  62987. 47:02:36integrating machine learning models into
  62988. 47:02:38web applications. And your question is,
  62989. 47:02:40you have developed a machine learning
  62990. 47:02:42model to predict real estate prices
  62991. 47:02:44based on various features like location,
  62992. 47:02:46size, and amenities. How would you
  62993. 47:02:48integrate this model into a web
  62994. 47:02:50application to allow users to get
  62995. 47:02:52real-time price predictions? Can you
  62996. 47:02:54provide a sample Python code snippet to
  62997. 47:02:56illustrate how you would prepare the
  62998. 47:02:57model for integration and handle user
  62999. 47:03:00request? So, starting with the approach
  63000. 47:03:03that is integrating a machine learning
  63001. 47:03:04model into a web application. This will
  63002. 47:03:07involve several steps to ensure the
  63003. 47:03:09model is accessible and perform well in
  63004. 47:03:11a live environment. So here's how you
  63005. 47:03:13can approach this task. We could divide
  63006. 47:03:15into steps and we'll start with number
  63007. 47:03:17one step that is model preparation.
  63008. 47:03:19We'll finalize and save the model. So
  63009. 47:03:22once your model is trained and
  63010. 47:03:23validated, save it using a format that
  63011. 47:03:26can be easily loaded into a web
  63012. 47:03:28application. So Python's pickle module
  63013. 47:03:30or TensorFlow's save model format are
  63014. 47:03:32commonly used for this purpose. Then we
  63015. 47:03:35can use web application backend setup.
  63016. 47:03:37For this, select a suitable web
  63017. 47:03:39framework. So, Flask is popularly known
  63018. 47:03:42for its simplicity and effectiveness in
  63019. 47:03:44integrating Python based machine
  63020. 47:03:46learning models. And after that, we'll
  63021. 47:03:48develop the API. After developing the
  63022. 47:03:51API within your Flask app that you can
  63023. 47:03:53receive user inputs for model features,
  63024. 47:03:56load the model, make prediction, and
  63025. 47:03:58return the result. And after this, we'll
  63026. 47:04:00develop the UI. We'll design a
  63027. 47:04:02user-friendly interface. We'll create a
  63028. 47:04:04simple and intuitive UI that lets users
  63029. 47:04:07input the feature like location, size
  63030. 47:04:09and submit them for prediction. And
  63031. 47:04:12after that we'll move to the deployment
  63032. 47:04:13phase. We'll use a cloud platform like
  63033. 47:04:15Heroku, AWS or Google Cloud to deploy
  63034. 47:04:18your Flask application. And then we have
  63035. 47:04:20the maintenance and updates. We'll
  63036. 47:04:22monitor and update regularly for the
  63037. 47:04:24performance and use the model as needed
  63038. 47:04:27based on user feedback. So now moving to
  63039. 47:04:30the Python code and see how this model
  63040. 47:04:33can be created. So here we'll start
  63041. 47:04:35importing the libraries and module and
  63042. 47:04:37we are using flask ple and jsonify and
  63043. 47:04:41we will start with app initialization.
  63044. 47:04:43We'll initialize a new flask web
  63045. 47:04:45application that would be a special
  63046. 47:04:47variable which gives python files a
  63047. 47:04:50unique name to differentiate between
  63048. 47:04:52them when they are important into other
  63049. 47:04:54scripts. And after that we'll load the
  63050. 47:04:56model. So loading a pretend machine
  63051. 47:04:58learning model from the file system. So
  63052. 47:05:00this model is assumed to be saved in the
  63053. 47:05:02same directory as this script. So the
  63054. 47:05:04model is loaded in RB mode which stands
  63055. 47:05:07for read binary. And after that we'll
  63056. 47:05:09move to API route and prediction
  63057. 47:05:11function. So we will define an API
  63058. 47:05:14endpoint at predict that listens for
  63059. 47:05:17post request. This is the URL that the
  63060. 47:05:20front end of the web application will
  63061. 47:05:21call to send data to the back end. And
  63062. 47:05:24after that we'll start with predicting
  63063. 47:05:25the function. And here we have extract
  63064. 47:05:27features that retrieves data sent into
  63065. 47:05:30the JSON format from the post request
  63066. 47:05:32that is request get_json and the force
  63067. 47:05:35we have set it as true here and
  63068. 47:05:38forcefully formats the request data into
  63069. 47:05:40JSON ensuring compatibility and then
  63070. 47:05:43we'll extract the relevant features that
  63071. 47:05:44is location size and amenities from the
  63072. 47:05:47JSON object and store them in a list as
  63073. 47:05:49expected by the model and after
  63074. 47:05:51preparing the features we'll make the
  63075. 47:05:53prediction we'll use the loaded model to
  63076. 47:05:55make a prediction based bas on the
  63077. 47:05:56provided feature and then we have the
  63078. 47:05:58return prediction method. Here we will
  63079. 47:06:00convert the prediction result into JSON
  63080. 47:06:02format using JSON and send it back to
  63081. 47:06:05the client and this will ensure that the
  63082. 47:06:06response can be easily handled by the
  63083. 47:06:08client application. So this was all
  63084. 47:06:10about the question number 17. Now moving
  63085. 47:06:12to the question number 18 that is based
  63086. 47:06:14on analyzing. And now we move to the
  63087. 47:06:16question number 18 that is based on
  63088. 47:06:18analyzing geospatial data. And your
  63089. 47:06:20question is you are tasked with
  63090. 47:06:21analyzing geospatial data to help a city
  63091. 47:06:24improve its public transportation
  63092. 47:06:25system. The data includes GPS
  63093. 47:06:27coordinates of bus stops, ridership
  63094. 47:06:30numbers and traffic patterns. What steps
  63095. 47:06:32would you take to analyze this data? And
  63096. 47:06:34can you provide a sample Python code
  63097. 47:06:35snippet to illustrate how you might
  63098. 47:06:37visualize bus stop location and
  63099. 47:06:39ridership? So you can start answering
  63100. 47:06:41this question that juice better data
  63101. 47:06:44analysis can provide critical insights
  63102. 47:06:46into how effectively a public
  63103. 47:06:47transportation system serves its city
  63104. 47:06:49and guide improvements and there's an
  63105. 47:06:52detailed approach for this and we can
  63106. 47:06:54start with data preparation and in this
  63107. 47:06:56we'll do data collection and data
  63108. 47:06:58cleaning and after this step we'll move
  63109. 47:07:00to the next step that is explorative
  63110. 47:07:02data analysis and in this we'll have
  63111. 47:07:04statistical summary we'll generate
  63112. 47:07:06descriptive statistics and then we have
  63113. 47:07:08correlation analysis
  63114. 47:07:10And after moving that we have geospatial
  63115. 47:07:13visualization that is mapping bus stop.
  63116. 47:07:16We'll plot the locations of bus stop on
  63117. 47:07:17a map to visually assess their
  63118. 47:07:19distribution across the city. And after
  63119. 47:07:22that we have heat maps that will create
  63120. 47:07:24ridership data to identify hot sports
  63121. 47:07:26and areas with potential service gaps.
  63122. 47:07:29And after geospatial visualization we'll
  63123. 47:07:31move with spatial analysis. We have
  63124. 47:07:33proximity analysis that will analyze the
  63125. 47:07:36proximity of bus stop to key areas like
  63126. 47:07:38commercial centers or residential areas.
  63127. 47:07:41And now moving to the fifth step that is
  63128. 47:07:43optimization and recommendation. So
  63129. 47:07:45we'll have a route optimization that
  63130. 47:07:47will suggest modifications to route
  63131. 47:07:50based on traffic patterns and ridership
  63132. 47:07:52demand and the policy recommendations
  63133. 47:07:54that will provide actionable
  63134. 47:07:55recommendations for improving bus
  63135. 47:07:57frequencies. Now move to the sample
  63136. 47:07:59Python code where we can define this
  63137. 47:08:02model and use it accordingly. And here
  63138. 47:08:05we will start importing the libraries
  63139. 47:08:06and modules. And here we'll start with
  63140. 47:08:09importing geopandas and m lib dotpipo.
  63141. 47:08:14And after importing we'll start with
  63142. 47:08:16data loading. So we will declare a
  63143. 47:08:19variable bus stops and load the bus
  63144. 47:08:22stops data from a shape file. So shape
  63145. 47:08:25files are popular geospatial vector data
  63146. 47:08:27formats for geographic information
  63147. 47:08:29system software and then we have the
  63148. 47:08:31wrership that will load wrership data
  63149. 47:08:33from a CSV file which includes columns
  63150. 47:08:35for longitude latitude and ridership
  63151. 47:08:38levels and after that we'll create geo
  63152. 47:08:40data frame that will convert the
  63153. 47:08:42wrership data frame into a geo data
  63154. 47:08:44frame and this step involves creating a
  63155. 47:08:46geometry column from the longitude and
  63156. 47:08:48latitude columns and then we have the
  63157. 47:08:51plotting one here we will plot the
  63158. 47:08:53graphs that would figures and axis and
  63159. 47:08:55create a figure for the single subplot
  63160. 47:08:58with a specified size that is 10 + 10
  63161. 47:09:00in. And then we have city map.plot. It
  63162. 47:09:04is assumed that there is a base map of
  63163. 47:09:06the city loaded as a geo data frame
  63164. 47:09:09named city map. This is plotted first
  63165. 47:09:12with a light gray color to serve as a
  63166. 47:09:14background for the other layers. So this
  63167. 47:09:16was all about question number 18. Now
  63168. 47:09:18moving to question number 19 that is
  63169. 47:09:20based on predictive maintenance using
  63170. 47:09:23machine learning and the question is you
  63171. 47:09:25are tasked with developing a predictive
  63172. 47:09:27maintenance system for a manufacturing
  63173. 47:09:29plant that relies heavily on automated
  63174. 47:09:31machinery. So the data available
  63175. 47:09:33includes machine operational parameters,
  63176. 47:09:35maintenance history and failure
  63177. 47:09:36incidents. What steps would you take to
  63178. 47:09:38develop a predictive model and can you
  63179. 47:09:40provide a sample Python code? So you can
  63180. 47:09:42start with predictive maintenance that
  63181. 47:09:44is essential in manufacturing as it
  63182. 47:09:46helps prevent equipment failures
  63183. 47:09:48reducing downtime and maintenance cost.
  63184. 47:09:51And here you would have a detail
  63185. 47:09:53approach or predictive model for this
  63186. 47:09:55starting with data collection and
  63187. 47:09:57integration. Then you can do EDA that is
  63188. 47:09:59exploratory data analysis and then we
  63189. 47:10:02can perform feature engineering and then
  63190. 47:10:04move to data prep-processing task and
  63191. 47:10:07then the selection model and training
  63192. 47:10:09and after that we have model evaluation
  63193. 47:10:11and deployment technique that we can do
  63194. 47:10:13for the model using appropriate metrics
  63195. 47:10:16such as precision, recall and F1 score.
  63196. 47:10:18So this was all about question number
  63197. 47:10:2019. So now move to question number 20
  63198. 47:10:22that is based on personalization using
  63199. 47:10:24machine learning and your question is
  63200. 47:10:26you are tasked with developing a machine
  63201. 47:10:28learning model to personalize content
  63202. 47:10:30recommendations for users on a media
  63203. 47:10:32streaming platform. The data available
  63204. 47:10:34includes user demographic retails
  63205. 47:10:36viewing history and ratings. So what
  63206. 47:10:38steps would you take to build a model
  63207. 47:10:40for personalized recommendations and can
  63208. 47:10:42you provide a sample Python code for
  63209. 47:10:44that? So you can start answering this
  63210. 47:10:46with creating a personalized
  63211. 47:10:47recommendation systems. This would be
  63212. 47:10:49essential for engaging users by
  63213. 47:10:51providing content that is relevant to
  63214. 47:10:54their interest. And there will be a
  63215. 47:10:55systematic approach or personalized
  63216. 47:10:57content recommendation. We'll start with
  63217. 47:11:00data collection and integration. And
  63218. 47:11:02after that, we'll perform EDA that is
  63219. 47:11:04explorative data analysis. And then we
  63220. 47:11:06have feature engineering. In this we'll
  63221. 47:11:08interact features and the temporal
  63222. 47:11:11features. We'll include time based
  63223. 47:11:13features to capture trends and
  63224. 47:11:15seasonality in viewing behavior. And
  63225. 47:11:17then we'll select the model that is by
  63226. 47:11:19collaborative filtering and hybrid
  63227. 47:11:21models. And then we'll train the model
  63228. 47:11:23and validation and implement and monitor
  63229. 47:11:26them. And after that we'll deploy the
  63230. 47:11:28model. So let's start with beginner
  63231. 47:11:30level questions. And number one is what
  63232. 47:11:32is machine learning? So machine learning
  63233. 47:11:34is a subset of artificial intelligence
  63234. 47:11:37that involves the use of algorithms and
  63235. 47:11:39statistical models to enable computers
  63236. 47:11:42to perform task without explicit
  63237. 47:11:44instructions. that is by relying on
  63238. 47:11:46patterns and interference. And now
  63239. 47:11:49moving to number second question that is
  63240. 47:11:51what are the different types of machine
  63241. 47:11:52learning. So the three main types of
  63242. 47:11:55machine learning are number one is
  63243. 47:11:57supervised learning and then comes
  63244. 47:11:59unsupervised learning and then there is
  63245. 47:12:01reinforcement learning. Now moving to
  63246. 47:12:04next question that is third that is what
  63247. 47:12:06is supervised learning. So supervised
  63248. 47:12:09learning involves training a model on a
  63249. 47:12:11label data set which means each training
  63250. 47:12:13example is paired with an output label.
  63251. 47:12:16The model learns to predict the output
  63252. 47:12:18from the input data. Now moving to the
  63253. 47:12:21fourth question that is what is
  63254. 47:12:22unsupervised learning. So unsupervised
  63255. 47:12:25involve training a model on data that
  63256. 47:12:27does not have labeled responses. The
  63257. 47:12:29model tries to learn the patterns and
  63258. 47:12:31the structure from the input data. So
  63259. 47:12:34guys, these are the beginner level
  63260. 47:12:35questions and now we'll move to the
  63261. 47:12:37fifth question that is what is
  63262. 47:12:38reinforcement learning. So reinforcement
  63263. 47:12:41learning is a type of machine learning
  63264. 47:12:43where an agent learns to make decisions
  63265. 47:12:45by performing actions and receiving
  63266. 47:12:47rewards or penalties. The goal is to
  63267. 47:12:50maximize the cumulative reward. So now
  63268. 47:12:52moving to the sixth question that is
  63269. 47:12:54what is a model in machine learning. So
  63270. 47:12:57a model in machine learning is a
  63271. 47:12:59mathematical representation of a real
  63272. 47:13:01world process. It is trained on data to
  63273. 47:13:04recognize patterns and make predictions
  63274. 47:13:06or decisions based on new data. So now
  63275. 47:13:09moving to seventh question that is what
  63276. 47:13:11is overfitting? So overfitting occurs
  63277. 47:13:13when a machine learning model performs
  63278. 47:13:15well on the training data but poorly on
  63279. 47:13:17new unseen data. It indicates that the
  63280. 47:13:20model has learned the noise and details
  63281. 47:13:22in the training data instead of the
  63282. 47:13:24actual patterns. So now coming to
  63283. 47:13:26question number eight that is what is
  63284. 47:13:28underfitting? So underfitting occurs
  63285. 47:13:30when a machine learning model is too
  63286. 47:13:32simple to capture the underlying
  63287. 47:13:34patterns in the data. It performs poorly
  63288. 47:13:37on both the training data and new data.
  63289. 47:13:39Now move to the next question that is
  63290. 47:13:41ninth question and the question is what
  63291. 47:13:43is a confusion matrix? So confusion
  63292. 47:13:45matrix is a table used to evaluate the
  63293. 47:13:48performance of a classification model.
  63294. 47:13:50It summarizes the number of correct and
  63295. 47:13:52incorrect predictions made by the model
  63296. 47:13:55and that is categorized by each class.
  63297. 47:13:58Now moving to the 10th question that is
  63298. 47:13:59what is cross validation? So cross
  63299. 47:14:02validation is a technique for assessing
  63300. 47:14:04how the results of a statistical
  63301. 47:14:06analysis will generalize to an
  63302. 47:14:08independent data set. It involves
  63303. 47:14:10partitioning the data into subsets.
  63304. 47:14:13Training the model on some subsets and
  63305. 47:14:15validating it on the remaining subsets.
  63306. 47:14:18This was all about that is the 10th
  63307. 47:14:21question or the overall 1 to 10
  63308. 47:14:23questions for beginner level. Now we'll
  63309. 47:14:25move to intermediate level and here
  63310. 47:14:27we'll cover 10 questions. So we'll start
  63311. 47:14:30with 11th question that is what is a ROC
  63312. 47:14:33curve. So ROC that is receiver operating
  63313. 47:14:37characteristic curve. It is a graphical
  63314. 47:14:39representation of a classifier's
  63315. 47:14:41performance across different thresholds.
  63316. 47:14:44It plots the true positive rate that is
  63317. 47:14:46TPR against a false positive rate that
  63318. 47:14:49is FPR. Now moving to 12th question that
  63319. 47:14:52is what is precision and recall. So
  63320. 47:14:54precision is the ratio of correctly
  63321. 47:14:56predicted positive observations to the
  63322. 47:14:58total predicted positives and recall is
  63323. 47:15:01the ratio of correctly predicted
  63324. 47:15:03positive observations to all actual
  63325. 47:15:06positives. So the formula is precision
  63326. 47:15:08equal to TP/TP
  63327. 47:15:11plus FP and the recall is TP/TP
  63328. 47:15:15+ F_sub_1. So now we'll move to the 13th
  63329. 47:15:19question that is what is the F1 score?
  63330. 47:15:22So the F1 score is the harmonic mean of
  63331. 47:15:25precision and recall. It provides a
  63332. 47:15:27balance between the two metrics and is
  63333. 47:15:29useful when you need to balance
  63334. 47:15:31precision and recall. F1 score is equal
  63335. 47:15:34to twice into precision into recall and
  63336. 47:15:38that is divided by precision plus
  63337. 47:15:40recall. Now we'll move to 14th question
  63338. 47:15:43and here we will cover regularization.
  63339. 47:15:46So the question is what is
  63340. 47:15:47regularization? So it is a technique
  63341. 47:15:50used to prevent overfitting by adding a
  63342. 47:15:52penalty to the model's complexity and
  63343. 47:15:55the common types of regularization
  63344. 47:15:57include L1 that is lasso and L2 ridge
  63345. 47:16:00regularization. Now we'll move to the
  63346. 47:16:0215th question that is what is the bias
  63347. 47:16:05variance tradeoff. So the bias variance
  63348. 47:16:07trade-off is a fundamental issue in
  63349. 47:16:10machine learning that involves balancing
  63350. 47:16:12the error introduced by the model's
  63351. 47:16:14assumptions and the error due to model
  63352. 47:16:17complexity. So a good model should have
  63353. 47:16:19low bias and low variance. Now we'll
  63354. 47:16:22move to the question number 16 that is
  63355. 47:16:24what is feature engineering. So feature
  63356. 47:16:26engineering is the process of creating
  63357. 47:16:28new features or modifying existing ones
  63358. 47:16:31to improve the performance of a machine
  63359. 47:16:33learning model. It involves techniques
  63360. 47:16:35like normalization and coding
  63361. 47:16:37categorical variables and creating
  63362. 47:16:40interaction terms. So now we'll move to
  63363. 47:16:42question number 17 and that is about
  63364. 47:16:45gradient descent. So the question is
  63365. 47:16:47what is gradient descent and your answer
  63366. 47:16:49is gradient descent is an optimization
  63367. 47:16:52algorithm used to minimize the cost
  63368. 47:16:54function in machine learning models and
  63369. 47:16:56it iteratively adjust the model
  63370. 47:16:58parameters in the direction of the
  63371. 47:17:00steepest descent of the coast function.
  63372. 47:17:04So with this we'll move to the 18th
  63373. 47:17:06question and that will cover with the
  63374. 47:17:08difference between bagging and boosting.
  63375. 47:17:11So the question is what is difference
  63376. 47:17:13between bagging and boosting and you
  63377. 47:17:14could answer this with starting with
  63378. 47:17:16bagging that is bootstrap aggregating
  63379. 47:17:20that involves training multiple models
  63380. 47:17:22on different subsets of the data and
  63381. 47:17:24averaging their predictions. Then comes
  63382. 47:17:26boosting that involves training models
  63383. 47:17:28sequentially with each new model
  63384. 47:17:30focusing on correcting the errors of the
  63385. 47:17:32previous ones. And then we have the
  63386. 47:17:35question number 19 that is what is a
  63387. 47:17:37decision tree? So a decision tree is a
  63388. 47:17:39nonparametric supervised learning
  63389. 47:17:41algorithm used for classification and
  63390. 47:17:44regression. It splits the data into
  63391. 47:17:46subsets based on the value of input
  63392. 47:17:48features resulting in a treel like
  63393. 47:17:50structure of decisions. Now we'll move
  63394. 47:17:52to question number 20 that is what is a
  63395. 47:17:54random forest. So random forest is an
  63396. 47:17:56ansemble learning method that combines
  63397. 47:17:59multiple decision trees to improve the
  63398. 47:18:01accuracy and robustness of the model. It
  63399. 47:18:04builds each tree using a random subset
  63400. 47:18:06of features and data points and then
  63401. 47:18:09averages their predictions. So these
  63402. 47:18:11were the questions that are for the
  63403. 47:18:13intermediate level and these are just
  63404. 47:18:15the basic questions or I will just say
  63405. 47:18:18the theoretical questions that can be
  63406. 47:18:20asked in an interview. So be prepared
  63407. 47:18:22for that. Now we'll move to the advanced
  63408. 47:18:24level interview questions and we'll
  63409. 47:18:26start with question number 21. And here
  63410. 47:18:28also we'll cover the 10 questions. So
  63411. 47:18:30number one question or that is 21
  63412. 47:18:33question and the question is what is a
  63413. 47:18:35support vector machine? So support
  63414. 47:18:37vector machine is a supervised learning
  63415. 47:18:39algorithm used for classification and
  63416. 47:18:42regression. It finds the optimal hyper
  63417. 47:18:44plane that maximizes the margin between
  63418. 47:18:46different classes in the feature space.
  63419. 47:18:49And then comes question number 22 that
  63420. 47:18:51is what is principal component analysis.
  63421. 47:18:54So principal component analysis is a
  63422. 47:18:56dimensionality reduction technique that
  63423. 47:18:58transforms highdimensional data into a
  63424. 47:19:01lower dimensional space by finding the
  63425. 47:19:03directions that is principal components
  63426. 47:19:06that maximize the variance in the data.
  63427. 47:19:08And then comes the question number 23
  63428. 47:19:10that is what is a neural network? So a
  63429. 47:19:12neural network is a series of algorithms
  63430. 47:19:14that attempt to recognize underlying
  63431. 47:19:17relationships in a set of data through a
  63432. 47:19:19process that mimics the way the human
  63433. 47:19:21brain operates. It consists of layer of
  63434. 47:19:24interconnected nodes or neurons. And
  63435. 47:19:27then comes the question number 24 that
  63436. 47:19:29is what is deep learning? So deep
  63437. 47:19:31learning is a subset of machine learning
  63438. 47:19:33that involves neural networks with many
  63439. 47:19:35layers that is deep neural networks and
  63440. 47:19:37it is particularly effective for task
  63441. 47:19:39like image and speech recognition. Now I
  63442. 47:19:42move to question number 25 that is what
  63443. 47:19:44is convolutional neural network that is
  63444. 47:19:48CNN. So we will start the answer by
  63445. 47:19:50answering the interviewer that a
  63446. 47:19:52convolutional neural network is a type
  63447. 47:19:54of deep learning model specifically
  63448. 47:19:56designed for processing structured grid
  63449. 47:19:58data like images. It uses convolutional
  63450. 47:20:01layers to extract special features or
  63451. 47:20:04the spatial features and patterns from
  63452. 47:20:06the input data. Now we move to the
  63453. 47:20:08question number 26 that is what is a
  63454. 47:20:10recurrent neural network or RNN. So a
  63455. 47:20:14recurrent neural network is a type of
  63456. 47:20:16neural network designed for sequential
  63457. 47:20:18data and it has connections that form
  63458. 47:20:20directed cycles allowing it to maintain
  63459. 47:20:23a memory of previous inputs and process
  63460. 47:20:26sequences of data. So this was all about
  63461. 47:20:28question number 26 and now we will cover
  63462. 47:20:30the question number 27 that is what is
  63463. 47:20:32the difference between batch gradient
  63464. 47:20:34descent and stoastic gradient descent.
  63465. 47:20:38So batch gradient descent computes the
  63466. 47:20:40gradient of the coast function using the
  63467. 47:20:43entire training data set while
  63468. 47:20:45stochastic gradient descent that is SGD
  63469. 47:20:48computes the gradient using only one
  63470. 47:20:50training example at a time. So SGD is
  63471. 47:20:54faster but noisier. Now we move to
  63472. 47:20:56question number 28 that is what is
  63473. 47:20:58dropout in neural networks. So dropout
  63474. 47:21:00is a regularization technique used in
  63475. 47:21:03neural networks to prevent overfitting
  63476. 47:21:05and it involves randomly setting a
  63477. 47:21:07fraction of the neurons to zero during
  63478. 47:21:10training forcing the network to learn
  63479. 47:21:12more robust features. And now we'll move
  63480. 47:21:15to question number 29 and that will be
  63481. 47:21:17about transfer learning. And your
  63482. 47:21:19question is what is transfer learning?
  63483. 47:21:22So we'll answer this to the interviewer
  63484. 47:21:23by starting that transfer learning is a
  63485. 47:21:26technique in machine learning where a
  63486. 47:21:28model developed for one task is reused
  63487. 47:21:30as the starting point for a model on a
  63488. 47:21:33second related task. It is particularly
  63489. 47:21:36useful when there is limited data
  63490. 47:21:37available for the second task. Now we'll
  63491. 47:21:40move to the last question and the 30th
  63492. 47:21:42question. So that is what is a
  63493. 47:21:44generative adversial network that is GN.
  63494. 47:21:48So you can start answering this. So
  63495. 47:21:50generative adversial network is a type
  63496. 47:21:52of deep learning model consisting of two
  63497. 47:21:55neural networks a generator and a
  63498. 47:21:57discriminator that are trained
  63499. 47:21:59simultaneously. The generator creates
  63500. 47:22:01fake data while the discriminator tries
  63501. 47:22:04to distinguish between real and fake
  63502. 47:22:06data leading to the generator producing
  63503. 47:22:08increasingly realistic data. And these
  63504. 47:22:11questions and answers are over and these
  63505. 47:22:14covers a wide range of topics in machine
  63506. 47:22:16learning and should help prepare for
  63507. 47:22:18interviews at various levels.
  63508. 47:22:19>> And with that we have reached the end of
  63509. 47:22:21the session on the AI and machine
  63510. 47:22:23learning engineer full course for
  63511. 47:22:25beginners. If you have any doubts or
  63512. 47:22:27questions about this video let us know
  63513. 47:22:29in the comment section below and a team
  63514. 47:22:30of experts will be happy to help you.
  63515. 47:22:32Until next time thank you and keep
  63516. 47:22:34learning. Stay tuned for more from
  63517. 47:22:35SimplyLearn.

About this transcript

This page contains the full transcript of AI And Machine Learning Full Course [FREE] | Learn AI And Machine Learning In 24 Hours | Simplilearn by Simplilearn, generated from the public captions YouTube serves with the video. The transcript has 423,981 words across 63,517 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.