YouTube2Text

4.2.7 An Introduction to Trees - Video 4: CART in R — Transcript

by MIT OpenCourseWare · 1,565 words · 290 segments · language en · Watch on YouTube

Full transcript

  1. 0:04In this video, we'll see how to build a
  2. 0:06CART model in R.
  3. 0:08Let's start by reading in the data file
  4. 0:10stevens.csv.
  5. 0:13We'll call our data frame stevens
  6. 0:16and use the read .csv function to read
  7. 0:20in the data file stevens
  8. 0:24.csv.
  9. 0:26Remember to navigate to the directory on
  10. 0:28your computer containing the file
  11. 0:30stevens.csv
  12. 0:31first.
  13. 0:33Now, let's take a look at our data using
  14. 0:35the str function.
  15. 0:40We have 566
  16. 0:42observations or Supreme Court cases and
  17. 0:46nine different variables.
  18. 0:48Docket is just a unique identifier for
  19. 0:50each case and term is the year of the
  20. 0:53case.
  21. 0:54Then, we have our six independent
  22. 0:56variables. The circuit court of origin,
  23. 1:00the issue area of the case,
  24. 1:02the type of petitioner,
  25. 1:04the type of respondent,
  26. 1:06the lower court direction,
  27. 1:09and whether or not the petitioner argued
  28. 1:11that a law or practice was
  29. 1:13unconstitutional.
  30. 1:15The last variable is our dependent
  31. 1:17variable, whether or not Justice Stevens
  32. 1:20voted to reverse the case.
  33. 1:22One for reverse and zero for affirm.
  34. 1:26Now, before building models, we need to
  35. 1:28split our data into a training set and a
  36. 1:31testing set.
  37. 1:32We'll do this using the sample.split
  38. 1:35function like we did last week for
  39. 1:37logistic regression.
  40. 1:39First, we need to load the package
  41. 1:41caTools
  42. 1:42with library
  43. 1:45caTools.
  44. 1:49Now, so that we all get the same split,
  45. 1:52we need to set the seed. Remember that
  46. 1:55this can be any number as long as we all
  47. 1:57use the same number.
  48. 1:59Let's set the seed
  49. 2:01to 3000.
  50. 2:06Now, let's create our split. We'll call
  51. 2:08it spl
  52. 2:11and we'll use the sample.split
  53. 2:13function
  54. 2:15where the first argument needs to be our
  55. 2:17outcome variable
  56. 2:19stevens
  57. 2:22reverse
  58. 2:24and then the second argument is the
  59. 2:26split ratio or the percentage of data
  60. 2:29that we want to put in the training set.
  61. 2:31In this case, we'll put 70% of the data
  62. 2:34in the training set.
  63. 2:37Now, let's create our training and
  64. 2:39testing sets using the subset function.
  65. 2:43We'll call our training set train
  66. 2:46and we'll take a subset of stevens
  67. 2:52only taking the observations for which
  68. 2:54spl is equal to true.
  69. 2:58We'll call our testing set test and here
  70. 3:01take a subset of stevens
  71. 3:04but this time taking the observations
  72. 3:06for which spl is equal to false.
  73. 3:12Now, we're ready to build our CART
  74. 3:13model.
  75. 3:14First, we need to install and load the
  76. 3:17rpart package and the rpart.plotting
  77. 3:20package.
  78. 3:21Remember that to install a new package,
  79. 3:24we use the install.packages
  80. 3:27function
  81. 3:29and then in parentheses and quotes, give
  82. 3:31the name of the package we want to
  83. 3:33install. In this case, rpart.
  84. 3:36After you hit enter, a CRAN mirror
  85. 3:38should pop up asking you to pick a
  86. 3:41location near you.
  87. 3:43Go ahead and pick the appropriate
  88. 3:45location. In my case, I'll pick
  89. 3:47Pennsylvania in the United States and
  90. 3:50hit okay.
  91. 3:52You should see some lines running your R
  92. 3:54console and then when you're back to the
  93. 3:56blinking cursor, load the package with
  94. 4:00library
  95. 4:01rpart.
  96. 4:04Now, let's install the package
  97. 4:12rpart .plot.
  98. 4:19Again, some lines should run in your R
  99. 4:20console and when you're back to the
  100. 4:22blinking cursor, load the package with
  101. 4:25library
  102. 4:27rpart.plot.
  103. 4:31Now, we can create our CART model using
  104. 4:33the rpart function.
  105. 4:35We'll call our model stevens.tree
  106. 4:39and we'll use the rpart function where
  107. 4:41the first argument is the same as if we
  108. 4:44were building a linear or logistic
  109. 4:45regression model.
  110. 4:47We give our dependent variable, in our
  111. 4:49case reverse,
  112. 4:51followed by a tilde sign and then the
  113. 4:53independent variables separated by plus
  114. 4:55signs. So, circuit plus
  115. 4:59issue plus
  116. 5:02petitioner
  117. 5:03plus
  118. 5:04respondent
  119. 5:06plus
  120. 5:08lower court
  121. 5:10plus unconstitutional.
  122. 5:13We also need to give our data set that
  123. 5:15should be used to build our model, which
  124. 5:17in our case is train.
  125. 5:20Now, we'll give two additional arguments
  126. 5:22here. The first one is method equals
  127. 5:25class.
  128. 5:27This tells rpart to build a
  129. 5:29classification tree instead of a
  130. 5:31regression tree.
  131. 5:32You'll see how we can create regression
  132. 5:34trees in recitation.
  133. 5:37The last argument we'll give is
  134. 5:39minbucket equals 25.
  135. 5:42This limits the tree so that it doesn't
  136. 5:45overfit to our training set.
  137. 5:47We selected a value of 25, but we could
  138. 5:49pick a smaller or larger value.
  139. 5:52We'll see another way to limit the tree
  140. 5:54later in this lecture.
  141. 5:57Now, let's plot our tree using the prp
  142. 6:00function where the only argument is the
  143. 6:02name of our model, stevens.tree.
  144. 6:07You should see the tree pop up in the
  145. 6:09graphics window.
  146. 6:11The first split of our tree is whether
  147. 6:13or not the lower court decision is
  148. 6:15liberal.
  149. 6:17If it is, then we move to the left in
  150. 6:19the tree and we check the respondent.
  151. 6:22If the respondent is a criminal
  152. 6:24defendant, injured person, politician,
  153. 6:28state, or the United States, we predict
  154. 6:31zero or affirm.
  155. 6:33You can see here that the prp function
  156. 6:36abbreviates the values of the
  157. 6:37independent variables.
  158. 6:39If you're not sure what the
  159. 6:41abbreviations are, you could create a
  160. 6:43table of the variable to see all of the
  161. 6:45possible values.
  162. 6:47prp will select the abbreviation so that
  163. 6:50they're uniquely identifiable.
  164. 6:52So, if you made a table, you could see
  165. 6:55that cri stands for criminal defendant,
  166. 6:58inj stands for injured person, etc.
  167. 7:02So, now moving on in our tree, if the
  168. 7:05respondent is not one of these types, we
  169. 7:08move on to the next split and we check
  170. 7:10the petitioner.
  171. 7:12If the petitioner is a city, employee,
  172. 7:15employer, government official, or
  173. 7:17politician, then we predict zero or
  174. 7:20affirm.
  175. 7:22If not, then we check the circuit court
  176. 7:24of origin.
  177. 7:25If it's the 10th, first, third, fourth,
  178. 7:30DC, or federal court, then we predict
  179. 7:33zero. Otherwise, we predict one or
  180. 7:36reverse.
  181. 7:38We can repeat the same process on the
  182. 7:40other side of the tree if the lower
  183. 7:42court decision is not liberal.
  184. 7:45Comparing this to a logistic regression
  185. 7:47model, we can see that it's very
  186. 7:49interpretable.
  187. 7:51A CART tree is a series of decision
  188. 7:53rules which can easily be explained.
  189. 7:57Now, let's see how well our CART model
  190. 7:59does at making predictions for the test
  191. 8:01set.
  192. 8:02So, back in our R console, we'll call
  193. 8:05our predictions predict cart
  194. 8:09and we'll use the predict function
  195. 8:12where the first argument is the name of
  196. 8:14our model, stevens
  197. 8:18tree.
  198. 8:20The second argument is the new data we
  199. 8:22want to make predictions for,
  200. 8:24test
  201. 8:26and we'll add a third argument here,
  202. 8:28which is type equals class.
  203. 8:32We need to give this argument when
  204. 8:34making predictions for our CART model if
  205. 8:36we want the majority class predictions.
  206. 8:40This is like using a threshold of .5.
  207. 8:43We'll see in a few minutes how we can
  208. 8:44leave this argument out and still get
  209. 8:46probabilities from our CART model.
  210. 8:49Now, let's compute the accuracy of our
  211. 8:51model by building a confusion matrix.
  212. 8:54So, we'll use the table function
  213. 8:57and first give the true outcome values
  214. 9:00test
  215. 9:01reverse
  216. 9:03and then our predictions predict cart.
  217. 9:09To compute the accuracy, we need to add
  218. 9:11up the observations we got correct, 41 +
  219. 9:1471 divided by the total number of
  220. 9:17observations in the table or the total
  221. 9:19number of observations in our test set.
  222. 9:23So, the accuracy of our CART model is
  223. 9:25.659.
  224. 9:27If you were to build a logistic
  225. 9:29regression model, you would get an
  226. 9:30accuracy of .665.
  227. 9:33And a baseline model that always
  228. 9:35predicts reverse, the most common
  229. 9:37outcome, has an accuracy of .547.
  230. 9:41So, our CART model significantly beats
  231. 9:43the baseline and is competitive with
  232. 9:46logistic regression.
  233. 9:47It's also much more interpretable than a
  234. 9:50logistic regression model would be.
  235. 9:53Lastly, to evaluate our model, let's
  236. 9:55generate an ROC curve for our CART model
  237. 9:58using the ROCR package.
  238. 10:01First, we need to load the package with
  239. 10:03the library function.
  240. 10:06And then, we need to generate our
  241. 10:08predictions again, this time without the
  242. 10:10type equals class argument. We'll call
  243. 10:13them predict ROC,
  244. 10:16and we'll use the predict function,
  245. 10:18giving just as the two arguments Stevens
  246. 10:21tree
  247. 10:22and new data equals test.
  248. 10:25Let's take a look at what this looks
  249. 10:27like by just typing predict ROC and
  250. 10:30hitting enter.
  251. 10:33For each observation in the test set, it
  252. 10:35gives two numbers, which can be thought
  253. 10:38of as the probability of outcome zero
  254. 10:40and the probability of outcome one.
  255. 10:43More concretely, each test set
  256. 10:45observation is classified into a subset
  257. 10:48or bucket of our CART tree.
  258. 10:50These numbers give the percentage of
  259. 10:52training set data in that subset with
  260. 10:55outcome zero and the percentage of data
  261. 10:58in the training set in that subset with
  262. 11:00outcome one.
  263. 11:02We'll use the second column as our
  264. 11:04probabilities to generate an ROC curve.
  265. 11:07So, just like we did last week for
  266. 11:09logistic regression, we'll start by
  267. 11:11using the prediction function. We'll
  268. 11:13call the output pred,
  269. 11:16and then use prediction,
  270. 11:18where the first argument is the second
  271. 11:20column of predict ROC, which we can
  272. 11:22access with square brackets,
  273. 11:25and the second argument is the true
  274. 11:27outcome values, test reverse.
  275. 11:30Now, we need to use the performance
  276. 11:32function,
  277. 11:35where the first argument is the outcome
  278. 11:38of the prediction function,
  279. 11:39and then the next two arguments are true
  280. 11:42positive rate and false positive rate,
  281. 11:45what we want on the X and Y axes of our
  282. 11:47ROC curve.
  283. 11:49Now, we can just plot our ROC curve by
  284. 11:51typing plot p e r f.
  285. 11:56If you switch back to your graphics
  286. 11:57window, you should see the ROC curve for
  287. 12:00our model.
  288. 12:01In the next quick question, we'll ask
  289. 12:04you to compute the test set AUC of this
  290. 12:06model.

About this transcript

This page contains the full transcript of 4.2.7 An Introduction to Trees - Video 4: CART in R by MIT OpenCourseWare, generated from the public captions YouTube serves with the video. The transcript has 1,565 words across 290 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.