4.2.9 An Introduction to Trees - Video 5: Random Forests — Transcript
Full transcript
- 0:03in this video we'll introduce a method
- 0:06that is similar to cart called random
- 0:09forests this method was designed to
- 0:12improve the prediction accuracy of cart
- 0:14and works by building a large number of
- 0:17cart trees unfortunately this makes the
- 0:20method less interpretable than cart so
- 0:23often you need to decide if you value
- 0:25the interpretability
- 0:26or the increase in accuracy more to make
- 0:30a prediction for a new observation each
- 0:33tree in the forest votes on the outcome
- 0:36and we pick the outcome that receives
- 0:38the majority of the votes so how does
- 0:42random forests build many cart trees we
- 0:45can't just run cart multiple times
- 0:47because it would create the same tree
- 0:49every time to prevent this random
- 0:53forests only allows each tree to split
- 0:56on a random subset of the available
- 0:58independent variables and each tree is
- 1:02built from what we call a bagged or
- 1:04bootstrapped sample of the data this
- 1:07just means that the data used as the
- 1:09training data for each tree is selected
- 1:11randomly with replacement let's look at
- 1:14an example suppose we have five data
- 1:17points in our training set we'll call
- 1:19them one two three four and five for the
- 1:23first tree will randomly pick five data
- 1:26points randomly sampled with replacement
- 1:29so the data could be two four five two
- 1:35and one each time we pick one of the
- 1:38five data points regardless of whether
- 1:40or not it's been selected already the
- 1:43these would be the five data points we
- 1:45would use when constructing the first
- 1:47cart tree then we repeat this process
- 1:50for the second tree this time the data
- 1:53set might be three five one five and two
- 1:57and we would use this data when building
- 1:59the second cart tree then we would
- 2:02repeat this process for each additional
- 2:04tree we want to create so since each
- 2:08tree sees a different set of variables
- 2:10and a different set of data we get
- 2:13what's called a forest of many different
- 2:15trees
- 2:17just like cart branda forests has some
- 2:21parameter values that need to be
- 2:22selected the first is the minimum number
- 2:25of observations in a subset or the min
- 2:28bucket parameter from cart when we
- 2:31create a random forest in our this will
- 2:33be called node size a smaller value of
- 2:37node size which leads to bigger trees
- 2:39may take longer in our random forests is
- 2:43much more computationally intensive than
- 2:46cart the second parameter is the number
- 2:49of trees to build which is called entry
- 2:52in our this should not be set to small
- 2:56but the larger it is the longer it will
- 2:58take a couple hundred trees is typically
- 3:01plenty a nice thing about random forests
- 3:04is that it's not as sensitive to the
- 3:06parameter values as car is in the next
- 3:09video we'll talk about a nice way to
- 3:11pick the cart parameter for random
- 3:14forests as long as the selection is
- 3:16reasonable it's okay let's switch to our
- 3:19and create a random forest model to
- 3:21predict the decisions of justice Stevens
- 3:26inner our consul let's start by
- 3:29installing and loading the package
- 3:31random forests we first need to install
- 3:34the package using the install dot
- 3:37packages function for the package random
- 3:41forests you should see a few lines run
- 3:45in your our console and then when you're
- 3:47back to the blinking cursor load the
- 3:50package at the library command now we're
- 3:56ready to build our random forest model
- 3:58we'll call it
- 4:00Stephens forest and use the random
- 4:03forest function first giving our
- 4:06dependent variable reverse followed by a
- 4:09tilde sign and then our independent
- 4:11variables separated by plus signs
- 4:13circuit issue petitioner respondent
- 4:20lower court an unconstitutional will use
- 4:27the data set train
- 4:30for random forests we need to give two
- 4:33additional arguments these are nodes
- 4:35size also known as min bucket for cart
- 4:39and we'll set this equal to 25 the same
- 4:42value we used for our cart model and
- 4:44then we need to set the parameter and
- 4:46tree this is the number of trees to
- 4:48build and we'll build 200 trees here
- 4:51then hit enter you should see an
- 4:55interesting warning message here in cart
- 4:58we added the argument method equals
- 5:00class so that it was clear that we're
- 5:02doing a classification problem as I
- 5:05mentioned earlier trees can also be used
- 5:07for regression problems which you'll see
- 5:09in the recitation the random forest
- 5:12function does not have a method argument
- 5:14so we'll may want to do a classification
- 5:16problem we need to make sure our outcome
- 5:19is a factor let's convert the variable
- 5:22reverse to a factor variable in both our
- 5:25training and our testing sets we do this
- 5:28by typing the name of the variable we
- 5:31want to convert in our case train
- 5:33reverse and then type a s dot factor and
- 5:37then in parentheses the variable name
- 5:40train reverse and just repeat this for
- 5:44the test set as well test reversed
- 5:47equals a s stop factor test dollar sign
- 5:52reverse now let's try creating a random
- 5:56forest again just use the up arrow to
- 5:59get back to the random forest line and
- 6:01hit enter we didn't get a warning
- 6:03message this time so our models ready to
- 6:06make predictions let's compute
- 6:09predictions on our test set we'll call
- 6:11our predictions predict forest and use
- 6:16the predict function to make predictions
- 6:18using our model Stephens forests and the
- 6:23new data set test let's look at the
- 6:28confusion matrix to compute our accuracy
- 6:31we'll use the table function and first
- 6:34give the true outcome test reverse and
- 6:37then our predictions predict forest
- 6:42our accuracy here is 40 plus 74 divided
- 6:48by 40 plus 37 plus 19 plus 74 so the
- 6:56accuracy of our random forest model is
- 6:59about 67% recall that our logistic
- 7:02regression model had an accuracy of
- 7:0566.5% and our cart model had an accuracy
- 7:08of 65.9% so a random forest model
- 7:12improved our accuracy a little bit over
- 7:15Kart sometimes you'll see a smaller
- 7:18improvement in accuracy and sometimes
- 7:20you'll see that random forests can
- 7:22significantly improve an accuracy over
- 7:24Kart we'll see this a lot in the
- 7:27recitation and the homework assignments
- 7:29keep in mind that random forests has a
- 7:32random component you may have gotten a
- 7:34different confusion matrix than me
- 7:36because there's a random component to
- 7:39this method
About this transcript
This page contains the full transcript of 4.2.9 An Introduction to Trees - Video 5: Random Forests by MIT OpenCourseWare, generated from the public captions YouTube serves with the video. The transcript has 1,051 words across 154 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.