YouTube transcript (wPJ1_Z8b0wk) — Transcript
Full transcript
- 0:02[Music]
- 0:07Hello, thanks for watching and welcome
- 0:10to the next video in my series on basic
- 0:12statistics. Now, as usual, a few things
- 0:15before we get started. Number one, if
- 0:17you're watching this video because you
- 0:19are struggling in a class right now, I
- 0:21want you to stay positive and keep your
- 0:23head up. If you're watching this, it
- 0:25means you've accomplished quite a bit
- 0:26already. You're very smart and talented,
- 0:29but you may have just hit a temporary
- 0:31rough patch. Now, I know with the right
- 0:33amount of hard work, practice, and
- 0:35patience, you can work through it. I
- 0:38have faith in you. Many other people
- 0:40around you have faith in you. So, so
- 0:43should you. Number two, please feel free
- 0:46to follow me here on YouTube, on
- 0:48Twitter, on Google+, or on LinkedIn.
- 0:52That way, when I upload a new video, you
- 0:54know about it. And it's always nice to
- 0:56connect with my viewers online. I feel
- 0:59that life is much too short and the
- 1:00world is much too large for us to miss
- 1:02the chance to connect when we can.
- 1:05Number three, if you like the video,
- 1:07please give it a thumbs up. Share it
- 1:10with classmates or colleagues or put it
- 1:12on a playlist. That does encourage me to
- 1:14keep making them for you. On the flip
- 1:17side, if you think there's something I
- 1:18can do better, please leave a
- 1:20constructive comment below the video,
- 1:21and I will take those ideas into account
- 1:24when I make new ones. And finally, just
- 1:27keep in mind that these videos are meant
- 1:29for individuals who are relatively new
- 1:31to stats. So, I'm just going over basic
- 1:34concepts, and I will be doing so in a
- 1:37slow, deliberate manner. Not only do I
- 1:40want you to know what is going on, but
- 1:43also why and how to apply it. So, all
- 1:46that being said, let's go ahead and get
- 1:49started.
- 1:51Hello and welcome to part three in our
- 1:53series on multiple regression. Now, I am
- 1:56going to assume you watched parts one
- 1:58and two of this series or you come to
- 2:00this video with that knowledge already
- 2:02in your head. So, if you need to go
- 2:05back, watch those two videos and then
- 2:07come back to this one. Otherwise, you
- 2:09may end up watching this one and become
- 2:11confused and frustrated. And neither of
- 2:13us want that to happen. So here in part
- 2:16three, we're going to talk about
- 2:17building regression models. So we will
- 2:20do that by introducing our independent
- 2:21variables into the model one at a time.
- 2:24Then we'll add them two at a time. And
- 2:26then in this case, we'll add all three.
- 2:28And then we'll see the models that come
- 2:30out of that process and then make a
- 2:32judgment call on which model is best. So
- 2:35we'll talk about all that criteria as
- 2:37far as determining the best model as we
- 2:39go forward. So all that being said,
- 2:41let's go ahead and get started.
- 2:46So as we discussed in parts one and two,
- 2:48conducting a multiple regression
- 2:50analysis requires a fair amount of
- 2:52pre-work before actually running the
- 2:54numbers in your software. So here are
- 2:56those steps real quickly. We generate a
- 2:59list of potential variables. So
- 3:01independence and the dependent variable.
- 3:03We collect data on those variables. We
- 3:06check the relationships between each
- 3:07independent variable and the dependent
- 3:09variable using scatter plots and
- 3:11correlations. So here we want to see
- 3:14which independent variables are actually
- 3:16related to the dependent variable in the
- 3:19first place because the ones that are
- 3:21not are probably going to be left out of
- 3:23the analysis.
- 3:25Then we'll check the relationships among
- 3:27the independent variables themselves
- 3:30using scatter plots and correlations. So
- 3:32here we're checking for multi-olinearity
- 3:36and then usually optional but I do it
- 3:38anyway and that is conduct simple linear
- 3:40regressions for each independent
- 3:42variable dependent variable pair and
- 3:45that's what we will do in this video
- 3:47because here we want to learn how to
- 3:48build models up and then tear them down
- 3:51to find the best one.
- 3:53Then we'll use the non-redundant
- 3:55independent variables in the analysis to
- 3:58find the best fitting model. And
- 4:00finally, most likely in a later video,
- 4:02we'll use the best fitting model to make
- 4:04predictions about the dependent
- 4:06variable. And we'll talk about
- 4:08confidence intervals and prediction
- 4:09intervals at that point.
- 4:14So remember, you are a small business
- 4:16owner that runs a delivery service like
- 4:18a courier service. You do same day
- 4:20deliveries of packages, letters, small
- 4:22cargo, etc. And you want to be able to
- 4:25predict the total travel time for each
- 4:29trip. So on a trip, you may have several
- 4:31deliveries. They may be close to your
- 4:33office, they may be far away, etc. So
- 4:36you go back and look at 10 past trips
- 4:38and record four pieces of information.
- 4:40The total miles traveled for that trip,
- 4:42the number of deliveries during that
- 4:44trip, the daily gas price, and finally
- 4:47the total travel time in hours, which is
- 4:50your dependent variable or the thing
- 4:51you're interested in predicting. So
- 4:54here's the data. We have miles traveled,
- 4:55that's our first independent variable,
- 4:57X1. number of deliveries during the
- 4:59trip, that's x2. And the gas price that
- 5:02day, that's x3. And then our dependent
- 5:05variable, travel time, is y there on the
- 5:07right.
- 5:11So remember, it's always a good idea to
- 5:13visualize the relationships. So we have
- 5:15our travel time dependent variable. Then
- 5:17we have three independent variables. So
- 5:19we have miles traveled, number of
- 5:22deliveries,
- 5:24and gas price. Now, of course, all three
- 5:26of those have some relationship to the
- 5:28dependent variable. We don't know that
- 5:29yet, but we will. But we also have to
- 5:31account for the relationships among the
- 5:33independent variables themselves, which
- 5:35are there in the dash line. So, we have
- 5:38six relationships we have to analyze and
- 5:40keep in mind as we do our analysis.
- 5:44Let's quickly review the scatter plots
- 5:46comparing the dependent variable and the
- 5:48independent variables individually. So
- 5:50the first scatter plot we have our
- 5:52travel time dependent variable versus
- 5:55our miles traveled or first independent
- 5:57variable and as you can see there we
- 5:59have a strong linear relationship starts
- 6:01at the lower left of the graph and goes
- 6:02to the upper right we have a correlation
- 6:04coefficient of 928 that's very very high
- 6:08we have a p value for the correlation of
- 6:090 which means it's less than 0.1 so we
- 6:14know that that variable miles traveled
- 6:17is related strongly related to our
- 6:19dependent variable able travel time. So
- 6:21we'll put a green check there.
- 6:24The second independent variable number
- 6:25of deliveries X2 that's also very
- 6:28strongly related. So the line starts in
- 6:30the lower left goes to the upper right.
- 6:32They fall along a rough line there. The
- 6:35correlation is 916. Again very high. The
- 6:38p value is less than 0001. So that is
- 6:41significant. So we'll put a green check
- 6:43there. So our first two independent
- 6:45variables do have strong linear
- 6:48relationships. very strong correlations
- 6:50with our dependent variable, which is
- 6:52good.
- 6:54Now, our third independent variable, gas
- 6:56price, does not. As you can see, the
- 6:59data points do not form any pattern.
- 7:01They're kind of all over the place. And
- 7:03we can see that in the correlation. It's
- 7:06267 with a p value of 0455.
- 7:10That is not significant. So, we'll put a
- 7:12red X there. So gas price X3 does not
- 7:17have any sort of linear relationship to
- 7:20the dependent variable right off the
- 7:22bat. So usually we would just go ahead
- 7:25and remove that from the model because
- 7:27if it doesn't have one to begin with,
- 7:29it's not going to contribute anything to
- 7:31the regression. But of course for now
- 7:34we're going to leave it in and see how
- 7:35it affects our numbers as we go forward.
- 7:40So now we have the scatter plots for the
- 7:42independent variable comparisons. So
- 7:44again, we're looking for
- 7:44multi-olinearity.
- 7:46So the first scatter plot, we have miles
- 7:48traveled versus number of deliveries.
- 7:51And here we have a problem. So our first
- 7:54two independent variables have a very
- 7:57very high correlation, a 0.956, a p
- 8:01value less than 0001. And I'll put a
- 8:04skull and crossbones there cuz this is a
- 8:06problem. two independent variables that
- 8:09are this highly correlated are going to
- 8:11cause some serious issues with our
- 8:13regression coefficients going forward.
- 8:16So, we'll keep that in mind as we go.
- 8:17I'll leave them in there for now so we
- 8:19can see how it affects things going
- 8:20forward.
- 8:22Now, miles traveled X1, gas price X3, no
- 8:25problems there. And then number of
- 8:28deliveries X2 and gas price X3, no
- 8:31problems there. So the only problem we
- 8:33have to worry about as far as
- 8:34multiolinearity goes is the correlation
- 8:37between the first two independent
- 8:38variables.
- 8:42So the quick summary. So the correlation
- 8:44analysis confirms the conclusions we
- 8:46reached by visual examination of the
- 8:48scatter plots. So we have some redundant
- 8:51multi-colinear variables. So miles
- 8:53traveled and number of deliveries are
- 8:55both highly correlated with each other
- 8:57and therefore are redundant. only one
- 9:00should be used in the regression
- 9:01analysis in the end. And we'll see how
- 9:04that all pans out as we build our model.
- 9:06We also have a non-contributing
- 9:08variable. So gas price is not correlated
- 9:11with the dependent variable at all and
- 9:14should probably be excluded right off
- 9:15the bat. Now again, for educational
- 9:19purposes, I'm going to leave all three
- 9:21variables in so we can see how putting
- 9:24them in there affects the regressions we
- 9:27do. But in the end, it'll all work out.
- 9:29You'll see.
- 9:32Okay. So, let's go ahead and get into
- 9:33some of our single variable regressions.
- 9:38Now, in this first step, we will perform
- 9:39a simple regression for each independent
- 9:42variable individually. The first will be
- 9:44conducted in Excel and then the rest in
- 9:46Mini Tab. That's what I prefer right
- 9:48now, but SPSS, SAS, Jump, R, etc. um or
- 9:52offline as well. You can get them all to
- 9:54generate basically the same output. So
- 9:57depending on which one you use, you
- 9:58should be fine. But I'll be focusing on
- 10:00a little bit of Excel and the rest
- 10:02MiniAB.
- 10:04We will discuss interpretations of the
- 10:06results we get from Mini Tab. And we
- 10:09will note how our results change. So
- 10:12we'll look at the coefficients in the
- 10:14regression. We'll look at the
- 10:16coefficient values. We'll look at their
- 10:19t statistics and we'll look at their p
- 10:21values.
- 10:22We'll look at the ANOVA table in the
- 10:24regression. So we'll look at the F value
- 10:27and the P value in that ANOVA table.
- 10:31We'll look and talk about the R 2, the R
- 10:342 adjusted and the R squar predicted and
- 10:37talk about what those mean. We'll also
- 10:39look at something called the VIF or the
- 10:41variance inflation factor and that's a
- 10:43statistic that MiniAB produces that will
- 10:45help us weed out multiolinearity.
- 10:49And we will also talk about something
- 10:51called malocp. That is a statistic that
- 10:54many tab outputs that will help us pick
- 10:56the best model in the end.
- 11:00So let's go ahead and look at our first
- 11:02regression. So we're going to regress
- 11:04travel time y which is our dependent
- 11:06variable on miles traveled which is our
- 11:08first independent variable. And here are
- 11:10our results from Excel. So we basically
- 11:12have three tables here. The first one in
- 11:14the top left are our regression
- 11:16statistics. Now in this case, multiple R
- 11:19is the same thing as our correlation.
- 11:21Since we only have one independent
- 11:23variable, they're the same thing. So
- 11:250.928, etc. is the same as the
- 11:28correlation we had a couple slides ago.
- 11:30Now R square is the proportion or
- 11:34percentage of variation in the dependent
- 11:37variable accounted for by the
- 11:39independent variable. So we can look at
- 11:42it as a percentage. So 86.15%
- 11:46of the variation in the dependent
- 11:48variable is accounted for by the
- 11:50independent variable. That's pretty
- 11:52high. Now the adjusted R square, that's
- 11:55the same thing as the R square.
- 11:57Obviously, it is just adjusted for the
- 11:59number of independent variables in our
- 12:01model, which in this case is one. So it
- 12:04will always be lower than the R squared.
- 12:07And how much lower really depends on the
- 12:09specific circumstances we're in. So the
- 12:12next number is the standard error of the
- 12:14regression. This is one of my favorite
- 12:16numbers, but unfortunately most people
- 12:19don't know how to use it or don't use it
- 12:21or just skip it or whatever else, but I
- 12:23think it's very helpful. So the standard
- 12:26error of the regression is the average
- 12:29distance of the data points from the
- 12:32regression line in dependent variable
- 12:35units. What we're saying here is that
- 12:37the data points are on average 342
- 12:43hours away from the regression line. It
- 12:46is in the units of the dependent
- 12:48variable and it gives us a measure of
- 12:51how tightly around the regression line
- 12:55our data points are. So it kind of forms
- 12:58a a channel or a band around the
- 13:03regression line. And the narrower that
- 13:05is, the more tightly our data points are
- 13:08around the regression. And the wider
- 13:10that band is, the more scattered they
- 13:13are from that regression line. So the
- 13:15standard error of the regression tells
- 13:17us relatively speaking how wide that
- 13:20band around the regression line is. And
- 13:23it's also helpful because it is in the
- 13:25units of the dependent variable. In this
- 13:27case, hours. And of course, we have 10
- 13:29observations. That's pretty
- 13:30self-explanatory. Now the ANOVA table
- 13:33that gives us the significance of the
- 13:35overall model. So we have an F statistic
- 13:38there of 49.768
- 13:40etc. with a p value of 0.00001
- 13:45that of course is significant. So the
- 13:48overall model here is significant. Now
- 13:51at the bottom we have some of our
- 13:53coefficient information. Now we're
- 13:55interested in the miles traveled
- 13:57coefficient. So under coefficients we
- 14:00have 0.0402
- 14:03that is the coefficient of our miles
- 14:05traveled and again that is in hours. So
- 14:08what we're saying there is that for
- 14:10every mile that's increased the time
- 14:14traveled increases by 042
- 14:18hours. Then we have the p value which is
- 14:210.1. So we know it is significant. Now,
- 14:24if you notice in this case, the p value
- 14:26for the miles traveled coefficient is
- 14:29the same as the p value for the ANOVA.
- 14:32And that's because we only have one
- 14:34independent variable. So, how can we use
- 14:36this information? Well, we can take our
- 14:39coefficient information at the bottom
- 14:40and generate our regression equation.
- 14:43So, we have an intercept of 3.1855, and
- 14:45again, that's rounded up here at the
- 14:47top. Plus 00403.
- 14:50That is our coefficient for miles
- 14:51traveled down here at the bottom. and
- 14:53then times the miles traveled. That's
- 14:55our independent variable. That's our x1.
- 14:58So we have 3.1856
- 15:00plus 0403
- 15:03x1 where x1 is miles traveled. So what
- 15:06does that mean? An increase in 1 mile.
- 15:10So miles traveled x1. So one mile will
- 15:13increase delivery time by 043
- 15:18hours. And that's how we can interpret
- 15:20the coefficient in this simple
- 15:22regression. Now, let's go ahead and make
- 15:24a rough prediction for miles traveled
- 15:27that's sort of within the range of our
- 15:29original data. So, we'll pick an 84 mile
- 15:32trip estimate. So, we go ahead and
- 15:35substitute 84 in for our x1. That gives
- 15:38us 6.5708
- 15:41hours. So, that's a very rough estimate
- 15:44of how long it would take for an 84 mile
- 15:47trip. So 6 hours and 34 minutes.
- 15:51But remember this is just an estimate.
- 15:53So it's going to have an interval around
- 15:55it. It's going to have some error around
- 15:57it. Now we can go ahead and find that
- 15:59prediction interval using some things we
- 16:02already know. And again what I've done
- 16:04here is a very rough sort of estimate.
- 16:07Mini tab can give us exact numbers but a
- 16:10rough estimate of the prediction
- 16:11interval for an 84 mile trip. So we have
- 16:146.5708
- 16:16plus or minus 2.31.
- 16:19Well, where do I get 2.31?
- 16:22That comes from our t distribution, our
- 16:24t table. So in this example, we have n
- 16:28minus 2 degrees of freedom. So that's 10
- 16:31minus 2 in this case because we have 10
- 16:33observations. So we go to our t table.
- 16:36We look at degrees of freedom of eight.
- 16:39Then we look down the table for a 95%
- 16:42interval and we have a critical t of
- 16:452.31. That's where that comes from. And
- 16:48then we have 3423.
- 16:51Where does that come from? Well, that's
- 16:54my magic number over here on the left.
- 16:56The standard error of the regression. So
- 16:58this is very much like any other
- 17:00interval we calculated back in interval
- 17:03estimation. So we have a point
- 17:05estimator. So 6.5708
- 17:08plus or minus the t alpha / 2 which is
- 17:122.31
- 17:14times the error which in this case
- 17:16is.3423.
- 17:19I don't expect you to sort of get this
- 17:20right now but I just want to show you
- 17:21sort of how we use what we have here. So
- 17:24that creates an interval of 5.7764
- 17:28to 7.3615
- 17:30hours or 5 hours 47 minutes to 7 hours
- 17:36and 22 minutes. That's our 95%
- 17:39prediction interval for an 84 mile trip.
- 17:43So again, that's sort of a step forward
- 17:46what we're going to do in future videos,
- 17:48but I just wanted to quickly show you
- 17:50how we can use the information we get
- 17:53from a regression to make some rough
- 17:55predictions for other values.
- 17:59So here is our second one toone
- 18:01regression. So we have our second
- 18:03independent variable number of
- 18:05deliveries and our dependent variable
- 18:07travel time. And this comes from Mini
- 18:09Tab. Now obviously it looks a little bit
- 18:11different. Now, I will say that one of
- 18:14the best skills to have if you're doing
- 18:16statistics or whatever else is to be
- 18:18able to use really any software package
- 18:21or at least look at the results of any
- 18:24software package and know how they
- 18:27correspond to each other. So, in this
- 18:29case, the F value of 41.96
- 18:34that's along the regression line is the
- 18:38same F value we had in Excel in the
- 18:40previous slide. So we look at the
- 18:42regression line going across. We have an
- 18:45F value of 41.96
- 18:48and a P value of 0. So we know that's
- 18:51significant. Now at the bottom we have
- 18:54the model summary. So S is the same
- 18:57thing as the standard error of the
- 18:59regression that we had in Excel. So here
- 19:02it's 368091.
- 19:04Then again we have R 2 we have R 2
- 19:08adjusted
- 19:09and then we have in this case R 2
- 19:13predicted. So this is an addition that
- 19:16MiniAB gives us and R 2 predicted is
- 19:21basically how well our model does at
- 19:26predicting
- 19:27additional data points. So we'll talk
- 19:29about that more as we go. But R squ is
- 19:32really about predictive power.
- 19:35So here are our coefficients. So again,
- 19:37this looks very similar to what we had
- 19:39in Excel. So we'll look at the number of
- 19:41deliveries uh row. So we have a
- 19:44coefficient of 4983. We have a t value
- 19:47of 6.48. Its p value is 0. So we know it
- 19:52is also significant. The vif we'll talk
- 19:55about later. And then the regression
- 19:57equation, which is really nice about
- 19:59many tab and other software packages, it
- 20:01actually gives us the regression
- 20:02equation. So we have 4.845.
- 20:05So that comes from the constant
- 20:07coefficient up there at the top.
- 20:09Plus4983
- 20:11that comes from the number of deliveries
- 20:13coefficient up there. And then we
- 20:15multiply that by the number of
- 20:16deliveries we actually have. So an
- 20:19increase in one delivery, one additional
- 20:22number of deliveries will increase
- 20:24delivery time by 4983 hours or almost a
- 20:30half an hour. Again, that's how we
- 20:32interpret a simple regression
- 20:34coefficient.
- 20:37So, let's go ahead and make a rough for
- 20:39delivery estimate. Now, we're not going
- 20:40to do the prediction interval again.
- 20:42We'll just do the four delivery
- 20:43estimate. So, let's say a trip has four
- 20:45deliveries. So, we can go ahead and
- 20:47substitute the four into our regression
- 20:49equation and we come up with an estimate
- 20:52of 6.838
- 20:54hours or 6 hours and 50 minutes. So if
- 20:59we come into work one day and we have a
- 21:01trip with four deliveries on it. So
- 21:04based on our data, the best estimate we
- 21:06have is a trip of 6 hours and 50
- 21:09minutes.
- 21:12So finally, here's our last single
- 21:14regression. So gas price X3 and travel
- 21:17time Y. So we go across the regression
- 21:20line here. We have an F value of
- 21:2462.
- 21:26That's very low. Then we have a p value
- 21:29of 0.455.
- 21:31That is not significant. That's
- 21:33obviously not below 0.05. So that
- 21:36basically confirms what we looked at in
- 21:38the scatter plots and correlations. And
- 21:40that is that gas price really has no
- 21:43relationship or no linear relationship
- 21:45to travel time. Now if we look down at
- 21:48the model summary, we have a standard
- 21:50error of regression of86
- 21:53hours. That is a huge standard error of
- 21:56the regression. We have an R squar
- 21:58that's 7.14%.
- 22:00You know why even bother? The R squ
- 22:02adjusted is literally zero and the R
- 22:04square predicted is literally zero. So
- 22:07we know that this variable is really a
- 22:10dud.
- 22:11So here are our coefficients. We have
- 22:14the gas price coefficient of 81. We have
- 22:18a T value of 78 very low obviously P
- 22:22value of 0455. Same thing as we have in
- 22:24the f value above the p value up there.
- 22:27We know that's not significant. The
- 22:30regression equation, why even bother?
- 22:32This is really not worth exploring any
- 22:34further. We know that gas price is
- 22:36really not a variable that contributes
- 22:38to travel time.
- 22:41Let's go ahead and summarize these three
- 22:43models. So in the first line, we have
- 22:45the case where we're only using our
- 22:47first independent variable. We had an F
- 22:50of 49.77,
- 22:52a p value that's very low, significant,
- 22:54less than 0001. We had a standard error
- 22:57of the regression of 34230.
- 23:01We had an R squar of 84.42
- 23:04and an R squar predicted of 79.07.
- 23:08Now I did that in many tabs, so we had
- 23:10it um but that's the actual R square
- 23:13predicted that was produced. And then in
- 23:16the second case where we had the second
- 23:18independent variable we had an f of
- 23:2041.96
- 23:22p value less than 0001.
- 23:25Now here we had a standard error of the
- 23:27regression of 36809.
- 23:31So remember what I said. What does that
- 23:33tell you? Well, in the first case in the
- 23:36um the top row, on average, our data
- 23:39points were 34230
- 23:42hours away from the regression line. In
- 23:46the second case, they were 36809 hours
- 23:50away from the regression line. So, which
- 23:53model has data points that are more
- 23:55clustered in around the regression line?
- 24:00Well, the first one does because it
- 24:01standard error of the regression is
- 24:03lower. Now, the R squ adjusted for the
- 24:05second one is 81.99.
- 24:08It's a little bit lower than the one at
- 24:09the top. And then the R square predicted
- 24:12is 70.27, which is quite a bit lower
- 24:15than the one at the top. And then
- 24:17finally, we sort of have our our dud,
- 24:20our no good one, where we're using X3.
- 24:22I'm not even going to go through that
- 24:24anymore because it's obviously um no
- 24:26good. So, if we look at this example
- 24:28where we're only using one independent
- 24:31variable, which one do you think is the
- 24:34best?
- 24:35Well, I would have to say that the first
- 24:36one's the best. We have the highest
- 24:39F-stistic. We have the lowest standard
- 24:42error of the regression. We have the
- 24:44highest R 2 adjusted. And we definitely
- 24:46have the highest R square predicted. So
- 24:49of these three, the first one is
- 24:52definitely I would say the best.
- 24:57So let's do the case where we have two
- 24:59variables. So now we're going to put two
- 25:01variables in at a time. So we'll put X1
- 25:04and X2 into the regression. We'll put X1
- 25:06and X3 into the regression. And we'll
- 25:08put X2 and X3 into the regression. So
- 25:11this is a subset of two variables but
- 25:13the interpretation is basically the
About this transcript
This page contains the full transcript of YouTube transcript (wPJ1_Z8b0wk) , generated from the public captions YouTube serves with the video. The transcript has 4,041 words across 607 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.