Paste This Into Claude, Never Hit a Token Limit Again — Transcript
Full transcript
- 0:00I'm going to break down how you can set
- 0:01up Claude's so you never hit a token
- 0:02limit again. The process is broken down
- 0:04into three parts that you can copy.
- 0:05Quick wins, system upgrades, and nuclear
- 0:07enhancements. But before we get to those
- 0:09fixes, you need to understand how AI
- 0:11token consumption works because a lot of
- 0:12people get this wrong. There are three
- 0:14key terms to understand: tokens, the
- 0:16model you're using, and your compute
- 0:17budget. A token is essentially how much
- 0:19text the model has to process. An AI
- 0:21model is what impacts how much compute
- 0:23is used per token. Better models require
- 0:25more compute per token. And your compute
- 0:28budget is how much compute is budgeted
- 0:30to your account. So Claude Pro, Max,
- 0:32these subscriptions all have an
- 0:34allocated budget you can draw down from.
- 0:36When you run out of tokens or hit a
- 0:38limit, that really has to do with the
- 0:39total compute associated with your
- 0:41account, not necessarily the amount of
- 0:42tokens that you consume. So with this
- 0:44foundational understanding, if we don't
- 0:46want to spend more money to increase our
- 0:47compute budget, there are two variables
- 0:49you can play with: the tokens we use and
- 0:51the model we use. A formula I'm going to
- 0:53be referencing throughout this video is
- 0:55compute budget used equals tokens
- 0:57consumed times model used. Every fix we
- 0:59cover in this video will optimize one of
- 1:01these two variables. So the first part,
- 1:03part one, which is quick wins, is all
- 1:05about token consumption. To identify
- 1:07some of these simple quick wins, the
- 1:08best thing to do is to audit your
- 1:10system. At the end of the day, you can't
- 1:11solve a problem if you're not sure
- 1:12what's causing it. So if you open Claude
- 1:14code and you type {slash} usage, you'll
- 1:16see a breakdown of how many tokens
- 1:18you've used, plus a section called
- 1:19what's using your limits, which is
- 1:21essentially where your tokens are going.
- 1:23This is Claude's way of helping you
- 1:24identify your problems. Now for me and
- 1:26for a lot of you, you'll likely see two
- 1:28different issues that jump out. The
- 1:29first for me is that 65% of my usage ran
- 1:32above 150K context. And number two is
- 1:35that 58% came from sub agent heavy
- 1:38usage. And we'll cover sub agents later,
- 1:39but for this part, we'll focus on
- 1:41optimizing our context. Context is
- 1:44essentially all of the additional
- 1:45information you provide to Claude to get
- 1:47a better response. The more context, the
- 1:50more tokens you use, the quicker you
- 1:52reach your limit. So to fix this, you
- 1:53need to set some new habits. So the
- 1:56first quick win is fix your contextual
- 1:58habits. As conversations progress, your
- 2:00context window will quickly fill up. To
- 2:02show that, on the left, I type {slash}
- 2:04context in a new chat. And then on the
- 2:06right is my best Bart Simpson
- 2:07impersonation where I sent a long
- 2:09message which had nothing but adding
- 2:11context a bunch of different times. And
- 2:12you can see just with this one message
- 2:14that my context more than doubled, which
- 2:16means that my token usage essentially
- 2:18doubled. Now, this is with one message,
- 2:19you can imagine how bad this gets over
- 2:21time. Some good habits to maintain your
- 2:23context is first, whenever you're
- 2:25switching tasks, run {slash} clear or
- 2:27start a new chat. The second is try and
- 2:29consistently work on projects versus
- 2:31coming back after 5 minutes. So, if
- 2:33you're waiting 10 minutes before
- 2:34responding, you don't get that caching
- 2:36benefit. The third is that you can
- 2:38actually adjust your effort. So, no
- 2:40matter what model you use, there is a
- 2:41default effort you can set. Simply put,
- 2:44the higher the effort, the more compute
- 2:45required per task. To adjust yours,
- 2:47click the bottom right corner on Claude
- 2:49Desktop, and then it may say low,
- 2:50medium, high, and then you can drag it
- 2:52to whatever effort you want. The fourth
- 2:54is that mid habit, as your context
- 2:55window fills up, let's say 60%, and you
- 2:58can see that in the bottom right corner,
- 2:59the circular icon, you type {slash}
- 3:01compact, which will compress everything
- 3:02into a short summary to help you manage
- 3:05your context. And if you are using
- 3:06Claude code directly in terminal, write
- 3:08{slash} status line so that you can see
- 3:10your context directly on screen. So,
- 3:12that first quick win is just improving
- 3:14your habits with all the strategies I
- 3:15just mentioned. The second is you need
- 3:16to run a contextual cleanup. If you
- 3:18start a new conversation and then type
- 3:20{slash} context in a completely fresh
- 3:22chat. This will show you everything that
- 3:24gets preloaded automatically to every
- 3:26conversation you have with Claude. So,
- 3:28for me, you'll see that I've already
- 3:29filled up 58.8k of my context window
- 3:32before typing a single word. The higher
- 3:34this number is, the higher your minimum
- 3:36token consumption will be for every new
- 3:38conversation. And so, we can actually
- 3:40clean this up so that we're not starting
- 3:41from such a high token count. Here's a
- 3:43prompt to help you do it, but going
- 3:44through it, the first part of the prompt
- 3:46will review any of your unused MCPs and
- 3:49delete them. If you want to do this
- 3:50manually, you can type {slash} MCP, and
- 3:52you'll see a full list of what you
- 3:53currently have set up. Next, it will
- 3:54clean up your skills. Every time you
- 3:56start a new session, your skills and
- 3:58their descriptions get loaded into
- 4:00context. So, if you have unused skills,
- 4:02just delete them. Or if your skill
- 4:04descriptions are extremely long, just
- 4:05shorten them. This part of the prompt
- 4:07will help you do exactly that. And then
- 4:08third, this will revise your Claude MD
- 4:10file. This file gets reread every
- 4:12message for the entire conversation. So,
- 4:14if you have a 4,000 token Claude MD, you
- 4:16start every conversation 4,000 tokens
- 4:19deep. The best practice in terms how to
- 4:21structure your Claude MD is you want it
- 4:22to tell Claude how to interact with your
- 4:24specific project. You don't necessarily
- 4:26want it to have extensive documentation.
- 4:28And a rule of thumb from Anthropic Docs
- 4:30directly is that you want to keep your
- 4:31Claude MD under 200 lines. Anything
- 4:33longer than that and you're paying a tax
- 4:35on every single message. So, that'll
- 4:37clean up your context you're starting at
- 4:38a lower point. The third quick win is
- 4:40reduce your output tokens. The more
- 4:43words that Claude says to you, the more
- 4:44tokens you spend, right? These are the
- 4:46response, these are the output tokens.
- 4:48And this is generally a relatively small
- 4:50amount of your overall consumption, but
- 4:52this is a quick win that I absolutely
- 4:54love. Just paste this or add this to
- 4:55your Claude MD. Be concise with all of
- 4:57your responses. Or you can go a step
- 4:59further and install a plugin that I love
- 5:00called Caveman. It makes Claude speak
- 5:02like a caveman, and so you can actually
- 5:04see on screen the before Caveman and
- 5:07after Caveman response. Those three
- 5:09quick wins are all foundational things
- 5:11that you need to do. But the next
- 5:12section are system upgrades that can
- 5:15make your system up to 60 to 90% more
- 5:17efficient. But before we get to that,
- 5:19one thing I'll spend a lot of tokens on
- 5:21unnecessarily is trying to make decks
- 5:23for my clients, which brings us to
- 5:24today's video sponsor Bolt, and
- 5:26specifically their new product Bolt
- 5:28Slides, which has changed how people
- 5:30make decks going forward. To use it,
- 5:31just go to bolt.new, click slide deck,
- 5:33and type something like build me a
- 5:35five-page deck about AI automation I can
- 5:37build for a client. Builds a real
- 5:39responsive web app styled and structured
- 5:41in seconds. Or you can connect it
- 5:43directly to Claude code and have it
- 5:44generate a deck for you without ever
- 5:46leaving the terminal. And so yes, Bolt
- 5:48can make a deck very quickly, but that's
- 5:50not why I like it. You guys know how
- 5:51much I stress the importance of using AI
- 5:53to improve the quality of your outputs,
- 5:55not just do things faster. And Bolt
- 5:57Slides enhances the deck creation
- 5:59process in three ways. First is you can
- 6:01embed whatever you want into your
- 6:03slides. For example, a live chart with
- 6:05updating data or a clickable diagram.
- 6:07Honestly, how bangers would it be if you
- 6:08had a product showing real-time usage
- 6:10and it's counting as you present. Now,
- 6:12this really unlocks a world of
- 6:14possibilities. Now, the second way is it
- 6:16looks native everywhere. So whether
- 6:17you're on a projector, a phone, a
- 6:19tablet, it doesn't matter because these
- 6:21slides are a native web app. One of the
- 6:22things that I personally hated when
- 6:24pitching brands like Google and Amazon
- 6:25at my last startup was whenever I shared
- 6:27a deck, I'd have to say, "Make sure to
- 6:28open this on desktop." And that was just
- 6:31annoying because if they viewed it on
- 6:32their phone, it would look terrible. But
- 6:34now with Bolt Slides, you can view it
- 6:35wherever, however. And the third is
- 6:37speed of iteration allows you to
- 6:38visualize decks as part of the creative
- 6:41process. Typically, when I'd make a
- 6:42sales or an investor deck, it started
- 6:44with a bullet point list. And then I
- 6:46take that bullet point list and then
- 6:47visualize it. And at that point, I'd
- 6:49often realize the concept didn't
- 6:51actually work. But by creating a
- 6:52visualization of these concepts earlier,
- 6:54you're able to streamline the entire
- 6:56process. Now, if you want access to
- 6:57Bolt's new feature, click the link below
- 6:59and differentiate yourself with the next
- 7:01generation of slide decks. Part two,
- 7:03system upgrade. So the first upgrade to
- 7:05your system is compress inputs
- 7:07>> [music]
- 7:07>> before AI sees them. So every piece of
- 7:09text that you provide Claude consumes
- 7:11tokens. So in an ideal world, you would
- 7:13compress the text so it's only sharing
- 7:15with Claude the text that actually
- 7:17matters. To help you visualize this,
- 7:19imagine you're working on a report for a
- 7:20client and you want Claude to review
- 7:22specific changes.
- 7:23>> [music]
- 7:23>> If you only made changes on the first
- 7:24page, would it make sense to share the
- 7:26entire 10-page report to Claude? No, it
- 7:28wouldn't, but Claude does this by
- 7:30default. So what you could do instead is
- 7:32use Claude hooks to pre-process these
- 7:34files to be more efficient with tokens.
- 7:36[music] This is pulled directly from
- 7:38Anthropic docs and it tells you to do
- 7:39exactly this. It says, "Offload
- 7:41processes to hooks and skills." And then
- 7:43it gives an example of reading a file
- 7:45and only pulling the information that
- 7:46has a prefix text error. Luckily, we
- 7:49don't have to figure out how to do this
- 7:50because there are some gigabrains who
- 7:52have helped us for free. There's a free
- 7:53open-source tool called RTK that will
- 7:56help make the text that's shared with
- 7:58Claude significantly more concise. To
- 8:00visualize where it sits, normally Claude
- 8:02will read the report, it'll reread the
- 8:04changes to check the work, and then the
- 8:06full raw output is dumped into Claude.
- 8:08And so let's say this process takes
- 8:10about 15,000 tokens. After RTK, Claude
- 8:13will edit the report, it'll reread the
- 8:15change to check work. RTK then uses
- 8:17deterministic logic to clean up the
- 8:19text, removes repeated text,
- 8:20boilerplate, formatting noise,
- 8:22compresses the text, and then shares
- 8:24that with Claude. And that would end up
- 8:25being about 1800 tokens. It's quite
- 8:28genius, and it works for both technical
- 8:30and non-technical tasks. I ran some
- 8:32tests on my computer across 13 commands,
- 8:34and I was able to save 92% of my tokens.
- 8:36Now, RTK claims 60 to 90% and the exact
- 8:40savings depend really on what you're
- 8:41doing. But to set this up, you literally
- 8:43can just paste their repo and then say
- 8:45set up RTK on my project, so it runs
- 8:47automatically in the background. Hit
- 8:48enter, and then it'll go through the
- 8:49process. The second system upgrade is
- 8:52leverage sub-agents with reduced models.
- 8:54Let's go back to the equation we had
- 8:56earlier. Compute budget used equals
- 8:58tokens consumed times model used. Thus
- 9:01far, we focused on tokens consumed. But
- 9:03what about the model that we actually
- 9:04use? The reality is if you can get a
- 9:06task done with Haiku instead of Fable,
- 9:08you save 90% of your compute budget. So
- 9:10the general rule of thumb in your brain
- 9:12is if AI could solve this task a year
- 9:14ago, you don't need a frontier model to
- 9:16solve it. Scraping, summarizing,
- 9:18formatting, fetching files, none of that
- 9:20needs the frontier models that are
- 9:22getting released. This framework is what
- 9:23I call using the minimum viable model.
- 9:26So how can we do this in practice
- 9:27without over-engineering our system
- 9:29pulling our hair out? What we do is we
- 9:31predefine the model we use inside Claude
- 9:33skills. Skills are predefined tasks that
- 9:35you use over and over again. So once you
- 9:37define the minimum viable model for that
- 9:39skill, it will use that model for all
- 9:41future runs. And when setting it up,
- 9:43there are two variables you can play
- 9:45with, context and model. Model, you
- 9:47specify the model that you want to use
- 9:49for this specific skill, and then the
- 9:50context, if you don't need to have
- 9:52contextual information, you can say
- 9:53context [music] fork, which will create
- 9:55a new thread for that skill to work on,
- 9:57in turn reducing the contextual
- 9:59information that's needed. If you do
- 10:01need previous info, you'll just leave
- 10:02that part blank in the skill itself.
- 10:05Here's a table on how to optimize skills
- 10:06based on specific needs. If you're still
- 10:08struggling to visualize why to do this,
- 10:11think of it like a team of lawyers. You
- 10:12want the partner at the law firm leading
- 10:14the case, but because his hourly is so
- 10:16expensive, you want junior lawyers doing
- 10:18a lot of the simple grunt work. It's the
- 10:20same idea here, and to find where this
- 10:22applies in your own setup, here's a
- 10:24prompt. Now, that's just one of the ways
- 10:25to upgrade your skills to be more
- 10:26efficient, but the next upgrade is one
- 10:29of my favorites. Upgrade number three,
- 10:30move your workflows into script-driven
- 10:32skills. After you've updated your skills
- 10:34to use a specific model, there is one
- 10:36more step. This one improves quality,
- 10:38consistency, and speed, while also
- 10:40reducing token consumption. So, a lot
- 10:42like RTK enhancement that uses computer
- 10:45logic to compress text, you can use
- 10:47computer logic or scripts to complete
- 10:49tasks that you're currently using AI
- 10:50for. So, a script will run the same way
- 10:52every time, call zero tokens, and never
- 10:54hallucinates. So, in an ideal world, we
- 10:56want to use AI for judgment, and then
- 10:58scripts for everything else that's
- 10:59repeatable. This way, AI doesn't have to
- 11:01refigure out something every time you do
- 11:02it. Here's a prompt that enhances your
- 11:04skills to leverage scripts and flags any
- 11:07skill your project is missing. Now,
- 11:08before we get to part three, which walks
- 11:10through nuclear enhancements, if this is
- 11:12your first video of mine, welcome to the
- 11:14channel. But, if this is your second or
- 11:15more, you know the drill, this is our
- 11:17anti-slop agreement. The visuals, the
- 11:19testing, the hours of research that went
- 11:21into this video, this is entirely built
- 11:22for humans, not for these token
- 11:24gobblers, okay? So, all that I ask is
- 11:27that you subscribe as part of this
- 11:28agreement, to help this content reach
- 11:29more people so that I can just keep
- 11:31doing this. Also, every video I give a
- 11:33Claude Mac subscription away. This
- 11:34video's winner is Chad N5N5X
- 11:37who's building a complete AI operating
- 11:40system. Shout out Chad, you are a
- 11:42legend. Now, if you want to enter the
- 11:43next giveaway, comment below with what
- 11:45you're building or any recent issues
- 11:47you've run into. And every video you
- 11:48comment on is another entry. Part three,
- 11:50nuclear enhancements. These are four big
- 11:53changes listed in the order that you
- 11:54should do them. And the last one is
- 11:55going viral right now, but I'll explain
- 11:57why I don't necessarily think it's worth
- 11:59your time. Nuclear enhancement one,
- 12:00route specific work to Codex. The
- 12:03reality is that certain models are more
- 12:05efficient than other models. But that
- 12:06also extends outside the model layer
- 12:08into the harness layer, the tool that's
- 12:10actually orchestrating using AI. So,
- 12:13when you prompt Claude code, the logic
- 12:14that interacts with the AI models is the
- 12:16harness. So, Codex ecosystem, OpenAI's
- 12:19product compared to Claude, on certain
- 12:21tasks it burns 4x less tokens than
- 12:23Claude does. And the reason for this, to
- 12:26simplify it, is that Claude is built to
- 12:28be thorough. It rereads, it verifies, it
- 12:30thinks before it acts, and you pay
- 12:32tokens for every step along the way.
- 12:34Whereas for Codex, it's built to be
- 12:35surgical. Get in, make the edit, get
- 12:37out. So, to use this to your advantage,
- 12:38you can intertwine Codex and Claude to
- 12:41get the most out of your subscriptions.
- 12:43The simplest way to do it, and my
- 12:44favorite, is there's a Claude code
- 12:45plugin for Codex. You can install that,
- 12:48then have Claude route any token-heavy
- 12:49execution task to Codex. This prompt
- 12:51will help you install that plugin, set
- 12:53it up, and then update your ClaudeMD to
- 12:55tell Claude to actually do this. Nuclear
- 12:57enhancement number two is use images
- 13:00instead of text. This one sounds fake,
- 13:02but it's so awesome that I had to
- 13:04include it. So, Claude processes images
- 13:06at a different rate than it processes
- 13:08text. So, you can convert text into
- 13:11images, [music]
- 13:11and then submit that to Claude to become
- 13:14more efficient with your tokens. Now,
- 13:15there's a lot going on there, but to
- 13:16visualize it, imagine you had a massive
- 13:18page of text with over 10,000 tokens,
- 13:21and then an image with that same text
- 13:23but in image form. In order for a Claude
- 13:25to actually process this image, it only
- 13:27takes 3,000 or 4,000 tokens, resulting
- 13:30in a 60 to 70% token reduction. These
- 13:32are based on claims from PX pipe, the
- 13:34open source tool on GitHub that will
- 13:36help you do exactly this. Now, there are
- 13:38some trade-offs like it won't read the
- 13:39text perfectly and there is a non-zero
- 13:41chance that Anthropic [music] patches
- 13:43this in the future, but I found this so
- 13:45damn interesting that I just had to
- 13:46include it. Nuclear enhancement number
- 13:48three, swap the engine out entirely.
- 13:50Most people aren't aware of this, but
- 13:51Claude code is the harness and the model
- 13:54that's inside of it and is actually used
- 13:56can be swapped. So, for example, you
- 13:58could use Claude code and only use
- 14:00OpenAI or other open source models. Now,
- 14:03to actually do this, it's pretty
- 14:04straightforward. There are some
- 14:05environment variables that Claude code
- 14:06will reference to know which model to
- 14:08use and [music] by default, it goes to
- 14:10Claude's models. Now, two potential
- 14:12options that I'm considering looking
- 14:13into are ZAI's GLM plan and then the
- 14:16Deep Seek plan. Based on research for
- 14:18these on a compute per dollar spent,
- 14:20these can provide you with more capacity
- 14:22than Claude's plans. There are
- 14:24trade-offs, right? Deep Seek is a
- 14:25Chinese model, you may be concerned with
- 14:27your data privacy and also when you're
- 14:29using these tools, you are getting
- 14:31slightly worse models. Here's a prompt
- 14:33you can use to go down this rabbit hole
- 14:34and learn a lot more and as part of
- 14:36that, it'll build you an implementation
- 14:37plan if you want to eject out of the
- 14:39Anthropic ecosystem. This prompt is also
- 14:41designed to help you identify current
- 14:43providers as offers are constantly
- 14:45changing. Now, the fourth nuclear
- 14:47enhancement is running your own model
- 14:48locally. This is an enhancement that I
- 14:50eventually will fully believe in. It's
- 14:52beautiful, right? You can run your own
- 14:54models locally and not rely on external
- 14:56data providers. So, for example, I have
- 14:58this Mac Mini, it's a corner here, you
- 15:00can't see it, but I could run a model on
- 15:02it and then route all of my requests
- 15:03directly to that instead of Claude. So,
- 15:05everything is staying within my office
- 15:07and now there are pros and cons to this.
- 15:09Some of the pros is you can essentially
- 15:11have limitless tokens. All I have to do
- 15:13is pay for the power to keep the
- 15:14hardware running. Next is that you own
- 15:16all of your data. If data privacy is a
- 15:18concern, this means nothing is exposed
- 15:20to external service providers. Third,
- 15:22[music] you own the entire AI stack. You
- 15:24aren't relying on anyone else for and
- 15:26this concept of running a model locally
- 15:27is why you will never actually be able
- 15:29to ban AI entirely. Now, the cons of
- 15:31this, first, my Mac Mini and then 99% of
- 15:34consumer hardware can't actually run any
- 15:36of the top-tier open-source models. And
- 15:38in order for you to get to those models,
- 15:40you're going to have to spend over
- 15:41$10,000. Next is that none of the
- 15:43frontier models like Fable are
- 15:45accessible to download locally, which
- 15:47means you're sacrificing quality in the
- 15:49short term. And then finally, you have
- 15:51to set up a computer server farm and as
- 15:53someone who turns his phone on airplane
- 15:54mode at night, I would not want all of
- 15:56that EMF radiation near me. Jokes aside
- 15:58with the EMF radiation, but you have to
- 16:00set up the server farm and you have to
- 16:01manage it. That's work. So, long-term, I
- 16:03do fully expect local models to be a
- 16:05thing as hardware gets better and the
- 16:07open-source models get better as well.
- 16:09And at that point, it may 100% be worth
- 16:12it, but right now, my recommendation for
- 16:14you is you can play around with local
- 16:15models. It's valuable for you to know
- 16:17that it exists, but I just wouldn't go
- 16:19all in on this. So, please don't go and
- 16:21buy expensive hardware. Maybe one day,
- 16:23just not today. So, with that being
- 16:24said, let's speed run the changes that
- 16:26you need to make today to get more out
- 16:28of your system. So, the first quick win
- 16:29is fix your contextual habits. Run
- 16:31{slash} clear every time you switch
- 16:33tasks, work in focused blocks so you can
- 16:35keep 90% cash discount, and set your
- 16:37effort to match the task. [music] And
- 16:39then use {slash} compact when you get to
- 16:4060% full on your context. The second
- 16:43quick win is run the cleanup prompt.
- 16:45Disconnect MCPs you don't use, archive
- 16:47unused skills, and shorten their
- 16:48descriptions, and turn your Claude MD
- 16:50into a directory instead of a document.
- 16:52And then make sure your Claude MD is
- 16:53less than 200 characters. The third
- 16:55[music] is cut your output tokens. Add
- 16:57be concise to your Claude MD, use the
- 16:59KMM plugin, or update your Claude MD to
- 17:01tell Claude to be concise with all their
- 17:03responses. Then for the system upgrades,
- 17:05install RTK so every command output gets
- 17:08crushed before Claude reads it. This
- 17:09could reduce your token input by 60 to
- 17:1190%. Second, enhance your skills so
- 17:14front work runs on minimum viable models
- 17:16instead of your most expensive ones. The
- 17:18third system update is turn every
- 17:20repeatable step into a script inside the
- 17:22skill. Computer code causes zero tokens
- 17:24to run. Then the nuclear enhancements.
- 17:26First, route any token-heavy execution
- 17:29to Codex with the plugin. This is where
- 17:30you could have two subscriptions with
- 17:32two different budgets and more total
- 17:33firepower. The second is use images
- 17:35instead of text. Keep an eye on this
- 17:37one. The third is that you can swap the
- 17:39engine out entirely if you want to work
- 17:40with different model providers. And the
- 17:42fourth, as I mentioned, you can run your
- 17:44own local model if you really want to.
- 17:46Now, once you apply this, you'll be able
- 17:47to build all day long instead of having
- 17:49to wait every 5 hours for your limits to
- 17:52reset. And if you like this video, you
- 17:53will love this video where I walk
- 17:55through my exact setup to leverage
- 17:57Claude's skills to build 10 times
- 17:59faster. This builds on a lot of what I
- 18:00covered in the skill optimization
- 18:02section of this video and you will love
- 18:04that. I'll see you over there. Peace.
About this transcript
This page contains the full transcript of Paste This Into Claude, Never Hit a Token Limit Again by Austin Marchese, generated from the public captions YouTube serves with the video. The transcript has 4,046 words across 586 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.