YouTube2Text

Paste This Into Claude, Never Hit a Token Limit Again — Transcript

by Austin Marchese · 4,046 words · 586 segments · language en · Watch on YouTube

Full transcript

  1. 0:00I'm going to break down how you can set
  2. 0:01up Claude's so you never hit a token
  3. 0:02limit again. The process is broken down
  4. 0:04into three parts that you can copy.
  5. 0:05Quick wins, system upgrades, and nuclear
  6. 0:07enhancements. But before we get to those
  7. 0:09fixes, you need to understand how AI
  8. 0:11token consumption works because a lot of
  9. 0:12people get this wrong. There are three
  10. 0:14key terms to understand: tokens, the
  11. 0:16model you're using, and your compute
  12. 0:17budget. A token is essentially how much
  13. 0:19text the model has to process. An AI
  14. 0:21model is what impacts how much compute
  15. 0:23is used per token. Better models require
  16. 0:25more compute per token. And your compute
  17. 0:28budget is how much compute is budgeted
  18. 0:30to your account. So Claude Pro, Max,
  19. 0:32these subscriptions all have an
  20. 0:34allocated budget you can draw down from.
  21. 0:36When you run out of tokens or hit a
  22. 0:38limit, that really has to do with the
  23. 0:39total compute associated with your
  24. 0:41account, not necessarily the amount of
  25. 0:42tokens that you consume. So with this
  26. 0:44foundational understanding, if we don't
  27. 0:46want to spend more money to increase our
  28. 0:47compute budget, there are two variables
  29. 0:49you can play with: the tokens we use and
  30. 0:51the model we use. A formula I'm going to
  31. 0:53be referencing throughout this video is
  32. 0:55compute budget used equals tokens
  33. 0:57consumed times model used. Every fix we
  34. 0:59cover in this video will optimize one of
  35. 1:01these two variables. So the first part,
  36. 1:03part one, which is quick wins, is all
  37. 1:05about token consumption. To identify
  38. 1:07some of these simple quick wins, the
  39. 1:08best thing to do is to audit your
  40. 1:10system. At the end of the day, you can't
  41. 1:11solve a problem if you're not sure
  42. 1:12what's causing it. So if you open Claude
  43. 1:14code and you type {slash} usage, you'll
  44. 1:16see a breakdown of how many tokens
  45. 1:18you've used, plus a section called
  46. 1:19what's using your limits, which is
  47. 1:21essentially where your tokens are going.
  48. 1:23This is Claude's way of helping you
  49. 1:24identify your problems. Now for me and
  50. 1:26for a lot of you, you'll likely see two
  51. 1:28different issues that jump out. The
  52. 1:29first for me is that 65% of my usage ran
  53. 1:32above 150K context. And number two is
  54. 1:35that 58% came from sub agent heavy
  55. 1:38usage. And we'll cover sub agents later,
  56. 1:39but for this part, we'll focus on
  57. 1:41optimizing our context. Context is
  58. 1:44essentially all of the additional
  59. 1:45information you provide to Claude to get
  60. 1:47a better response. The more context, the
  61. 1:50more tokens you use, the quicker you
  62. 1:52reach your limit. So to fix this, you
  63. 1:53need to set some new habits. So the
  64. 1:56first quick win is fix your contextual
  65. 1:58habits. As conversations progress, your
  66. 2:00context window will quickly fill up. To
  67. 2:02show that, on the left, I type {slash}
  68. 2:04context in a new chat. And then on the
  69. 2:06right is my best Bart Simpson
  70. 2:07impersonation where I sent a long
  71. 2:09message which had nothing but adding
  72. 2:11context a bunch of different times. And
  73. 2:12you can see just with this one message
  74. 2:14that my context more than doubled, which
  75. 2:16means that my token usage essentially
  76. 2:18doubled. Now, this is with one message,
  77. 2:19you can imagine how bad this gets over
  78. 2:21time. Some good habits to maintain your
  79. 2:23context is first, whenever you're
  80. 2:25switching tasks, run {slash} clear or
  81. 2:27start a new chat. The second is try and
  82. 2:29consistently work on projects versus
  83. 2:31coming back after 5 minutes. So, if
  84. 2:33you're waiting 10 minutes before
  85. 2:34responding, you don't get that caching
  86. 2:36benefit. The third is that you can
  87. 2:38actually adjust your effort. So, no
  88. 2:40matter what model you use, there is a
  89. 2:41default effort you can set. Simply put,
  90. 2:44the higher the effort, the more compute
  91. 2:45required per task. To adjust yours,
  92. 2:47click the bottom right corner on Claude
  93. 2:49Desktop, and then it may say low,
  94. 2:50medium, high, and then you can drag it
  95. 2:52to whatever effort you want. The fourth
  96. 2:54is that mid habit, as your context
  97. 2:55window fills up, let's say 60%, and you
  98. 2:58can see that in the bottom right corner,
  99. 2:59the circular icon, you type {slash}
  100. 3:01compact, which will compress everything
  101. 3:02into a short summary to help you manage
  102. 3:05your context. And if you are using
  103. 3:06Claude code directly in terminal, write
  104. 3:08{slash} status line so that you can see
  105. 3:10your context directly on screen. So,
  106. 3:12that first quick win is just improving
  107. 3:14your habits with all the strategies I
  108. 3:15just mentioned. The second is you need
  109. 3:16to run a contextual cleanup. If you
  110. 3:18start a new conversation and then type
  111. 3:20{slash} context in a completely fresh
  112. 3:22chat. This will show you everything that
  113. 3:24gets preloaded automatically to every
  114. 3:26conversation you have with Claude. So,
  115. 3:28for me, you'll see that I've already
  116. 3:29filled up 58.8k of my context window
  117. 3:32before typing a single word. The higher
  118. 3:34this number is, the higher your minimum
  119. 3:36token consumption will be for every new
  120. 3:38conversation. And so, we can actually
  121. 3:40clean this up so that we're not starting
  122. 3:41from such a high token count. Here's a
  123. 3:43prompt to help you do it, but going
  124. 3:44through it, the first part of the prompt
  125. 3:46will review any of your unused MCPs and
  126. 3:49delete them. If you want to do this
  127. 3:50manually, you can type {slash} MCP, and
  128. 3:52you'll see a full list of what you
  129. 3:53currently have set up. Next, it will
  130. 3:54clean up your skills. Every time you
  131. 3:56start a new session, your skills and
  132. 3:58their descriptions get loaded into
  133. 4:00context. So, if you have unused skills,
  134. 4:02just delete them. Or if your skill
  135. 4:04descriptions are extremely long, just
  136. 4:05shorten them. This part of the prompt
  137. 4:07will help you do exactly that. And then
  138. 4:08third, this will revise your Claude MD
  139. 4:10file. This file gets reread every
  140. 4:12message for the entire conversation. So,
  141. 4:14if you have a 4,000 token Claude MD, you
  142. 4:16start every conversation 4,000 tokens
  143. 4:19deep. The best practice in terms how to
  144. 4:21structure your Claude MD is you want it
  145. 4:22to tell Claude how to interact with your
  146. 4:24specific project. You don't necessarily
  147. 4:26want it to have extensive documentation.
  148. 4:28And a rule of thumb from Anthropic Docs
  149. 4:30directly is that you want to keep your
  150. 4:31Claude MD under 200 lines. Anything
  151. 4:33longer than that and you're paying a tax
  152. 4:35on every single message. So, that'll
  153. 4:37clean up your context you're starting at
  154. 4:38a lower point. The third quick win is
  155. 4:40reduce your output tokens. The more
  156. 4:43words that Claude says to you, the more
  157. 4:44tokens you spend, right? These are the
  158. 4:46response, these are the output tokens.
  159. 4:48And this is generally a relatively small
  160. 4:50amount of your overall consumption, but
  161. 4:52this is a quick win that I absolutely
  162. 4:54love. Just paste this or add this to
  163. 4:55your Claude MD. Be concise with all of
  164. 4:57your responses. Or you can go a step
  165. 4:59further and install a plugin that I love
  166. 5:00called Caveman. It makes Claude speak
  167. 5:02like a caveman, and so you can actually
  168. 5:04see on screen the before Caveman and
  169. 5:07after Caveman response. Those three
  170. 5:09quick wins are all foundational things
  171. 5:11that you need to do. But the next
  172. 5:12section are system upgrades that can
  173. 5:15make your system up to 60 to 90% more
  174. 5:17efficient. But before we get to that,
  175. 5:19one thing I'll spend a lot of tokens on
  176. 5:21unnecessarily is trying to make decks
  177. 5:23for my clients, which brings us to
  178. 5:24today's video sponsor Bolt, and
  179. 5:26specifically their new product Bolt
  180. 5:28Slides, which has changed how people
  181. 5:30make decks going forward. To use it,
  182. 5:31just go to bolt.new, click slide deck,
  183. 5:33and type something like build me a
  184. 5:35five-page deck about AI automation I can
  185. 5:37build for a client. Builds a real
  186. 5:39responsive web app styled and structured
  187. 5:41in seconds. Or you can connect it
  188. 5:43directly to Claude code and have it
  189. 5:44generate a deck for you without ever
  190. 5:46leaving the terminal. And so yes, Bolt
  191. 5:48can make a deck very quickly, but that's
  192. 5:50not why I like it. You guys know how
  193. 5:51much I stress the importance of using AI
  194. 5:53to improve the quality of your outputs,
  195. 5:55not just do things faster. And Bolt
  196. 5:57Slides enhances the deck creation
  197. 5:59process in three ways. First is you can
  198. 6:01embed whatever you want into your
  199. 6:03slides. For example, a live chart with
  200. 6:05updating data or a clickable diagram.
  201. 6:07Honestly, how bangers would it be if you
  202. 6:08had a product showing real-time usage
  203. 6:10and it's counting as you present. Now,
  204. 6:12this really unlocks a world of
  205. 6:14possibilities. Now, the second way is it
  206. 6:16looks native everywhere. So whether
  207. 6:17you're on a projector, a phone, a
  208. 6:19tablet, it doesn't matter because these
  209. 6:21slides are a native web app. One of the
  210. 6:22things that I personally hated when
  211. 6:24pitching brands like Google and Amazon
  212. 6:25at my last startup was whenever I shared
  213. 6:27a deck, I'd have to say, "Make sure to
  214. 6:28open this on desktop." And that was just
  215. 6:31annoying because if they viewed it on
  216. 6:32their phone, it would look terrible. But
  217. 6:34now with Bolt Slides, you can view it
  218. 6:35wherever, however. And the third is
  219. 6:37speed of iteration allows you to
  220. 6:38visualize decks as part of the creative
  221. 6:41process. Typically, when I'd make a
  222. 6:42sales or an investor deck, it started
  223. 6:44with a bullet point list. And then I
  224. 6:46take that bullet point list and then
  225. 6:47visualize it. And at that point, I'd
  226. 6:49often realize the concept didn't
  227. 6:51actually work. But by creating a
  228. 6:52visualization of these concepts earlier,
  229. 6:54you're able to streamline the entire
  230. 6:56process. Now, if you want access to
  231. 6:57Bolt's new feature, click the link below
  232. 6:59and differentiate yourself with the next
  233. 7:01generation of slide decks. Part two,
  234. 7:03system upgrade. So the first upgrade to
  235. 7:05your system is compress inputs
  236. 7:07>> [music]
  237. 7:07>> before AI sees them. So every piece of
  238. 7:09text that you provide Claude consumes
  239. 7:11tokens. So in an ideal world, you would
  240. 7:13compress the text so it's only sharing
  241. 7:15with Claude the text that actually
  242. 7:17matters. To help you visualize this,
  243. 7:19imagine you're working on a report for a
  244. 7:20client and you want Claude to review
  245. 7:22specific changes.
  246. 7:23>> [music]
  247. 7:23>> If you only made changes on the first
  248. 7:24page, would it make sense to share the
  249. 7:26entire 10-page report to Claude? No, it
  250. 7:28wouldn't, but Claude does this by
  251. 7:30default. So what you could do instead is
  252. 7:32use Claude hooks to pre-process these
  253. 7:34files to be more efficient with tokens.
  254. 7:36[music] This is pulled directly from
  255. 7:38Anthropic docs and it tells you to do
  256. 7:39exactly this. It says, "Offload
  257. 7:41processes to hooks and skills." And then
  258. 7:43it gives an example of reading a file
  259. 7:45and only pulling the information that
  260. 7:46has a prefix text error. Luckily, we
  261. 7:49don't have to figure out how to do this
  262. 7:50because there are some gigabrains who
  263. 7:52have helped us for free. There's a free
  264. 7:53open-source tool called RTK that will
  265. 7:56help make the text that's shared with
  266. 7:58Claude significantly more concise. To
  267. 8:00visualize where it sits, normally Claude
  268. 8:02will read the report, it'll reread the
  269. 8:04changes to check the work, and then the
  270. 8:06full raw output is dumped into Claude.
  271. 8:08And so let's say this process takes
  272. 8:10about 15,000 tokens. After RTK, Claude
  273. 8:13will edit the report, it'll reread the
  274. 8:15change to check work. RTK then uses
  275. 8:17deterministic logic to clean up the
  276. 8:19text, removes repeated text,
  277. 8:20boilerplate, formatting noise,
  278. 8:22compresses the text, and then shares
  279. 8:24that with Claude. And that would end up
  280. 8:25being about 1800 tokens. It's quite
  281. 8:28genius, and it works for both technical
  282. 8:30and non-technical tasks. I ran some
  283. 8:32tests on my computer across 13 commands,
  284. 8:34and I was able to save 92% of my tokens.
  285. 8:36Now, RTK claims 60 to 90% and the exact
  286. 8:40savings depend really on what you're
  287. 8:41doing. But to set this up, you literally
  288. 8:43can just paste their repo and then say
  289. 8:45set up RTK on my project, so it runs
  290. 8:47automatically in the background. Hit
  291. 8:48enter, and then it'll go through the
  292. 8:49process. The second system upgrade is
  293. 8:52leverage sub-agents with reduced models.
  294. 8:54Let's go back to the equation we had
  295. 8:56earlier. Compute budget used equals
  296. 8:58tokens consumed times model used. Thus
  297. 9:01far, we focused on tokens consumed. But
  298. 9:03what about the model that we actually
  299. 9:04use? The reality is if you can get a
  300. 9:06task done with Haiku instead of Fable,
  301. 9:08you save 90% of your compute budget. So
  302. 9:10the general rule of thumb in your brain
  303. 9:12is if AI could solve this task a year
  304. 9:14ago, you don't need a frontier model to
  305. 9:16solve it. Scraping, summarizing,
  306. 9:18formatting, fetching files, none of that
  307. 9:20needs the frontier models that are
  308. 9:22getting released. This framework is what
  309. 9:23I call using the minimum viable model.
  310. 9:26So how can we do this in practice
  311. 9:27without over-engineering our system
  312. 9:29pulling our hair out? What we do is we
  313. 9:31predefine the model we use inside Claude
  314. 9:33skills. Skills are predefined tasks that
  315. 9:35you use over and over again. So once you
  316. 9:37define the minimum viable model for that
  317. 9:39skill, it will use that model for all
  318. 9:41future runs. And when setting it up,
  319. 9:43there are two variables you can play
  320. 9:45with, context and model. Model, you
  321. 9:47specify the model that you want to use
  322. 9:49for this specific skill, and then the
  323. 9:50context, if you don't need to have
  324. 9:52contextual information, you can say
  325. 9:53context [music] fork, which will create
  326. 9:55a new thread for that skill to work on,
  327. 9:57in turn reducing the contextual
  328. 9:59information that's needed. If you do
  329. 10:01need previous info, you'll just leave
  330. 10:02that part blank in the skill itself.
  331. 10:05Here's a table on how to optimize skills
  332. 10:06based on specific needs. If you're still
  333. 10:08struggling to visualize why to do this,
  334. 10:11think of it like a team of lawyers. You
  335. 10:12want the partner at the law firm leading
  336. 10:14the case, but because his hourly is so
  337. 10:16expensive, you want junior lawyers doing
  338. 10:18a lot of the simple grunt work. It's the
  339. 10:20same idea here, and to find where this
  340. 10:22applies in your own setup, here's a
  341. 10:24prompt. Now, that's just one of the ways
  342. 10:25to upgrade your skills to be more
  343. 10:26efficient, but the next upgrade is one
  344. 10:29of my favorites. Upgrade number three,
  345. 10:30move your workflows into script-driven
  346. 10:32skills. After you've updated your skills
  347. 10:34to use a specific model, there is one
  348. 10:36more step. This one improves quality,
  349. 10:38consistency, and speed, while also
  350. 10:40reducing token consumption. So, a lot
  351. 10:42like RTK enhancement that uses computer
  352. 10:45logic to compress text, you can use
  353. 10:47computer logic or scripts to complete
  354. 10:49tasks that you're currently using AI
  355. 10:50for. So, a script will run the same way
  356. 10:52every time, call zero tokens, and never
  357. 10:54hallucinates. So, in an ideal world, we
  358. 10:56want to use AI for judgment, and then
  359. 10:58scripts for everything else that's
  360. 10:59repeatable. This way, AI doesn't have to
  361. 11:01refigure out something every time you do
  362. 11:02it. Here's a prompt that enhances your
  363. 11:04skills to leverage scripts and flags any
  364. 11:07skill your project is missing. Now,
  365. 11:08before we get to part three, which walks
  366. 11:10through nuclear enhancements, if this is
  367. 11:12your first video of mine, welcome to the
  368. 11:14channel. But, if this is your second or
  369. 11:15more, you know the drill, this is our
  370. 11:17anti-slop agreement. The visuals, the
  371. 11:19testing, the hours of research that went
  372. 11:21into this video, this is entirely built
  373. 11:22for humans, not for these token
  374. 11:24gobblers, okay? So, all that I ask is
  375. 11:27that you subscribe as part of this
  376. 11:28agreement, to help this content reach
  377. 11:29more people so that I can just keep
  378. 11:31doing this. Also, every video I give a
  379. 11:33Claude Mac subscription away. This
  380. 11:34video's winner is Chad N5N5X
  381. 11:37who's building a complete AI operating
  382. 11:40system. Shout out Chad, you are a
  383. 11:42legend. Now, if you want to enter the
  384. 11:43next giveaway, comment below with what
  385. 11:45you're building or any recent issues
  386. 11:47you've run into. And every video you
  387. 11:48comment on is another entry. Part three,
  388. 11:50nuclear enhancements. These are four big
  389. 11:53changes listed in the order that you
  390. 11:54should do them. And the last one is
  391. 11:55going viral right now, but I'll explain
  392. 11:57why I don't necessarily think it's worth
  393. 11:59your time. Nuclear enhancement one,
  394. 12:00route specific work to Codex. The
  395. 12:03reality is that certain models are more
  396. 12:05efficient than other models. But that
  397. 12:06also extends outside the model layer
  398. 12:08into the harness layer, the tool that's
  399. 12:10actually orchestrating using AI. So,
  400. 12:13when you prompt Claude code, the logic
  401. 12:14that interacts with the AI models is the
  402. 12:16harness. So, Codex ecosystem, OpenAI's
  403. 12:19product compared to Claude, on certain
  404. 12:21tasks it burns 4x less tokens than
  405. 12:23Claude does. And the reason for this, to
  406. 12:26simplify it, is that Claude is built to
  407. 12:28be thorough. It rereads, it verifies, it
  408. 12:30thinks before it acts, and you pay
  409. 12:32tokens for every step along the way.
  410. 12:34Whereas for Codex, it's built to be
  411. 12:35surgical. Get in, make the edit, get
  412. 12:37out. So, to use this to your advantage,
  413. 12:38you can intertwine Codex and Claude to
  414. 12:41get the most out of your subscriptions.
  415. 12:43The simplest way to do it, and my
  416. 12:44favorite, is there's a Claude code
  417. 12:45plugin for Codex. You can install that,
  418. 12:48then have Claude route any token-heavy
  419. 12:49execution task to Codex. This prompt
  420. 12:51will help you install that plugin, set
  421. 12:53it up, and then update your ClaudeMD to
  422. 12:55tell Claude to actually do this. Nuclear
  423. 12:57enhancement number two is use images
  424. 13:00instead of text. This one sounds fake,
  425. 13:02but it's so awesome that I had to
  426. 13:04include it. So, Claude processes images
  427. 13:06at a different rate than it processes
  428. 13:08text. So, you can convert text into
  429. 13:11images, [music]
  430. 13:11and then submit that to Claude to become
  431. 13:14more efficient with your tokens. Now,
  432. 13:15there's a lot going on there, but to
  433. 13:16visualize it, imagine you had a massive
  434. 13:18page of text with over 10,000 tokens,
  435. 13:21and then an image with that same text
  436. 13:23but in image form. In order for a Claude
  437. 13:25to actually process this image, it only
  438. 13:27takes 3,000 or 4,000 tokens, resulting
  439. 13:30in a 60 to 70% token reduction. These
  440. 13:32are based on claims from PX pipe, the
  441. 13:34open source tool on GitHub that will
  442. 13:36help you do exactly this. Now, there are
  443. 13:38some trade-offs like it won't read the
  444. 13:39text perfectly and there is a non-zero
  445. 13:41chance that Anthropic [music] patches
  446. 13:43this in the future, but I found this so
  447. 13:45damn interesting that I just had to
  448. 13:46include it. Nuclear enhancement number
  449. 13:48three, swap the engine out entirely.
  450. 13:50Most people aren't aware of this, but
  451. 13:51Claude code is the harness and the model
  452. 13:54that's inside of it and is actually used
  453. 13:56can be swapped. So, for example, you
  454. 13:58could use Claude code and only use
  455. 14:00OpenAI or other open source models. Now,
  456. 14:03to actually do this, it's pretty
  457. 14:04straightforward. There are some
  458. 14:05environment variables that Claude code
  459. 14:06will reference to know which model to
  460. 14:08use and [music] by default, it goes to
  461. 14:10Claude's models. Now, two potential
  462. 14:12options that I'm considering looking
  463. 14:13into are ZAI's GLM plan and then the
  464. 14:16Deep Seek plan. Based on research for
  465. 14:18these on a compute per dollar spent,
  466. 14:20these can provide you with more capacity
  467. 14:22than Claude's plans. There are
  468. 14:24trade-offs, right? Deep Seek is a
  469. 14:25Chinese model, you may be concerned with
  470. 14:27your data privacy and also when you're
  471. 14:29using these tools, you are getting
  472. 14:31slightly worse models. Here's a prompt
  473. 14:33you can use to go down this rabbit hole
  474. 14:34and learn a lot more and as part of
  475. 14:36that, it'll build you an implementation
  476. 14:37plan if you want to eject out of the
  477. 14:39Anthropic ecosystem. This prompt is also
  478. 14:41designed to help you identify current
  479. 14:43providers as offers are constantly
  480. 14:45changing. Now, the fourth nuclear
  481. 14:47enhancement is running your own model
  482. 14:48locally. This is an enhancement that I
  483. 14:50eventually will fully believe in. It's
  484. 14:52beautiful, right? You can run your own
  485. 14:54models locally and not rely on external
  486. 14:56data providers. So, for example, I have
  487. 14:58this Mac Mini, it's a corner here, you
  488. 15:00can't see it, but I could run a model on
  489. 15:02it and then route all of my requests
  490. 15:03directly to that instead of Claude. So,
  491. 15:05everything is staying within my office
  492. 15:07and now there are pros and cons to this.
  493. 15:09Some of the pros is you can essentially
  494. 15:11have limitless tokens. All I have to do
  495. 15:13is pay for the power to keep the
  496. 15:14hardware running. Next is that you own
  497. 15:16all of your data. If data privacy is a
  498. 15:18concern, this means nothing is exposed
  499. 15:20to external service providers. Third,
  500. 15:22[music] you own the entire AI stack. You
  501. 15:24aren't relying on anyone else for and
  502. 15:26this concept of running a model locally
  503. 15:27is why you will never actually be able
  504. 15:29to ban AI entirely. Now, the cons of
  505. 15:31this, first, my Mac Mini and then 99% of
  506. 15:34consumer hardware can't actually run any
  507. 15:36of the top-tier open-source models. And
  508. 15:38in order for you to get to those models,
  509. 15:40you're going to have to spend over
  510. 15:41$10,000. Next is that none of the
  511. 15:43frontier models like Fable are
  512. 15:45accessible to download locally, which
  513. 15:47means you're sacrificing quality in the
  514. 15:49short term. And then finally, you have
  515. 15:51to set up a computer server farm and as
  516. 15:53someone who turns his phone on airplane
  517. 15:54mode at night, I would not want all of
  518. 15:56that EMF radiation near me. Jokes aside
  519. 15:58with the EMF radiation, but you have to
  520. 16:00set up the server farm and you have to
  521. 16:01manage it. That's work. So, long-term, I
  522. 16:03do fully expect local models to be a
  523. 16:05thing as hardware gets better and the
  524. 16:07open-source models get better as well.
  525. 16:09And at that point, it may 100% be worth
  526. 16:12it, but right now, my recommendation for
  527. 16:14you is you can play around with local
  528. 16:15models. It's valuable for you to know
  529. 16:17that it exists, but I just wouldn't go
  530. 16:19all in on this. So, please don't go and
  531. 16:21buy expensive hardware. Maybe one day,
  532. 16:23just not today. So, with that being
  533. 16:24said, let's speed run the changes that
  534. 16:26you need to make today to get more out
  535. 16:28of your system. So, the first quick win
  536. 16:29is fix your contextual habits. Run
  537. 16:31{slash} clear every time you switch
  538. 16:33tasks, work in focused blocks so you can
  539. 16:35keep 90% cash discount, and set your
  540. 16:37effort to match the task. [music] And
  541. 16:39then use {slash} compact when you get to
  542. 16:4060% full on your context. The second
  543. 16:43quick win is run the cleanup prompt.
  544. 16:45Disconnect MCPs you don't use, archive
  545. 16:47unused skills, and shorten their
  546. 16:48descriptions, and turn your Claude MD
  547. 16:50into a directory instead of a document.
  548. 16:52And then make sure your Claude MD is
  549. 16:53less than 200 characters. The third
  550. 16:55[music] is cut your output tokens. Add
  551. 16:57be concise to your Claude MD, use the
  552. 16:59KMM plugin, or update your Claude MD to
  553. 17:01tell Claude to be concise with all their
  554. 17:03responses. Then for the system upgrades,
  555. 17:05install RTK so every command output gets
  556. 17:08crushed before Claude reads it. This
  557. 17:09could reduce your token input by 60 to
  558. 17:1190%. Second, enhance your skills so
  559. 17:14front work runs on minimum viable models
  560. 17:16instead of your most expensive ones. The
  561. 17:18third system update is turn every
  562. 17:20repeatable step into a script inside the
  563. 17:22skill. Computer code causes zero tokens
  564. 17:24to run. Then the nuclear enhancements.
  565. 17:26First, route any token-heavy execution
  566. 17:29to Codex with the plugin. This is where
  567. 17:30you could have two subscriptions with
  568. 17:32two different budgets and more total
  569. 17:33firepower. The second is use images
  570. 17:35instead of text. Keep an eye on this
  571. 17:37one. The third is that you can swap the
  572. 17:39engine out entirely if you want to work
  573. 17:40with different model providers. And the
  574. 17:42fourth, as I mentioned, you can run your
  575. 17:44own local model if you really want to.
  576. 17:46Now, once you apply this, you'll be able
  577. 17:47to build all day long instead of having
  578. 17:49to wait every 5 hours for your limits to
  579. 17:52reset. And if you like this video, you
  580. 17:53will love this video where I walk
  581. 17:55through my exact setup to leverage
  582. 17:57Claude's skills to build 10 times
  583. 17:59faster. This builds on a lot of what I
  584. 18:00covered in the skill optimization
  585. 18:02section of this video and you will love
  586. 18:04that. I'll see you over there. Peace.

About this transcript

This page contains the full transcript of Paste This Into Claude, Never Hit a Token Limit Again by Austin Marchese, generated from the public captions YouTube serves with the video. The transcript has 4,046 words across 586 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.