Paste This Into Claude, Never Hit a Token Limit Again — Transcript
Full transcript
- 0:00You keep running out of Claude or Codex
- 0:02or or chat GPT or Kimmy or whatever you
- 0:04want and you don't do anything
- 0:05unreasonable to run out of tokens. You
- 0:07asked a handful of questions and it told
- 0:09you to come back in 3 hours or 5 hours
- 0:11or next week. On one working day, my
- 0:14tracker recorded 3.77 billion tokens
- 0:17moving through my Codex workspace. Of
- 0:20that, 3.59
- 0:23billion were reused input, almost 96%.
- 0:27And look, I push these tools really
- 0:28hard. That was 143 separate Codex
- 0:31threads in one day.
- 0:32But I did not type 3.77 billion tokens
- 0:36and neither did you. So, if the answer
- 0:38fails, the retry carries most of all of
- 0:41that a second time and a third time and
- 0:42a fourth time while you're telling
- 0:44Claude or you're telling Codex, you got
- 0:45to fix this. So,
- 0:48this is the core idea that every single
- 0:51rule that I'm about to give you comes
- 0:53back to. The message you typed is the
- 0:55tiniest part of the overall call. I'm
- 0:58Nate B. Jones. I'm here to help you
- 1:00build the life you want with AI. That's
- 1:02what I'm going to show you how to do
- 1:03today. We're going to go through all 15
- 1:05and by the way, if you're a beginner,
- 1:06this is for you. You're going to go
- 1:08through things you can do initially. I
- 1:10built a skill so you don't have to
- 1:11remember all of it and if you're a
- 1:13little bit more technical, we are going
- 1:14to get to a really cool multi-agent
- 1:16automated solution at the end. So, stay
- 1:17tuned for that. Your 10th message, your
- 1:1920th message, your 100th message in the
- 1:22chat cost many more times than your
- 1:25first and nobody ever raised the price
- 1:27on you. It's just that every time you
- 1:28hit enter, the way LLMs work, the entire
- 1:32conversation gets wrapped up in a bow
- 1:34and sent again from the top. That's how
- 1:36LLMs pretend to have memory right now.
- 1:38Watch how fast that compounds. Your
- 1:41first message costs exactly what you
- 1:44typed. That's the easy part. Your second
- 1:46costs what you typed plus the answer
- 1:49plus what you originally typed. Your
- 1:5010th costs what you typed plus all the
- 1:53nine previous conversations you already
- 1:55paid for. And by message 30,
- 1:58the thing you actually wrote just now is
- 2:00a tiny tiny rounding error. That old
- 2:03material, it has a name. It's called
- 2:05reused input. The part of every request
- 2:09the model has already seen. So, two
- 2:11things to call out. One,
- 2:13this is not getting fixed by the labs.
- 2:17The labs are not going to magically fix
- 2:19this tomorrow. It is up to us.
- 2:21The labs aren't going to fix it because
- 2:22frankly, they have an incentive to get
- 2:24us using the product. And if we're using
- 2:26that's great up to a point, right? If
- 2:28they run out of compute, they're going
- 2:29to sort of rein it back. But
- 2:30fundamentally, the labs want us to use
- 2:34their tokens. So, it's up to us to make
- 2:36the most of the limits they give us. AI,
- 2:39it's like this desk. And there's a
- 2:41belief going around
- 2:43that this is going to fix itself right
- 2:44now. And that our desks can clean
- 2:46themselves and the models are going to
- 2:48get bigger context windows and they're
- 2:50going to run longer and have more agents
- 2:52getting more tools. And soon you can
- 2:53point the thing at your work. And it's
- 2:55just going to run, right? We In other
- 2:57words, we believe we're going to be able
- 2:59to have a messy desk and it's going to
- 3:01be fine. But if I take all the stuff off
- 3:03of my shelves and put it on the desk,
- 3:05I'm not going to be very productive. I'm
- 3:07going to be overloaded and stressed and
- 3:09it's not going to work. It just doesn't
- 3:10work that way. I have to pick one Lego
- 3:12set for my desk. The more capable the
- 3:15tool, the more tools you hand it, the
- 3:17more material it puts in and the faster
- 3:20you hit the wall. So, ironically, more
- 3:21capability may end up giving you more
- 3:24cleanup problems. And I talk about all
- 3:26three in this video. But none of it
- 3:29really takes the responsibility of
- 3:31organizing your AI desk for you. And
- 3:34reducing your token consumption is not
- 3:36what any of these companies is graded
- 3:38on.
- 3:39It's your desk. You have to own it. Is
- 3:42it fun to clean your desk? It can be
- 3:43kind of boring, right? It's almost
- 3:45always the boring stuff though that
- 3:47gives you the superpowers. So, level one
- 3:49is for everybody. It's how you keep your
- 3:51desk clean. Nine habits, nothing that
- 3:54you need to install, will work on any
- 3:56AI. And some of it is advice we've known
- 3:59for a while, and I'll tell you when it
- 4:01is, but guess what? Old advice may still
- 4:03be something you need to hear if it's
- 4:04not something you're following right
- 4:06now.
- 4:07Then we're going to get to level two.
- 4:09Level two is kind of like hiring someone
- 4:11to keep your desk clean. I built a
- 4:12skill, it's called Token Saver. It
- 4:15installs into Codex, it installs into
- 4:16Cloud Code, and it does most of level
- 4:19one for you while you keep working
- 4:21normally, which is kind of handy. Level
- 4:24three is stopping the mess before it
- 4:26ever gets to your desk. And that's a
- 4:28software piece that I've been building
- 4:30around my Ringer Multi-Agent Framework.
- 4:32It's the most powerful option by far,
- 4:35and I will tell you exactly where the
- 4:36edges are and how it works when you are
- 4:38ready at the end of the video. Level
- 4:40one, clean your own desk. Rule number
- 4:44one, you've got to edit your mistakes.
- 4:46So, if you're writing something in AI
- 4:48and you're like, "Oh, man, I didn't mean
- 4:49to write that. I had a typo. I asked the
- 4:51wrong thing." Do not say that was wrong
- 4:53in the next chat. Just edit it, which
- 4:56you can do, and resend the message. And
- 4:58you may have heard of this, but almost
- 5:00nobody does it. The model will give you
- 5:02something wrong because your request was
- 5:03unclear, and your instinct is to just
- 5:05say, "No, that's not it." Well,
- 5:08instead, hit that little edit button,
- 5:11and make sure that you are actually
- 5:13correcting the unclear request you had
- 5:15before, and get it right. Rule number
- 5:17two, ask related questions together, and
- 5:21say how you want the answer to appear.
- 5:24Not new advice, again, but absolutely
- 5:26worth a minute for you to realize. If
- 5:28you have multiple questions from the
- 5:30same document set, put them all into one
- 5:33query, and then specify what you want at
- 5:36the end. What does it look like? Is it a
- 5:37one-pager you want? Is it 150 words? Is
- 5:41it just give me the bullets? Is it give
- 5:42me the headline? Name it and say it.
- 5:45Because that way you are reducing the
- 5:48ambiguity that the AI is going to spend
- 5:50tokens on to give you answers on. So,
- 5:52just name what your questions are all at
- 5:54once. Don't just string them along as
- 5:55multiple questions and name the answer
- 5:58you want in advance. Rule number three,
- 6:00start a clean task when the job changes.
- 6:04And this is this is a huge win and
- 6:06people seem to resist this. I know, in
- 6:08fact, a lot of people who believe they
- 6:10are in romantic relationships with AI
- 6:12because they didn't bother to do this.
- 6:14Long conversations are really good while
- 6:17you're focused on the same problem, but
- 6:18they're really terrible for trying to
- 6:21deal with token usage when you're trying
- 6:25to get specific questions answered. And
- 6:27we get away with this more now cuz
- 6:28models are smarter. And so, you can
- 6:31disambiguate more in that long context
- 6:33window, but it's really token heavy and
- 6:35you hit your usage limits faster. And
- 6:37this drove the biggest measured change
- 6:38when I was testing all of these out on
- 6:41my own Codex and Claude installs. When
- 6:43you are carrying a working conversation,
- 6:46you carry so many reused tokens with
- 6:49you. Like, you can carry 50,000,
- 6:51100,000, a million, or as I've been
- 6:53saying, if you're getting into the
- 6:54billions, it's hundreds of millions of
- 6:55tokens with you. Start from scratch.
- 6:59Now, a clean task doesn't actually
- 7:01report zero every time because, of
- 7:03course, Codex and Claude will still send
- 7:05their own starting instructions.
- 7:07But what it does is it stops the old
- 7:09conversation from writing along.
- 7:12And you're not throwing the old task
- 7:13away. You can keep it. Just stop making
- 7:15the next job that you that you want it
- 7:17to do carry that task. Rule four is
- 7:19related. Carry the answer that you care
- 7:22about, not the argument. Let's say you
- 7:24have a multi-step process. You have a
- 7:25research piece, you want to have a
- 7:27document writing piece. This was big for
- 7:29me. Make sure that you separate out the
- 7:33stages of the job and that when you have
- 7:35an artifact that's produced after a
- 7:37stage, like a research report, that's
- 7:39what you carry forward into the next
- 7:42step so that you only have that to give
- 7:45to your AI to start. You don't have
- 7:4810 million, 100 million tokens of
- 7:50research that you're handing, much of
- 7:52which is irrelevant because you had to
- 7:54guide the research along the way. Be
- 7:56precise.
- 7:58You don't need to stack your bad first
- 8:00draft and your three rounds of criticism
- 8:02and the sources you rejected and the
- 8:04model's reasoning all into the next
- 8:05step. You can actually just take the
- 8:08result you got and move on to the next.
- 8:10Keep your desk clean. Rule five, ask for
- 8:13only the answer you need. Input gets a
- 8:15lot of attention because it gets so big,
- 8:17but output costs you twice. It costs you
- 8:19once when it's written very expensively
- 8:22and again when it becomes part of the
- 8:23input on the next turn and again when
- 8:25there's a turn after that and again when
- 8:26there's a turn after that, you keep
- 8:27getting billed on it. So, remember, when
- 8:30when we first got like the capability to
- 8:33write a lot from AI, I saw a lot of
- 8:35people who would go tell deep research
- 8:37to write them 50-page papers. But, what
- 8:40I see a lot more of now is people
- 8:42saying, "Can I get it in 50 words? Can
- 8:44Can you please write it precisely in a
- 8:46way that I can understand it and give me
- 8:48what matters?" We need more of that
- 8:50because when you do that, you're saving
- 8:52not just on the output tokens, you're
- 8:54saving on every single response that
- 8:56comes after that. And so, ask exactly
- 8:59for what you need. If you need a
- 9:00paragraph, ask for that. If you need
- 9:02JSON, ask for that. If you need five
- 9:04bullets, get five bullets. Rule six is
- 9:06really simple and it's not something
- 9:07people understand a lot. You want to
- 9:10search the file yourself whenever you
- 9:12can and don't make the model search the
- 9:13file. Yes, the model can search the file
- 9:16now. It's a great convenience, it's also
- 9:18a massive token burner.
- 9:20You want to be in a position to say,
- 9:23"Hey, I searched it. I found these
- 9:25things. This is what you should focus
- 9:27on. This is what I'm including in this
- 9:28snippet for you. You don't need to go
- 9:30read this entire file." Rule seven is
- 9:32related. You want to send the lightest
- 9:34useful form of the source. In other
- 9:37words,
- 9:38if the words matter and the layout
- 9:40doesn't, just convert it to text. I've
- 9:43talked about this before. You don't have
- 9:45to send a PDF just cuz the source came
- 9:48to you in a PDF, convert it to markdown,
- 9:51convert it to text, and just paste the
- 9:54text in. It's so much more efficient.
- 9:57Don't just be lazy and say, I can throw
- 10:00a PDF and 18 screenshots in and it's all
- 10:02fine and it will sort it out because as
- 10:04tempting as it is, cuz it does sort it
- 10:06out, eat your token bill so so fast.
- 10:09Make sure that you take the time to
- 10:12actually get your sources in order.
- 10:15Clean your desk. It's common theme. Rule
- 10:17nine ties into those of you who have
- 10:19open brain. Keep your answers around
- 10:21somewhere you can find them. And like if
- 10:23you have open brain, it may be in your
- 10:24open brain database or you may have your
- 10:26own database that you use, but really
- 10:29really important. If you are working on
- 10:32a particular problem set and you're
- 10:33working around the edges of it, you have
- 10:35multiple conversations. The more you
- 10:37keep information about that in a place
- 10:39you can look it up easily, the more the
- 10:42AI is going to be able to look it up and
- 10:44not have to go recreate it and not have
- 10:46to go dig for it in the sources and
- 10:47recalculate it. It's just going to be
- 10:49able to get that particular piece of
- 10:50data out of the database and be done
- 10:52with it and that is so so much more
- 10:53token efficient. So, if you're not using
- 10:56open brain, you can check it out. I have
- 10:58a lot of videos on it. I can link it
- 11:00here. It's super easy to get started on,
- 11:02but make sure that you have a system
- 11:04that allows you to go and retrieve data
- 11:07for stuff that you would look up
- 11:08multiple times otherwise because the
- 11:10cheaper you can make that, the more
- 11:12you're saving tokens. Let's say you're
- 11:14like, "Nate, I love this. I'm lazy. I
- 11:17don't do this. I don't remember to do
- 11:19this. Please help me." That's what I
- 11:21built the skill for. I built the skill
- 11:23called token saver. One command will
- 11:25install it. You can get it into Codex,
- 11:27you can get it into Claude Code,
- 11:28wherever you do your skills, and then it
- 11:30will keep working with you the way you
- 11:33already work. And so, you can just say
- 11:34use the token saver skill for this job
- 11:36and it will handle a lot of the tedious
- 11:38parts of level one for you. It'll search
- 11:40before opening large sources. It will
- 11:43send selected passages instead of whole
- 11:45files. It will run exact work as code
- 11:47wherever it can. It saves the version
- 11:49you accepted and builds your next
- 11:51request from that result plus your
- 11:52change. It keeps answers to the length
- 11:55that you asked for and it stops
- 11:57pointless retries over and over. So,
- 11:59there's a good things about it and I
- 12:02want you to use it and find ways to
- 12:04continue to improve it. And that's
- 12:05something that I love about our
- 12:06community in the slack is that we find
- 12:08ways to improve each other's skills. So,
- 12:10I'm launching this. I've tested it. My
- 12:12team tested it. We love it. I'll also
- 12:14call out that this helps you with the
- 12:16next three rules I'm about to give you
- 12:18that are really difficult to do by hand.
- 12:20And so, I want you to like listen to
- 12:22these, understand them, but realize the
- 12:24skill is going to help you get there.
- 12:26So, rule 10, load only the tools the job
- 12:29can use. Now, this is something where I
- 12:31quite frankly in six or eight months
- 12:34expect the models to be good enough to
- 12:36solve this problem, but they're not
- 12:38reliably good enough at it today. Every
- 12:40tool you connect carries a description
- 12:42right now. What it does, when to use it,
- 12:45what arguments it takes. That
- 12:47description is model input before the
- 12:49model does anything at all.
- 12:51Now, Anthropic has published some work
- 12:54that makes this whole problem space
- 12:56very, very real. A typical setup with
- 12:58several tool servers connected like say
- 13:00GitHub and Slack and Sentry and Grafana.
- 13:03It burns roughly 55,000
- 13:06tokens in tool definition
- 13:08before Claude does anything. Now, we're
- 13:10starting on more advanced rules. This is
- 13:13stuff where the skill can be supported,
- 13:14but you also have to use your head.
- 13:16Sometimes starting clean on a thread
- 13:18like I recommended in the sort of the
- 13:20beginner section, it's not very
- 13:21practical. Let's say you're debugging a
- 13:23particular system and the model needs a
- 13:25decision from you and you can't restart
- 13:28it because it's in the middle of the
- 13:29task. This is an area where compaction
- 13:32and context editing become really
- 13:33important. OpenAI supports compaction
- 13:36for long-running work, and it carries
- 13:37forward the state, and it turns out it
- 13:40needs fewer tokens as a result, but
- 13:41you're depending on their native
- 13:42compaction capabilities. Anthropic
- 13:45supports context editing, which clears
- 13:47old tool results out and and clears
- 13:50thinking blocks before the next request,
- 13:52and that's also super helpful. And so,
- 13:54the skill helps with it with that, but I
- 13:55also need to be honest with you that
- 13:56this is just something you as an
- 13:58advanced AI user should be aware of. You
- 14:00should know where your context window
- 14:01is, and you should recognize that when
- 14:05you clear out old material, you are
- 14:07depending on a approximated version of
- 14:12the initial prompt, approximated
- 14:14versions of the initial responses that
- 14:17the model makers are using to enable you
- 14:20to continue work on a long-running
- 14:21thread.
- 14:23That is not perfect, but it's a whole
- 14:24lot better than just hitting a wall,
- 14:26which is what we used to do. So, the
- 14:28takeaway for you is use the skill,
- 14:30understand where your context window is,
- 14:32and then make sure that you actually
- 14:35anticipate in your work
- 14:38the consequences of hitting that wall.
- 14:40Put the signal early. The skill also
- 14:43helps when you want to figure out what
- 14:45model to use because sometimes a smaller
- 14:48model helps you get the job done
- 14:49cheaper, but you have to trade off the
- 14:51context, etc. So, what the skill does is
- 14:53what it looks at the at the question or
- 14:55the problem you're tackling, and it
- 14:57comes up with, at least an initial
- 14:59opinion on whether the model you're
- 15:01using is correct or not. You can
- 15:02obviously disagree, you can move on, you
- 15:04can say, "No, I want to use this model,"
- 15:05but at least you have a first blush
- 15:07approximation at what a token-efficient
- 15:10model solution is. I I like to say, like
- 15:13if you're doing a serious task, use the
- 15:16absolute dumbest model that will still
- 15:19get the work done for you. And the more
- 15:21you work with AI, the more you have a
- 15:23feel for what that line of dumbest model
- 15:25is. And so, the skill just kind of helps
- 15:27you get there, helps you take a guess at
- 15:29that if you're new at that. I hear a lot
- 15:31about prompt caching. That's what we're
- 15:32talking about with this rule. Prompt
- 15:34caching is really, really important if
- 15:37you're doing repeated work. It's
- 15:38especially important with API work. It's
- 15:40not something that I would recommend for
- 15:42people who are doing just initial desk
- 15:44work. It's not a clear your desk feature
- 15:46if you're a regular knowledge worker.
- 15:47It's something where it's an API
- 15:49feature. You're going to be in a
- 15:50position where you can cache a prompt or
- 15:52cache part of what you're sending, and
- 15:54that makes it much, much more efficient
- 15:56to send because you're not sending the
- 15:57whole message back and forth. And that's
- 15:59really all you need to know. And if
- 16:00you're someone who's already diving into
- 16:01this, you're like, "Yeah, Nate, I do
- 16:03prompt caching, or at least I know what
- 16:04prompt caching is." And you're off to
- 16:06the races. And if you're someone who's
- 16:07like, "What is an API and what is prompt
- 16:09caching?" Well, by definition, you don't
- 16:11need it. And you can focus on all the
- 16:13other good habits that I just talked
- 16:14about to keep your desk clean and make
- 16:15sure that you're not running into your
- 16:17AI limits. And this is where the Ringer
- 16:19multi-agent framework comes in.
- 16:20Everything so far shares a single
- 16:23ceiling.
- 16:24And that ceiling is that a skill cannot
- 16:27make the call it is inside of any
- 16:30smaller. Like, if you're in the middle
- 16:31of a conversation and you're using the
- 16:33skill,
- 16:34it's going to do its very best, but it's
- 16:36going to sort of work on things going
- 16:38forward. By the time the model reads the
- 16:40skill, the request initially has already
- 16:43been sent, right? The conversation, the
- 16:44standing instructions, the tool
- 16:46definitions, the hidden setup I talked
- 16:47about. A lot of that is in the envelope
- 16:50before the skill ever gets invoked.
- 16:53And so, what I want to do is think about
- 16:55that initial request and how we hook
- 16:57into that and make that cleaner. And
- 16:59that's the problem I'm solving with
- 17:01Ringer and level three. And yes, it is
- 17:03absolutely a bit more of an advanced
- 17:04solution. Don't don't be scared of it.
- 17:06You can absolutely do it. I'm just
- 17:08telling you honestly what it's actually
- 17:10going to take. Ringer runs locally
- 17:12between your AI and the model provider.
- 17:16And before the request goes up to the
- 17:18model provider,
- 17:19it can return an answer without a model
- 17:22call in some instances. It can run a
- 17:24fixed local recipe with no model call in
- 17:26some instances. It can select only the
- 17:28useful passages to send through in some
- 17:30instances, or it can forward a small
- 17:32request under hard limits, or it can
- 17:34even stop it entirely. It is not another
- 17:36chat window. You don't go somewhere else
- 17:38to use it. Instead, it is effectively an
- 17:42intermediary that hooks in, and it helps
- 17:45to constrain the size of what you're
- 17:48sending with all of these tricks built
- 17:50in. Rule number 14, the the
- 17:52second-to-last one, you want to make
- 17:55sure that you are able to enforce hard
- 17:57limits if you're serious about token
- 17:59usage. So, if you want to say, "I only
- 18:02want to send packets of a certain size
- 18:05out, or I only want to get back packets
- 18:08of a certain size, or I want to make
- 18:09sure I have a hard limit on my call, so
- 18:11I'm never sending 10 million tokens."
- 18:14That's something you can enforce with an
- 18:16in-between intermediary like Ringer.
- 18:18It's not something you can really do
- 18:20without that.
- 18:21The other thing that's important is that
- 18:23Ringer allows you to take advantage of
- 18:25what I talked about with OpenBrain. So,
- 18:27OpenBrain has, "Oh, if it's got an
- 18:29answer, we can go get it." Well, Ringer
- 18:31can go hit OpenBrain and come back, or
- 18:33your database of choice and come back,
- 18:35and say, "We've already had an accepted
- 18:37answer here. We already talked about
- 18:38this last week. We've got this response.
- 18:40Is this what you mean?"
- 18:42And that saves you the call, right? It
- 18:43saves you the entire 100% of the call.
- 18:45We have our desk here.
- 18:47The nine things that we're talking about
- 18:49initially are you picking up your pen
- 18:51and your paper and your LEGOs and
- 18:53keeping the desk clean. And then the
- 18:54skill is kind of like someone who comes
- 18:56in and cleans your desk for you every
- 18:58night. And then Ringer is really a
- 19:01magical system that keeps your desk
- 19:04clean for you
- 19:05before the desk ever gets messy. I'm
- 19:08always telling you to think big on this
- 19:09channel, and I don't want token limits
- 19:11to be the thing that holds you back. And
- 19:13so this video is all about making sure
- 19:15that you can literally 10x the value of
- 19:17your tokens and get where you want to
- 19:19go. I'm saying 10x for a reason and I
- 19:20was able to audit down all of the tokens
- 19:24that I've been using and say where they
- 19:26actually going, what am I wasting and
- 19:29how do I make sure that I'm putting my
- 19:30tokens toward their maximum value so I'm
- 19:33not wasting the dollars I'm spending on
- 19:34subscriptions. So better tools are
- 19:36coming and they will get better but
- 19:39you're still going to have to keep your
- 19:40desk clean.
- 19:41And everything is linked below. The 15
- 19:43rules are written up in full. I have the
- 19:45skill for you over on Substack and
- 19:48Ringer is there and if you want a whole
- 19:50introduction to Ringer, I have a whole
- 19:51video on that. And if you tuned out and
- 19:53you're the kind of person that wants
- 19:54Ringer, you can get that started. If
- 19:56you've tried any of the tools I've
- 19:57mentioned, if you tried the skill, if
- 19:59you've tried Ringer, let me know below.
- 20:01If you're just trying it for the first
- 20:02time, let me know below, too. Let me
- 20:04know what your experience is and if you
- 20:06know someone who's running out of AI,
- 20:08share this video with them. We don't
- 20:10want them to run out of AI.
About this transcript
This page contains the full transcript of Paste This Into Claude, Never Hit a Token Limit Again by AI News & Strategy Daily | Nate B Jones, generated from the public captions YouTube serves with the video. The transcript has 4,126 words across 579 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.