How To Save 90% of Claude Code Token Usage — Transcript
Full transcript
- 0:00Everyone how's it going in this video i'm going to teach
- 0:02you how to reduce your cloud code token
- 0:05usage by up to 90
- 0:07in
- 0:08four different strategies it's easy and free to do and you
- 0:11can do it right away and in this video i'm going to show
- 0:13you step by step and also a quick comparison on like
- 0:16visually being able to see difference in tokens that you'll
- 0:19be ending up using now all of these strategies actually do
- 0:22have some trade-offs and i'll also go over those trade-offs
- 0:26as well so be sure to stick around through the entire video
- 0:29so you don't miss out on exactly how to use these things
- 0:31properly as usual i made some slides so there's some
- 0:34visual representation of what i'm talking about so let's get
- 0:37right into the four different strategies now the first
- 0:40strategy is actually to index the code essentially we're
- 0:43going to create an index of the code now if you don't
- 0:46know what indexing is it's essentially like search actually is
- 0:52a very famous usage of an index. Essentially what Google
- 0:56does behind the scenes is it creates a map of like
- 0:59keywords that point to different websites as a very
- 1:02dumbed down example. So the core idea is that we want
- 1:06to essentially create a graph map of our code base.
- 1:09So here's a visual representation. Now on the left side,
- 1:12it's kind of how clock code already does
- 1:15the grep and reading and scanning of the code base.
- 1:18It will go one by one and deep dive into the code base.
- 1:21If you're asking about some particular, you know,
- 1:24server implementation or some validations like file or an
- 1:28auth file, we'll try to go and grep it using regular file grep
- 1:31and they will go and find it and load those into the
- 1:34token context. So it ends up spending a lot of time
- 1:37because sometimes to find the files that it's looking for,
- 1:40it would have read a lot of things that it doesn't need.
- 1:42Now, when you index the entire code base into essentially
- 1:46a curable graph, You can see this representation here.
- 1:51You essentially use like natural language to try to find the
- 1:54code base via a search. And then you essentially do this
- 1:57indexing ahead of time once so that Kako can leverage
- 2:01this code graph so that it can find a lot of these things
- 2:04under the hood without having to read all of this other
- 2:08files just to get to where it needs to go. Now the way to
- 2:11get this to work is essentially use this repo
- 2:14called CodeGraph. It's very easy to install. All you need to
- 2:18do is just copy this
- 2:20mpx command and then paste it in or just install using the
- 2:24npm global and then you just init the project here.
- 2:27So for example, we could do CodeGraph init-i and then it
- 2:30would initialize it. I already initialized it over here.
- 2:34And there's a bunch of things that you can do with this.
- 2:36So for example, you could do CodeGraph status
- 2:39and they will kind of give you a metadata about
- 2:42the CodeGraph, like what is in here, what kind of nodes
- 2:45are there, like import. Is there a component?
- 2:48You know, the common things that you kind of
- 2:50would need. And then as an example, we could do
- 2:53Curie. So let's do CodeGraph Curie LLM, for example.
- 2:57And then here it found two search results where it says it's
- 3:02HTML to text and active session.
- 3:05So essentially just use like semantic natural language to
- 3:08be able to find the relative code.
- 3:11Now you wouldn't normally
- 3:13use like
- 3:14this manually, but Clock Code essentially learns how to
- 3:18use this CLI tool under the hood and then it will do all of
- 3:22these like functions for you. Now, if you look at the docs,
- 3:26this is the CLI reference right here. It has all of these
- 3:28commands and then essentially Clock Code will read the
- 3:31CLI tool,
- 3:32understand how to use it and leverage the CLI to do
- 3:35the Curies. So that's the first strategy is essentially to
- 3:38create a graph.
- 3:40So that instead
- 3:42of Clock Code reading files one by one, kind of using grep,
- 3:46you do some work ahead of time.
- 3:48And then you index a bunch of things so you could quickly
- 3:50get to the file that you're looking for more semantically.
- 3:54Now this is a very common strategy and essentially how
- 3:56like a lot of search based products work, right?
- 4:00Or graph recommendations like in Meta. But there are
- 4:02trade-offs to this strategy. So the number one trade-off
- 4:06with this
- 4:07graph strategies, this pre-indexing strategy
- 4:10is that now there is two source of truth, which is that one,
- 4:15you have your code base as a source of truth. But now
- 4:18when Clock Code is looking and researching things,
- 4:21there is the index that is also the source of truth.
- 4:24That means that some system, ideally Clock Code,
- 4:27will have to spend time syncing.
- 4:30Now this repo actually has a method of syncing, but if you
- 4:33forget to do it, or it just hasn't been done since some time
- 4:37period that you've been coding, something could go out
- 4:40of sync. And then now Clock Code will confidently be
- 4:43wrong that, hey, this function may not exist, or this file
- 4:47doesn't exist. Or it does exist, even though it
- 4:49doesn't exist.
- 4:51So there's this tax of like having to keep it synced.
- 4:54The other kind of major trade-off, which is not as big of
- 4:58a problem, but essentially for compound engineering,
- 5:01when you're adding context files into big modules,
- 5:04for example, the way it systematically does code search,
- 5:08it may miss some of that. It might find the actual file that it
- 5:13needs to, but it might miss that module's Cloud.md
- 5:17that it's supposed to load normally. Because the way
- 5:20normal Cloud.md works
- 5:23is if you put one in some directory, if the code search
- 5:27goes into that file and reads that file, it'll actually go up
- 5:30and read that Cloud.md in that module. But because
- 5:33you're using a query version,
- 5:35it may miss it.
- 5:36So there is this inherent issue. And there's also
- 5:38probably some
- 5:40issue with accuracy, especially for
- 5:43deep dives. And at the end of the day, you still need to
- 5:46read the accuracy. Actual file once it files where the file is,
- 5:50right? So it does save a lot of tokens in the retrieval of
- 5:54the file, but you still will have to spend tokens on reading
- 5:57the file. But yeah, so those are kind of the
- 5:59high-level trade-offs. All right, strategy number two is to
- 6:03compress the outputs.
- 6:05So what does this mean? And I think this is a really
- 6:08good example. So if you take a look at this
- 6:10video animation,
- 6:12on the left side is essentially an NPM run. And in the
- 6:15NPM run, anyone who is programmed for a while knows
- 6:19that server logs or just CLI logs or just logs in general is
- 6:23very noisy. And oftentimes it's information that
- 6:27ClockCode doesn't actually need. So there's this open
- 6:30source library called RTK that actually takes a lot of these
- 6:34noisy logs and then compresses them. So on the right
- 6:37side is something like, oh, 43 test paths, rather than listing
- 6:41out all the 43 tests. And then like, it'll also like bundle all
- 6:44the warnings if it's suppressed them, for example.
- 6:47So there could be huge savings here. And before that,
- 6:50all of the tokens get read back by ClockCode, it just like
- 6:53shrinks all of it. So you can imagine how much savings it
- 6:56would have, especially
- 6:58if your workflow is heavy in logs or CLI usage, or like a
- 7:01bunch of like noisy things that ClockCode doesn't actually
- 7:04need to know. This open source library will save a ton
- 7:08of tokens. So as you can see, this is the open source
- 7:11library right here. And it's very easy to install as well. I just
- 7:15use homebrew and then you could just have it installed
- 7:18into like a specific directory or globally. And again,
- 7:22there's a bunch of ways you can like manually just trigger it
- 7:25using the CLI commands. And this is pretty
- 7:27common pattern. A lot of people have been realizing that
- 7:30CLIs are just kind of a better form factor than MCPs.
- 7:34And if they can help it, they would build a CLI over like an
- 7:38MCP because ClockCode is very good at juggling CLIs.
- 7:41Okay, so we'll do like a very simple example of doing like a
- 7:45Git log, for
- 7:46example. Git log.
- 7:47And you can see like, if you actually do Git log,
- 7:50there's so much, you know,
- 7:52information about the Git log, it's just it goes on for
- 7:55a while.
- 7:56So let's use RTK now.
- 7:59And then you can see that all of these things
- 8:01were compressed, and it shows how many lines it
- 8:03was omitted. And it just saves a ton of tokens. And again,
- 8:07like CodeGraph, it's ClockCode will leverage this
- 8:10tooling for you on your behalf. So you will get a ton
- 8:13of savings. But
- 8:14there is definitely a trade off here. And it's pretty obvious.
- 8:18And essentially, the trade off is accuracy, because the
- 8:22compression is lossy. By doing this method, you definitely
- 8:25have a chance of like dropping a very important
- 8:29log or some message that you are hoping for. And the
- 8:32system tries its best to, you know, compress
- 8:35and just remove things that it doesn't need.
- 8:37But if you're trying to debug bunch of server logs, it might
- 8:41be worth to turn this off when you're doing that
- 8:44specific work. So like, you should really know when to turn
- 8:47it on and when to turn it off. And at what point do you
- 8:50need the extra logs for the improved like self healing,
- 8:54for example, if you want to self correct, you may want to
- 8:57read the system logs to make sure that everything is
- 8:59happening step by step in the way that you want it. But in
- 9:02most of the time, you don't actually need all of the logs.
- 9:05So this is a really good strategy to compress a lot of the
- 9:09output before it comes to you. All right, strategy number
- 9:12three is,
- 9:13in fact, my favorite strategy.
- 9:16And honestly, I think they misnamed this, but essentially is
- 9:19to use caveman, which is another open source library.
- 9:23Now at a high level,
- 9:25it's to get cloud code to talk
- 9:28less.
- 9:29So in this diagram, you can see on the left side, there's a
- 9:32bunch of like long outputs on the right side is a very like
- 9:35aggressive caveman example.
- 9:38And
- 9:38it will just compress everything. And I want to show this
- 9:42video because I think it's the most accurate assessment
- 9:45to what this is doing.
- 9:47I waste time say lot word when few word do track.
- 9:50That's probably one of my favorite clips from the office.
- 9:53The office was ahead of its times, honestly.
- 9:57And you know, Kevin was ahead of this time, and definitely
- 10:01a clock code Maxer. But yeah, caveman, like I said,
- 10:04probably should have been named Kevin, in my opinion,
- 10:06but at a high level, it's really that it's just to shorten the
- 10:10amount of words that clock code says to basically say the
- 10:14same thing. It's very easy to install cave man, you can just
- 10:17do NPM install here, just copy this command and paste it
- 10:20into your bash command. There's many different ways.
- 10:23And essentially, it's a skill that sets
- 10:26your clock code session into a specific mode. So there's
- 10:29like light full ultra, and then when
- 10:33essentially it's per session, so you could remove it at any
- 10:36time or like at it. So it really depends. So let's give it a shot
- 10:40real quick. Tell me about this project. Alright, so on the left
- 10:43is without caveman. That's the caveman Ultra. And then
- 10:46here is came in and saw and it's ready for the next project.
- 10:49Tell me about this project. Alright, so as you can see
- 10:52right here, it's about like half of the output. And I would
- 10:56say you still get about like the same quality.
- 10:59It's pretty accurate. So that's caveman. I use it quite a bit.
- 11:02I keep it on ultra
- 11:04until I like need to get away from a specific trade off.
- 11:07And we're gonna go into the trade off now. So the obvious
- 11:10trade off here is that yes, you save tokens, but the
- 11:14correctness is at risk here. Because essentially, clock code
- 11:18really needs to leverage this like feedback loop of
- 11:21messages in and out to get as much context that it needs
- 11:25to answer things questions better. But when you
- 11:28essentially shrink the outputs dramatically,
- 11:30those messages get resent back to clock code, right?
- 11:34The way the context window works is the context that you
- 11:37have essentially clock code leverages it in this
- 11:40agentic loop. And that's how it knows what you've been
- 11:43working on. So your context history that session if the
- 11:46quality is bad, then the output may lead to like
- 11:50bad answers. Essentially, I'm working on a video right now
- 11:53on system design for a clock code. If you're interested in
- 11:56that subscribe,
- 11:57I think it's very important to understand like
- 12:00system design, system design is probably the most
- 12:02important thing in this new world of AI. So I'm doing a
- 12:06whole series
- 12:08on system design for like agentic tools agentic,
- 12:11like building chat GPT, building clock code, building like a
- 12:14rack system, like all of that kind of stuff. So if
- 12:17you're interested, don't forget to subscribe
- 12:19to this channel and my newsletter. Alright, the final
- 12:21strategy is kind of boring, but it's good old just managing
- 12:26the clock code session properly using like compact,
- 12:29changing the models, the docs have pretty good
- 12:32recommendations on managing it. Here is kind of the
- 12:36main ways to main built in ways to save your
- 12:41tokens. Alright, so obviously, there is
- 12:44such context, right here. And this kind of shows you
- 12:47what's in the current context. And I look at this every now
- 12:51and then to make sure that my
- 12:52usage just makes sense. You know, sometimes if you have
- 12:56like a really large cloud.md, you won't realize it until you
- 13:00look at it. And every time you turn on the quad code,
- 13:03it's using like 40k tokens, and you're like, what is going on?
- 13:06How is it that every time. I just opened my session is
- 13:09using 4%? Well, it could just be that you have a 50k
- 13:13token cloud.md.
- 13:15It's like a pretty small cloud.md that like just is already
- 13:19indexed and like points to different
- 13:21things. So this is a really good way to kind of audit what's
- 13:25using the most tokens. And sometimes you'll find that like
- 13:29MCPs end up using a lot in this case, it's not there's a
- 13:32bunch of things that's loaded, it's ready to go. And these
- 13:35are my MCPs that I have. But essentially, this is a way to
- 13:39debug the current context. There's also slash clear,
- 13:42which clears
- 13:43the context. And.
- 13:45I will say that if you're working on like new tasks, let's say
- 13:48you have a you just finished a task and you want to do
- 13:51another new task, then it's good to clear the task if you
- 13:54need a previous context or not. There's also slash model
- 13:57to switch the models. And some people ask me if like
- 14:01haiku or sauna is ever worth using. And my answer is
- 14:04100% Yes, I think one of the best use cases of haiku is
- 14:09actually using it in slash Chrome, or just navigating like
- 14:13Chrome haiku does a really good job. And it's actually
- 14:15better in most use cases, because it's faster to navigate
- 14:19because the model itself is fast. For sauna, I actually
- 14:22leverage sauna for a lot of like my scheduling jobs. If I have
- 14:26like a slash schedule, I will try to see if that schedule job
- 14:31that I want to run every like day or whatever, let's say I,
- 14:35I have one that like just cleans up my desktop, like it's just
- 14:40like looks at my desktop and moves everything into
- 14:43folders that I could probably just use sauna or haiku.
- 14:46Because I know that that specific thing can be done
- 14:49with sauna. If I know that that task
- 14:52is repeatedly successful, we're using sauna or haiku,
- 14:55I'll always end up using that. So don't forget, try switching
- 14:59the models. And you don't have to be on the million
- 15:024.7 context. Now, pro tip here is that if you're doing any
- 15:06significant programming tasks,
- 15:08I do recommend just being on
- 15:10the the most, like reliable, especially for planning and
- 15:14doing deep dives. The quality self, in my opinion, is a token
- 15:17savings versus like trying to save money with sauna and
- 15:21getting a bad quality result. And the last few things is I
- 15:25would recommend using plan
- 15:27mode first to plan out an execution rather than just
- 15:31like having
- 15:32clock code and go do just a bunch of things. And then for
- 15:35programming in general, I would say use like
- 15:38pencil or Figma
- 15:40to design something, run off designs,
- 15:43rather than just like coding something right.
- 15:45So yeah, so those are kind of the high level tips, I would
- 15:48say that I have for just managing your
- 15:51session. Now, a lot of what I said actually is in this stock,
- 15:56I'll link this below. But you know, clock code also has like
- 15:59advanced strategies. Most of this is already like kind of
- 16:02covered in my talk, not the repos, they won't cover the
- 16:06open source projects, but more about like the
- 16:08strategies that. I just mentioned about managing
- 16:10your session. So feel free to take a look at that.
- 16:13For example,
- 16:15they said agent teams cost a lot,
- 16:17which makes sense. So should you use all of these things?
- 16:21And in my opinion, I think the answer is yes.
- 16:24But again, there's trade offs.
- 16:27And in my opinion,
- 16:28the most the largest trade off here is really costs
- 16:33versus
- 16:33quality. And also just like the
- 16:38quality and and and also the complexity, there's a
- 16:40complexity aspect, right? There's
- 16:43the code graph, it can get stale.
- 16:45The RTX proxy is lossy for sure. So you're straight up just
- 16:48missing messages. And sometimes you'll miss a log or
- 16:51something that you need. And caveman may be over
- 16:54trimming things so that
- 16:56the message history is not great. But
- 16:59it all of these things dramatically save the tokens.
- 17:03So really, it's you understanding the how to use these
- 17:07tools and essentially being like, oh, okay, in this
- 17:10particular application, I may not need all of the logs. So I
- 17:14could just have RTX be on. And then oh, in this case,
- 17:17caveman is great, because I don't really need it right now.
- 17:21I don't need this. I'm not doing like a crazy plan mode
- 17:24right now or something like that.
- 17:26So yeah, but all of these things add complexities,
- 17:28add multiple layers, multiple things for you to juggle
- 17:31multiple failure points, essentially for clock code, right?
- 17:35So yeah, you got your own risk,
- 17:38but I guarantee you, you will save
- 17:40tokens
- 17:41using these strategies. So yeah, I hope you enjoy
- 17:44this video. I make a ton of videos on clock code and AI
- 17:48coding agent decoding. So feel free to check them
- 17:51out here.
- 17:52And don't forget to subscribe to my newsletter, I
- 17:55have a bunch of content there that I don't really post here,
- 17:59I actually also just launched a new course with bye bye go.
- 18:02And I'll be teaching you over there. So if you're interested,
- 18:05it is a paid live course. It's a two day thing, you could build
- 18:09a bunch of portfolio projects.
- 18:11And I teach you basically everything I know about clock
- 18:14code. So if you're interested, sign up to a bye bye goes
- 18:17newsletter or look at this website, the course may or may
- 18:20not have already happened. So if you missed it, you might
- 18:24have to join the next course. But yeah, I hope you guys
- 18:26enjoyed this video. And until I see you guys on the
- 18:29next one.
About this transcript
This page contains the full transcript of How To Save 90% of Claude Code Token Usage by John Kim, generated from the public captions YouTube serves with the video. The transcript has 3,396 words across 366 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.