YouTube2Text

How To Save 90% of Claude Code Token Usage — Transcript

by John Kim · 3,396 words · 366 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Everyone how's it going in this video i'm going to teach
  2. 0:02you how to reduce your cloud code token
  3. 0:05usage by up to 90
  4. 0:07in
  5. 0:08four different strategies it's easy and free to do and you
  6. 0:11can do it right away and in this video i'm going to show
  7. 0:13you step by step and also a quick comparison on like
  8. 0:16visually being able to see difference in tokens that you'll
  9. 0:19be ending up using now all of these strategies actually do
  10. 0:22have some trade-offs and i'll also go over those trade-offs
  11. 0:26as well so be sure to stick around through the entire video
  12. 0:29so you don't miss out on exactly how to use these things
  13. 0:31properly as usual i made some slides so there's some
  14. 0:34visual representation of what i'm talking about so let's get
  15. 0:37right into the four different strategies now the first
  16. 0:40strategy is actually to index the code essentially we're
  17. 0:43going to create an index of the code now if you don't
  18. 0:46know what indexing is it's essentially like search actually is
  19. 0:52a very famous usage of an index. Essentially what Google
  20. 0:56does behind the scenes is it creates a map of like
  21. 0:59keywords that point to different websites as a very
  22. 1:02dumbed down example. So the core idea is that we want
  23. 1:06to essentially create a graph map of our code base.
  24. 1:09So here's a visual representation. Now on the left side,
  25. 1:12it's kind of how clock code already does
  26. 1:15the grep and reading and scanning of the code base.
  27. 1:18It will go one by one and deep dive into the code base.
  28. 1:21If you're asking about some particular, you know,
  29. 1:24server implementation or some validations like file or an
  30. 1:28auth file, we'll try to go and grep it using regular file grep
  31. 1:31and they will go and find it and load those into the
  32. 1:34token context. So it ends up spending a lot of time
  33. 1:37because sometimes to find the files that it's looking for,
  34. 1:40it would have read a lot of things that it doesn't need.
  35. 1:42Now, when you index the entire code base into essentially
  36. 1:46a curable graph, You can see this representation here.
  37. 1:51You essentially use like natural language to try to find the
  38. 1:54code base via a search. And then you essentially do this
  39. 1:57indexing ahead of time once so that Kako can leverage
  40. 2:01this code graph so that it can find a lot of these things
  41. 2:04under the hood without having to read all of this other
  42. 2:08files just to get to where it needs to go. Now the way to
  43. 2:11get this to work is essentially use this repo
  44. 2:14called CodeGraph. It's very easy to install. All you need to
  45. 2:18do is just copy this
  46. 2:20mpx command and then paste it in or just install using the
  47. 2:24npm global and then you just init the project here.
  48. 2:27So for example, we could do CodeGraph init-i and then it
  49. 2:30would initialize it. I already initialized it over here.
  50. 2:34And there's a bunch of things that you can do with this.
  51. 2:36So for example, you could do CodeGraph status
  52. 2:39and they will kind of give you a metadata about
  53. 2:42the CodeGraph, like what is in here, what kind of nodes
  54. 2:45are there, like import. Is there a component?
  55. 2:48You know, the common things that you kind of
  56. 2:50would need. And then as an example, we could do
  57. 2:53Curie. So let's do CodeGraph Curie LLM, for example.
  58. 2:57And then here it found two search results where it says it's
  59. 3:02HTML to text and active session.
  60. 3:05So essentially just use like semantic natural language to
  61. 3:08be able to find the relative code.
  62. 3:11Now you wouldn't normally
  63. 3:13use like
  64. 3:14this manually, but Clock Code essentially learns how to
  65. 3:18use this CLI tool under the hood and then it will do all of
  66. 3:22these like functions for you. Now, if you look at the docs,
  67. 3:26this is the CLI reference right here. It has all of these
  68. 3:28commands and then essentially Clock Code will read the
  69. 3:31CLI tool,
  70. 3:32understand how to use it and leverage the CLI to do
  71. 3:35the Curies. So that's the first strategy is essentially to
  72. 3:38create a graph.
  73. 3:40So that instead
  74. 3:42of Clock Code reading files one by one, kind of using grep,
  75. 3:46you do some work ahead of time.
  76. 3:48And then you index a bunch of things so you could quickly
  77. 3:50get to the file that you're looking for more semantically.
  78. 3:54Now this is a very common strategy and essentially how
  79. 3:56like a lot of search based products work, right?
  80. 4:00Or graph recommendations like in Meta. But there are
  81. 4:02trade-offs to this strategy. So the number one trade-off
  82. 4:06with this
  83. 4:07graph strategies, this pre-indexing strategy
  84. 4:10is that now there is two source of truth, which is that one,
  85. 4:15you have your code base as a source of truth. But now
  86. 4:18when Clock Code is looking and researching things,
  87. 4:21there is the index that is also the source of truth.
  88. 4:24That means that some system, ideally Clock Code,
  89. 4:27will have to spend time syncing.
  90. 4:30Now this repo actually has a method of syncing, but if you
  91. 4:33forget to do it, or it just hasn't been done since some time
  92. 4:37period that you've been coding, something could go out
  93. 4:40of sync. And then now Clock Code will confidently be
  94. 4:43wrong that, hey, this function may not exist, or this file
  95. 4:47doesn't exist. Or it does exist, even though it
  96. 4:49doesn't exist.
  97. 4:51So there's this tax of like having to keep it synced.
  98. 4:54The other kind of major trade-off, which is not as big of
  99. 4:58a problem, but essentially for compound engineering,
  100. 5:01when you're adding context files into big modules,
  101. 5:04for example, the way it systematically does code search,
  102. 5:08it may miss some of that. It might find the actual file that it
  103. 5:13needs to, but it might miss that module's Cloud.md
  104. 5:17that it's supposed to load normally. Because the way
  105. 5:20normal Cloud.md works
  106. 5:23is if you put one in some directory, if the code search
  107. 5:27goes into that file and reads that file, it'll actually go up
  108. 5:30and read that Cloud.md in that module. But because
  109. 5:33you're using a query version,
  110. 5:35it may miss it.
  111. 5:36So there is this inherent issue. And there's also
  112. 5:38probably some
  113. 5:40issue with accuracy, especially for
  114. 5:43deep dives. And at the end of the day, you still need to
  115. 5:46read the accuracy. Actual file once it files where the file is,
  116. 5:50right? So it does save a lot of tokens in the retrieval of
  117. 5:54the file, but you still will have to spend tokens on reading
  118. 5:57the file. But yeah, so those are kind of the
  119. 5:59high-level trade-offs. All right, strategy number two is to
  120. 6:03compress the outputs.
  121. 6:05So what does this mean? And I think this is a really
  122. 6:08good example. So if you take a look at this
  123. 6:10video animation,
  124. 6:12on the left side is essentially an NPM run. And in the
  125. 6:15NPM run, anyone who is programmed for a while knows
  126. 6:19that server logs or just CLI logs or just logs in general is
  127. 6:23very noisy. And oftentimes it's information that
  128. 6:27ClockCode doesn't actually need. So there's this open
  129. 6:30source library called RTK that actually takes a lot of these
  130. 6:34noisy logs and then compresses them. So on the right
  131. 6:37side is something like, oh, 43 test paths, rather than listing
  132. 6:41out all the 43 tests. And then like, it'll also like bundle all
  133. 6:44the warnings if it's suppressed them, for example.
  134. 6:47So there could be huge savings here. And before that,
  135. 6:50all of the tokens get read back by ClockCode, it just like
  136. 6:53shrinks all of it. So you can imagine how much savings it
  137. 6:56would have, especially
  138. 6:58if your workflow is heavy in logs or CLI usage, or like a
  139. 7:01bunch of like noisy things that ClockCode doesn't actually
  140. 7:04need to know. This open source library will save a ton
  141. 7:08of tokens. So as you can see, this is the open source
  142. 7:11library right here. And it's very easy to install as well. I just
  143. 7:15use homebrew and then you could just have it installed
  144. 7:18into like a specific directory or globally. And again,
  145. 7:22there's a bunch of ways you can like manually just trigger it
  146. 7:25using the CLI commands. And this is pretty
  147. 7:27common pattern. A lot of people have been realizing that
  148. 7:30CLIs are just kind of a better form factor than MCPs.
  149. 7:34And if they can help it, they would build a CLI over like an
  150. 7:38MCP because ClockCode is very good at juggling CLIs.
  151. 7:41Okay, so we'll do like a very simple example of doing like a
  152. 7:45Git log, for
  153. 7:46example. Git log.
  154. 7:47And you can see like, if you actually do Git log,
  155. 7:50there's so much, you know,
  156. 7:52information about the Git log, it's just it goes on for
  157. 7:55a while.
  158. 7:56So let's use RTK now.
  159. 7:59And then you can see that all of these things
  160. 8:01were compressed, and it shows how many lines it
  161. 8:03was omitted. And it just saves a ton of tokens. And again,
  162. 8:07like CodeGraph, it's ClockCode will leverage this
  163. 8:10tooling for you on your behalf. So you will get a ton
  164. 8:13of savings. But
  165. 8:14there is definitely a trade off here. And it's pretty obvious.
  166. 8:18And essentially, the trade off is accuracy, because the
  167. 8:22compression is lossy. By doing this method, you definitely
  168. 8:25have a chance of like dropping a very important
  169. 8:29log or some message that you are hoping for. And the
  170. 8:32system tries its best to, you know, compress
  171. 8:35and just remove things that it doesn't need.
  172. 8:37But if you're trying to debug bunch of server logs, it might
  173. 8:41be worth to turn this off when you're doing that
  174. 8:44specific work. So like, you should really know when to turn
  175. 8:47it on and when to turn it off. And at what point do you
  176. 8:50need the extra logs for the improved like self healing,
  177. 8:54for example, if you want to self correct, you may want to
  178. 8:57read the system logs to make sure that everything is
  179. 8:59happening step by step in the way that you want it. But in
  180. 9:02most of the time, you don't actually need all of the logs.
  181. 9:05So this is a really good strategy to compress a lot of the
  182. 9:09output before it comes to you. All right, strategy number
  183. 9:12three is,
  184. 9:13in fact, my favorite strategy.
  185. 9:16And honestly, I think they misnamed this, but essentially is
  186. 9:19to use caveman, which is another open source library.
  187. 9:23Now at a high level,
  188. 9:25it's to get cloud code to talk
  189. 9:28less.
  190. 9:29So in this diagram, you can see on the left side, there's a
  191. 9:32bunch of like long outputs on the right side is a very like
  192. 9:35aggressive caveman example.
  193. 9:38And
  194. 9:38it will just compress everything. And I want to show this
  195. 9:42video because I think it's the most accurate assessment
  196. 9:45to what this is doing.
  197. 9:47I waste time say lot word when few word do track.
  198. 9:50That's probably one of my favorite clips from the office.
  199. 9:53The office was ahead of its times, honestly.
  200. 9:57And you know, Kevin was ahead of this time, and definitely
  201. 10:01a clock code Maxer. But yeah, caveman, like I said,
  202. 10:04probably should have been named Kevin, in my opinion,
  203. 10:06but at a high level, it's really that it's just to shorten the
  204. 10:10amount of words that clock code says to basically say the
  205. 10:14same thing. It's very easy to install cave man, you can just
  206. 10:17do NPM install here, just copy this command and paste it
  207. 10:20into your bash command. There's many different ways.
  208. 10:23And essentially, it's a skill that sets
  209. 10:26your clock code session into a specific mode. So there's
  210. 10:29like light full ultra, and then when
  211. 10:33essentially it's per session, so you could remove it at any
  212. 10:36time or like at it. So it really depends. So let's give it a shot
  213. 10:40real quick. Tell me about this project. Alright, so on the left
  214. 10:43is without caveman. That's the caveman Ultra. And then
  215. 10:46here is came in and saw and it's ready for the next project.
  216. 10:49Tell me about this project. Alright, so as you can see
  217. 10:52right here, it's about like half of the output. And I would
  218. 10:56say you still get about like the same quality.
  219. 10:59It's pretty accurate. So that's caveman. I use it quite a bit.
  220. 11:02I keep it on ultra
  221. 11:04until I like need to get away from a specific trade off.
  222. 11:07And we're gonna go into the trade off now. So the obvious
  223. 11:10trade off here is that yes, you save tokens, but the
  224. 11:14correctness is at risk here. Because essentially, clock code
  225. 11:18really needs to leverage this like feedback loop of
  226. 11:21messages in and out to get as much context that it needs
  227. 11:25to answer things questions better. But when you
  228. 11:28essentially shrink the outputs dramatically,
  229. 11:30those messages get resent back to clock code, right?
  230. 11:34The way the context window works is the context that you
  231. 11:37have essentially clock code leverages it in this
  232. 11:40agentic loop. And that's how it knows what you've been
  233. 11:43working on. So your context history that session if the
  234. 11:46quality is bad, then the output may lead to like
  235. 11:50bad answers. Essentially, I'm working on a video right now
  236. 11:53on system design for a clock code. If you're interested in
  237. 11:56that subscribe,
  238. 11:57I think it's very important to understand like
  239. 12:00system design, system design is probably the most
  240. 12:02important thing in this new world of AI. So I'm doing a
  241. 12:06whole series
  242. 12:08on system design for like agentic tools agentic,
  243. 12:11like building chat GPT, building clock code, building like a
  244. 12:14rack system, like all of that kind of stuff. So if
  245. 12:17you're interested, don't forget to subscribe
  246. 12:19to this channel and my newsletter. Alright, the final
  247. 12:21strategy is kind of boring, but it's good old just managing
  248. 12:26the clock code session properly using like compact,
  249. 12:29changing the models, the docs have pretty good
  250. 12:32recommendations on managing it. Here is kind of the
  251. 12:36main ways to main built in ways to save your
  252. 12:41tokens. Alright, so obviously, there is
  253. 12:44such context, right here. And this kind of shows you
  254. 12:47what's in the current context. And I look at this every now
  255. 12:51and then to make sure that my
  256. 12:52usage just makes sense. You know, sometimes if you have
  257. 12:56like a really large cloud.md, you won't realize it until you
  258. 13:00look at it. And every time you turn on the quad code,
  259. 13:03it's using like 40k tokens, and you're like, what is going on?
  260. 13:06How is it that every time. I just opened my session is
  261. 13:09using 4%? Well, it could just be that you have a 50k
  262. 13:13token cloud.md.
  263. 13:15It's like a pretty small cloud.md that like just is already
  264. 13:19indexed and like points to different
  265. 13:21things. So this is a really good way to kind of audit what's
  266. 13:25using the most tokens. And sometimes you'll find that like
  267. 13:29MCPs end up using a lot in this case, it's not there's a
  268. 13:32bunch of things that's loaded, it's ready to go. And these
  269. 13:35are my MCPs that I have. But essentially, this is a way to
  270. 13:39debug the current context. There's also slash clear,
  271. 13:42which clears
  272. 13:43the context. And.
  273. 13:45I will say that if you're working on like new tasks, let's say
  274. 13:48you have a you just finished a task and you want to do
  275. 13:51another new task, then it's good to clear the task if you
  276. 13:54need a previous context or not. There's also slash model
  277. 13:57to switch the models. And some people ask me if like
  278. 14:01haiku or sauna is ever worth using. And my answer is
  279. 14:04100% Yes, I think one of the best use cases of haiku is
  280. 14:09actually using it in slash Chrome, or just navigating like
  281. 14:13Chrome haiku does a really good job. And it's actually
  282. 14:15better in most use cases, because it's faster to navigate
  283. 14:19because the model itself is fast. For sauna, I actually
  284. 14:22leverage sauna for a lot of like my scheduling jobs. If I have
  285. 14:26like a slash schedule, I will try to see if that schedule job
  286. 14:31that I want to run every like day or whatever, let's say I,
  287. 14:35I have one that like just cleans up my desktop, like it's just
  288. 14:40like looks at my desktop and moves everything into
  289. 14:43folders that I could probably just use sauna or haiku.
  290. 14:46Because I know that that specific thing can be done
  291. 14:49with sauna. If I know that that task
  292. 14:52is repeatedly successful, we're using sauna or haiku,
  293. 14:55I'll always end up using that. So don't forget, try switching
  294. 14:59the models. And you don't have to be on the million
  295. 15:024.7 context. Now, pro tip here is that if you're doing any
  296. 15:06significant programming tasks,
  297. 15:08I do recommend just being on
  298. 15:10the the most, like reliable, especially for planning and
  299. 15:14doing deep dives. The quality self, in my opinion, is a token
  300. 15:17savings versus like trying to save money with sauna and
  301. 15:21getting a bad quality result. And the last few things is I
  302. 15:25would recommend using plan
  303. 15:27mode first to plan out an execution rather than just
  304. 15:31like having
  305. 15:32clock code and go do just a bunch of things. And then for
  306. 15:35programming in general, I would say use like
  307. 15:38pencil or Figma
  308. 15:40to design something, run off designs,
  309. 15:43rather than just like coding something right.
  310. 15:45So yeah, so those are kind of the high level tips, I would
  311. 15:48say that I have for just managing your
  312. 15:51session. Now, a lot of what I said actually is in this stock,
  313. 15:56I'll link this below. But you know, clock code also has like
  314. 15:59advanced strategies. Most of this is already like kind of
  315. 16:02covered in my talk, not the repos, they won't cover the
  316. 16:06open source projects, but more about like the
  317. 16:08strategies that. I just mentioned about managing
  318. 16:10your session. So feel free to take a look at that.
  319. 16:13For example,
  320. 16:15they said agent teams cost a lot,
  321. 16:17which makes sense. So should you use all of these things?
  322. 16:21And in my opinion, I think the answer is yes.
  323. 16:24But again, there's trade offs.
  324. 16:27And in my opinion,
  325. 16:28the most the largest trade off here is really costs
  326. 16:33versus
  327. 16:33quality. And also just like the
  328. 16:38quality and and and also the complexity, there's a
  329. 16:40complexity aspect, right? There's
  330. 16:43the code graph, it can get stale.
  331. 16:45The RTX proxy is lossy for sure. So you're straight up just
  332. 16:48missing messages. And sometimes you'll miss a log or
  333. 16:51something that you need. And caveman may be over
  334. 16:54trimming things so that
  335. 16:56the message history is not great. But
  336. 16:59it all of these things dramatically save the tokens.
  337. 17:03So really, it's you understanding the how to use these
  338. 17:07tools and essentially being like, oh, okay, in this
  339. 17:10particular application, I may not need all of the logs. So I
  340. 17:14could just have RTX be on. And then oh, in this case,
  341. 17:17caveman is great, because I don't really need it right now.
  342. 17:21I don't need this. I'm not doing like a crazy plan mode
  343. 17:24right now or something like that.
  344. 17:26So yeah, but all of these things add complexities,
  345. 17:28add multiple layers, multiple things for you to juggle
  346. 17:31multiple failure points, essentially for clock code, right?
  347. 17:35So yeah, you got your own risk,
  348. 17:38but I guarantee you, you will save
  349. 17:40tokens
  350. 17:41using these strategies. So yeah, I hope you enjoy
  351. 17:44this video. I make a ton of videos on clock code and AI
  352. 17:48coding agent decoding. So feel free to check them
  353. 17:51out here.
  354. 17:52And don't forget to subscribe to my newsletter, I
  355. 17:55have a bunch of content there that I don't really post here,
  356. 17:59I actually also just launched a new course with bye bye go.
  357. 18:02And I'll be teaching you over there. So if you're interested,
  358. 18:05it is a paid live course. It's a two day thing, you could build
  359. 18:09a bunch of portfolio projects.
  360. 18:11And I teach you basically everything I know about clock
  361. 18:14code. So if you're interested, sign up to a bye bye goes
  362. 18:17newsletter or look at this website, the course may or may
  363. 18:20not have already happened. So if you missed it, you might
  364. 18:24have to join the next course. But yeah, I hope you guys
  365. 18:26enjoyed this video. And until I see you guys on the
  366. 18:29next one.

About this transcript

This page contains the full transcript of How To Save 90% of Claude Code Token Usage by John Kim, generated from the public captions YouTube serves with the video. The transcript has 3,396 words across 366 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.