YouTube2Text

Stop Prompting Claude. Use Karpathy's Method Instead. — Transcript

by Austin Marchese · 2,909 words · 425 segments · language en · Watch on YouTube

Full transcript

  1. 0:00I just listened to Andrej Karpathy speak
  2. 0:01at AISN 2026, and I learned something
  3. 0:04that I wasn't expecting. Almost everyone
  4. 0:06is prompting Claude wrong. So, I decided
  5. 0:08to dig deeper and see exactly how
  6. 0:09Karpathy, the former head of AI at
  7. 0:11Tesla, uses AI in 2026. And it turns out
  8. 0:14that Karpathy's method for building 10
  9. 0:16times faster can be broken down into
  10. 0:18three simple layers. So, in today's
  11. 0:20video, I'll be breaking down each layer
  12. 0:22so that anybody can apply them. And then
  13. 0:23I'll show you the one thing that
  14. 0:25Karpathy said to focus on in the age of
  15. 0:27AI. So, layer one is the spec. AI models
  16. 0:29are incredibly smart, but they're still
  17. 0:31missing something. To showcase their
  18. 0:33current limitation, Karpathy explained a
  19. 0:35simple question AI will get wrong.
  20. 0:37>> I want to go to a car wash to wash my
  21. 0:39car, and it's 50 m away. Should I drive
  22. 0:42or should I walk? And state-of-the-art
  23. 0:45models today will tell you to walk
  24. 0:46because it's so close.
  25. 0:48>> At first, I actually didn't believe
  26. 0:49this, so I went to Claude, Gemini, Grok,
  27. 0:51and ChatGPT, asked them the same
  28. 0:52question, and they all gave me the same
  29. 0:54answer. And it reveals the whole
  30. 0:55foundation of this video. AI is
  31. 0:57brilliant at what can be measured, but
  32. 1:00for context-driven things like needing a
  33. 1:02car for a car wash, it has no signal to
  34. 1:05act on. So, how do you bridge this gap
  35. 1:07between your understanding and your
  36. 1:08contextual information and AI's
  37. 1:10computational power? That's where the
  38. 1:11spec comes in. And a spec is how you
  39. 1:13deliver your understanding to Claude in
  40. 1:15a format it can use. A term you may have
  41. 1:17heard is Claude's plan mode, which
  42. 1:19essentially can be used to help you
  43. 1:20create a plan before building anything.
  44. 1:22But Karpathy thinks that this is too
  45. 1:25high-level.
  46. 1:25>> I actually don't even like the plan
  47. 1:27mode. I I would
  48. 1:28I mean, obviously it's very useful, but
  49. 1:30I think there's something more general
  50. 1:31here where you have to work with your
  51. 1:32agent to design a spec that is very
  52. 1:34detailed.
  53. 1:34>> Now, Karpathy isn't telling you that
  54. 1:36plan mode is bad. What he's actually
  55. 1:38saying is you have to go deeper, work
  56. 1:40with these AI tools to design the actual
  57. 1:42spec. So, how do you create a spec that
  58. 1:44Claude can successfully use to build
  59. 1:46what you're trying to build? The first
  60. 1:48step is you have to uncover your goal.
  61. 1:50If you just say, "Create a end-of-month
  62. 1:52report," that's a task, but the actual
  63. 1:54goal is a conclusion you're trying to
  64. 1:56draw, the decision the report drives.
  65. 1:59And what the goal actually is is
  66. 2:00something AI will literally never be
  67. 2:02able to decide. So, to help you do this,
  68. 2:04we'll tell Claude to interview me to
  69. 2:07identify the goal of this project. This
  70. 2:09is the way to get the information out of
  71. 2:11you and into the spec. Now, step two is
  72. 2:13be agile with how you work. There are
  73. 2:15two methods of completing any task. The
  74. 2:17first is waterfall, and the other is
  75. 2:19agile. Waterfall is you take a big task
  76. 2:22and you complete the entire thing, and
  77. 2:24then you show the final product. Agile
  78. 2:26on the other hand is you break that same
  79. 2:27task into small buckets, and you show
  80. 2:30the result throughout the entire process
  81. 2:32to make sure you're going in the right
  82. 2:33direction. And people are extremely
  83. 2:35susceptible to using AI agents in a
  84. 2:37waterfall manner because they want to
  85. 2:38give them everything to do at once. The
  86. 2:41better move is agile specking. You want
  87. 2:43to have a tight scope, a clear
  88. 2:45checkpoint, you want to review the
  89. 2:46output, adjust it, and then repeat. To
  90. 2:48help with this, we'll tell Claude to
  91. 2:50bias towards smaller and more
  92. 2:51compartmentalized specs. Step three is
  93. 2:54you want to be precise and use your
  94. 2:55brain. The more precise you are, the
  95. 2:57less AI has to assume. And every
  96. 2:59assumption that AI makes is a chance for
  97. 3:01it to drift from the final product you
  98. 3:03actually want. And when you have AI
  99. 3:05create a spec for you, you have to use
  100. 3:07your brain to think critically about
  101. 3:09what that spec actually says. So, to
  102. 3:12help you use your brain, you can say
  103. 3:13"Make me verify key decisions explicitly
  104. 3:16to ensure nothing is missed." And when
  105. 3:18you put these three pieces together, we
  106. 3:20have a final prompt we can use in Claude
  107. 3:22to help create a tightly scoped,
  108. 3:24well-thought-out
  109. 3:26aligns with our actual goal. This is a
  110. 3:27process that I call modern engineering,
  111. 3:29which every successful person has to
  112. 3:31become. Now, layer two is the verifier.
  113. 3:34Layer two sits on top of the spec. This
  114. 3:36is the verification process. One of the
  115. 3:38most frustrating things about AI is
  116. 3:40reviewing and verifying the output. And
  117. 3:42unlike a human, it can't grasp
  118. 3:44non-measurable things. So, how can we
  119. 3:47help AI verify its own outputs? Well,
  120. 3:49first you need to understand the mental
  121. 3:50model behind this. And Karpathy explains
  122. 3:53it as animals versus ghosts. Here's him
  123. 3:55getting asked a question about this in a
  124. 3:57recent interview. And if it sounds
  125. 3:59confusing, don't worry, I will simplify
  126. 4:00it after.
  127. 4:01>> And the idea is that we're not building
  128. 4:03animals, we are summoning ghosts. Why
  129. 4:05does that framing matter? And what does
  130. 4:07it actually change about how you build
  131. 4:09and deploy and evaluate or even trust
  132. 4:12them?
  133. 4:12>> Yeah, I think the reason I wrote about
  134. 4:14this is because I'm trying to wrap my
  135. 4:15head around what these things are,
  136. 4:16right? Because if you have a good model
  137. 4:18of what they are or are not, then you're
  138. 4:19going to be more competent at uh using
  139. 4:21them. I think it's just um
  140. 4:23coming to terms with the fact that these
  141. 4:24things are not, you know, animal
  142. 4:26intelligences. Like if you yell at them,
  143. 4:27they're not going to work better or
  144. 4:29worse or doesn't have any impact. Um
  145. 4:32and uh
  146. 4:33it's all just kind of like these
  147. 4:34statistical simulation circuits. It's
  148. 4:36more just being suspicious of it and um
  149. 4:38figuring it out over time.
  150. 4:39>> Now, that's some gigabrain stuff, but
  151. 4:41let me simplify it. People, me and you,
  152. 4:43are used to interacting with people,
  153. 4:45which Karpathy is calling animals. These
  154. 4:48animals are driven by different
  155. 4:49motivators and emotions, which help
  156. 4:51produce the final product and output
  157. 4:53within a team setting. And if you say to
  158. 4:55a person, become an expert at SEO
  159. 4:57marketing in the next 14 days or you're
  160. 4:59fired, they're going to figure it out.
  161. 5:01That's because they have these intrinsic
  162. 5:03motivations. But AI is not that.
  163. 5:05Karpathy describes it as a ghost, but in
  164. 5:07my eyes that's a little too confusing,
  165. 5:09so throw it out the window. Instead,
  166. 5:10think of it like a robot librarian. If
  167. 5:12you ask it that same SEO question, the
  168. 5:14librarian will only suggest resources
  169. 5:17and answers based on the books in its
  170. 5:19library. If it doesn't have a book, it
  171. 5:21can't help you. And part of the
  172. 5:22challenge here is that the librarian
  173. 5:24doesn't know when it's missing a
  174. 5:26specific book. So, it may just
  175. 5:28confidently make something up. And
  176. 5:29that's what's happening when AI nails
  177. 5:31math and fumbles things with context.
  178. 5:33It's brilliant because the library has
  179. 5:35the clear answers. But if it doesn't,
  180. 5:37then it's confidently wrong or
  181. 5:40uncertain. Which means interacting with
  182. 5:42it like it's an animal, i.e. a human,
  183. 5:44doesn't help, right? Yelling at it,
  184. 5:46pleading, just saying, "Make this
  185. 5:47better." doesn't necessarily work.
  186. 5:49Really, the only lever you have, which
  187. 5:50most people don't even think to use, is
  188. 5:53the verification lever. Because by
  189. 5:55optimizing this, it makes it so that
  190. 5:56you're playing within the actual rules
  191. 5:58that the AI follows. So, how do you help
  192. 6:00AI verify the output so it's up to the
  193. 6:02standard you want? Well, there are three
  194. 6:04places to focus on. First, you want to
  195. 6:06set the evaluation criteria up front.
  196. 6:08Before Claude touches a single thing,
  197. 6:10whether that's technical or
  198. 6:11non-technical tasks, define what good
  199. 6:13looks like with precision. For example,
  200. 6:16a vague way to evaluate an output is,
  201. 6:17"Make this report look good." Whereas, a
  202. 6:20precise way would say, "The report must
  203. 6:22have three sections, each ends with a
  204. 6:24recommendation." And if you're making
  205. 6:25the connection, this is very similar to
  206. 6:27what we covered in layer one. The more
  207. 6:29precise you are up front, the less room
  208. 6:30Claude will have to make mistakes. To
  209. 6:32help enforce this, we'll add this to our
  210. 6:34verification Claude prompt. Outline the
  211. 6:37evaluation criteria you will use to
  212. 6:38ensure a high-quality final product. Be
  213. 6:40precise. The second step is use a second
  214. 6:43AI model as the critic. Think of this
  215. 6:45like a second robot librarian from a
  216. 6:47different library. You use that
  217. 6:49librarian to grade the output of the
  218. 6:51first librarian. This other librarian
  219. 6:53has a whole different set of books, and
  220. 6:55that may give them insight into why this
  221. 6:57first librarian is right or wrong. Now,
  222. 6:59a tactical way to do this, if you use
  223. 7:01Claude code, you can install the Codex
  224. 7:03plugin, which will allow you to directly
  225. 7:05ask Codex questions within your Claude
  226. 7:07code session. So, you could say
  227. 7:09something like, "If this turns into a
  228. 7:10complex build, run the final output by
  229. 7:13Codex to ensure both systems agree." And
  230. 7:15step three is pull external signal where
  231. 7:18possible. The question here is, how can
  232. 7:19you bring in additional context that
  233. 7:21will help you verify an output? Here are
  234. 7:23two concrete examples. Let's say you're
  235. 7:25deploying an app and you're not sure if
  236. 7:26it's successfully deployed. What you can
  237. 7:28do instead is connect your Claude
  238. 7:29session with your system where it's
  239. 7:31deployed, so it can verify that it has
  240. 7:33been deployed successfully. We are
  241. 7:35making a connection to pull external
  242. 7:36data to enhance our verification layer.
  243. 7:40And now, if it says that the deployment
  244. 7:41was successful, we know for certainty
  245. 7:43that it actually was. In a non-technical
  246. 7:45example, let's say you're working on a
  247. 7:47monthly report. You could bring in your
  248. 7:48historical reports to use as reference
  249. 7:50for the exact format that the final
  250. 7:52output should be in, pulling in data and
  251. 7:54empowering the verification process.
  252. 7:56Now, bringing in this concept with the
  253. 7:57first two points, combining this third
  254. 7:59point with the first two points, here is
  255. 8:01a prompt that you can run in Claude,
  256. 8:03which will help ensure that you are
  257. 8:04adding a proper evaluation layer where
  258. 8:06it makes sense. I can't stress how
  259. 8:07important this is. The creator of Claude
  260. 8:09code, Boris Cherney, said it best. If
  261. 8:11Claude has a feedback loop, it will two
  262. 8:13to three x quality of the final result.
  263. 8:15So, layer one and layer two are about
  264. 8:16creating specs and evaluating the
  265. 8:18output. The third layer, however, is
  266. 8:20where we build a foundation that can't
  267. 8:21be replicated. But, before we get to
  268. 8:23that, if this is your first video of
  269. 8:25mine, welcome to the channel. If it's
  270. 8:26your second or more, here is our
  271. 8:28anti-slop agreement. The visuals, the
  272. 8:30testing, the hours of research that went
  273. 8:32into this video, this is entirely built
  274. 8:34for humans, not for AI clunkers. So, all
  275. 8:37that I ask is that you subscribe as part
  276. 8:38of this agreement because it helps it
  277. 8:39reach more people so that I can keep
  278. 8:41making videos like this. Also, every
  279. 8:43couple of weeks, I give away a Claude
  280. 8:44Max subscription, so comment below with
  281. 8:46whatever you're building to enter. Layer
  282. 8:48three, the environment. So, layer one
  283. 8:50and layer two need somewhere to live,
  284. 8:51and that's layer three, which is the
  285. 8:53environment that you build in. Think of
  286. 8:54this layer as a workshop. The spec is a
  287. 8:56blueprint pinned to the wall, the
  288. 8:58verifier is the quality check station by
  289. 9:00the door, and then the environment is
  290. 9:02the workshop itself. You need to create
  291. 9:04the proper tooling and the proper system
  292. 9:06so that the whole thing can function at
  293. 9:07a high level. Now, the problem here is
  294. 9:09that most people use the workshop from
  295. 9:10scratch every time they use AI. And no,
  296. 9:12if you have a single chat with your
  297. 9:14entire conversation history, that is not
  298. 9:16what I'm talking about. So, how do you
  299. 9:17create a proper workspace that improves
  300. 9:20over time? First is you need to set up a
  301. 9:22proper Claude MD file. Every time you
  302. 9:24prompt Claude, your Claude.md file gets
  303. 9:26injected automatically. It's essentially
  304. 9:28the first thing that Claude reads to
  305. 9:30help determine how it should operate.
  306. 9:31For example, you can add to your Claude
  307. 9:33MD before building anything multi-step,
  308. 9:35include a verification plan. Now,
  309. 9:37verification is forced into every build,
  310. 9:39not something that you have to remember
  311. 9:41to say. This is just one of the ways
  312. 9:42that you can improve this Claude MD, and
  313. 9:44here's actually mine on the screen, and
  314. 9:46I'm going to call out a couple of
  315. 9:47sections. The first is I outline how
  316. 9:49this repo works. So, think of my repo as
  317. 9:51my workspace. It gives high-level to the
  318. 9:53details around it. I then tell it the
  319. 9:55custom skills and how they're routed,
  320. 9:57how to use them. I then outline the
  321. 9:59architecture of the training data or
  322. 10:01knowledge architecture so that the AI
  323. 10:03knows where to look for certain
  324. 10:04information. And then I have key working
  325. 10:06rules that it should follow no matter
  326. 10:08what. Make this your environment. It's
  327. 10:10your world, and AI is living in it. It
  328. 10:12should not feel like the other way
  329. 10:13around. The second step is you need to
  330. 10:15build your LLM knowledge base. Karpathy
  331. 10:17went viral for this concept on Twitter
  332. 10:19that he calls his LLM knowledge base.
  333. 10:21And this is essentially creating a
  334. 10:22folder system on your machine that
  335. 10:24you're able to ingest your own training
  336. 10:26data in a way that makes it really easy
  337. 10:28for Claude to understand where
  338. 10:30information is. This is so important
  339. 10:32because your data is your moat. And this
  340. 10:35begins the process of building out your
  341. 10:37own intellectual data property. And step
  342. 10:40three is you have to start building out
  343. 10:42your skill set. A general rule of thumb
  344. 10:43that I have is if you plan on doing
  345. 10:45something repeatedly, create a custom
  346. 10:47skill for that. Think of this like a
  347. 10:49handbook to complete a specific task.
  348. 10:50And the more you use these skills, the
  349. 10:52better they'll become. I have a saying
  350. 10:53that I tell my team, the best way to
  351. 10:55find a leak in a hose is to run water
  352. 10:57through it. And it's the same with
  353. 10:58skills. The more you use them, the more
  354. 11:00you'll realize where you need to fix
  355. 11:01them and where they're really good. Keep
  356. 11:03running water through it and your
  357. 11:05system's going to compound over time.
  358. 11:07Step four is create rules for what the
  359. 11:09AI can and can't work on. Depending on
  360. 11:11the cost of getting something wrong, you
  361. 11:13need to establish different AI
  362. 11:14guardrails. So, here's how to think of
  363. 11:16this, right? So, take the Claude.md file
  364. 11:18that I mentioned earlier. You could add
  365. 11:20a line that says, "Don't make up
  366. 11:21information," but that's a guide, not
  367. 11:23necessarily a hard rule. So, at the end
  368. 11:25of the day, AI can still ignore it. So,
  369. 11:27if you have things that are critical not
  370. 11:29to get wrong, then you need to introduce
  371. 11:31rule-based guardrails to ensure that the
  372. 11:34AI can't bypass them. To help you
  373. 11:36visualize this, imagine you have a
  374. 11:37folder called "Important, Don't Edit."
  375. 11:40You could have a rule in Claude MD that
  376. 11:41says, "Don't touch anything in the
  377. 11:43/important, don't edit folder." And that
  378. 11:45might get you 80% of the way there, but
  379. 11:48it's essentially a request, not a rule.
  380. 11:51Claude can still touch those files. So,
  381. 11:53instead, you add a pre-tool use hook
  382. 11:56before Claude uses the write or edit
  383. 11:58tool, and it checks to see the file that
  384. 12:00it's trying to edit. Now, Claude
  385. 12:02literally can't make the edit, and it's
  386. 12:04enforced at the tool level, not the
  387. 12:06prompt level. And as a result of this,
  388. 12:08this is now a concrete rule that the
  389. 12:10agent can't bypass. So, with this in
  390. 12:12mind, bucket things into three groups.
  391. 12:14The first is always do. This is things
  392. 12:16that AI should run on autopilot. The
  393. 12:18second is ask first. So, this is
  394. 12:20anything that you want to double-check.
  395. 12:22And then the third is never do. These
  396. 12:24are lines that can't be crossed that are
  397. 12:26absolutely critical not to get wrong.
  398. 12:28Here's a prompt that brings all of these
  399. 12:30four points that I mentioned to help
  400. 12:32audit your system and create an
  401. 12:33optimized environment for Claude to
  402. 12:35interact with. That's the Karpathy
  403. 12:36method end-to-end, the spec, the
  404. 12:38verifier, and the environment. But
  405. 12:40there's a question that needs to be
  406. 12:41answered. What's the one thing that
  407. 12:42Karpathy thinks we should focus on in
  408. 12:44the age of AI? Here's him getting asked
  409. 12:46this in an interview.
  410. 12:47>> What still remains worth learning deeply
  411. 12:50when intelligence gets cheap as we move
  412. 12:53into the next eight era of AI?
  413. 12:55>> You can outsource your thinking, but you
  414. 12:57can't outsource your understanding. And
  415. 12:58the thing with everything we covered
  416. 12:59here is that the three layers are
  417. 13:01centered around your understanding of
  418. 13:03the bigger picture. You need to
  419. 13:04understand your goals and what's needed
  420. 13:06to direct AI to start working for you.
  421. 13:08Now, if you like this video, you will
  422. 13:09love this one where I do a deep dive
  423. 13:11into four Claude projects that you need
  424. 13:13to build today using these three layers.
  425. 13:16I'll see you over there. Peace.

About this transcript

This page contains the full transcript of Stop Prompting Claude. Use Karpathy's Method Instead. by Austin Marchese, generated from the public captions YouTube serves with the video. The transcript has 2,909 words across 425 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.