Stop Prompting Claude. Use Karpathy's Method Instead. — Transcript
Full transcript
- 0:00I just listened to Andrej Karpathy speak
- 0:01at AISN 2026, and I learned something
- 0:04that I wasn't expecting. Almost everyone
- 0:06is prompting Claude wrong. So, I decided
- 0:08to dig deeper and see exactly how
- 0:09Karpathy, the former head of AI at
- 0:11Tesla, uses AI in 2026. And it turns out
- 0:14that Karpathy's method for building 10
- 0:16times faster can be broken down into
- 0:18three simple layers. So, in today's
- 0:20video, I'll be breaking down each layer
- 0:22so that anybody can apply them. And then
- 0:23I'll show you the one thing that
- 0:25Karpathy said to focus on in the age of
- 0:27AI. So, layer one is the spec. AI models
- 0:29are incredibly smart, but they're still
- 0:31missing something. To showcase their
- 0:33current limitation, Karpathy explained a
- 0:35simple question AI will get wrong.
- 0:37>> I want to go to a car wash to wash my
- 0:39car, and it's 50 m away. Should I drive
- 0:42or should I walk? And state-of-the-art
- 0:45models today will tell you to walk
- 0:46because it's so close.
- 0:48>> At first, I actually didn't believe
- 0:49this, so I went to Claude, Gemini, Grok,
- 0:51and ChatGPT, asked them the same
- 0:52question, and they all gave me the same
- 0:54answer. And it reveals the whole
- 0:55foundation of this video. AI is
- 0:57brilliant at what can be measured, but
- 1:00for context-driven things like needing a
- 1:02car for a car wash, it has no signal to
- 1:05act on. So, how do you bridge this gap
- 1:07between your understanding and your
- 1:08contextual information and AI's
- 1:10computational power? That's where the
- 1:11spec comes in. And a spec is how you
- 1:13deliver your understanding to Claude in
- 1:15a format it can use. A term you may have
- 1:17heard is Claude's plan mode, which
- 1:19essentially can be used to help you
- 1:20create a plan before building anything.
- 1:22But Karpathy thinks that this is too
- 1:25high-level.
- 1:25>> I actually don't even like the plan
- 1:27mode. I I would
- 1:28I mean, obviously it's very useful, but
- 1:30I think there's something more general
- 1:31here where you have to work with your
- 1:32agent to design a spec that is very
- 1:34detailed.
- 1:34>> Now, Karpathy isn't telling you that
- 1:36plan mode is bad. What he's actually
- 1:38saying is you have to go deeper, work
- 1:40with these AI tools to design the actual
- 1:42spec. So, how do you create a spec that
- 1:44Claude can successfully use to build
- 1:46what you're trying to build? The first
- 1:48step is you have to uncover your goal.
- 1:50If you just say, "Create a end-of-month
- 1:52report," that's a task, but the actual
- 1:54goal is a conclusion you're trying to
- 1:56draw, the decision the report drives.
- 1:59And what the goal actually is is
- 2:00something AI will literally never be
- 2:02able to decide. So, to help you do this,
- 2:04we'll tell Claude to interview me to
- 2:07identify the goal of this project. This
- 2:09is the way to get the information out of
- 2:11you and into the spec. Now, step two is
- 2:13be agile with how you work. There are
- 2:15two methods of completing any task. The
- 2:17first is waterfall, and the other is
- 2:19agile. Waterfall is you take a big task
- 2:22and you complete the entire thing, and
- 2:24then you show the final product. Agile
- 2:26on the other hand is you break that same
- 2:27task into small buckets, and you show
- 2:30the result throughout the entire process
- 2:32to make sure you're going in the right
- 2:33direction. And people are extremely
- 2:35susceptible to using AI agents in a
- 2:37waterfall manner because they want to
- 2:38give them everything to do at once. The
- 2:41better move is agile specking. You want
- 2:43to have a tight scope, a clear
- 2:45checkpoint, you want to review the
- 2:46output, adjust it, and then repeat. To
- 2:48help with this, we'll tell Claude to
- 2:50bias towards smaller and more
- 2:51compartmentalized specs. Step three is
- 2:54you want to be precise and use your
- 2:55brain. The more precise you are, the
- 2:57less AI has to assume. And every
- 2:59assumption that AI makes is a chance for
- 3:01it to drift from the final product you
- 3:03actually want. And when you have AI
- 3:05create a spec for you, you have to use
- 3:07your brain to think critically about
- 3:09what that spec actually says. So, to
- 3:12help you use your brain, you can say
- 3:13"Make me verify key decisions explicitly
- 3:16to ensure nothing is missed." And when
- 3:18you put these three pieces together, we
- 3:20have a final prompt we can use in Claude
- 3:22to help create a tightly scoped,
- 3:24well-thought-out
- 3:26aligns with our actual goal. This is a
- 3:27process that I call modern engineering,
- 3:29which every successful person has to
- 3:31become. Now, layer two is the verifier.
- 3:34Layer two sits on top of the spec. This
- 3:36is the verification process. One of the
- 3:38most frustrating things about AI is
- 3:40reviewing and verifying the output. And
- 3:42unlike a human, it can't grasp
- 3:44non-measurable things. So, how can we
- 3:47help AI verify its own outputs? Well,
- 3:49first you need to understand the mental
- 3:50model behind this. And Karpathy explains
- 3:53it as animals versus ghosts. Here's him
- 3:55getting asked a question about this in a
- 3:57recent interview. And if it sounds
- 3:59confusing, don't worry, I will simplify
- 4:00it after.
- 4:01>> And the idea is that we're not building
- 4:03animals, we are summoning ghosts. Why
- 4:05does that framing matter? And what does
- 4:07it actually change about how you build
- 4:09and deploy and evaluate or even trust
- 4:12them?
- 4:12>> Yeah, I think the reason I wrote about
- 4:14this is because I'm trying to wrap my
- 4:15head around what these things are,
- 4:16right? Because if you have a good model
- 4:18of what they are or are not, then you're
- 4:19going to be more competent at uh using
- 4:21them. I think it's just um
- 4:23coming to terms with the fact that these
- 4:24things are not, you know, animal
- 4:26intelligences. Like if you yell at them,
- 4:27they're not going to work better or
- 4:29worse or doesn't have any impact. Um
- 4:32and uh
- 4:33it's all just kind of like these
- 4:34statistical simulation circuits. It's
- 4:36more just being suspicious of it and um
- 4:38figuring it out over time.
- 4:39>> Now, that's some gigabrain stuff, but
- 4:41let me simplify it. People, me and you,
- 4:43are used to interacting with people,
- 4:45which Karpathy is calling animals. These
- 4:48animals are driven by different
- 4:49motivators and emotions, which help
- 4:51produce the final product and output
- 4:53within a team setting. And if you say to
- 4:55a person, become an expert at SEO
- 4:57marketing in the next 14 days or you're
- 4:59fired, they're going to figure it out.
- 5:01That's because they have these intrinsic
- 5:03motivations. But AI is not that.
- 5:05Karpathy describes it as a ghost, but in
- 5:07my eyes that's a little too confusing,
- 5:09so throw it out the window. Instead,
- 5:10think of it like a robot librarian. If
- 5:12you ask it that same SEO question, the
- 5:14librarian will only suggest resources
- 5:17and answers based on the books in its
- 5:19library. If it doesn't have a book, it
- 5:21can't help you. And part of the
- 5:22challenge here is that the librarian
- 5:24doesn't know when it's missing a
- 5:26specific book. So, it may just
- 5:28confidently make something up. And
- 5:29that's what's happening when AI nails
- 5:31math and fumbles things with context.
- 5:33It's brilliant because the library has
- 5:35the clear answers. But if it doesn't,
- 5:37then it's confidently wrong or
- 5:40uncertain. Which means interacting with
- 5:42it like it's an animal, i.e. a human,
- 5:44doesn't help, right? Yelling at it,
- 5:46pleading, just saying, "Make this
- 5:47better." doesn't necessarily work.
- 5:49Really, the only lever you have, which
- 5:50most people don't even think to use, is
- 5:53the verification lever. Because by
- 5:55optimizing this, it makes it so that
- 5:56you're playing within the actual rules
- 5:58that the AI follows. So, how do you help
- 6:00AI verify the output so it's up to the
- 6:02standard you want? Well, there are three
- 6:04places to focus on. First, you want to
- 6:06set the evaluation criteria up front.
- 6:08Before Claude touches a single thing,
- 6:10whether that's technical or
- 6:11non-technical tasks, define what good
- 6:13looks like with precision. For example,
- 6:16a vague way to evaluate an output is,
- 6:17"Make this report look good." Whereas, a
- 6:20precise way would say, "The report must
- 6:22have three sections, each ends with a
- 6:24recommendation." And if you're making
- 6:25the connection, this is very similar to
- 6:27what we covered in layer one. The more
- 6:29precise you are up front, the less room
- 6:30Claude will have to make mistakes. To
- 6:32help enforce this, we'll add this to our
- 6:34verification Claude prompt. Outline the
- 6:37evaluation criteria you will use to
- 6:38ensure a high-quality final product. Be
- 6:40precise. The second step is use a second
- 6:43AI model as the critic. Think of this
- 6:45like a second robot librarian from a
- 6:47different library. You use that
- 6:49librarian to grade the output of the
- 6:51first librarian. This other librarian
- 6:53has a whole different set of books, and
- 6:55that may give them insight into why this
- 6:57first librarian is right or wrong. Now,
- 6:59a tactical way to do this, if you use
- 7:01Claude code, you can install the Codex
- 7:03plugin, which will allow you to directly
- 7:05ask Codex questions within your Claude
- 7:07code session. So, you could say
- 7:09something like, "If this turns into a
- 7:10complex build, run the final output by
- 7:13Codex to ensure both systems agree." And
- 7:15step three is pull external signal where
- 7:18possible. The question here is, how can
- 7:19you bring in additional context that
- 7:21will help you verify an output? Here are
- 7:23two concrete examples. Let's say you're
- 7:25deploying an app and you're not sure if
- 7:26it's successfully deployed. What you can
- 7:28do instead is connect your Claude
- 7:29session with your system where it's
- 7:31deployed, so it can verify that it has
- 7:33been deployed successfully. We are
- 7:35making a connection to pull external
- 7:36data to enhance our verification layer.
- 7:40And now, if it says that the deployment
- 7:41was successful, we know for certainty
- 7:43that it actually was. In a non-technical
- 7:45example, let's say you're working on a
- 7:47monthly report. You could bring in your
- 7:48historical reports to use as reference
- 7:50for the exact format that the final
- 7:52output should be in, pulling in data and
- 7:54empowering the verification process.
- 7:56Now, bringing in this concept with the
- 7:57first two points, combining this third
- 7:59point with the first two points, here is
- 8:01a prompt that you can run in Claude,
- 8:03which will help ensure that you are
- 8:04adding a proper evaluation layer where
- 8:06it makes sense. I can't stress how
- 8:07important this is. The creator of Claude
- 8:09code, Boris Cherney, said it best. If
- 8:11Claude has a feedback loop, it will two
- 8:13to three x quality of the final result.
- 8:15So, layer one and layer two are about
- 8:16creating specs and evaluating the
- 8:18output. The third layer, however, is
- 8:20where we build a foundation that can't
- 8:21be replicated. But, before we get to
- 8:23that, if this is your first video of
- 8:25mine, welcome to the channel. If it's
- 8:26your second or more, here is our
- 8:28anti-slop agreement. The visuals, the
- 8:30testing, the hours of research that went
- 8:32into this video, this is entirely built
- 8:34for humans, not for AI clunkers. So, all
- 8:37that I ask is that you subscribe as part
- 8:38of this agreement because it helps it
- 8:39reach more people so that I can keep
- 8:41making videos like this. Also, every
- 8:43couple of weeks, I give away a Claude
- 8:44Max subscription, so comment below with
- 8:46whatever you're building to enter. Layer
- 8:48three, the environment. So, layer one
- 8:50and layer two need somewhere to live,
- 8:51and that's layer three, which is the
- 8:53environment that you build in. Think of
- 8:54this layer as a workshop. The spec is a
- 8:56blueprint pinned to the wall, the
- 8:58verifier is the quality check station by
- 9:00the door, and then the environment is
- 9:02the workshop itself. You need to create
- 9:04the proper tooling and the proper system
- 9:06so that the whole thing can function at
- 9:07a high level. Now, the problem here is
- 9:09that most people use the workshop from
- 9:10scratch every time they use AI. And no,
- 9:12if you have a single chat with your
- 9:14entire conversation history, that is not
- 9:16what I'm talking about. So, how do you
- 9:17create a proper workspace that improves
- 9:20over time? First is you need to set up a
- 9:22proper Claude MD file. Every time you
- 9:24prompt Claude, your Claude.md file gets
- 9:26injected automatically. It's essentially
- 9:28the first thing that Claude reads to
- 9:30help determine how it should operate.
- 9:31For example, you can add to your Claude
- 9:33MD before building anything multi-step,
- 9:35include a verification plan. Now,
- 9:37verification is forced into every build,
- 9:39not something that you have to remember
- 9:41to say. This is just one of the ways
- 9:42that you can improve this Claude MD, and
- 9:44here's actually mine on the screen, and
- 9:46I'm going to call out a couple of
- 9:47sections. The first is I outline how
- 9:49this repo works. So, think of my repo as
- 9:51my workspace. It gives high-level to the
- 9:53details around it. I then tell it the
- 9:55custom skills and how they're routed,
- 9:57how to use them. I then outline the
- 9:59architecture of the training data or
- 10:01knowledge architecture so that the AI
- 10:03knows where to look for certain
- 10:04information. And then I have key working
- 10:06rules that it should follow no matter
- 10:08what. Make this your environment. It's
- 10:10your world, and AI is living in it. It
- 10:12should not feel like the other way
- 10:13around. The second step is you need to
- 10:15build your LLM knowledge base. Karpathy
- 10:17went viral for this concept on Twitter
- 10:19that he calls his LLM knowledge base.
- 10:21And this is essentially creating a
- 10:22folder system on your machine that
- 10:24you're able to ingest your own training
- 10:26data in a way that makes it really easy
- 10:28for Claude to understand where
- 10:30information is. This is so important
- 10:32because your data is your moat. And this
- 10:35begins the process of building out your
- 10:37own intellectual data property. And step
- 10:40three is you have to start building out
- 10:42your skill set. A general rule of thumb
- 10:43that I have is if you plan on doing
- 10:45something repeatedly, create a custom
- 10:47skill for that. Think of this like a
- 10:49handbook to complete a specific task.
- 10:50And the more you use these skills, the
- 10:52better they'll become. I have a saying
- 10:53that I tell my team, the best way to
- 10:55find a leak in a hose is to run water
- 10:57through it. And it's the same with
- 10:58skills. The more you use them, the more
- 11:00you'll realize where you need to fix
- 11:01them and where they're really good. Keep
- 11:03running water through it and your
- 11:05system's going to compound over time.
- 11:07Step four is create rules for what the
- 11:09AI can and can't work on. Depending on
- 11:11the cost of getting something wrong, you
- 11:13need to establish different AI
- 11:14guardrails. So, here's how to think of
- 11:16this, right? So, take the Claude.md file
- 11:18that I mentioned earlier. You could add
- 11:20a line that says, "Don't make up
- 11:21information," but that's a guide, not
- 11:23necessarily a hard rule. So, at the end
- 11:25of the day, AI can still ignore it. So,
- 11:27if you have things that are critical not
- 11:29to get wrong, then you need to introduce
- 11:31rule-based guardrails to ensure that the
- 11:34AI can't bypass them. To help you
- 11:36visualize this, imagine you have a
- 11:37folder called "Important, Don't Edit."
- 11:40You could have a rule in Claude MD that
- 11:41says, "Don't touch anything in the
- 11:43/important, don't edit folder." And that
- 11:45might get you 80% of the way there, but
- 11:48it's essentially a request, not a rule.
- 11:51Claude can still touch those files. So,
- 11:53instead, you add a pre-tool use hook
- 11:56before Claude uses the write or edit
- 11:58tool, and it checks to see the file that
- 12:00it's trying to edit. Now, Claude
- 12:02literally can't make the edit, and it's
- 12:04enforced at the tool level, not the
- 12:06prompt level. And as a result of this,
- 12:08this is now a concrete rule that the
- 12:10agent can't bypass. So, with this in
- 12:12mind, bucket things into three groups.
- 12:14The first is always do. This is things
- 12:16that AI should run on autopilot. The
- 12:18second is ask first. So, this is
- 12:20anything that you want to double-check.
- 12:22And then the third is never do. These
- 12:24are lines that can't be crossed that are
- 12:26absolutely critical not to get wrong.
- 12:28Here's a prompt that brings all of these
- 12:30four points that I mentioned to help
- 12:32audit your system and create an
- 12:33optimized environment for Claude to
- 12:35interact with. That's the Karpathy
- 12:36method end-to-end, the spec, the
- 12:38verifier, and the environment. But
- 12:40there's a question that needs to be
- 12:41answered. What's the one thing that
- 12:42Karpathy thinks we should focus on in
- 12:44the age of AI? Here's him getting asked
- 12:46this in an interview.
- 12:47>> What still remains worth learning deeply
- 12:50when intelligence gets cheap as we move
- 12:53into the next eight era of AI?
- 12:55>> You can outsource your thinking, but you
- 12:57can't outsource your understanding. And
- 12:58the thing with everything we covered
- 12:59here is that the three layers are
- 13:01centered around your understanding of
- 13:03the bigger picture. You need to
- 13:04understand your goals and what's needed
- 13:06to direct AI to start working for you.
- 13:08Now, if you like this video, you will
- 13:09love this one where I do a deep dive
- 13:11into four Claude projects that you need
- 13:13to build today using these three layers.
- 13:16I'll see you over there. Peace.
About this transcript
This page contains the full transcript of Stop Prompting Claude. Use Karpathy's Method Instead. by Austin Marchese, generated from the public captions YouTube serves with the video. The transcript has 2,909 words across 425 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.