Anthropic Engineer Explains: What to Build Instead of AI Agents — Transcript
Full transcript
- 0:00So, Enthropic Engineers just said that
- 0:01they stopped building agents and they
- 0:03started building something completely
- 0:04different. So, if you're still focused
- 0:05on building agents, you're probably
- 0:07wasting hours on something that'll never
- 0:08work the way that you actually want it
- 0:09to. But if you focus on what these
- 0:11engineers are actually building, you'll
- 0:12have a system that improves on its own.
- 0:14So, in this video, I'm going to show you
- 0:15what they said to build instead and the
- 0:17four things that actually make it work
- 0:18so that you can start getting the same
- 0:19quality results that Enthropic engineers
- 0:21are getting. So, let's get into it. All
- 0:23right. So, Enthropic isn't saying that
- 0:24agents are dead. Barry Jien and Mahesh
- 0:26Marog, the two people who created agent
- 0:28skills adanthropic, said that they
- 0:29basically stopped rebuilding a separate
- 0:31agent for every single job because the
- 0:33agent underneath had become way more
- 0:35general purpose than they expected. So,
- 0:36real quick, the easiest way to
- 0:37understand this is to look at your
- 0:39phone. Your phone has a processor, an
- 0:41operating system, and then all the apps
- 0:43that you actually use every day. And a
- 0:44few massive companies build the
- 0:46processor and the operating system. You
- 0:48probably aren't changing either one of
- 0:49those things, but you can choose the
- 0:51apps, and each app gives that same phone
- 0:53a very specific capability. And the
- 0:55whole AI stack is starting to look
- 0:56pretty similar to that. The model is
- 0:58kind of like the processor. The agent
- 0:59runtime is like the operating [music]
- 1:01system. And then skills are the apps. So
- 1:03cloud code can already read files, write
- 1:05code, call tools, and work through a
- 1:06task. You don't need to necessarily
- 1:08rebuild all of that every time that you
- 1:09want help creating a presentation or
- 1:11researching a company or writing a
- 1:12LinkedIn post or doing whatever kind of
- 1:14things you need to do like that. You
- 1:15give the same general purpose agent a
- 1:17skill that contains the process, the
- 1:19context, the scripts, and the examples
- 1:21for that specific job. So now let me
- 1:23talk about the four practical ways to
- 1:24make those skills work better. So the
- 1:26first one is simple. Stop making Claude
- 1:28solve the same technical problem over
- 1:30and over. And team kept watching Claude
- 1:32write basically the exact same Python
- 1:33script every time it needed to apply
- 1:35styling to a slide deck. It would spend
- 1:37tons of tokens just recreating the code
- 1:39that has already been written. And
- 1:40because it was rebuilding it from
- 1:41scratch, the result could change from
- 1:43one run to the next. So it wasn't super
- 1:44consistent. So what they did is they had
- 1:46Claude save that script inside the skill
- 1:48as in their own words a tool for its
- 1:50future self. So now the next time it
- 1:52needs to style a presentation, it can
- 1:53just run the version that already was
- 1:55proven to work. And developers have
- 1:56followed this idea forever. It's called
- 1:58dry or dr, which stands for don't repeat
- 2:01yourself. If you solve a problem in
- 2:02code, you can just save the solution and
- 2:04reuse it instead of rewriting the code
- 2:05every single time. And you can do the
- 2:06exact same thing with Claude. I've got
- 2:08skills in my own AI operating system
- 2:09that use the same renderers, the same
- 2:11templates, scripts, and stuff like that
- 2:13every time. My carousel workflow doesn't
- 2:15ask Claude to reinvent how a slide gets
- 2:16rendered on every run. the skill just
- 2:18points it to the files that already work
- 2:19and then Claude can just focus on the
- 2:20new content. Like literally just
- 2:22changing the text. So the next time
- 2:23Claude writes a script that gives you a
- 2:24result that you really like, don't leave
- 2:26that code trapped inside the chat just
- 2:28to get lost later when you start up new
- 2:29sessions. Tell it something like save
- 2:31the script you just used inside the
- 2:32skills script folder. Update the
- 2:34skill.mmd so that future runs execute
- 2:36that file instead of trying to rewrite
- 2:37it again. Then I want you to run the
- 2:39skill again and verify the result. And
- 2:40real quick, you still obviously need to
- 2:42test it. You run the same type of task
- 2:44twice. You compare the important parts
- 2:45and make sure the skill is actually
- 2:47calling the saved file. The surrounding
- 2:49AI output may still vary a little bit,
- 2:50but you've replaced one fresh guess with
- 2:52a proven piece of code. So, moral of the
- 2:54story, don't pay Claude to rediscover a
- 2:56solution that you already have. But once
- 2:58you start building a bunch of these
- 2:59skills, Claude needs to know which one
- 3:01belongs to the job in front of it. So,
- 3:02let's move on to number two. So, think
- 3:04about a mechanic for a second. A
- 3:05mechanic might own hundreds of tools,
- 3:07but he doesn't dump every single one
- 3:08onto the bench before he starts changing
- 3:10a tire. He basically just identifies the
- 3:12job and grabs the few tools that he
- 3:14needs and leaves everything else inside
- 3:15the toolbox. And skills work in a very
- 3:17similar way because of something that
- 3:18Enthropic calls progressive disclosure.
- 3:21So when Claude starts, it doesn't read
- 3:22the full instructions and all the
- 3:24examples and all the scripts from every
- 3:25single skill you have. That would be a
- 3:27huge waste of time and tokens. It starts
- 3:29with the name and the description of
- 3:30each skill. And this is called the YAML
- 3:32front matter. Then when your specific
- 3:33prompt matches a description of a skill,
- 3:35that's when it reads the full skill. MD
- 3:37file. and any larger references or
- 3:38scripts can just stay in the folder
- 3:40until the task actually needs them.
- 3:41Enthropic describes this as letting
- 3:43Claude load information only as needed.
- 3:46That basically keeps irrelevant
- 3:47instructions out of the working context,
- 3:49which helps you avoid the bloat and
- 3:50confusion which sometimes gets called
- 3:52context rod. But this depends on your
- 3:53description being really really clear.
- 3:54If one skill says help with content and
- 3:56another one says create marketing
- 3:57assets, then cla basically guessing
- 3:59because those descriptions sort of
- 4:00overlap and they don't tell the agent
- 4:02when either of those skills should
- 4:03actually run. So, a stronger description
- 4:05might say something like, "This skill
- 4:06creates LinkedIn carousels from a topic,
- 4:08from a transcript, or an outline. Use
- 4:10this when the user asks for a carousel,
- 4:12carousel slides, or a LinkedIn document
- 4:13post." Now, Claude knows what the skill
- 4:15does and exactly when to use it. So,
- 4:17just keep each skill focused on one
- 4:18specific job. Put the words a real
- 4:20person would use inside the description
- 4:22and make sure two skills aren't
- 4:23competing for the same request. And you
- 4:25can actually have Claude Code audit this
- 4:26for you. [music] Just say, "Hey, review
- 4:28all my skill descriptions. For each one,
- 4:30tell me what it does, when it should
- 4:31trigger, and where it overlaps with any
- 4:33other skills. Rewrite only descriptions
- 4:35that are [music] ambiguous. And then you
- 4:36can test these three things. The first
- 4:38one is an obvious request that should
- 4:39trigger it. The second one is a
- 4:40differently worded request that still
- 4:41should trigger it. And the third one is
- 4:43an unrelated request that definitely
- 4:44should not trigger it. And just remember
- 4:46that a skill that Claude can't find is
- 4:47basically a skill that you don't have.
- 4:48So finding the correct skill handles
- 4:50today's tasks. But the next step is
- 4:52actually making sure that the skill gets
- 4:54better every single time you use it. So,
- 4:55number three, every time you correct
- 4:57Claude and then you close the chat,
- 4:58there's a pretty good chance you just
- 4:59threw that lesson away. Maybe [music] it
- 5:00used the wrong tone or it skipped a
- 5:02validation step or maybe it formatted
- 5:04the final output in a way that you don't
- 5:06want to see again. And if all you say to
- 5:07Claude is fix it, then it probably will
- 5:09fix it, but the process stays broken.
- 5:10And this is something I've had to work
- 5:11through inside my own AI operating
- 5:13system. If an agent tells me it can't
- 5:14find a file that I know exists, I don't
- 5:16just hand it the path and keep moving. I
- 5:17ask it to backtrack. I basically ask it
- 5:19to show me its work, show me where it
- 5:21searched, figure out why it missed the
- 5:22file, and then update the routing or the
- 5:24skill so that next run starts in the
- 5:26correct place. Enthropic designs skills
- 5:28as a step toward this kind of continuous
- 5:30learning. Their guarantee is that
- 5:32anything that Claude writes down can be
- 5:33used efficiently by a future version of
- 5:35itself. Now, here's the thing. Skills
- 5:37don't remember every single thing, and
- 5:38they aren't a recording of every
- 5:39conversation that you've ever had. They
- 5:40basically store the procedural knowledge
- 5:42that Claude needs to do the [music]
- 5:44specific job. So, when a process is
- 5:45wrong, update the instructions in the
- 5:47skill.mmd. When Claude is missing your
- 5:49voice or your brand or your examples,
- 5:51then add a reference file. And when the
- 5:52same mistake keeps happening, then add a
- 5:54clear rule that explicitly prevents
- 5:55that. And then rerun the same task. You
- 5:57could use a prompt like this. Review
- 5:59what went wrong during this run. Decide
- 6:00whether the cause was the process or
- 6:03missing context or a weak rule or
- 6:05unreliable code. And then update the
- 6:06skill in the smallest durable place.
- 6:08Then rerun the same task and verify
- 6:10[music] the fix. Over time, the skill
- 6:11becomes a living record of how you want
- 6:13the job done. And I do want to be
- 6:14precise about the phrase model proof. No
- 6:16skill can just force a weaker model to
- 6:18perform exactly like the stronger one.
- 6:20Different models will obviously still
- 6:21produce different results and they
- 6:23interpret skills in different ways. But
- 6:24your process can be portable. Agent
- 6:26skills are an open format. So the same
- 6:28core like skills folder can work across
- 6:30compatible agent harnesses like codeex
- 6:32or Hermes agent or anything else. So
- 6:34test an important skill with another
- 6:35compatible agent. If the result falls
- 6:37apart, then look for hidden assumptions,
- 6:38missing examples or instructions that
- 6:40only one model understands and then
- 6:42tighten up the skill and keep testing.
- 6:43All right, so the first three were very
- 6:44important, but this one is probably the
- 6:46most important. A skill shouldn't hand
- 6:48you its first attempt and call the job
- 6:49done. This is probably the biggest gap
- 6:51that I see with a lot of AI workflows.
- 6:53You run the skill, it creates the thing,
- 6:54saves the file, and then comes back to
- 6:56you and says, "Hey, I'm done." And then
- 6:57you open it and you realize that the
- 6:58formatting is broken, the sources don't
- 7:00support the claims, or that the script
- 7:01falls completely flat for the person
- 7:03you're trying to reach. So, the AI did
- 7:05maybe 70 or 80% of the job, and now
- 7:07you're the human, you know, manually
- 7:09doing the last 20 or 30%. But if you
- 7:11already know how to personally check
- 7:12that work, then just bake those checks
- 7:14into the skill and let the AI close more
- 7:16of that gap for you. So, for example,
- 7:17for a slide deck, have the skill render
- 7:19each slide as an image, inspect the
- 7:21screenshots, fix anything that's cropped
- 7:22or hard to read or out of bounds, and
- 7:24rerender it. For a research report, make
- 7:26it open the primary sources, match the
- 7:27claims [music] to the evidence, and
- 7:29remove anything that it couldn't
- 7:30actually verify. And if you're creating
- 7:31a script or an ad or something more
- 7:33subjective, then have a few different
- 7:35personas, like a few different sub
- 7:36aents, review it and discuss it. A
- 7:38beginner agent can tell you where
- 7:39they're confused. [music] A skeptical
- 7:40buyer agent can tell you what they don't
- 7:42believe. Someone from your actual
- 7:43audience can tell you where they would
- 7:45probably click away. Now, you don't have
- 7:46to accept every single piece of feedback
- 7:48that you get from these different agent
- 7:49personas. That would probably make the
- 7:50output worse, but the skill can find the
- 7:52issues that show up more than once, make
- 7:54the strongest revisions, and then run
- 7:55the review again. And sometimes just
- 7:57hearing those different perspectives is
- 7:58really helpful. And real quick,
- 7:59verification isn't clawed reading its
- 8:01own work and saying, "Looks good to me."
- 8:02It needs some kind of evidence outside
- 8:04of that first draft, whether that's a
- 8:05screenshot, a test result, [music] a
- 8:07source, a reference example, or like I
- 8:08said, feedback from a few different
- 8:09agent perspectives. And you can add
- 8:11something like this into almost any
- 8:12skill. Say something like, "Before
- 8:14returning the final output, define the
- 8:15acceptance criteria. Create the first
- 8:17version, inspect it using the relevant
- 8:18verification method, fix every issue you
- 8:20find, and then run another pass. Return
- 8:22the output only after it meets the
- 8:23criteria with a short summary of what
- 8:25you checked. And if something can't be
- 8:27verified, then tell me exactly what
- 8:28remains. It's even better if there's
- 8:30some sort of objective success metric
- 8:32that you can set and just have the
- 8:33agents keep working until it hits that
- 8:35objectively. But anyways, now the first
- 8:37output is basically an internal draft.
- 8:38Claude reviews it, catches the obvious
- 8:40problems, and then improves it before
- 8:41you see any of it, before you waste any
- 8:43of your time and attention on it.
- 8:45Because you're ultimately still going to
- 8:46be the final judge, especially when
- 8:47there's things like taste or strategy or
- 8:49business judgment involved in the
- 8:51process. The goal, though, is just to
- 8:52stop spending your time catching
- 8:54problems that the AI could have caught
- 8:55on its own. Your first look should not
- 8:57be the agent's first look. It should be
- 8:58the agent's, you know, fourth or fifth
- 9:00or maybe even sixth look. So now you
- 9:01know what enthropic engineers are
- 9:03actually building. They save proven code
- 9:04instead of rewriting it. They give every
- 9:06skill a precise description so Claude
- 9:08loads only the correct one. They turn
- 9:10corrections into durable instructions
- 9:11that keep improving. And they verify the
- 9:13work before it ever reaches you. That's
- 9:15how you take a general purpose agent and
- 9:16you teach it how you work specifically.
- 9:18And if you want to see how I organize
- 9:19the agents, the context, connections,
- 9:21capabilities, and cadence around all of
- 9:22this, then I'll put my full AI operating
- 9:24system course on screen right up here. I
- 9:27also have free courses, templates, and
- 9:29resources inside my free community that
- 9:30will help you build agents and workflows
- 9:32from scratch. You can join with the link
- 9:33in the description. And if you want to
- 9:34go deeper [music] on turning these
- 9:35skills into a career, then you can check
- 9:37out my plus community. But anyways, that
- 9:38is going to do it for this one. So, if
- 9:40you guys enjoyed, you learned something
- 9:41new, please give it a like. It helps me
- 9:42out a ton. And as always, I appreciate
- 9:44you guys making it to the end of the
- 9:45video. I'll see you on the next one.
- 9:46Thanks everyone.
About this transcript
This page contains the full transcript of Anthropic Engineer Explains: What to Build Instead of AI Agents by Nate Herk | AI Automation, generated from the public captions YouTube serves with the video. The transcript has 2,515 words across 369 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.