YouTube2Text

Anthropic Engineer Explains: What to Build Instead of AI Agents — Transcript

by Nate Herk | AI Automation · 2,515 words · 369 segments · language en · Watch on YouTube

Full transcript

  1. 0:00So, Enthropic Engineers just said that
  2. 0:01they stopped building agents and they
  3. 0:03started building something completely
  4. 0:04different. So, if you're still focused
  5. 0:05on building agents, you're probably
  6. 0:07wasting hours on something that'll never
  7. 0:08work the way that you actually want it
  8. 0:09to. But if you focus on what these
  9. 0:11engineers are actually building, you'll
  10. 0:12have a system that improves on its own.
  11. 0:14So, in this video, I'm going to show you
  12. 0:15what they said to build instead and the
  13. 0:17four things that actually make it work
  14. 0:18so that you can start getting the same
  15. 0:19quality results that Enthropic engineers
  16. 0:21are getting. So, let's get into it. All
  17. 0:23right. So, Enthropic isn't saying that
  18. 0:24agents are dead. Barry Jien and Mahesh
  19. 0:26Marog, the two people who created agent
  20. 0:28skills adanthropic, said that they
  21. 0:29basically stopped rebuilding a separate
  22. 0:31agent for every single job because the
  23. 0:33agent underneath had become way more
  24. 0:35general purpose than they expected. So,
  25. 0:36real quick, the easiest way to
  26. 0:37understand this is to look at your
  27. 0:39phone. Your phone has a processor, an
  28. 0:41operating system, and then all the apps
  29. 0:43that you actually use every day. And a
  30. 0:44few massive companies build the
  31. 0:46processor and the operating system. You
  32. 0:48probably aren't changing either one of
  33. 0:49those things, but you can choose the
  34. 0:51apps, and each app gives that same phone
  35. 0:53a very specific capability. And the
  36. 0:55whole AI stack is starting to look
  37. 0:56pretty similar to that. The model is
  38. 0:58kind of like the processor. The agent
  39. 0:59runtime is like the operating [music]
  40. 1:01system. And then skills are the apps. So
  41. 1:03cloud code can already read files, write
  42. 1:05code, call tools, and work through a
  43. 1:06task. You don't need to necessarily
  44. 1:08rebuild all of that every time that you
  45. 1:09want help creating a presentation or
  46. 1:11researching a company or writing a
  47. 1:12LinkedIn post or doing whatever kind of
  48. 1:14things you need to do like that. You
  49. 1:15give the same general purpose agent a
  50. 1:17skill that contains the process, the
  51. 1:19context, the scripts, and the examples
  52. 1:21for that specific job. So now let me
  53. 1:23talk about the four practical ways to
  54. 1:24make those skills work better. So the
  55. 1:26first one is simple. Stop making Claude
  56. 1:28solve the same technical problem over
  57. 1:30and over. And team kept watching Claude
  58. 1:32write basically the exact same Python
  59. 1:33script every time it needed to apply
  60. 1:35styling to a slide deck. It would spend
  61. 1:37tons of tokens just recreating the code
  62. 1:39that has already been written. And
  63. 1:40because it was rebuilding it from
  64. 1:41scratch, the result could change from
  65. 1:43one run to the next. So it wasn't super
  66. 1:44consistent. So what they did is they had
  67. 1:46Claude save that script inside the skill
  68. 1:48as in their own words a tool for its
  69. 1:50future self. So now the next time it
  70. 1:52needs to style a presentation, it can
  71. 1:53just run the version that already was
  72. 1:55proven to work. And developers have
  73. 1:56followed this idea forever. It's called
  74. 1:58dry or dr, which stands for don't repeat
  75. 2:01yourself. If you solve a problem in
  76. 2:02code, you can just save the solution and
  77. 2:04reuse it instead of rewriting the code
  78. 2:05every single time. And you can do the
  79. 2:06exact same thing with Claude. I've got
  80. 2:08skills in my own AI operating system
  81. 2:09that use the same renderers, the same
  82. 2:11templates, scripts, and stuff like that
  83. 2:13every time. My carousel workflow doesn't
  84. 2:15ask Claude to reinvent how a slide gets
  85. 2:16rendered on every run. the skill just
  86. 2:18points it to the files that already work
  87. 2:19and then Claude can just focus on the
  88. 2:20new content. Like literally just
  89. 2:22changing the text. So the next time
  90. 2:23Claude writes a script that gives you a
  91. 2:24result that you really like, don't leave
  92. 2:26that code trapped inside the chat just
  93. 2:28to get lost later when you start up new
  94. 2:29sessions. Tell it something like save
  95. 2:31the script you just used inside the
  96. 2:32skills script folder. Update the
  97. 2:34skill.mmd so that future runs execute
  98. 2:36that file instead of trying to rewrite
  99. 2:37it again. Then I want you to run the
  100. 2:39skill again and verify the result. And
  101. 2:40real quick, you still obviously need to
  102. 2:42test it. You run the same type of task
  103. 2:44twice. You compare the important parts
  104. 2:45and make sure the skill is actually
  105. 2:47calling the saved file. The surrounding
  106. 2:49AI output may still vary a little bit,
  107. 2:50but you've replaced one fresh guess with
  108. 2:52a proven piece of code. So, moral of the
  109. 2:54story, don't pay Claude to rediscover a
  110. 2:56solution that you already have. But once
  111. 2:58you start building a bunch of these
  112. 2:59skills, Claude needs to know which one
  113. 3:01belongs to the job in front of it. So,
  114. 3:02let's move on to number two. So, think
  115. 3:04about a mechanic for a second. A
  116. 3:05mechanic might own hundreds of tools,
  117. 3:07but he doesn't dump every single one
  118. 3:08onto the bench before he starts changing
  119. 3:10a tire. He basically just identifies the
  120. 3:12job and grabs the few tools that he
  121. 3:14needs and leaves everything else inside
  122. 3:15the toolbox. And skills work in a very
  123. 3:17similar way because of something that
  124. 3:18Enthropic calls progressive disclosure.
  125. 3:21So when Claude starts, it doesn't read
  126. 3:22the full instructions and all the
  127. 3:24examples and all the scripts from every
  128. 3:25single skill you have. That would be a
  129. 3:27huge waste of time and tokens. It starts
  130. 3:29with the name and the description of
  131. 3:30each skill. And this is called the YAML
  132. 3:32front matter. Then when your specific
  133. 3:33prompt matches a description of a skill,
  134. 3:35that's when it reads the full skill. MD
  135. 3:37file. and any larger references or
  136. 3:38scripts can just stay in the folder
  137. 3:40until the task actually needs them.
  138. 3:41Enthropic describes this as letting
  139. 3:43Claude load information only as needed.
  140. 3:46That basically keeps irrelevant
  141. 3:47instructions out of the working context,
  142. 3:49which helps you avoid the bloat and
  143. 3:50confusion which sometimes gets called
  144. 3:52context rod. But this depends on your
  145. 3:53description being really really clear.
  146. 3:54If one skill says help with content and
  147. 3:56another one says create marketing
  148. 3:57assets, then cla basically guessing
  149. 3:59because those descriptions sort of
  150. 4:00overlap and they don't tell the agent
  151. 4:02when either of those skills should
  152. 4:03actually run. So, a stronger description
  153. 4:05might say something like, "This skill
  154. 4:06creates LinkedIn carousels from a topic,
  155. 4:08from a transcript, or an outline. Use
  156. 4:10this when the user asks for a carousel,
  157. 4:12carousel slides, or a LinkedIn document
  158. 4:13post." Now, Claude knows what the skill
  159. 4:15does and exactly when to use it. So,
  160. 4:17just keep each skill focused on one
  161. 4:18specific job. Put the words a real
  162. 4:20person would use inside the description
  163. 4:22and make sure two skills aren't
  164. 4:23competing for the same request. And you
  165. 4:25can actually have Claude Code audit this
  166. 4:26for you. [music] Just say, "Hey, review
  167. 4:28all my skill descriptions. For each one,
  168. 4:30tell me what it does, when it should
  169. 4:31trigger, and where it overlaps with any
  170. 4:33other skills. Rewrite only descriptions
  171. 4:35that are [music] ambiguous. And then you
  172. 4:36can test these three things. The first
  173. 4:38one is an obvious request that should
  174. 4:39trigger it. The second one is a
  175. 4:40differently worded request that still
  176. 4:41should trigger it. And the third one is
  177. 4:43an unrelated request that definitely
  178. 4:44should not trigger it. And just remember
  179. 4:46that a skill that Claude can't find is
  180. 4:47basically a skill that you don't have.
  181. 4:48So finding the correct skill handles
  182. 4:50today's tasks. But the next step is
  183. 4:52actually making sure that the skill gets
  184. 4:54better every single time you use it. So,
  185. 4:55number three, every time you correct
  186. 4:57Claude and then you close the chat,
  187. 4:58there's a pretty good chance you just
  188. 4:59threw that lesson away. Maybe [music] it
  189. 5:00used the wrong tone or it skipped a
  190. 5:02validation step or maybe it formatted
  191. 5:04the final output in a way that you don't
  192. 5:06want to see again. And if all you say to
  193. 5:07Claude is fix it, then it probably will
  194. 5:09fix it, but the process stays broken.
  195. 5:10And this is something I've had to work
  196. 5:11through inside my own AI operating
  197. 5:13system. If an agent tells me it can't
  198. 5:14find a file that I know exists, I don't
  199. 5:16just hand it the path and keep moving. I
  200. 5:17ask it to backtrack. I basically ask it
  201. 5:19to show me its work, show me where it
  202. 5:21searched, figure out why it missed the
  203. 5:22file, and then update the routing or the
  204. 5:24skill so that next run starts in the
  205. 5:26correct place. Enthropic designs skills
  206. 5:28as a step toward this kind of continuous
  207. 5:30learning. Their guarantee is that
  208. 5:32anything that Claude writes down can be
  209. 5:33used efficiently by a future version of
  210. 5:35itself. Now, here's the thing. Skills
  211. 5:37don't remember every single thing, and
  212. 5:38they aren't a recording of every
  213. 5:39conversation that you've ever had. They
  214. 5:40basically store the procedural knowledge
  215. 5:42that Claude needs to do the [music]
  216. 5:44specific job. So, when a process is
  217. 5:45wrong, update the instructions in the
  218. 5:47skill.mmd. When Claude is missing your
  219. 5:49voice or your brand or your examples,
  220. 5:51then add a reference file. And when the
  221. 5:52same mistake keeps happening, then add a
  222. 5:54clear rule that explicitly prevents
  223. 5:55that. And then rerun the same task. You
  224. 5:57could use a prompt like this. Review
  225. 5:59what went wrong during this run. Decide
  226. 6:00whether the cause was the process or
  227. 6:03missing context or a weak rule or
  228. 6:05unreliable code. And then update the
  229. 6:06skill in the smallest durable place.
  230. 6:08Then rerun the same task and verify
  231. 6:10[music] the fix. Over time, the skill
  232. 6:11becomes a living record of how you want
  233. 6:13the job done. And I do want to be
  234. 6:14precise about the phrase model proof. No
  235. 6:16skill can just force a weaker model to
  236. 6:18perform exactly like the stronger one.
  237. 6:20Different models will obviously still
  238. 6:21produce different results and they
  239. 6:23interpret skills in different ways. But
  240. 6:24your process can be portable. Agent
  241. 6:26skills are an open format. So the same
  242. 6:28core like skills folder can work across
  243. 6:30compatible agent harnesses like codeex
  244. 6:32or Hermes agent or anything else. So
  245. 6:34test an important skill with another
  246. 6:35compatible agent. If the result falls
  247. 6:37apart, then look for hidden assumptions,
  248. 6:38missing examples or instructions that
  249. 6:40only one model understands and then
  250. 6:42tighten up the skill and keep testing.
  251. 6:43All right, so the first three were very
  252. 6:44important, but this one is probably the
  253. 6:46most important. A skill shouldn't hand
  254. 6:48you its first attempt and call the job
  255. 6:49done. This is probably the biggest gap
  256. 6:51that I see with a lot of AI workflows.
  257. 6:53You run the skill, it creates the thing,
  258. 6:54saves the file, and then comes back to
  259. 6:56you and says, "Hey, I'm done." And then
  260. 6:57you open it and you realize that the
  261. 6:58formatting is broken, the sources don't
  262. 7:00support the claims, or that the script
  263. 7:01falls completely flat for the person
  264. 7:03you're trying to reach. So, the AI did
  265. 7:05maybe 70 or 80% of the job, and now
  266. 7:07you're the human, you know, manually
  267. 7:09doing the last 20 or 30%. But if you
  268. 7:11already know how to personally check
  269. 7:12that work, then just bake those checks
  270. 7:14into the skill and let the AI close more
  271. 7:16of that gap for you. So, for example,
  272. 7:17for a slide deck, have the skill render
  273. 7:19each slide as an image, inspect the
  274. 7:21screenshots, fix anything that's cropped
  275. 7:22or hard to read or out of bounds, and
  276. 7:24rerender it. For a research report, make
  277. 7:26it open the primary sources, match the
  278. 7:27claims [music] to the evidence, and
  279. 7:29remove anything that it couldn't
  280. 7:30actually verify. And if you're creating
  281. 7:31a script or an ad or something more
  282. 7:33subjective, then have a few different
  283. 7:35personas, like a few different sub
  284. 7:36aents, review it and discuss it. A
  285. 7:38beginner agent can tell you where
  286. 7:39they're confused. [music] A skeptical
  287. 7:40buyer agent can tell you what they don't
  288. 7:42believe. Someone from your actual
  289. 7:43audience can tell you where they would
  290. 7:45probably click away. Now, you don't have
  291. 7:46to accept every single piece of feedback
  292. 7:48that you get from these different agent
  293. 7:49personas. That would probably make the
  294. 7:50output worse, but the skill can find the
  295. 7:52issues that show up more than once, make
  296. 7:54the strongest revisions, and then run
  297. 7:55the review again. And sometimes just
  298. 7:57hearing those different perspectives is
  299. 7:58really helpful. And real quick,
  300. 7:59verification isn't clawed reading its
  301. 8:01own work and saying, "Looks good to me."
  302. 8:02It needs some kind of evidence outside
  303. 8:04of that first draft, whether that's a
  304. 8:05screenshot, a test result, [music] a
  305. 8:07source, a reference example, or like I
  306. 8:08said, feedback from a few different
  307. 8:09agent perspectives. And you can add
  308. 8:11something like this into almost any
  309. 8:12skill. Say something like, "Before
  310. 8:14returning the final output, define the
  311. 8:15acceptance criteria. Create the first
  312. 8:17version, inspect it using the relevant
  313. 8:18verification method, fix every issue you
  314. 8:20find, and then run another pass. Return
  315. 8:22the output only after it meets the
  316. 8:23criteria with a short summary of what
  317. 8:25you checked. And if something can't be
  318. 8:27verified, then tell me exactly what
  319. 8:28remains. It's even better if there's
  320. 8:30some sort of objective success metric
  321. 8:32that you can set and just have the
  322. 8:33agents keep working until it hits that
  323. 8:35objectively. But anyways, now the first
  324. 8:37output is basically an internal draft.
  325. 8:38Claude reviews it, catches the obvious
  326. 8:40problems, and then improves it before
  327. 8:41you see any of it, before you waste any
  328. 8:43of your time and attention on it.
  329. 8:45Because you're ultimately still going to
  330. 8:46be the final judge, especially when
  331. 8:47there's things like taste or strategy or
  332. 8:49business judgment involved in the
  333. 8:51process. The goal, though, is just to
  334. 8:52stop spending your time catching
  335. 8:54problems that the AI could have caught
  336. 8:55on its own. Your first look should not
  337. 8:57be the agent's first look. It should be
  338. 8:58the agent's, you know, fourth or fifth
  339. 9:00or maybe even sixth look. So now you
  340. 9:01know what enthropic engineers are
  341. 9:03actually building. They save proven code
  342. 9:04instead of rewriting it. They give every
  343. 9:06skill a precise description so Claude
  344. 9:08loads only the correct one. They turn
  345. 9:10corrections into durable instructions
  346. 9:11that keep improving. And they verify the
  347. 9:13work before it ever reaches you. That's
  348. 9:15how you take a general purpose agent and
  349. 9:16you teach it how you work specifically.
  350. 9:18And if you want to see how I organize
  351. 9:19the agents, the context, connections,
  352. 9:21capabilities, and cadence around all of
  353. 9:22this, then I'll put my full AI operating
  354. 9:24system course on screen right up here. I
  355. 9:27also have free courses, templates, and
  356. 9:29resources inside my free community that
  357. 9:30will help you build agents and workflows
  358. 9:32from scratch. You can join with the link
  359. 9:33in the description. And if you want to
  360. 9:34go deeper [music] on turning these
  361. 9:35skills into a career, then you can check
  362. 9:37out my plus community. But anyways, that
  363. 9:38is going to do it for this one. So, if
  364. 9:40you guys enjoyed, you learned something
  365. 9:41new, please give it a like. It helps me
  366. 9:42out a ton. And as always, I appreciate
  367. 9:44you guys making it to the end of the
  368. 9:45video. I'll see you on the next one.
  369. 9:46Thanks everyone.

About this transcript

This page contains the full transcript of Anthropic Engineer Explains: What to Build Instead of AI Agents by Nate Herk | AI Automation, generated from the public captions YouTube serves with the video. The transcript has 2,515 words across 369 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.