YouTube2Text

Building Great Agent Skills: The Missing Manual — Transcript

by AI Engineer · 4,204 words · 597 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Hello friends, I was dearly hoping to be
  2. 0:01able to come to the Air Engineer World's
  3. 0:03Fair, but family matters have intruded
  4. 0:06and I'm not able to make it. However, I
  5. 0:08will not be leaving you empty-handed.
  6. 0:09I'm going to give you the talk that I
  7. 0:11would have given in San Francisco. This
  8. 0:13talk is called The Missing Manual, How
  9. 0:15to Write Great Skills, and I think that
  10. 0:18the ability to distinguish good skills
  11. 0:20from bad skills is only getting more
  12. 0:22important. As developers, we seem to be
  13. 0:24pretty talented at finding different
  14. 0:25forms of hell for us to go to. In like a
  15. 0:29few years ago, we had tutorial hell,
  16. 0:31which is where you would go into a bunch
  17. 0:32of tutorials trying to learn something,
  18. 0:34not be able to piece it together, and
  19. 0:36sort of just get into this cycle you
  20. 0:38couldn't get out of. We had framework
  21. 0:41hell, where every other 10 minutes there
  22. 0:43was a JavaScript framework being
  23. 0:44announced, and you know, you had to
  24. 0:45learn the hot new thing all the time.
  25. 0:47And now, I think we have another version
  26. 0:49of hell, which is skill hell. Skill hell
  27. 0:52is where you have all of these skills
  28. 0:54available, freely available, that you
  29. 0:55can download, contribute to, you can
  30. 0:57figure out on your own, but you don't
  31. 0:59really know how the pieces all work
  32. 1:01together. You can't tell a good skill
  33. 1:03from a bad skill. And this means that
  34. 1:04people are trying to piece together
  35. 1:06these frameworks, trying to try
  36. 1:07everything that's out there all at once.
  37. 1:10And they sort of can't, or rather, they
  38. 1:12don't get the results that the skills
  39. 1:14themselves promise. This is true at an
  40. 1:15individual level, but it's also true at
  41. 1:17an organization level, too.
  42. 1:19Organizations have no way or no
  43. 1:21understanding on how to build good
  44. 1:23skills, how to take their operating
  45. 1:25procedures and turn them into things
  46. 1:27that an agent can do. you don't do that,
  47. 1:29then it's hard to get the bounty that
  48. 1:31skills can offer. Just one more skill,
  49. 1:34bro. That's kind of seems like what
  50. 1:36we're saying. And I feel a bit of guilt
  51. 1:38here, too, because we have Matt Pocock
  52. 1:40skills, which is my skills repo, which
  53. 1:42is one of the most popular engineering
  54. 1:44skill sets out there. And so, I feel
  55. 1:46like I want to help the people who use
  56. 1:48my skills get out of skill hell. So, how
  57. 1:50do we do it? How do we get out of this?
  58. 1:51Well, what is actually missing here?
  59. 1:54Well, in my opinion, the thing that
  60. 1:55we're missing is we don't know what
  61. 1:57makes a skill great. We can't yet look
  62. 1:59at a skill and go, "Okay, this skill is
  63. 2:01doing these good things and these bad
  64. 2:03things." There's no shared rubric, no
  65. 2:06framework for looking at a skill and
  66. 2:07making it better. And so, that's what
  67. 2:09I'm going to give you in this talk. I'm
  68. 2:10going to give you a skill checklist, a
  69. 2:12checklist of things you can look at
  70. 2:14inside the skill to make sure that it's
  71. 2:16doing what it says it's doing and ways
  72. 2:18you can improve it, ways you can write
  73. 2:20skills. This checklist looks like this.
  74. 2:22We start with the trigger of the skill,
  75. 2:24how the skill is invoked, and the
  76. 2:26decisions that you need to design there.
  77. 2:29Then the internal structure of the
  78. 2:31skill, how the skill is actually
  79. 2:32composed and laid out internally. Then
  80. 2:35number three is how do you actually
  81. 2:36steer using the skill? How do you get
  82. 2:39the skill to tell the agent what to do?
  83. 2:42Then four, how do you make the skill as
  84. 2:44small as possible? Because once we've
  85. 2:46got a working skill, we then need to
  86. 2:49basically maximize it, prune out all of
  87. 2:51the irrelevant stuff, prune out all of
  88. 2:52the no ops. There's one handy advantage
  89. 2:54of me not being in the room with you,
  90. 2:56which is you can immediately go and try
  91. 2:57this out because I've encoded all of
  92. 2:59this into a new skill in my repo called
  93. 3:01writing great skills. So, if you've got
  94. 3:03an immediate use case for this, then
  95. 3:05just go to this my skills repo, you
  96. 3:06know, just close this browser, get out
  97. 3:08of here, and go and use this skill to
  98. 3:11either improve your skills or write
  99. 3:13great new ones. But, let's now go
  100. 3:14through the checklist then. We have
  101. 3:16number one, the trigger, the way the
  102. 3:18skill is invoked. And in order to talk
  103. 3:19about this, I'm actually going to do a
  104. 3:21bit of comparison here, which is that my
  105. 3:23skills are often compared to another set
  106. 3:25of extremely popular engineering skills
  107. 3:27called superpowers. And I'm really often
  108. 3:29asked the question, "How do your skills
  109. 3:31compare to superpowers? What's the
  110. 3:33difference between them?" To understand
  111. 3:34that, we need to understand the
  112. 3:35difference between user invoked and
  113. 3:37model invoked skills. Anytime you have a
  114. 3:40skill, you can always invoke it
  115. 3:42manually. So, the skill sits on your
  116. 3:44file system, the agent will just be able
  117. 3:46to pull up the skill and understand
  118. 3:48what's in there. And you can always do
  119. 3:50that by communicating that to the agent.
  120. 3:52Doesn't always look like this forward
  121. 3:53slash depending on the harness, but that
  122. 3:55you can always use or invoke your
  123. 3:57skills. Another way that skills can be
  124. 3:59invoked is by the agent itself. These
  125. 4:01are called model invocable skills or
  126. 4:03model invoked skills. You can take a
  127. 4:05description, so the description of the
  128. 4:08skill always ends up in the agent's
  129. 4:10context, and the agent can look in that
  130. 4:13and go, "Okay, based on that
  131. 4:14description, I'm going to invoke the
  132. 4:16skill and I end up reading the skill.md
  133. 4:19file, which is where the meat of the
  134. 4:21skill is, into my context window. That's
  135. 4:23how you invoke a skill. That's what
  136. 4:24happens when a skill is invoked. So,
  137. 4:26this description serves as a kind of
  138. 4:28context pointer. It sits in the agent's
  139. 4:30context pointing to another file where
  140. 4:33the agent can go if it wants more
  141. 4:35context. But, that context pointer, you
  142. 4:36don't need to put it into the agent's
  143. 4:39context. It can just be invisible from
  144. 4:42the agent, and that is what we call a
  145. 4:43user invocable skill. So, some skills
  146. 4:46can only be invoked by the user because
  147. 4:48they don't have this context pointer.
  148. 4:50It's optional. For instance, we can see
  149. 4:52in my code base design here, this is a
  150. 4:54model invocable skill. It has a
  151. 4:56description that ends up in the agent's
  152. 4:57context window. But, if we look at my
  153. 4:59grill me skill instead, we can see it
  154. 5:01has disable model invocation true. This
  155. 5:04means that this little description here
  156. 5:05will only show to the user. It won't be
  157. 5:08visible to the agent. So, this then is
  158. 5:09tip number one. Decide if your skill is
  159. 5:12user invoked or model invoked. Now, you
  160. 5:14might think that model invoked skills
  161. 5:16are better, right? Because either the
  162. 5:17model can invoke it itself or the user
  163. 5:20can invoke it. It's more flexible. But,
  164. 5:22every time you add a model invoked skill
  165. 5:24into your agent's environment, it
  166. 5:27increases what I'm going to call the
  167. 5:28context load on that agent. It adds a
  168. 5:31new description, which is costing you
  169. 5:34tokens on every request, but also adding
  170. 5:36a different thing for the agent to think
  171. 5:39about. So, if you have a hundred model
  172. 5:41invoked skills, that's going to be a
  173. 5:42hundred descriptions inside the context
  174. 5:45for your agent. So, it seems to make
  175. 5:46sense then to either tamp down the
  176. 5:48number of model invoked skills or to
  177. 5:50just use all user invoked skills. But,
  178. 5:53user invoked skills have a different
  179. 5:54load, which is the more user invoked
  180. 5:57skills you have, the higher cognitive
  181. 5:59load on the user. In other words, the
  182. 6:00more things the user needs to keep in
  183. 6:03their head, the more skill you require
  184. 6:05from the pilot. And so, if we compare
  185. 6:07Matt Percot skills to Superpowers,
  186. 6:10Superpowers is primarily model invoked
  187. 6:12skills. It gives the agent superpowers.
  188. 6:16Whereas my skills, I much prefer to be
  189. 6:18in full control. That means I get to
  190. 6:20keep the context load on the agent as
  191. 6:22small as possible, but it does impose
  192. 6:25more of a cognitive load on me. So, I
  193. 6:27need to understand the skills really
  194. 6:29deeply in order to get the most use out.
  195. 6:31So, why have I done this? Why did I
  196. 6:33prefer user invoked skills? Well, every
  197. 6:36time you have a model invoked skill, it
  198. 6:38basically you get a cost in
  199. 6:40unpredictability. Because every time you
  200. 6:42have a context pointer pointing from one
  201. 6:44resource to another, the model may just
  202. 6:46choose not to follow it, you know, even
  203. 6:49if it's absolutely perfect for the task,
  204. 6:51it may just choose not to invoke the
  205. 6:54skill. I much prefer removing that level
  206. 6:57of unpredictability, imposing a bit more
  207. 6:59cognitive load on the user, and what you
  208. 7:01get is just you're removing a class of
  209. 7:04problem from even being a problem.
  210. 7:06Because this unpredictability leaves
  211. 7:07people to need to eval their skills to
  212. 7:10make sure they're being called at the
  213. 7:11right time, which is really nasty and
  214. 7:14it's a problem I prefer to avoid. But,
  215. 7:16what I'm hoping to show you here is that
  216. 7:17model invoked skills and user invoked
  217. 7:19skills both have their same costs. So,
  218. 7:22it's not an easy decision which one you
  219. 7:24choose. So, that then is the trigger,
  220. 7:26how the skill gets invoked. Now, let's
  221. 7:28talk about the structure, the internal
  222. 7:31layout of the skill. I think of there as
  223. 7:32being two main units that you need to
  224. 7:34put into most skills. These two units
  225. 7:37are the steps and the reference. The
  226. 7:40steps are the step-by-step procedure
  227. 7:42that the skill is going to walk through
  228. 7:45and the reference is any supporting
  229. 7:46information that helps it walk through
  230. 7:48those steps. You can have skills that
  231. 7:51have no steps and are only reference and
  232. 7:53you can have skills that are no
  233. 7:55reference and only a set of simple steps
  234. 7:57to walk through. But if you start
  235. 7:58thinking of skills as composed of these
  236. 8:00two units, it really helps just break
  237. 8:02them down a lot more. If we look at an
  238. 8:04example, one of my skills called 2 PRD
  239. 8:06creates a product requirements document
  240. 8:09out of the current context window. It's
  241. 8:10got three steps in it. So it finds the
  242. 8:13relevant context, it confirms the test
  243. 8:16seams with the user. So there's like a
  244. 8:18little human in the loop checkpoint
  245. 8:19there just to make sure we're not doing
  246. 8:21anything weird with the testing, which I
  247. 8:22find really important. And then we write
  248. 8:25the product requirements document. To
  249. 8:26handle those three steps, we've got two
  250. 8:29bits of reference material. We've got a
  251. 8:31little bit of reference on what is a
  252. 8:33test seam and then we've got a product
  253. 8:36requirements document template. So just
  254. 8:38a literal markdown template which is
  255. 8:41used to write the PRD. So this is a
  256. 8:42great way to write a skill from scratch.
  257. 8:45You work out if you need some steps,
  258. 8:47then you write those steps and you work
  259. 8:49out what reference material those steps
  260. 8:50need and you put it in a separate little
  261. 8:52spot in the skill, which is for
  262. 8:54reference material. However, there's a
  263. 8:55really important constraint that we need
  264. 8:57to think about, which is tip number
  265. 8:59three, we want to make the main skill.md
  266. 9:02file as small as possible. Every skill
  267. 9:05is composed of its description and then
  268. 9:07a skill.md file and then any reference
  269. 9:09material that branches off that. And
  270. 9:11this skill.md file, if we make it small,
  271. 9:14then we're saving in a bunch of
  272. 9:15different ways. Smaller skills are just
  273. 9:17easier to maintain, easier to audit,
  274. 9:19fewer words to think about. And every
  275. 9:22time you shave off a word, that is a
  276. 9:24token shaved, that multiple tokens
  277. 9:26shaved from your skills cost. So I do
  278. 9:28believe that small skills are really
  279. 9:31important both for maintainers and for
  280. 9:33users. One really useful way you can
  281. 9:35make your skill smaller is by thinking
  282. 9:37about the different branches of the
  283. 9:39skill, the different ways the skill can
  284. 9:41be used. Because if you have reference
  285. 9:43material that's only used in one branch,
  286. 9:45then that's a candidate for being
  287. 9:46removed from the main skill.md. For
  288. 9:49instance, if we look at my 2PRD here, we
  289. 9:51have two pieces of reference material,
  290. 9:53what is a test seam and the PRD
  291. 9:55template. Well, we need the PRD template
  292. 9:57every single time because we are always
  293. 10:00creating a PRD and we probably also need
  294. 10:02the what is a test seam information
  295. 10:04every time because we're always asking
  296. 10:06about the test seams. So, 2PRD, there's
  297. 10:08only one branch and all the reference
  298. 10:10material belongs on that branch, so it
  299. 10:12probably also belongs in the skill.md
  300. 10:15file. However, if we look at a different
  301. 10:16skill of mine, which is domain modeling,
  302. 10:19domain modeling does two things. It
  303. 10:21updates a local glossary called
  304. 10:23context.md and then it also creates
  305. 10:25architectural decision records. In other
  306. 10:27words, it's doing two different things
  307. 10:30or it might actually choose to do
  308. 10:31neither of these, in which case it
  309. 10:33doesn't need the template and it doesn't
  310. 10:34need the ADR template, either. So, in
  311. 10:37other words, domain modeling has two or
  312. 10:39maybe three branches and this means that
  313. 10:41we don't need to include the ADR
  314. 10:43template or the context.md template into
  315. 10:46the main skill. They can be moved into
  316. 10:49separate zones. The way you do that is
  317. 10:51you have the skill.md file, then you put
  318. 10:54it behind a context pointer and you
  319. 10:56point the context template to a separate
  320. 10:59markdown file inside the skills folder.
  321. 11:01That context pointer literally just
  322. 11:03says, "If you need the template or if
  323. 11:04you need to update the context.md file,
  324. 11:07go to this file." And I call that an
  325. 11:09external reference. It's a reference
  326. 11:12that's external to the skill.md that you
  327. 11:14can just easily reference the agent can
  328. 11:16pull in very easily because it's bundled
  329. 11:18along with the skill. So, this is a
  330. 11:19technique you can use for making the
  331. 11:22skill.md as small as possible, which is
  332. 11:24has so many benefits.
  333. 11:26Hide branching reference material behind
  334. 11:28context pointers. In other words, if you
  335. 11:30feel like your skill is going to be used
  336. 11:32in lots of different ways, then take the
  337. 11:34reference material that's relevant for
  338. 11:35those branches and hide them behind
  339. 11:37context pointers. So, that is structure.
  340. 11:39We need to think about making the
  341. 11:41skill.md super duper small. We need to
  342. 11:43think about the branches in our skill
  343. 11:45moving material out behind context
  344. 11:47pointers. And we need to think about
  345. 11:49steps and reference, which are the two
  346. 11:51main units inside a skill. Let's go next
  347. 11:54to steering, the actual ways we get the
  348. 11:57agent to do what we want it to do. And
  349. 11:58for me, steering comes down to one
  350. 12:01really cool technique, which is the kind
  351. 12:03of main thing I want you to get from
  352. 12:05this talk. This technique fixes this
  353. 12:07issue, which is the agent doesn't do
  354. 12:10what I want. In other words, I specify
  355. 12:12something in the skill, I think that
  356. 12:14I've been clear, and then it just
  357. 12:15doesn't do the thing. Now, I think the
  358. 12:17main reason this happens is because
  359. 12:19you're not using a technique called
  360. 12:21leading words. The idea of leading
  361. 12:23words, or light vert if you like
  362. 12:26literary theory, I suppose, is that
  363. 12:28there are certain words that pack in a
  364. 12:30bunch of meaning into a very small
  365. 12:32space. These leading words are really
  366. 12:35powerful with agents because you put the
  367. 12:37leading word in the skill itself in the
  368. 12:40text, and then the agent will repeat the
  369. 12:42leading word back to itself as part of
  370. 12:44its operations, as part of its thinking
  371. 12:46tokens, and as part of its output to
  372. 12:48you. And then, because it's
  373. 12:50re-emphasizing that word and that word
  374. 12:52hopefully describes what you want from
  375. 12:54the agent, that then goes and changes
  376. 12:56its behavior. Let's make this more
  377. 12:58concrete with an example. So, let's
  378. 13:00imagine that we have a problem, which is
  379. 13:02a classic problem with agents, which is
  380. 13:03that they code layer by layer. In other
  381. 13:06words, if you give them a big tranche of
  382. 13:07work to do, they will generally code up
  383. 13:09all of the database layer, then all of
  384. 13:11the schemas, then all of the API
  385. 13:13endpoints, then all of the front end.
  386. 13:15They don't do the sort of typical um
  387. 13:18human thing, which is to seek feedback
  388. 13:20early on, get something small working,
  389. 13:23and then expand out from there. Now, we
  390. 13:25can try to encourage the agent to do
  391. 13:26that by just saying, you know, don't
  392. 13:28code layer by layer, Um make sure that
  393. 13:30you create a small slice first and then
  394. 13:32go from there. But what if instead we
  395. 13:34use the leading word? We said, "vertical
  396. 13:36slice" is our leading word. We want to
  397. 13:39slice up the work instead of horizontal
  398. 13:41slices into vertical slices. A vertical
  399. 13:44slice is a pretty well-known terminology
  400. 13:46in development, and so this will
  401. 13:48hopefully trigger the agent's priors and
  402. 13:50it will understand what we mean. We
  403. 13:52don't just have to like have a two-word
  404. 13:54skill where it just says vertical slice.
  405. 13:56What we're doing is we're packing lots
  406. 13:57of meaning into a relatively short
  407. 14:00phrase that we then repeat throughout
  408. 14:02the skill. The cool thing about this
  409. 14:03technique is you can know if it's worked
  410. 14:05because you say vertical slice in your
  411. 14:07skill, and then you'll notice in the
  412. 14:09reasoning traces that it's saying,
  413. 14:10"Okay, we're going to do this as a thin
  414. 14:12vertical slice." Then you should get
  415. 14:14better implementation plans. Everyone
  416. 14:15I've explained this technique to sort of
  417. 14:17feels like, "Oh yeah, I've been doing
  418. 14:19that for a while. I've been using these
  419. 14:21little phrases to try to encourage the
  420. 14:23agent to do what I want." All I'm asking
  421. 14:26you now is to use those consistently
  422. 14:28within your skills and watch in the
  423. 14:31thinking traces as the agent adopts your
  424. 14:33way of doing it. So often if the agent
  425. 14:35isn't doing what you want, you need to
  426. 14:37make your leading words more consistent,
  427. 14:40more powerful, and look for others
  428. 14:42because, you know, English is a pretty
  429. 14:45wide API in terms of different functions
  430. 14:47you can call, different things you can
  431. 14:49experiment with, and there are many
  432. 14:50leading word candidates out there. And
  433. 14:52agents are actually pretty good at
  434. 14:53helping you think of them. Another
  435. 14:55little lever you can use with agents is
  436. 14:58sometimes the agent just doesn't do
  437. 15:00enough legwork. What I mean by this is
  438. 15:02that, okay, we're on a step, let's say,
  439. 15:05and maybe the step is to ask clarifying
  440. 15:07questions or to explore the code base,
  441. 15:09and the agent just doesn't do enough of
  442. 15:12it. It doesn't put enough effort into
  443. 15:14that particular step. A real classic
  444. 15:16case of this and something that I have
  445. 15:18found almost everywhere it
  446. 15:20is plan mode. Because in plan mode, we
  447. 15:23have two steps. We have ask clarifying
  448. 15:25questions and then create a plan. And
  449. 15:28what I have found in every single
  450. 15:29implementation of plan mode I've tried
  451. 15:31is that ask clarifying questions just,
  452. 15:34you know, it doesn't ever do enough
  453. 15:36legwork. It sees that its ultimate goal
  454. 15:38is to create a plan and so it just does
  455. 15:40a small amount of legwork with ask
  456. 15:42clarifying questions, ask you a couple
  457. 15:43of things, and then eagerly creates the
  458. 15:46plan. So, what was my solution here?
  459. 15:48Instead of doing plan mode, I instead
  460. 15:50have a skill called grill with docs,
  461. 15:52which is kind of my ask clarifying
  462. 15:54questions phase. And then, I split that
  463. 15:57up into a separate skill. So, I split
  464. 15:59the planning into its own skill. So,
  465. 16:02grill with docs now is its own skill
  466. 16:04where the agent only sees that part of
  467. 16:07the process. And then, after grill with
  468. 16:09docs completes, we then go and do 2 PRD.
  469. 16:12In other words, we have step one and
  470. 16:14step two, but the agent only sees one
  471. 16:16step at a time. So, this is a really
  472. 16:18cool technique for increasing legwork on
  473. 16:21the step that you're on by hiding the
  474. 16:23future goal, hiding the future steps.
  475. 16:25It's not always necessary to split
  476. 16:27skills into individual steps, but in
  477. 16:31particular cases where you really want
  478. 16:33an extra chunk of legwork, it really
  479. 16:36there's no technique like it. It works
  480. 16:37very, very well. So, that is steering
  481. 16:39using leading words to capture what you
  482. 16:41want in small reusable tokens and then
  483. 16:44making sure that it's doing the right
  484. 16:46amount of legwork per step. So, let's
  485. 16:48head now into pruning. Now, pruning
  486. 16:50really is just a quick fire set of
  487. 16:52failure modes, different things that you
  488. 16:54can get wrong. And the first is fairly
  489. 16:56obvious is we do not want massive
  490. 16:59skills. Massive skills are usually a
  491. 17:02kind of symptom of something else going
  492. 17:04wrong. So, a symptom of one of these
  493. 17:05other failure modes. And the first one
  494. 17:07is pretty simple. Don't repeat yourself.
  495. 17:10You need to make sure you're watching
  496. 17:12out for duplication. And in general, I
  497. 17:14like to have every part of the skill to
  498. 17:17have a single source of truth. In other
  499. 17:19words, if you have a piece of reference
  500. 17:21material like the PRD template, let's
  501. 17:22say, or something even smaller like what
  502. 17:25is a test seam, you make sure that you
  503. 17:27don't repeat that in several places or
  504. 17:29like cover multiple steps in multiple
  505. 17:31places. Just make sure each part has a
  506. 17:34single source of truth and you're not
  507. 17:35repeating yourself even across reference
  508. 17:38material, too. The next way that skills
  509. 17:39get big is via sediment. And sediment is
  510. 17:43just a classic thing when people are
  511. 17:46working on the same set of docs, really,
  512. 17:48which is that everyone starts
  513. 17:50contributing to a shared markdown file.
  514. 17:52People add their own stuff. They don't
  515. 17:54feel brave enough to delete and modify
  516. 17:56anyone else's. And so you just end up
  517. 17:57with this huge amount of sediment with
  518. 18:00often irrelevant material for the skill,
  519. 18:02especially stuff that hasn't been laid
  520. 18:04out properly. With a skill with a lot of
  521. 18:06sediments, you really need to look at
  522. 18:07structure. That's the first thing you
  523. 18:09need to do. You need to make sure that
  524. 18:10the stuff that's been added is relevant
  525. 18:12for all branches. If it's not, then move
  526. 18:15it into the correct branches. Or if it's
  527. 18:16just totally irrelevant, maybe just
  528. 18:18remove it or kill it. Or maybe there's
  529. 18:21stuff in there that's totally stale, in
  530. 18:22which case you just need to kill it
  531. 18:24dead. The next failure mode is really
  532. 18:25common when an agent writes your skills,
  533. 18:28which are no-ops. So things inside the
  534. 18:31skill that appear to do something but
  535. 18:34don't actually influence the agent's
  536. 18:36behavior inside the context of the
  537. 18:37skill. Let's imagine we have an
  538. 18:38implement skill and we have an entire
  539. 18:40paragraph of the skill that tells the
  540. 18:42agent to write a long detailed commit
  541. 18:44message. What would happen if you just
  542. 18:46deleted that paragraph? Well, the agent
  543. 18:49would probably still write a decent like
  544. 18:51long commit message. People ask me a lot
  545. 18:53how I get my skills so small, and it's
  546. 18:56just using these techniques, using
  547. 18:58deletion tests, using um making sure
  548. 19:00that I compact things into leading
  549. 19:02words, I don't have anything irrelevant
  550. 19:04in there, and I don't have any sediment.
  551. 19:06And that finally brings us to the full
  552. 19:08sweep of things. Number one, we check
  553. 19:10the trigger. We make sure that it's
  554. 19:12firing at the right times. We check
  555. 19:14whether we're imposing context load or
  556. 19:16cognitive load. With structure, we think
  557. 19:18about branches. We think about
  558. 19:21structuring things into steps and
  559. 19:23reference. And we make sure that
  560. 19:25material that's only relevant for one
  561. 19:26branch is outside of the main skill.md.
  562. 19:29With steering, we're thinking about
  563. 19:30condensing text down into leading words
  564. 19:33and watching those leading words appear
  565. 19:35in the reasoning traces. And we're also
  566. 19:37thinking about legwork. Should we break
  567. 19:39this skill down further to increase its
  568. 19:42focus on the current phase by hiding the
  569. 19:44future phase phase from it. And with
  570. 19:46pruning, we're doing a final pruning
  571. 19:48pass over the entire skill, watching out
  572. 19:50for sediments, watching out for crud,
  573. 19:52and watching out especially for no-ops.
  574. 19:54Now, all of this stuff, the best way to
  575. 19:57get started with this framework is
  576. 19:58inside this skill, inside the writing
  577. 20:00great skills skill. You can check it out
  578. 20:03from my papago skills, download it, use
  579. 20:05it to improve your own skills, and maybe
  580. 20:07even use it to run over some community
  581. 20:10authored skills so you can check that
  582. 20:12the skills that you're actually pulling
  583. 20:14in are any good. If you want to follow
  584. 20:15along with my stuff, then I have a
  585. 20:17newsletter up on aihero.dev. And my
  586. 20:20plans for the next few months are to
  587. 20:21release an AI coding crash course, which
  588. 20:23is an intro to a lot of the stuff I've
  589. 20:25been talking about and how you get off
  590. 20:27the ground working with engineering and
  591. 20:29AI. I hope that what I've given you is
  592. 20:32enough to help you escape from skill
  593. 20:34hell or at least try to make the bitter
  594. 20:36journey out of there. I'm so sorry not
  595. 20:38to be able to attend in person, but
  596. 20:40thanks for watching. I'll see you very
  597. 20:42soon.

About this transcript

This page contains the full transcript of Building Great Agent Skills: The Missing Manual by AI Engineer, generated from the public captions YouTube serves with the video. The transcript has 4,204 words across 597 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.