YouTube2Text

What is an AI harness? I build one live in less than 30 minutes — Transcript

by How I AI · 4,324 words · 621 segments · language en · Watch on YouTube

Full transcript

  1. 0:00A harness is some code around an AI
  2. 0:03agent that makes it more effective. Why
  3. 0:06we've seen people build these specific
  4. 0:08use case harnesses is sometimes with the
  5. 0:11specific job, you just want to
  6. 0:14micromanage a little bit. You just want
  7. 0:15to be more prescriptive about how that
  8. 0:18job gets done. [music] I'm going to show
  9. 0:19you how it works and then we will talk
  10. 0:22about how I built it. So the interface I
  11. 0:25built for my harness is a terminal UI.
  12. 0:28The harness core is run on cloud
  13. 0:30agent SDK and then it's connected to
  14. 0:32real tools. So it's connected to [music]
  15. 0:34Sentry, Vercel, and then it's connected
  16. 0:36to linear and GitHub in terms of getting
  17. 0:38[music] tasks done. I think we all have
  18. 0:40done good work, but then now I've
  19. 0:42realized that these agents can help us
  20. 0:44solve very very specific problems by
  21. 0:46constraining that work. It's really like
  22. 0:48changed my mind about how work gets
  23. 0:51done.
  24. 0:52>> [music]
  25. 0:54>> Everybody's saying it's not the model,
  26. 0:56it's the harness. But you know what not
  27. 0:59everybody is saying?
  28. 1:00What is [music] a harness? In today's
  29. 1:03how I AI episode, I'm going to demystify
  30. 1:07the idea of a harness, write my own
  31. 1:09harness and show you how you can do the
  32. 1:11same, and explain to you why a custom
  33. 1:14harness makes sense and could be better
  34. 1:16than using Claude code or Codex alone.
  35. 1:19Let's get to it. This episode is brought
  36. 1:21to you by Bolt.new, the AI app builder
  37. 1:24for people who have ideas and want to
  38. 1:26ship them.
  39. 1:27Most AI tools spit out code that looks
  40. 1:30great in the demo and falls apart the
  41. 1:31second [music] you try to do anything
  42. 1:33real with it. Or they lock you into
  43. 1:36their own platform with no real way out.
  44. 1:38Bolt is different. [music] You describe
  45. 1:40what you want to build, a startup MVP, a
  46. 1:42landing page, an internal tool, a side
  47. 1:44project, and Bolt generates
  48. 1:46production-ready code in minutes.
  49. 1:49Connect Stripe or any other MCP, hook up
  50. 1:51your domain, and deploy it live.
  51. 1:54Founders are using Bolt to build
  52. 1:55businesses doing real [music] revenue.
  53. 1:58Product managers are shipping prototypes
  54. 2:00their teams actually use. Designers and
  55. 2:03marketers are launching campaigns
  56. 2:05without waiting in line. Anyone can
  57. 2:07build, engineering can ship, everyone
  58. 2:10[music] wins. You just need an idea and
  59. 2:12a weekend. Check it out at
  60. 2:14bolt.new/howiai.
  61. 2:18Before I get into how to build a
  62. 2:20harness, let's talk about what a harness
  63. 2:22is and I'm going to make it as simple as
  64. 2:26I can for all of you. A harness is some
  65. 2:29code around an AI agent. Yes, you heard
  66. 2:32it here first. A harness is just code
  67. 2:36around an AI agent that makes it more
  68. 2:38effective. Can that code have AI in it?
  69. 2:41Sure. Does that code have to have AI in
  70. 2:44it? Not necessarily. What is the goal of
  71. 2:47a harness? To make the AI better. It is
  72. 2:49so simple and I feel like the way that
  73. 2:52people have been talking about this have
  74. 2:54made it such a mystery that I wanted to
  75. 2:56make it just very clear to you all. It
  76. 2:59is just writing more code around your AI
  77. 3:01to make it more useful for a specific
  78. 3:04use case.
  79. 3:05So, what are the parts of a harness?
  80. 3:07Well, a harness is going to have
  81. 3:09specific context, it's going to be able
  82. 3:11to take specific actions, and it's going
  83. 3:13to have a goal of specific outcomes.
  84. 3:16It's just as simple as that. And I want
  85. 3:20to talk about when it makes sense to
  86. 3:22build a harness and when it doesn't. And
  87. 3:26I think you want to build a harness when
  88. 3:28the same workflow needs the same setup
  89. 3:32and the same outcomes. And so, it's kind
  90. 3:35of similar to when you would build an AI
  91. 3:37agent. In fact, harness agent, sometimes
  92. 3:39you can interchange some of these
  93. 3:41concepts, but really it's when there is
  94. 3:44a sort of combination of deterministic
  95. 3:47and non-deterministic workflow,
  96. 3:48step-by-step process tools use cases,
  97. 3:52you want your AI to follow up to do a
  98. 3:55specific job. Usually those jobs are
  99. 3:57like slightly more complex. And this is
  100. 4:00why you've seen these coding harnesses
  101. 4:02come out like coding is a job to be
  102. 4:04done. It needs specific tools. It
  103. 4:06typically goes through kind of a
  104. 4:07standard workflow. And so coding
  105. 4:09harnesses are very popular, but you
  106. 4:12could also do things like managing
  107. 4:14production incidents where you need to
  108. 4:15go through a specific process, getting
  109. 4:17PRs ready for release, of handling
  110. 4:20support escalations, managing
  111. 4:22migrations.
  112. 4:24Even non-technical use cases like doing
  113. 4:26research in a very specific way or
  114. 4:28consolidating docs in a very specific
  115. 4:30way. That's how you and why you would
  116. 4:31use a harness.
  117. 4:34So, how did I decide what kind of
  118. 4:37harness I would build? Well, I looked
  119. 4:39across my business at Chat Priority and
  120. 4:41I thought, "What am I doing sort of
  121. 4:44repeatedly and consistently that I think
  122. 4:46AI could be good at? That I think we
  123. 4:48could be doing better if we were more
  124. 4:50structured about the AI and how we used
  125. 4:52it." And I thought that fixing bugs, you
  126. 4:57all if you've listened to this podcast,
  127. 4:58look, I ship code so I ship bugs. Fixing
  128. 5:00bugs is a very specific workflow where
  129. 5:04we've built some custom internal tools
  130. 5:06that I've been generally doing with
  131. 5:07Claude Coder Codex, but I had a had this
  132. 5:09hypothesis that I could do a better job
  133. 5:12of triaging bugs if I built my own
  134. 5:14harness. And so, I picked Sentry
  135. 5:18debugging and sorry for the Claude slot
  136. 5:21content here. Um
  137. 5:23Sentry debugging and debugging Sentry
  138. 5:25issues, really figuring out the issue,
  139. 5:27using some of our custom internal tools,
  140. 5:30and then doing all the follow-up actions
  141. 5:33we do when we close bugs was like a good
  142. 5:35first harness. It had coding in it. It
  143. 5:38needed custom content and custom
  144. 5:40context. There were like specific
  145. 5:42outcomes I wanted to make sure that we
  146. 5:44followed like tracking everything in
  147. 5:46linear and writing follow-up docs that
  148. 5:48the rest of the engineering team could
  149. 5:49use. And so, we chose uh debugging our
  150. 5:53Sentry bugs, by we I mean me and Codex,
  151. 5:55chose debugging Sentry as a good use
  152. 5:58case to demonstrate how to build a
  153. 6:00harness.
  154. 6:01Now, why wouldn't I just use an AI
  155. 6:03coding tool directly? Well, I have been
  156. 6:05using AI coding tools directly. And I
  157. 6:08think the problem with using a
  158. 6:09general-purpose coding tool and why
  159. 6:12we've seen people build these specific
  160. 6:14use case harnesses is sometimes with a
  161. 6:17specific job, you just want to
  162. 6:20micromanage a little bit. You just want
  163. 6:21to be more prescriptive about how that
  164. 6:24job gets done. And so, if you can
  165. 6:26identify the right workflows, you can
  166. 6:28actually be more efficient, more
  167. 6:30consistent, and have better outcomes if
  168. 6:34you build a harness. So, for this
  169. 6:36specific use case, you know, with a
  170. 6:39direct AI tool like Claude Code, um I
  171. 6:42would have to explain what I want the
  172. 6:45the agent to do. So, I'd say like, "Dear
  173. 6:47agent, please fix this bug. Here it is."
  174. 6:50and send a link. Instead of this
  175. 6:51harness, I can literally just paste in
  176. 6:53the link and the agent already knows my
  177. 6:55intent, already knows what the job to be
  178. 6:57done.
  179. 6:58A second thing that I wasn't that
  180. 6:59worried about, but is interesting when
  181. 7:01you build your harnesses, you can be
  182. 7:02really prescriptive about what tools
  183. 7:04it's allowed to do and what it's allowed
  184. 7:05to execute and not. So, for example, if
  185. 7:08you wanted to build an investigate-only
  186. 7:11harness, you could make sure that your
  187. 7:15harness, your code editor, never
  188. 7:17actually wrote code. It only explored
  189. 7:20and explained root cause. You can also
  190. 7:22repeat the same process over time if you
  191. 7:25encode it in a harness. And so, if you
  192. 7:27want like a very precise step-by-step
  193. 7:30flow, including outcomes, so for us,
  194. 7:32every time we fixed a
  195. 7:35Sentry bug, we want it documented in
  196. 7:36Linear, we want a very specific report,
  197. 7:39we might even want to follow up with
  198. 7:40customers that it was impacted with.
  199. 7:42You could encode that in a skill, but
  200. 7:44then again, you have to babysit it. When
  201. 7:46we build this harness, we knew it would
  202. 7:47happen every time. And then, from a
  203. 7:51model perspective, you can do multimodal
  204. 7:53routing and all sorts of interesting
  205. 7:54things and ways that you couldn't with a
  206. 7:56general-purpose AI model. So, I'm going
  207. 7:59to show you how it works, and then we
  208. 8:01will talk about how I built it. Okay, so
  209. 8:05the interface I built for my harness is
  210. 8:07a terminal UI, again, like Claude Code
  211. 8:09or Code Geeks, something you would run
  212. 8:10in your eye in a UI.
  213. 8:13And just so you know,
  214. 8:14your harness does not have to be a TUI.
  215. 8:17It doesn't have to be a CLI. It doesn't
  216. 8:19even have to have letters. It could be a
  217. 8:21web app.
  218. 8:22I did it in a TUI one because I haven't
  219. 8:24built one in a while. I thought it would
  220. 8:25be fun. And two, I just want to show
  221. 8:27that building your own custom harness
  222. 8:29means you can build your own custom
  223. 8:31interface into these AI agents as well.
  224. 8:34So, the harness is the whole experience,
  225. 8:37including the human experience that
  226. 8:39makes it more useful and easier to use.
  227. 8:42And so, um this TUI is pretty easy to
  228. 8:45invoke. I just run TUI. You can see it
  229. 8:48here. It's kind of cute. It's been made
  230. 8:50cute. Um I use this library called Ink,
  231. 8:53which helps you make cute TUIs. I don't
  232. 8:55think they would say cute, but I'm going
  233. 8:56to say cute. And you can see here that
  234. 8:59this terminal UI really reflects the
  235. 9:01structure of the harness itself. So, you
  236. 9:03see all the runs um that it's done so
  237. 9:06far,
  238. 9:08errors and how it's fixed things. And
  239. 9:10then, sort of our process, which is it
  240. 9:12gathers evidence, it streams in
  241. 9:14activities, and then it builds some
  242. 9:16artifacts. And so, I'm going to actually
  243. 9:18have it investigate this Sentry error
  244. 9:20over here. It's one where our edit um
  245. 9:24operations are getting dropped sometime
  246. 9:26by the agents. And that has now kicked
  247. 9:30off our specific harness. So, what it's
  248. 9:33going to do is it's going to start this
  249. 9:34investigation
  250. 9:36run. It's going to kick off a Claude SDK
  251. 9:38session, which is a fundamental part of
  252. 9:40how I built this. It's going to go ahead
  253. 9:42and start gathering evidence and coming
  254. 9:44up with a root cause hypothesis of
  255. 9:47what's causing this issue and how we
  256. 9:49might fix it. Now, as you can see, I
  257. 9:51chose I investigate, not F fix. So, the
  258. 9:55investigation should not touch and
  259. 9:58modify files. And again, this is
  260. 10:00something that I would have had to like
  261. 10:02prompt to the agent and say, I only want
  262. 10:04you to investigate. I do not want you to
  263. 10:06ship a fix. But instead, I can just
  264. 10:08click I, paste in that Sentry issue, and
  265. 10:12it's off to the races. This episode is
  266. 10:16brought to you by Customer.io.
  267. 10:18You're here because you'd rather use AI
  268. 10:20than talk about it. With Customer.io,
  269. 10:23you describe the campaign you want to
  270. 10:25build, and the AI agent creates for you
  271. 10:28the audience, the messages, and the
  272. 10:30timing. You review it, make any changes
  273. 10:33you want, and launch.
  274. 10:35Instead of spending hours stitching
  275. 10:36together tools and workflows, you can
  276. 10:38focus on the work that actually drives
  277. 10:40growth. Every campaign is tied back to
  278. 10:43results, so you can see what's working
  279. 10:45and what to do next. [music]
  280. 10:47More than 9,000 brands use Customer.io
  281. 10:50to turn the data they already have into
  282. 10:52messages customers remember.
  283. 10:54Visit customer.io/howiai
  284. 10:59to try it today.
  285. 11:01Customer.io, more impact from every
  286. 11:04message. While this is running, I'm
  287. 11:06going to just go and show you a little
  288. 11:08bit about how this works and how I have
  289. 11:11actually built it. Okay, so this is the
  290. 11:13high-level architecture of the app. So,
  291. 11:16the front end is a terminal UI or a C
  292. 11:19CLI. Each invocation of the harness we
  293. 11:23call a run, so it's running a task. Each
  294. 11:26task has a specific input, usually
  295. 11:28that's a Sentry issue. And then there
  296. 11:31are specific flags I put on the harness
  297. 11:33that allow it to edit the source, modify
  298. 11:38the inputs, or even message customers
  299. 11:40only if I flag and approve it. So again,
  300. 11:43this is just a little bit more control
  301. 11:44over how the agent works. The harness
  302. 11:47core is run on Claude agent SDK and so
  303. 11:51all the agentic planning is run through
  304. 11:53the Claude agent SDK which has some of
  305. 11:56the primitives of Claude code including
  306. 11:58wrapping files and writing files and all
  307. 12:00those sorts of things that we find
  308. 12:01useful.
  309. 12:02And then what's really interesting about
  310. 12:04this harness, and you've seen in other
  311. 12:06harnesses like open claw, is it can
  312. 12:09create its own artifacts in its file
  313. 12:12store. And so we have this artifact
  314. 12:14store, I will show it to you in a
  315. 12:15minute, and it basically saves all the
  316. 12:18evidence from these runs to the file
  317. 12:20system for the agent to use in the
  318. 12:23future. And then it's connected to real
  319. 12:24tools, so it's connected to century, for
  320. 12:26sale, the Claude SDK, it's running
  321. 12:29Sonnet 46. I think that's the right the
  322. 12:31right model for the job. And then it's
  323. 12:32connected to linear and GitHub in terms
  324. 12:34of getting tasks done. Now what's really
  325. 12:38interesting as well is you can prompt
  326. 12:42this in a custom way. So instead of the
  327. 12:44general like you are Claude code, make
  328. 12:46no mistakes, you are our, you know, sort
  329. 12:49of model genius.
  330. 12:51I'm saying specifically that you're
  331. 12:53working inside the Chat Purity
  332. 12:54engineering harness, it's Chat Purity
  333. 12:56specific, it's not an open-ended coding
  334. 12:58system. We want to use these artifacts
  335. 13:02as a source of truth and here's the plan
  336. 13:04to attack a very specific problem. And
  337. 13:08what I want you to return is X, Y, and
  338. 13:10Z. And again, I don't have to copy and
  339. 13:12paste this, I don't even have to put it
  340. 13:13in a skill where hopefully it will get
  341. 13:15invoked in the right way. I've actually
  342. 13:17encoded this in a very specific step in
  343. 13:20the harness to make sure that the model
  344. 13:22falls it every time. And so
  345. 13:25there's several of these types of custom
  346. 13:27prompts inside my harness. There is um
  347. 13:30the artifacts that get generated. There
  348. 13:32are tool policies around like what tools
  349. 13:34can be called and which ones can't. And
  350. 13:36then um I have just decided again to use
  351. 13:39Claude Sonnet 4 6, um which is really I
  352. 13:42think the right model for this
  353. 13:43particular workflow.
  354. 13:45Okay, I want to talk a little bit about
  355. 13:47the code and how you generate this. And
  356. 13:49then just like a peek behind the scenes.
  357. 13:51I actually ran dueling Claude code and
  358. 13:54Codex sessions and essentially said like
  359. 13:57help me build a harness. I think I want
  360. 13:59to use the Claude agent SDK. Here's what
  361. 14:01I would like it to do. And then like
  362. 14:03closed my eyes and tried to get it done.
  363. 14:06Honestly, it was not a one-shot. I don't
  364. 14:08know if it was my prompting or the
  365. 14:11models were being funky. It was GPT 5.5
  366. 14:14and Opus, but both of them really wanted
  367. 14:16to build something super deterministic.
  368. 14:19So, they like really resisted putting
  369. 14:21any AI in the harness and I had had to
  370. 14:25really prompt it very, very specifically
  371. 14:28to get what I want. So, I would say if
  372. 14:30you were trying to do this, I would be
  373. 14:32very specific about the workflow. I
  374. 14:35would be very specific about the tools.
  375. 14:37I would be very specific about where
  376. 14:39custom prompts make sense. And then I
  377. 14:42would suggest using an agent SDK either
  378. 14:45from Claude or from OpenAI to run most
  379. 14:49of it because without that prompting, I
  380. 14:52just did not get what I wanted out of
  381. 14:54these models. The second thing I will
  382. 14:55say, funnily enough, Codex did the best
  383. 14:58job at building the agent, but it used
  384. 15:01Claude agent's SDK to actually implement
  385. 15:04the agent. So, we are spanning across
  386. 15:05models and spanning across coding agents
  387. 15:08here.
  388. 15:09But the actual harness itself is pretty
  389. 15:11simple. It's got sort of a high-level
  390. 15:15index of how you get to the TUI. And
  391. 15:17then it's got like, I don't know, eight
  392. 15:19files of specific things it can do. So,
  393. 15:23it can hunt for bugs in Sentry. Um it
  394. 15:27has a Sentry adapter to effectively use
  395. 15:29the Sentry API in a very specific way.
  396. 15:32So, instead of using the MCP generally,
  397. 15:35instead of like having your coding agent
  398. 15:37wander through all these traces, I'm
  399. 15:39just very precise about exactly what I
  400. 15:41think you need to pull from a bug report
  401. 15:42perspective, what's useful, what's not,
  402. 15:44and made that connector really
  403. 15:46opinionated. It's got similar a linear
  404. 15:49integration and a Vercel integration and
  405. 15:51a GitHub integration. So, again, not
  406. 15:53like generally how you can use these
  407. 15:55tools, but specifically how you would
  408. 15:58use these tools when you are searching
  409. 15:59for a bug. And then, after those tools
  410. 16:03and data sources are used, the bug is
  411. 16:05identified and triaged, then there is
  412. 16:08this artifact file here that outputs and
  413. 16:12spits out the specific artifact I want
  414. 16:14to see after a bug run is done.
  415. 16:17And that artifact bundle looks something
  416. 16:19like this. So, it's literally just uh
  417. 16:23the task run, um which is all the
  418. 16:25messages, the report, so what was the
  419. 16:28Sentry issue, here's a brief on what we
  420. 16:30discovered, here's any logs that we
  421. 16:33think are relevant, what the Cloud
  422. 16:35Worker ended up doing, and then the
  423. 16:37summary of the output. And then, we also
  424. 16:40output this beautiful HTML file um that
  425. 16:42I can show you that shows you
  426. 16:45what happened and how it all worked, as
  427. 16:47well as a worker report. So, I will show
  428. 16:49you those outcomes, as well. Just
  429. 16:51pulling up this code for you again,
  430. 16:56it's pretty straightforward. It's giving
  431. 16:58me all the instructions on where to put
  432. 17:00my specific API keys, and then, I can
  433. 17:03just run it in this very opinionated
  434. 17:06way. So, in addition to running the TUI,
  435. 17:09which lets me sort of like navigate
  436. 17:10through the UI and use this harness,
  437. 17:12something I might want to do as a human,
  438. 17:14it also has built these really easy
  439. 17:16command line tools, where if I just
  440. 17:18quickly want to run this harness against
  441. 17:21specific issues with specific flags on
  442. 17:24tool use, I can definitely do that. And
  443. 17:26what's kind of interesting about this is
  444. 17:28yes, I built this harness and you can
  445. 17:30see here I built this like fun UI so
  446. 17:33that I could use it in a fun way and it
  447. 17:35makes for a better demo, but really this
  448. 17:38harness is a structured way to give
  449. 17:41agents the job of running these
  450. 17:44investigations on an on a simpler basis.
  451. 17:46And so you can imagine while I design
  452. 17:48the TUI for human, actually giving a
  453. 17:52kind of
  454. 17:53all intelligent agent a specific harness
  455. 17:57to solve a specific problem with agents
  456. 17:59in that,
  457. 18:00I think that's how you're going to get
  458. 18:01real leverage and really custom outcomes
  459. 18:04out of things like coding agents like
  460. 18:07Claude Code. And so going through this
  461. 18:09process has really opened my mind to
  462. 18:12we've gotten so used to like the open
  463. 18:14chat field. Like if I just type in, the
  464. 18:17agent will do good work. And I think we
  465. 18:18all have done good work. But then now
  466. 18:20I've realized that these agents can help
  467. 18:22us solve very, very specific problems
  468. 18:25using other agents and by constraining
  469. 18:27that work, we can actually get specific
  470. 18:29jobs done really efficiently and then
  471. 18:32use the general purpose agent to sort of
  472. 18:35orchestrate it. So it's really like
  473. 18:36changed my mind about how work gets
  474. 18:39done.
  475. 18:40As you can see here again, it's just a
  476. 18:42couple files. It's really not too much.
  477. 18:45The adapters to the data sources, um a
  478. 18:49couple workflows. In particular, this
  479. 18:52bug hunter workflow which just goes
  480. 18:54through exactly how we want to hunt
  481. 18:56bugs, including how we want to put
  482. 18:59together summaries of bug reports and
  483. 19:02then some files here in terms of running
  484. 19:05the TUI or the CLI. And then as I said,
  485. 19:08we have this artifacts folder that gets
  486. 19:10updated every time a run happens where I
  487. 19:13can click in and actually see exactly
  488. 19:17what happened out of a run. So, let's go
  489. 19:19and see if this run happened well and
  490. 19:24what I can find out. So, now I have the
  491. 19:26full context. Here's the investigation
  492. 19:28brief and I can go look for it. So, this
  493. 19:31is Bug Hunter C7. Let's see if I can
  494. 19:35find this one.
  495. 19:36Here it is. Here's the investigation
  496. 19:38brief on that edit document operations
  497. 19:40dropped. I have confirmed evidence. So,
  498. 19:43it's saying, "Yes, there was definitely
  499. 19:45a Sentry warning. It's impacted 150
  500. 19:47users. It's still happening hourly. Um
  501. 19:50it's a warning, so it's not an actual
  502. 19:52error."
  503. 19:54And the Versel logs were unavailable and
  504. 19:56so we weren't able to use that data. And
  505. 19:58then it found likely root causes. So,
  506. 20:01invalid original range or overlapping
  507. 20:04original range. And so, it's identified
  508. 20:06a couple potential root causes as well
  509. 20:09as a blind spot in this particular
  510. 20:10function. It's told me exactly where in
  511. 20:13the product surface um the issue is and
  512. 20:17then how I would actually verify this by
  513. 20:20fetching a raw Sentry event to see if
  514. 20:23the issues that they've identified are
  515. 20:25correct. It's identifying should it um
  516. 20:29issue a linear issue and it says, "Yes,
  517. 20:31we should definitely make a linear issue
  518. 20:32to fix this." And so, this should get
  519. 20:34assigned to somebody. And then it
  520. 20:36doesn't recommend turning on patch mode
  521. 20:39and actually fixing this. So, again,
  522. 20:41this is like a very specific outcome I
  523. 20:43wanted. I wanted to say like, "What's
  524. 20:45all the evidence?
  525. 20:46Priority rank the root causes. Make a
  526. 20:48suggestion on the next step if we need
  527. 20:51to verify this more. Tell me if I need
  528. 20:53to assign it to somebody in Linear and
  529. 20:55then tell me if you can fix it." And
  530. 20:57they're saying, "No, I don't think I can
  531. 20:58fix it yet. I need a little bit more
  532. 21:00information."
  533. 21:02And all of that is built because I have
  534. 21:05done this like very specific workflow
  535. 21:08and encoded that in
  536. 21:11what we're calling a harness, which is
  537. 21:12just code around an agent. So, how would
  538. 21:16you you build your own harness? I feel
  539. 21:18like hopefully you're still with me not
  540. 21:20too much of that went over your head.
  541. 21:22Just to reiterate, I just identified a
  542. 21:25specific workflow. I determined what the
  543. 21:29run against the task would look like. I
  544. 21:32made very opinionated calls to tools or
  545. 21:35data sources, so I didn't just say like
  546. 21:37use an MCP, although that could be part
  547. 21:39of your harness. But what I did is I
  548. 21:41made adapters that made the calls to
  549. 21:44these external APIs and tools very
  550. 21:45specific. I thought about what the
  551. 21:48structured artifacts out of that
  552. 21:50workflow might be. I decided what rules
  553. 21:53and permissions I wanted to give this
  554. 21:55harness and which ones I didn't. I
  555. 21:57decided whether I wanted to use Claude
  556. 21:59Code or Codex or a model router to
  557. 22:02actually run these things. And then I
  558. 22:04built a surface to interact with this
  559. 22:06agent. So, I built a TUI so I could
  560. 22:09actually look and work with this harness
  561. 22:12in a way. It could be a TUI, it could be
  562. 22:14a CLI, it could be a web app, but I
  563. 22:16built some way to interact with this.
  564. 22:18So, this is what you need to do.
  565. 22:20Identify a workflow. Uh really write it
  566. 22:22down on, you know, proverbial paper,
  567. 22:24HTML or markdown. Figure out what
  568. 22:27sources of data you want and then plug
  569. 22:29it all into Claude Code or into Codex as
  570. 22:32I did and have it build your own harness
  571. 22:35and then test it against real data. So,
  572. 22:38that's it. I just I really hope that you
  573. 22:41walk away from this realizing that these
  574. 22:44mystery terms like harness are not that
  575. 22:47mysterious. A harness is simply putting
  576. 22:50some structure around how AI works. Yes,
  577. 22:53Cursor is like a really complex harness.
  578. 22:56Yes, Codex and Claude Code are very
  579. 22:58complex coding harnesses. But at the end
  580. 23:01of the day, they're code that wraps
  581. 23:03these AI agents and these AI calls to
  582. 23:05make them more efficient in doing a very
  583. 23:08specific job. And so whether you're
  584. 23:10doing that in a very prescriptive way
  585. 23:12like I just showed where I want to show
  586. 23:14you how I triage sentry bugs, do the
  587. 23:16investigation and pass it on to the
  588. 23:18team, or you're doing it a broad way
  589. 23:20like these general purpose coding agents
  590. 23:22that just have access to tools and
  591. 23:25context and methods that make the coding
  592. 23:27workflow better.
  593. 23:29That's all harnesses. You can think of
  594. 23:31harnesses that you can build. You can
  595. 23:33build them in the terminal. You can
  596. 23:35build them for CLIs. You can even build
  597. 23:37them as web apps. I'm starting to
  598. 23:39hypothesize that a wrapper is just a
  599. 23:42harness and that is going to upgrade
  600. 23:44everything that I've vibe coded over the
  601. 23:46last 3 years.
  602. 23:48This has been totally a learning
  603. 23:50experience for me here on How AI. This
  604. 23:52is my very first harness that I've built
  605. 23:55live on the show. I hope it's useful for
  606. 23:57you. And if you're interested in me
  607. 23:59building other and demystifying AI
  608. 24:01terms, let me know in the comments.
  609. 24:04Thanks for joining How AI.
  610. 24:07Thanks so much for watching. If you
  611. 24:09enjoyed the show, please like and
  612. 24:11subscribe here on YouTube or even
  613. 24:13better, leave us a comment with your
  614. 24:14thoughts. You can also find this podcast
  615. 24:17on Apple Podcasts, Spotify, or your
  616. 24:19favorite podcast app. Please consider
  617. 24:22leaving us a rating and review which
  618. 24:24will help others find the show. You can
  619. 24:26see all our episodes and learn more
  620. 24:28about the show at howiaiipod.com.
  621. 24:32See you next time.

About this transcript

This page contains the full transcript of What is an AI harness? I build one live in less than 30 minutes by How I AI, generated from the public captions YouTube serves with the video. The transcript has 4,324 words across 621 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.