YouTube2Text

No Vibes Allowed: Solving Hard Problems in Complex Codebases – Dex Horthy, HumanLayer — Transcript

by AI Engineer · 4,538 words · 624 segments · language en · Watch on YouTube

Full transcript

  1. 0:13[music]
  2. 0:20Hi everybody. How y'all doing?
  3. 0:24>> It's exciting. I'm Dex. Uh, as they did
  4. 0:26in the great intro, I've been hacking on
  5. 0:27agents for a while. Um, our talk 12
  6. 0:30factor agents at AI engineer in June was
  7. 0:32one of the top talks of all time. Uh, I
  8. 0:35think top eight or something. One of the
  9. 0:36best ones from from AI engineer in June.
  10. 0:38May or may not have said something about
  11. 0:39context engineering. Um, why am I here
  12. 0:42today? What am I here to talk about? Um,
  13. 0:44I want to talk about one of my favorite
  14. 0:45talks from AI engineer in June. And I
  15. 0:47know we all got the update from Eigor
  16. 0:48yesterday, but they wouldn't let me
  17. 0:50change my slides. So, this is going to
  18. 0:51be about what Eigor talked about in
  19. 0:53June. uh basically that they surveyed a
  20. 0:55100,000 developers across all company
  21. 0:57sizes and they found that most of the
  22. 0:59time you use AI for software engineering
  23. 1:01you're doing a lot of rework a lot of
  24. 1:02codebase churn uh and it doesn't really
  25. 1:05work well for complex tasks brownfield
  26. 1:07code bases um and you can see in the
  27. 1:10chart basically you are shipping a lot
  28. 1:11more but a lot of it is just reworking
  29. 1:13the slop that you shipped last week so
  30. 1:16uh and then the other side right was
  31. 1:18that uh if you're doing green field
  32. 1:20little versel dashboard something like
  33. 1:22this then it's going to work great. Uh
  34. 1:25if you're going to go in a 10-year-old
  35. 1:27Java at codebase, maybe not so much. And
  36. 1:29this matched my experience personally
  37. 1:31and talking to a lot of founders and
  38. 1:32great engineers, too much slop uh tech
  39. 1:35debt factory. It's just it's not going
  40. 1:36to work from our codebase. Like maybe
  41. 1:37someday when the models get better, but
  42. 1:40that's what context engineering is all
  43. 1:41about. How can we get the most out of
  44. 1:43today's models? How do we manage our
  45. 1:45context window? So we talked about this
  46. 1:47in August. Um I have to confess
  47. 1:49something. The first time I used cloud
  48. 1:51code, I was not impressed. It was like,
  49. 1:53okay, this is a little bit better. I get
  50. 1:54it. I like the UX. Um, but since then,
  51. 1:57we as a team figured something out. Um,
  52. 2:00that we were actually able to get, you
  53. 2:01know, 2 to 3x more throughput. And we
  54. 2:03were shipping so much that we had no
  55. 2:05choice but to change the way we
  56. 2:07collaborated. We rewired everything
  57. 2:09about how we build software. Uh, it was
  58. 2:11a team of three. It took eight weeks. It
  59. 2:13was really freaking hard. Uh, but now
  60. 2:15that we solved it, we're we're never
  61. 2:16going back. This is the whole no slop
  62. 2:18thing. I think I think we got somewhere
  63. 2:20with this went super viral on HackerNews
  64. 2:22in September. Uh we have thousands of
  65. 2:24folks who have gone on to GitHub and
  66. 2:25grabbed our you know research plan
  67. 2:26implement prompt system. Um so the goals
  68. 2:29here which we kind of backed our way
  69. 2:31into we need AI that can work well in
  70. 2:34brownfield code bases that can solve
  71. 2:36complex problems. No slop, right? No
  72. 2:39more slop. Uh and we had to maintain
  73. 2:42mental alignment. I'll talk a little bit
  74. 2:43more about what that means in a minute.
  75. 2:44And of course we want to spend with
  76. 2:46everything we want to spend as many
  77. 2:47tokens as possible. what we can offload
  78. 2:48meaningfully to the AI is really really
  79. 2:51important. Um, super high leverage. So,
  80. 2:53this is advanced context engineering for
  81. 2:55coding agents. Um, I'll start with kind
  82. 2:57of like framing this. The most naive way
  83. 3:00to use a coding agent is to ask it for
  84. 3:02something and then tell it why it's
  85. 3:03wrong and resteere it and ask and ask
  86. 3:05and ask until you run out of context or
  87. 3:07you give up or you cry. Um, we can be a
  88. 3:10little bit smarter about this. Most
  89. 3:11people discover this pretty early on in
  90. 3:13their AI like exploration. uh is that it
  91. 3:17might be better if you start a
  92. 3:18conversation and you're off track that
  93. 3:22uh you just start a new context window.
  94. 3:24You say, "Okay, we went down that path.
  95. 3:25Let's start again. Same prompt, same
  96. 3:26task, but this time we're going to go
  97. 3:28down this path and like don't go over
  98. 3:29there cuz that doesn't work." So, uh how
  99. 3:32do you know when it's time to start
  100. 3:34over?
  101. 3:35If you see this,
  102. 3:39it's probably time to start over, right?
  103. 3:41This is what Claude says when you tell
  104. 3:43it it's screwing up.
  105. 3:46Um, so we can be even smarter about
  106. 3:47this. We can do what I call intentional
  107. 3:49compaction. Um, and this is basically
  108. 3:51whether you're on track or not, you can
  109. 3:53take uh your existing context window and
  110. 3:56ask the agent to compress it down into a
  111. 3:58markdown file. You can review this, you
  112. 4:00can tag it, and then when the new agent
  113. 4:01starts, it gets straight to work instead
  114. 4:03of having to do all that searching and
  115. 4:04codebase understanding and getting
  116. 4:06caught up. Um, what goes into
  117. 4:08compaction? Well, the question is like
  118. 4:10what takes up space in your context
  119. 4:12window. So, um, it's looking for files,
  120. 4:15it's understanding code flow, it's
  121. 4:17editing files, it's test and build
  122. 4:19output. And if you have one of those
  123. 4:20MCPs that's dumping JSON and a bunch of
  124. 4:22UU ids into your context window, you
  125. 4:25know, God help you. Uh, so what should
  126. 4:28we compact? I'll get more specifics
  127. 4:29here, but this is a really good
  128. 4:30compaction. This is exactly what we're
  129. 4:32working on. The exact files and line
  130. 4:34numbers that matter to the problem that
  131. 4:35we're solving. Um, why are we so
  132. 4:38obsessed with context? Because LMS are
  133. 4:41actually got roasted on YouTube for this
  134. 4:42one. And they're not pure functions cuz
  135. 4:43they're nondeterministic, but they are
  136. 4:45stateless. And the only way to get
  137. 4:46better better performance out of an LLM
  138. 4:49is to put better tokens in and then you
  139. 4:51get better tokens out. And so every turn
  140. 4:53of the loop when Claude is picking the
  141. 4:54next tool or any coding agent is picking
  142. 4:56the next and there could be hundreds of
  143. 4:57right next steps and hundreds of wrong
  144. 4:59next steps. But the only thing that
  145. 5:01influences what comes out next is what
  146. 5:03is in the conversation so far. So we're
  147. 5:05going to optimize this context window
  148. 5:07for correctness, completeness, size, and
  149. 5:10a little bit of trajectory. And the
  150. 5:11trajectory one is interesting because a
  151. 5:13lot of people say, "Well, I I told the
  152. 5:14agent to do something and it did
  153. 5:16[clears throat] something wrong. So, I
  154. 5:17corrected it and I yelled at it and then
  155. 5:19it did something wrong again and then I
  156. 5:20yelled at it." And then the LM is
  157. 5:21looking at this conversation says,
  158. 5:23"Okay, cool. I did something wrong. The
  159. 5:24human yelled at me and then I did
  160. 5:25something wrong and the human yelled at
  161. 5:26me." So, the next most likely conver
  162. 5:27token in this conversation is I better
  163. 5:30do something wrong so the human can yell
  164. 5:31at me again. So, mind be mindful of your
  165. 5:34trajectory. If you were going to invert
  166. 5:36this, the worst thing you can have is
  167. 5:37incorrect information, then missing
  168. 5:39information, and then just too much
  169. 5:41noise. Um, if you like equations,
  170. 5:43there's a dumb equation if you want to
  171. 5:45think about it this way. Um, Jeff
  172. 5:48Huntley uh did a lot of research on
  173. 5:49coding agents. Uh, he put it really
  174. 5:51well. Just the more you use the context
  175. 5:53window, the worse outcomes you'll get.
  176. 5:55This leads to a concept I'm in a very
  177. 5:56very academic concept called the dumb
  178. 5:58zone. So, you have your context window.
  179. 6:01You have 168,000 tokens roughly. Some
  180. 6:03are reserved for output and compaction.
  181. 6:05This varies by model. Um, but we'll use
  182. 6:07cloud code as an example here. Around
  183. 6:09the 40% line is where you're going to
  184. 6:10start to see some diminishing returns
  185. 6:12depending on your task. Um, if you have
  186. 6:15too many MCPs in your coding agent, you
  187. 6:17are doing all your work in the dumb zone
  188. 6:18and you're never going to get good
  189. 6:20results. People talked about this. I'm
  190. 6:22not going to talk about that one. Your
  191. 6:23mileage may vary. 40% is like it depends
  192. 6:24on how complex the task is, but this is
  193. 6:26kind of a good guideline. Um so back to
  194. 6:29compaction or as I will call it from now
  195. 6:31on cleverly avoiding the dumb zone. Um
  196. 6:35we can do sub agents. Um if you have a
  197. 6:37front-end sub aent and a backend sub
  198. 6:39aent and a QA sub aent and a data data
  199. 6:41scientist sub aent
  200. 6:43please stop. Sub aents are not for
  201. 6:46anthropomorphizing roles. They are for
  202. 6:47controlling context. And so what you can
  203. 6:49do is if you want to go find how
  204. 6:51something works in a large codebase um
  205. 6:53you can steer the coding agent to do
  206. 6:55this if it supports sub agents or you
  207. 6:56can build your own sub agent system. But
  208. 6:58basically you say hey go find how this
  209. 7:00works and it can fork out a new context
  210. 7:02window that is going to go do all that
  211. 7:04reading and searching and finding and
  212. 7:06reading entire files and understanding
  213. 7:08the codebase and then just return a
  214. 7:11really really succinct message back up
  215. 7:13to the parent agent of just like hey the
  216. 7:15file you want is here. parent agent can
  217. 7:18read that one file and get straight to
  218. 7:20work. And so this is really powerful. If
  219. 7:22you wield these correctly, you can get
  220. 7:24good responses like this and then you
  221. 7:26can manage your context really, really
  222. 7:27well. Um, what works even better than
  223. 7:29sub agents or like a layer on top of sub
  224. 7:31aents is a workflow I call frequent
  225. 7:33intentional compaction. We're going to
  226. 7:35talk about research plan implement in a
  227. 7:37minute, but like the point is you're
  228. 7:38constantly st keeping your context
  229. 7:40window small. You're building your
  230. 7:42entire workflow around context
  231. 7:44management. So comes in three phases.
  232. 7:46research, plan, implement. Um, and we're
  233. 7:49going to try to stay in the smart zone
  234. 7:50the whole time. So, the research is all
  235. 7:52about understanding how the system
  236. 7:53works, finding the right files, staying
  237. 7:55objective. Here's a prompt you can use
  238. 7:57to do research. Here's the output of um,
  239. 7:59a research prompt. These are all open
  240. 8:01source. You can go grab them and play
  241. 8:02with them yourself. Um, planning, you're
  242. 8:05going to outline the exact steps. You're
  243. 8:06going to include file names and line
  244. 8:07snippets. You're going to be very
  245. 8:08explicit about how we're going to test
  246. 8:09things after every change. Here's a good
  247. 8:11planning prompt. Here's one of our
  248. 8:13plans. It's got actual code snippets in
  249. 8:14it. Um, and then we're gonna implement.
  250. 8:16And if you read one of these plans, you
  251. 8:18can see very easily how the dumbest
  252. 8:20model in the world is probably not going
  253. 8:21to screw this up. Um, so we just go
  254. 8:23through and we run the plan and we keep
  255. 8:25the context low. As a planning prompt,
  256. 8:27like I said, it's the least exciting
  257. 8:28part of the process. Um, I wanted to put
  258. 8:30this into practice. So, working for us,
  259. 8:32uh, I do a podcast with my buddy uh,
  260. 8:34Vibv, who's the CEO of a company called
  261. 8:35Boundary ML. Uh, and I said, "Hey, I'm
  262. 8:38going to try to oneshot a fix to your
  263. 8:39300,000line Rust codebase for a
  264. 8:41programming language."
  265. 8:43Um, and the whole episode goes in, it's
  266. 8:45like an hour and a half. Uh, I'm not
  267. 8:46going to talk through it right now, but
  268. 8:47we built a bunch of research and then we
  269. 8:48threw them out because they were bad.
  270. 8:49And then we made a plan and we made a
  271. 8:50plan without research and with research
  272. 8:52and compared all the results. It's a fun
  273. 8:53time. Uh, by that was Monday night. By
  274. 8:56Tuesday morning, we were on the show and
  275. 8:57the CTO had like seen the PR and like
  276. 9:00didn't realize I was doing it as a bit
  277. 9:01for a podcast and basically was like,
  278. 9:03"Yeah, this looks good. We'll get in the
  279. 9:04next release." He I think he was a
  280. 9:06little confused. Um, here's the the
  281. 9:08plan. But anyways, uh, yeah, confirmed
  282. 9:11works in brownfield code bases and no
  283. 9:14slop. But I wanted to see if we could
  284. 9:16solve complex problems. So, Vib was
  285. 9:18still a little skeptical. I sat down, we
  286. 9:20sat down for like 7 hours on a Saturday
  287. 9:21and we shipped 35,000 lines of code to
  288. 9:24BAML. One of the PRs got merged like a
  289. 9:26week later. I will say some of this is
  290. 9:28codegen. You know, you update your
  291. 9:29behavior. All the golden files update
  292. 9:31and stuff, but we shipped a lot of code
  293. 9:32that day. Um, he estimates it was about
  294. 9:341 to 2 weeks and 7 hours. And uh, so
  295. 9:37cool. We can solve complex problems.
  296. 9:40There are limits to this. I sat down
  297. 9:41with my buddy Blake. We tried to remove
  298. 9:43Hadoop dependencies from Parket Java. If
  299. 9:46you know what Paret Java is, I'm sorry
  300. 9:49uh for whatever happened to you to get
  301. 9:51you to this point in your career. Uh it
  302. 9:53did not go well. Uh here's the plans,
  303. 9:55here's the research. Uh at a certain
  304. 9:57point, we threw everything out and we
  305. 9:58actually went back to the whiteboard. We
  306. 10:00had to actually once we had learned
  307. 10:01where were the where all the foot guns
  308. 10:03were, we we went back to okay, how is
  309. 10:05this actually going to fit together? Um,
  310. 10:07and this brings me to a really
  311. 10:08interesting point that Jake's going to
  312. 10:09talk about later. Uh, do not outsource
  313. 10:12the thinking. AI cannot replace
  314. 10:14thinking. It can only amplify the
  315. 10:16thinking you have done or the lack of
  316. 10:17thinking you have done. So people ask,
  317. 10:20so Dex, this is specri development,
  318. 10:22right? No, specri development is broken.
  319. 10:28Not the idea, but the phrase. Um,
  320. 10:32it's not well defined. This is Brietta
  321. 10:34from Thought Works. Um, and a lot of
  322. 10:36people just say spec and they mean a
  323. 10:37more detailed prompt. Does anyone
  324. 10:40remember this picture? Does anyone know
  325. 10:41what this is from? All right, that's a
  326. 10:43deep cut. Uh, there will never be a year
  327. 10:45of agents because of semantic diffusion.
  328. 10:47Martin Fowler said this in 2006. We come
  329. 10:49up with a good term with a good
  330. 10:51definition and then everybody gets
  331. 10:53excited and everybody starts meaning it
  332. 10:55to mean a hundred things to a 100
  333. 10:56different people and it becomes useless.
  334. 10:59We had an agent is a person, an agent is
  335. 11:01a micros service. An agent is a chatbot.
  336. 11:03An agent is a workflow. And thank you,
  337. 11:05Simon. We're back to the beginning. An
  338. 11:07agent is just tools in a loop. Um, this
  339. 11:10is happening to spec driven dev. I used
  340. 11:12to have Sean's uh slide in the beginning
  341. 11:14of this talk, but it caused a bunch of
  342. 11:15people to focus on the wrong things. His
  343. 11:17thing of like, forget the code. It's
  344. 11:18like assembly now and you just focus on
  345. 11:20the markdown. Very cool idea, but people
  346. 11:23say Spectrum Dev is writing a better
  347. 11:24prompt, a product requirements document.
  348. 11:26Sometimes it's using like verifiable
  349. 11:28feedback loops and back pressure. Maybe
  350. 11:30it is treating the code like assembly
  351. 11:32like Sean taught us. Um, but a lot of
  352. 11:34people is just using a bunch of markdown
  353. 11:35files while you're coding. Or my
  354. 11:37favorite, I just stumbled upon this last
  355. 11:39week. Uh, a spec is documentation for an
  356. 11:42open source library. So it's gone. It's
  357. 11:44as specri dev is overhyped. It's useless
  358. 11:47now. It's semantically diffused.
  359. 11:50Um, so I want to talk about like four
  360. 11:52things that actually work today. The
  361. 11:54tactical and practical steps that we
  362. 11:55found working internally and with a
  363. 11:57bunch of users. Um, we do the research,
  364. 11:59we figure out how the system works. Um,
  365. 12:02remember Momento? This is the best the
  366. 12:04best movie on context engineering, as
  367. 12:06Peter says it. Guy wakes up, he has no
  368. 12:08memory. He has to like read his own
  369. 12:10tattoos to figure out who he is and what
  370. 12:12he's up to. If you don't onboard your
  371. 12:14agents, they will make stuff up. And so,
  372. 12:17if this is your team, this is very
  373. 12:18simplified for most of you. Most of you
  374. 12:20have much bigger orgs than this. But
  375. 12:21let's say you want to do some work over
  376. 12:22here. Um, one thing you could do is you
  377. 12:25could put onboarding into every repo.
  378. 12:27You put a bunch of context. Here's the
  379. 12:28repo. Here's how it works. This is a
  380. 12:30compression of all the context in the
  381. 12:32codebase that the agent can see ahead of
  382. 12:34time before actually getting to work.
  383. 12:36This is challenging because
  384. 12:38sometimes it gets too long. As your
  385. 12:40codebase gets really big, you either
  386. 12:41have to make this longer or you have to
  387. 12:43leave information out. And so as you uh
  388. 12:47are reading through this, you're going
  389. 12:49to read the context of this big 5
  390. 12:50million line monor repo and you're going
  391. 12:52to use all the smart zone just to learn
  392. 12:54how it works. And you're not going to be
  393. 12:55able to do any good tool calling in the
  394. 12:56dumb zone. So that's uh you can
  395. 13:01you can shard this down the stack. You
  396. 13:02can do they're just talking about
  397. 13:03progressive disclosure. You could split
  398. 13:04this up, right? You could just put a
  399. 13:06file in the root of every repo and then
  400. 13:08like at every level you have like
  401. 13:10additional context based on if you're
  402. 13:12working here, this is what you need to
  403. 13:14know. Uh we don't document the files
  404. 13:16themselves cuz they're the source of
  405. 13:17truth. But then as your agent is
  406. 13:19working, you know, you pull in the root
  407. 13:20context and then you pull in the
  408. 13:21subcontext. We won't talk about any
  409. 13:22specific like you could use cloudd for
  410. 13:24this, you can use hooks for this,
  411. 13:25whatever it is. Um, but then you still
  412. 13:27have plenty of room in the smart zone
  413. 13:28because you're only pulling in what you
  414. 13:29need to know. Um, the problem with this
  415. 13:32is that it gets out of date. And so
  416. 13:34every time you ship a new feature, you
  417. 13:36need to kind of like cache and validate
  418. 13:38and rebuild large parts of this internal
  419. 13:40documentation. And you could use a lot
  420. 13:43of AI and make it part of your process
  421. 13:44to update this. Um, but I want to ask a
  422. 13:48question between the actual code, the
  423. 13:49function names, the comments, and the
  424. 13:50documentation. Does anyone want to guess
  425. 13:52what is on the y-axis of this chart?
  426. 13:57slop
  427. 13:57>> slop. It's actually the amount of lies
  428. 14:00you can find in any one part of your
  429. 14:02codebase.
  430. 14:04Um, so you could make a part of your
  431. 14:05process to update this, but you probably
  432. 14:07shouldn't cuz you probably won't. What
  433. 14:08we prefer is on demand compressed
  434. 14:10context. So if I'm building a feature
  435. 14:12that relates to SCM providers and Jira
  436. 14:14and Linear, um, I would just give it a
  437. 14:16little bit of steering. I would say,
  438. 14:17hey, we're going over in like this like
  439. 14:19part of the codebase over here. Um, and
  440. 14:22a good research uh prompt or or slash
  441. 14:24command might take you or skill even uh
  442. 14:27launch a bunch of sub aents to take
  443. 14:29these vertical slices through the
  444. 14:30codebase and then build up a research
  445. 14:32document that is just a snapshot of the
  446. 14:34actually true based on the code itself
  447. 14:37parts of the codebase that matter. We
  448. 14:39are compressing truth. Um, planning is
  449. 14:42leverage. Planning is about compression
  450. 14:44of intent. Um, and in plan we're going
  451. 14:46to outline the exact steps. We take our
  452. 14:49research and our PRD or our bug ticket
  453. 14:50or our whatever it is and we create a
  454. 14:52plan and we create a plan file. So we're
  455. 14:54compacting again. And I want to pause to
  456. 14:56talk about mental alignment. Um does
  457. 14:58anyone know what code review is for?
  458. 15:03>> Mental alignment. Mental alignment is it
  459. 15:06is about finding making sure things are
  460. 15:07correct and stuff but the most important
  461. 15:09thing is how do we keep everybody on the
  462. 15:10team on the same page about how the
  463. 15:12codebase is changing and why. And I can
  464. 15:14read a thousand lines of Golang every
  465. 15:16week. Uh sorry I can't read a thousand.
  466. 15:18It's hard. I can do it. I don't want to.
  467. 15:20Um, and as our team grows, I all the
  468. 15:22code gets reviewed. We don't not read
  469. 15:24the code, but I, as you know, a
  470. 15:25technical leader in the in on the team,
  471. 15:27I can read the plans and I can keep up
  472. 15:29to date and I can that's enough. I can
  473. 15:31catch some problems early and I maintain
  474. 15:33understanding of how the system is
  475. 15:34evolving. Um, Mitchell had this really
  476. 15:36good post about how he's been putting
  477. 15:37his AMP threads on his pull requests so
  478. 15:39that you can see not just, hey, here's a
  479. 15:41wall of green text in GitHub, but here's
  480. 15:43the exact steps, here's the prompts, and
  481. 15:44hey, I ran the build at the end and it
  482. 15:46passed. This takes the reviewer on a
  483. 15:48journey in a way that a GitHub PR just
  484. 15:50can't. And as you're shipping more and
  485. 15:52more in two to three times as much code,
  486. 15:54it's really on you to find ways to keep
  487. 15:57your team on the same page and show them
  488. 15:59here's the steps I did and here's how we
  489. 16:00tested it manually. Um, your goal is
  490. 16:03leverage. So you want high confidence
  491. 16:04that the model will actually do the
  492. 16:05right thing. I can't read this plan and
  493. 16:07know what actually is going to happen
  494. 16:09and what code changes are going to
  495. 16:10happen. So we've over time iterated
  496. 16:12towards our plans include actual code
  497. 16:15snippets of what's going to change. So
  498. 16:17your goal is leverage. You want
  499. 16:18compression of intent and you want
  500. 16:20reliable execution. Um and so I don't
  501. 16:22know I have a physics background. We
  502. 16:23like to draw lines through the center of
  503. 16:26peaks and curves. Uh as your plans get
  504. 16:28longer, reliability goes up, readability
  505. 16:30goes down. There's a sweet spot for you
  506. 16:32and your team and your codebase. you
  507. 16:34should try to find it because when we
  508. 16:35review the research and the plans, if
  509. 16:37they're good, then we can get mental
  510. 16:39alignment. Um, don't outsource the
  511. 16:42thinking. I've said this before, this is
  512. 16:44not magic. There is no perfect prompt.
  513. 16:46You still will not work if you do not
  514. 16:49read the plan. So, we built our entire
  515. 16:51process around you, the builder, are in
  516. 16:53back and forth with the agent reading
  517. 16:55the plans as they're created. And then
  518. 16:57if you need peer review, you can send it
  519. 16:58to someone and say, "Hey, does this plan
  520. 16:59look right? Is this the right approach?
  521. 17:00Is this the right order to look at these
  522. 17:02things?" Um Jake again wrote a really
  523. 17:04good blog post about like the thing that
  524. 17:05makes research plan implementing
  525. 17:06valuable is you the human in the loop
  526. 17:09making sure it's correct. So if you take
  527. 17:12one thing away from this talk it should
  528. 17:14be that a bad line of code is a bad line
  529. 17:16of code and a bad part of a plan is
  530. 17:20could be a hundred bad lines of code and
  531. 17:22a bad line of research like a
  532. 17:24misunderstanding of how the system works
  533. 17:26and where things are your whole thing is
  534. 17:28going to be hosed. You're going to be
  535. 17:29telling sending the model off in the
  536. 17:30wrong direction. And so when we're
  537. 17:32working internally and with users, we're
  538. 17:34constantly trying to move human effort
  539. 17:36and focus to the highest leverage parts
  540. 17:38of this pipeline. Um, don't outsource
  541. 17:40the thinking. Watch out for tools that
  542. 17:42just spew out a bunch of markdown files
  543. 17:44just to make you feel good. I'm not
  544. 17:46going to name names here. Uh, sometimes
  545. 17:48this is overkill. And the way I like to
  546. 17:49think about this is like, yeah, you
  547. 17:51don't always need a full research plan
  548. 17:53implement. Sometimes you need more,
  549. 17:55sometimes you need less. If you're
  550. 17:56changing the color of a button, just
  551. 17:58talk to the agent and tell it what to
  552. 17:59do. Um, if you're doing like a simple
  553. 18:02plan and as a small feature, if you're
  554. 18:04doing medium features across multiple
  555. 18:06repos, then do one research, then build
  556. 18:08a plan. Basically, the hardest problem
  557. 18:10you can solve, the ceiling goes up the
  558. 18:12more of this context engineering
  559. 18:14compaction you're willing to do. Um, and
  560. 18:16so if you're in the top right corner,
  561. 18:18you're probably going to have to do
  562. 18:19more. A lot of people ask me, "How do I
  563. 18:21know how much context engineering to
  564. 18:22use?" It takes reps. You will get it
  565. 18:25wrong. You have to get it wrong over and
  566. 18:26over and over again. Sometimes you'll go
  567. 18:27too big. Sometimes you go too small.
  568. 18:29Pick one tool and get some reps. I
  569. 18:32recommend against minmaxing across cloud
  570. 18:34and codeex and all these different
  571. 18:35tools. Um, so I'm not a big acronym guy.
  572. 18:39Uh, we said specri dev was broken. Uh,
  573. 18:42research plan and implement I don't
  574. 18:44think will be the steps. The important
  575. 18:45part is compaction and context
  576. 18:46engineering and staying in the smart
  577. 18:47zone. But people are calling this RPI
  578. 18:50and there's nothing I can do about it.
  579. 18:52So, uh, just be wary. There is no
  580. 18:54perfect prompt. There is no silver
  581. 18:55bullet. Um, if you really want a hypy
  582. 18:58word, you can call this harness harness
  583. 19:00engineering, which is part of context
  584. 19:01engineering, and it's how you integrate
  585. 19:02with the integration points on codeex,
  586. 19:04claude, cursor, whatever. How you
  587. 19:06customize your codebase. Um, so what's
  588. 19:09next? I think the coding agent stuff is
  589. 19:12actually going to be commoditized.
  590. 19:13People are going to learn how to do this
  591. 19:14and get better at it. And the hard part
  592. 19:16is going to be how do you adapt your
  593. 19:17team and your workflow and the SDLC to
  594. 19:20work in a world where 99% of your code
  595. 19:22is shipped by AI. Uh, and if you can't
  596. 19:25figure this out, you're hosed because
  597. 19:26there's kind of a rift growing where
  598. 19:28like staff engineers don't adopt AI
  599. 19:29because it doesn't make them that much
  600. 19:30faster. And then junior mid-levels
  601. 19:32engineers use a lot because it fills in
  602. 19:34skill gaps and then it also produces
  603. 19:36some slop. And then the senior engineers
  604. 19:37hate it more and more every week because
  605. 19:39they're cleaning up slop that was
  606. 19:40shipped by cursor the week before. Uh,
  607. 19:43this is not AI's fault. This is not the
  608. 19:44mid-level engineers fault. Like if
  609. 19:46cultural change is really hard and it
  610. 19:48needs to come from the top if it's going
  611. 19:49to work. So if you're a technical leader
  612. 19:51at your company, pick one tool and get
  613. 19:53some reps. If you want to help, we are
  614. 19:56hiring. We're building an Aentic IDE to
  615. 19:58help teams of all sizes speedrun the
  616. 20:00journey to 99% AI generated code. Uh if
  617. 20:04we'd love to we'd love to talk if you
  618. 20:05want to work with us. Uh go go hit our
  619. 20:07website, send us an email, come find me
  620. 20:08in the hallway. Uh thank you all so much
  621. 20:10for your energy.
  622. 20:14[music]
  623. 20:16Heat.
  624. 20:30[music]

About this transcript

This page contains the full transcript of No Vibes Allowed: Solving Hard Problems in Complex Codebases – Dex Horthy, HumanLayer by AI Engineer, generated from the public captions YouTube serves with the video. The transcript has 4,538 words across 624 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.