YouTube2Text

Everything We Got Wrong About Research-Plan-Implement - Dexter Horthy — Transcript

by AAIF Live · 6,100 words · 909 segments · language en · Watch on YouTube

Full transcript

  1. 0:04All right, before I bring up Dex, I got
  2. 0:07to say I met James and James, yeah, can
  3. 0:11you stand up real fast? James told me
  4. 0:13this morning that he is going We can't
  5. 0:15see the QR code. You got to hit it up.
  6. 0:17He's going to be having dinner tonight
  7. 0:19and everybody's invited. So,
  8. 0:22if you want to go,
  9. 0:24just scan that QR code real fast and go
  10. 0:27hang out with James. That's awesome. And
  11. 0:30that's what we're going for here. That
  12. 0:31is great.
  13. 0:32Yeah, you rock. Yeah, that's I I like
  14. 0:34it. I like it, too. Okay, so I'm going
  15. 0:37to bring up Dex.
  16. 0:38Earlier somebody said we need to have a
  17. 0:40mustache competition. I don't think
  18. 0:42that's going to happen, but I will say
  19. 0:44that he gave us 200 slides that he's
  20. 0:47about to present.
  21. 0:48>> Woah, woah, 100 158.
  22. 0:50>> 158. We're going to keep him honest on
  23. 0:53the timing, all right? So, let it start
  24. 0:55now.
  25. 0:56>> Amazing. Let's do it. What's up,
  26. 0:58everybody?
  27. 1:02Uh
  28. 1:02I am Dex. Uh this is a talk with a very
  29. 1:05long title that I'm not going to read
  30. 1:06because we're on the clock now.
  31. 1:08Uh I have been talking about coding
  32. 1:09agents for quite a while, basically
  33. 1:11since like August. Um we did a long talk
  34. 1:13in November. Um there's this methodology
  35. 1:16that we've been talking about a lot uh
  36. 1:17called research plan implement. Um
  37. 1:20lot of upvotes on Hacker News. There's
  38. 1:21probably 10,000 people who have gone to
  39. 1:23our open source and grabbed our prompts
  40. 1:25and are using them internally from small
  41. 1:26startups up to the enterprise.
  42. 1:28Um it all started with this guy. Um this
  43. 1:30guy
  44. 1:31Has anyone seen this talk?
  45. 1:33Yes? Okay, so Igor went in and he said
  46. 1:35like, "Okay, cool. We're using a lot of
  47. 1:36tokens. We're spending a lot of money to
  48. 1:37get AI developer productivity." But what
  49. 1:39he found was that it actually tends to
  50. 1:41lead to a lot of rework. Like you are
  51. 1:42shipping 50% more, but half of that is
  52. 1:45just cleaning up the slop from last
  53. 1:47week. And the other thing they found,
  54. 1:48and these are last year's numbers, so
  55. 1:50this is not account for Opus 4.5. So, I
  56. 1:52would I would inflate this a little bit,
  57. 1:53but like it's great for low complexity
  58. 1:56greenfield tasks, not great for
  59. 1:59high complexity brownfield tasks. Um,
  60. 2:02and so I could give you a talk about RPI
  61. 2:04and why it's great. Uh, but that would
  62. 2:06be boring. And there's other talks. So
  63. 2:08if you haven't seen them, go watch them.
  64. 2:09They'll give more context, but uh
  65. 2:12I'm going to tell you everything we got
  66. 2:13wrong about RPI today. Um, we thought we
  67. 2:15had this AI thing thing figured out. I
  68. 2:18uh am am uh humble enough to admit when
  69. 2:20I was wrong.
  70. 2:21Uh, so we got a couple things wrong.
  71. 2:23One thing that's very relevant if you've
  72. 2:24been on Twitter today is uh I don't
  73. 2:26think it's okay to not read the code.
  74. 2:28Uh,
  75. 2:29I also don't think you should read
  76. 2:30really long plan files. Uh, well those
  77. 2:32two are related.
  78. 2:34Uh, and no, Claude should not be allowed
  79. 2:36to have If you If you're writing
  80. 2:37production code that is used by users
  81. 2:39and you're going to get paged at 3:00
  82. 2:40a.m. if it's broken, uh
  83. 2:42no slop. This is the year, 2026, no more
  84. 2:45slop. So we're all on this journey.
  85. 2:46We're all figuring this out. We all are
  86. 2:48wrong all the time. We did get a couple
  87. 2:50things right. Uh, there is no magic
  88. 2:51prompt. Uh, do not outsource the
  89. 2:53thinking. You the engineer are an
  90. 2:55important part of this process, and seek
  91. 2:57leverage. There's a lot of code being
  92. 2:59written. Find ways to make sure it's
  93. 3:01correct without having to read all of it
  94. 3:02and re-steer after the fact. Um
  95. 3:05So lots of people I'm sure have heard of
  96. 3:06research plan implement. Has anyone
  97. 3:08actually run this Claude command,
  98. 3:10research code base?
  99. 3:11Okay, cool. Leave your hand up if you've
  100. 3:13run it like this. Tell me how this
  101. 3:15system works.
  102. 3:16Has anyone run it like this of like,
  103. 3:18"Hey, I want to build this thing. Go do
  104. 3:18the research."
  105. 3:20Okay.
  106. 3:21Um, or maybe go fetch a ticket or
  107. 3:22something. What about the create plan uh
  108. 3:25prompt? Okay, couple hands. Uh,
  109. 3:28how many of you run it like that? You
  110. 3:29know, "Hey, we got to go build this
  111. 3:30thing." Yes?
  112. 3:32Has anyone run it like this?
  113. 3:33Work back and forth with me starting
  114. 3:35with your open questions and outline
  115. 3:36before writing the plan.
  116. 3:37Okay, some of you found out about the
  117. 3:39magic words. A lot of people didn't
  118. 3:41though. Uh, we'll get into why that's a
  119. 3:42problem. Um, so since October we've
  120. 3:45basically worked with thousands of
  121. 3:46engineers from tiny startups all the way
  122. 3:48up to Fortune 500s.
  123. 3:49Um, and we would find over and over
  124. 3:51again we would give these tools to an
  125. 3:52expert uh and they would get great
  126. 3:54results. They would go sit and talk to
  127. 3:55Claude for 70 hours a week and they
  128. 3:56would start shipping like crazy.
  129. 3:58And then they would go give it to their
  130. 3:59team
  131. 4:00and the results were not always so good.
  132. 4:02And so people weren't getting good
  133. 4:04results. And so we got in the trenches
  134. 4:05with our users and we went to go figure
  135. 4:07out what was going wrong. And the first
  136. 4:09thing that was going wrong was people
  137. 4:10were not getting good research.
  138. 4:12So we talked about this in November.
  139. 4:14This is
  140. 4:15one of the only slides I'm ever using,
  141. 4:16but you would pick a zone of your code
  142. 4:17base. You would say, "Oh, we're going to
  143. 4:18build something over here." And then you
  144. 4:20would launch it in coding agent session
  145. 4:21to go send these sub agents through
  146. 4:23these deep vertical slices through the
  147. 4:24code base for just the context
  148. 4:26compressed context about what is the
  149. 4:28thing we're about to go build.
  150. 4:30Right? And we said, "Keep things
  151. 4:32objective. Discourage opinions. Don't
  152. 4:35actually put any implementation details
  153. 4:36in there. You just want to compress the
  154. 4:38truth. What is true about how the code
  155. 4:40works today?"
  156. 4:42And a skilled engineer was really good
  157. 4:43at taking, "Okay, here's my ticket. Let
  158. 4:45me write some questions that will cause
  159. 4:47the model to go touch all the parts of
  160. 4:48the code base that matter." So if it
  161. 4:50was, you know, add a new endpoint to
  162. 4:51reticulate splines across tenants, we
  163. 4:54would say something like, "Okay, tell me
  164. 4:55how endpoint endpoints work and trace
  165. 4:57the logic flow for everything that
  166. 4:58touches splines and go find the workers
  167. 5:00that do all the reticulation."
  168. 5:02So if this is your ticket, a lot of
  169. 5:04people would run it like this. They
  170. 5:06would just say, "Hey, research code
  171. 5:07base. Here's what I'm building."
  172. 5:08And the problem is that good research is
  173. 5:10all facts, but if you tell the model
  174. 5:11what you're building, then you get
  175. 5:12opinions. And we don't We'll get into
  176. 5:14why the model shouldn't have opinions
  177. 5:15later. Comes back to this thing that
  178. 5:17Jake from Netflix came up with, which is
  179. 5:19do not outsource the thinking.
  180. 5:21The other thing that wasn't working is
  181. 5:23people were getting not great plans.
  182. 5:25And basically there were these steps
  183. 5:27built into this planning prompt
  184. 5:29that was this single giant like
  185. 5:31monolithic thing with 85 or more
  186. 5:33instructions. And it had these steps in
  187. 5:36it of like, "Cool, present design
  188. 5:37options to the user, get feedback on the
  189. 5:39structure before you actually go write
  190. 5:41the plan." And so a good planning
  191. 5:43session would look something like, you
  192. 5:45know, you have your Claude system tools
  193. 5:47in your prompt and then you say, "Hey,
  194. 5:49create plan."
  195. 5:50Loads the skill, looks at your ticket,
  196. 5:52loads your research doc, launch a bunch
  197. 5:54of sub agents to go find a bunch of
  198. 5:55things that are true about the code
  199. 5:56base, just confirm some stuff that
  200. 5:58wasn't maybe in the research. This is
  201. 5:59all one big context window, by the way.
  202. 6:00Usually, I'll use these columns to mean
  203. 6:02separate context windows, but today this
  204. 6:04is all one session. I just slides are
  205. 6:06sideways, so I had to put them on next
  206. 6:08to each other.
  207. 6:09Um but the agent would come and ask
  208. 6:11questions. Say, "Okay, here's our
  209. 6:12options for question one." User would
  210. 6:13pick an option. User would pick an
  211. 6:15option. And then eventually, it would
  212. 6:16say, "Cool, here's the order we're going
  213. 6:17to do the things. Um what do you think?"
  214. 6:19And the user could say, "Well, we need
  215. 6:20to add a testing step up front and I
  216. 6:21want to swap phases three and four."
  217. 6:23Assistant would give the new outline of
  218. 6:25the phases.
  219. 6:26Then the user would approve it. And only
  220. 6:28then
  221. 6:29would we write our plan file.
  222. 6:30Um
  223. 6:32complex process of aligning with the
  224. 6:33user on what was what was going to be
  225. 6:35built.
  226. 6:36Um but for about 50% of people, maybe
  227. 6:38more, if you didn't prompt it with this
  228. 6:40work back and forth with me or Opus was
  229. 6:42just feeling dumb for that that
  230. 6:44particular hour of the day, um
  231. 6:47it would just take the stuff and it
  232. 6:48would just immediately go and write the
  233. 6:50plan out. And so, you would get this and
  234. 6:52they'd be like, "Cool, I wrote the
  235. 6:53plan." Didn't ask me any questions, made
  236. 6:55all the decisions for me. Yikes.
  237. 6:57So, we give the tools to people and some
  238. 7:00people got good results and some people
  239. 7:01didn't. And we dug in and we were like,
  240. 7:02"What's the difference?"
  241. 7:04Uh
  242. 7:05and people would literally say this to
  243. 7:06me. They'd be like, "Well, you have to
  244. 7:07say the magic words." And I found myself
  245. 7:09in workshops full of enterprise
  246. 7:10engineers saying, "Well, guys, guys,
  247. 7:12guys, guys, yeah, here's the software,
  248. 7:13but don't forget to say the magic
  249. 7:14words." It was um quite frankly, it was
  250. 7:16embarrassing.
  251. 7:17But if you said this, work back and
  252. 7:18forth with me starting with your open
  253. 7:19questions and outline before writing the
  254. 7:21plan, then the agent would actually ask
  255. 7:22you the questions. And this isn't the
  256. 7:24user's fault. If you built a tool that
  257. 7:26requires hours and hours of training and
  258. 7:28reps to get like good results from, go
  259. 7:31fix the tool. And so, I'll talk about
  260. 7:33how we did that. Um but why these steps
  261. 7:35were getting skipped, the one of the big
  262. 7:36takeaways I'll give you today is like
  263. 7:38you have an instruction budget. Uh
  264. 7:41my co-founder Kyle is somewhere over
  265. 7:42here. He wrote this really good blog
  266. 7:44post in December or November, I guess
  267. 7:46technically, um that basically cited
  268. 7:47this archive paper, which again, this is
  269. 7:49from last year, so the number is
  270. 7:51probably a little bit higher now, but
  271. 7:52that frontier LLMs could only follow
  272. 7:54about 150 to 200 instructions with like
  273. 7:57good consistency. Anything more than
  274. 7:58that and it's kind of half attending to
  275. 8:00all of them and you're rolling the dice.
  276. 8:02So, if you have a prompt with 85
  277. 8:04instructions and your Claude MD and your
  278. 8:06system prompt and your tools and your
  279. 8:08MCP, um yeah, you're not likely to get
  280. 8:12full adherence to the workflow. So, more
  281. 8:14on how we fix this later.
  282. 8:16The other thing that I think, um really
  283. 8:18wasn't working for people was like we
  284. 8:20advocated for reading the plans that
  285. 8:21were output. This is me on stage in
  286. 8:24November telling people, you have to
  287. 8:25read the plan, otherwise it won't work.
  288. 8:28Um some people even would PR their plans
  289. 8:30and code review them together. But a
  290. 8:32thousand line plan tends to be about a
  291. 8:34thousand lines of code within 10% or so,
  292. 8:37and plans can have surprises. So, you
  293. 8:38would go and you would review the plan
  294. 8:40and then you would go right to code and
  295. 8:42it would be different. And so, you're
  296. 8:43telling you're asking one of your
  297. 8:44co-workers like, okay, you go spend an
  298. 8:46hour reading this and tell me what's
  299. 8:47wrong with it, and then you would go
  300. 8:48implement it and it would be different.
  301. 8:49They'd have to go read the code again
  302. 8:50and see what the surprises were and what
  303. 8:52changed. Um and so, this isn't leverage.
  304. 8:55Leverage is about like do less work to
  305. 8:57get more output. So, the new advice, uh
  306. 9:01don't read the plans.
  307. 9:02Please, read the code.
  308. 9:04Uh just cuz it's it's the same amount of
  309. 9:06work and like look for leverage
  310. 9:07elsewhere and I'll talk about how we
  311. 9:08found better leverage. Um and you may
  312. 9:10say, "Hey Dex, in August you said don't
  313. 9:12read the code. You said that the plans
  314. 9:14are enough. Just don't just just go just
  315. 9:15ship and let Claude do its thing."
  316. 9:17I was wrong. I am humble enough to admit
  317. 9:20when I was wrong. Uh this is actually a
  318. 9:21very big conversation right now. Please,
  319. 9:24please read the code. We tried not
  320. 9:25reading the code for like 6 months.
  321. 9:27Uh it did not end well. We had to rip
  322. 9:28out and replace large parts of that
  323. 9:30system. Um
  324. 9:31and you may say, "Hey Dex, but other
  325. 9:33people don't read the code." Beats,
  326. 9:34300,000 lines and counting. Uh no one's
  327. 9:37read that code, allegedly.
  328. 9:39Uh open claw, Pete's like, "Okay, you
  329. 9:40know, I know the structure and the
  330. 9:42pieces and how they fit together, but I
  331. 9:43don't read every line of every PR."
  332. 9:46Um these are OSS projects. They don't
  333. 9:48charge money.
  334. 9:49Nobody gets paged at 3:00 a.m. if it's
  335. 9:51broken, and no one gets fined millions
  336. 9:53of dollars if it's done wrong. I will
  337. 9:55also say though,
  338. 9:56these are OSS. They are very, very cool
  339. 9:58projects. I am humbled, deeply humbled
  340. 10:01by the accomplishments of the
  341. 10:02maintainers, and the stakes are still
  342. 10:04high. Like, if you break open claw, a
  343. 10:06lot of people are going to be upset. But
  344. 10:08they are different than if you were,
  345. 10:09say, working in a regulated industry
  346. 10:11shipping production SAS code.
  347. 10:13Um so, if you have people who depend on
  348. 10:14your code,
  349. 10:16please, I'm begging you, please read it.
  350. 10:19Please read it. We have a profession to
  351. 10:20uphold. 2026 is supposed to be the year
  352. 10:23of no more slop. Uh literally everyone
  353. 10:26is talking about the difference between
  354. 10:27slop and craft.
  355. 10:29Uh this is why I'm a little mid on agent
  356. 10:31swarms and the whole gas town thing
  357. 10:33because you still need to be able to
  358. 10:35ensure quality, and like going 10 times
  359. 10:37faster doesn't matter if you're going to
  360. 10:38throw it all away in 6 months. So, shoot
  361. 10:41for 2 to 3x. That's actually another
  362. 10:42talk of like how you measure this and
  363. 10:44how you actually get there and maintain
  364. 10:45like a near human level of quality. Um
  365. 10:48but I'll talk about the goals and like
  366. 10:49what you should think about if you want
  367. 10:50to get there is you should have high
  368. 10:51leverage planning.
  369. 10:53You should not outsource the thinking.
  370. 10:55Read and own the code.
  371. 10:56And ideally we will avoid uh
  372. 10:59magic words.
  373. 11:00So, uh
  374. 11:01we got better research, we got better
  375. 11:02plans, we got better leverage. I'm going
  376. 11:04to talk about each of those um as far as
  377. 11:06like, in general what we in like it's
  378. 11:08specifically what we did, and also some
  379. 11:10general concepts as you're building
  380. 11:11workflows and systems around coding
  381. 11:13agents, what you can do.
  382. 11:14So, we talked about a skilled This is
  383. 11:15the least exciting one, but talked about
  384. 11:17how a skilled engineer could detangle
  385. 11:18the ticket to the questions to the
  386. 11:20research,
  387. 11:21uh and then the research would be very
  388. 11:22objective.
  389. 11:23Um
  390. 11:24basically we just hide the ticket from
  391. 11:26the context window that's doing
  392. 11:27research, and we do it
  393. 11:28deterministically. So, basically you
  394. 11:30have one context window to generate
  395. 11:31questions, and then a fresh context
  396. 11:33window with no knowledge of what we're
  397. 11:34building to go make your research doc.
  398. 11:37Um this is pretty trivial. If you're
  399. 11:38familiar with the concept of query
  400. 11:39planning,
  401. 11:40um
  402. 11:41it's
  403. 11:41similar in concept but for, you know,
  404. 11:43LLMs reading through codebases.
  405. 11:46Um so, I've been hacking on agents for a
  406. 11:48while and before we did the coding agent
  407. 11:49stuff, I wrote this paper called 12
  408. 11:50factor agents, which was uh allegedly
  409. 11:53the first time anyone was like talking a
  410. 11:54lot about context engineering. Uh
  411. 11:57there's two ways to read context
  412. 11:59engineering and most people jumped in.
  413. 12:01Is anyone building like rag pipelines?
  414. 12:02Is it Raise your hand if you built a rag
  415. 12:04pipeline.
  416. 12:05Okay, some people are feeling uh not
  417. 12:07like not raising their hands today. Um
  418. 12:10But it's like, okay, put more
  419. 12:11information in, the model can't make
  420. 12:13sense of it. I actually think the more
  421. 12:15interesting read of context engineering
  422. 12:17is like better instructions and simpler
  423. 12:19tasks and smaller context windows. Of
  424. 12:21course, we all know Jeff now. I don't
  425. 12:22have to introduce him anymore. He used
  426. 12:23to have to I used to have to tell people
  427. 12:25who Jeff was when I was talking. Um we
  428. 12:28talked about this like context window
  429. 12:29thing as the idea of the dumb zone,
  430. 12:31which is, you know, you have about
  431. 12:33168,000 tokens and 200,000 but some of
  432. 12:36them are reserved for output. You have
  433. 12:38various things that they're for and
  434. 12:39around like 40% on average depending on
  435. 12:41what you're doing and how much of your
  436. 12:42context is user messages versus files
  437. 12:44and all of this stuff, you hit this
  438. 12:46point where you have degrading results.
  439. 12:48And obviously sometimes you can get
  440. 12:49still get good enough for you results at
  441. 12:5160% but the less of the context window
  442. 12:54you use, the better results you will
  443. 12:55get. Um our friends at Databricks were
  444. 12:57just talking about you have too many
  445. 12:58MCPs. The whole context window is full
  446. 13:00of instructions about how to use a bunch
  447. 13:01of tools that you don't care about and
  448. 13:03then by the time you're writing code,
  449. 13:04the model's like not good at following
  450. 13:05your instructions.
  451. 13:06So, you're not just giving the model too
  452. 13:07much information, you're also probably
  453. 13:10giving it too many instructions.
  454. 13:12And so, the idea of what we're doing was
  455. 13:14this thing like makes a lot of sense,
  456. 13:16use prompts for control flow. This is a
  457. 13:18customer support example but, you know,
  458. 13:19if it's a complaint, go do this. If it's
  459. 13:21product feedback, go do this. If it's a
  460. 13:23billing issue, go do this.
  461. 13:25Um
  462. 13:25And what you could do instead is you can
  463. 13:27instead of using prompts for control
  464. 13:28flow,
  465. 13:29you can kind of classify the input and
  466. 13:31then feed it to a series of smaller,
  467. 13:33more focused prompts where there are far
  468. 13:35fewer instructions and far fewer actions
  469. 13:37to choose from. I'm sure many of us have
  470. 13:38already done things like this to improve
  471. 13:40the performance of pipelines. Um so this
  472. 13:42was a single mega prompt with 85
  473. 13:44instructions.
  474. 13:45Um and if you did it right, you would go
  475. 13:46through all these different steps. All
  476. 13:48these different phases were part of
  477. 13:49that. And if any of the instructions
  478. 13:50didn't get followed, you would skip the
  479. 13:52things that made this really high
  480. 13:53leverage.
  481. 13:54Um so we split it across several
  482. 13:55prompts.
  483. 13:56And so like before it was research,
  484. 13:57plan, implement, now it's questions,
  485. 13:59research, design, structure, plan, work
  486. 14:00tree, implement, PR. We're not actually
  487. 14:01not going to have time to talk about the
  488. 14:03implement side of the thing today. But
  489. 14:05um
  490. 14:06we split up the planning into a design
  491. 14:07discussion, an outline, and a plan.
  492. 14:10And before it was 85 instructions, now
  493. 14:13they're all less than 40, which is
  494. 14:14really exciting. And I think some of
  495. 14:15them could actually be even smaller.
  496. 14:17We're still iterating on them. The
  497. 14:18lesson is don't use prompts for control
  498. 14:20flow if you can use control flow for
  499. 14:21control flow. Like the if statement is
  500. 14:23really, really powerful and LLMs are
  501. 14:25really good at classifying things. This
  502. 14:26is not just true for coding agents. This
  503. 14:28is any AI LLM-based system you're
  504. 14:29building.
  505. 14:30Um and it's really funny cuz we were
  506. 14:32writing all this stuff and we got on
  507. 14:33stage and we said like full fat agents
  508. 14:35don't work. Don't just call tools in a
  509. 14:37loop, do context engineering and build
  510. 14:38workflows and graphs and micro agents.
  511. 14:40We told everybody don't do this. And
  512. 14:42then we turned around in August and
  513. 14:43we're like, "Oh,
  514. 14:44all right, but this Claude code thing is
  515. 14:46pretty good." And we turned around and
  516. 14:47we wrote this giant monolithic prompt.
  517. 14:49So we figured it was time to actually go
  518. 14:50drink our own Kool-Aid.
  519. 14:52Um
  520. 14:53mind your instruction budget.
  521. 14:56How do we get better leverage?
  522. 14:57So we split things up to get better
  523. 14:58instruction following, right?
  524. 15:00These three different phases. But we
  525. 15:02also got more leverage. I'm going to
  526. 15:03talk about why. Because even if the plan
  527. 15:05is a thousand lines and the code is a
  528. 15:06thousand lines, your design discussion
  529. 15:08might only be 200 lines. And you get a
  530. 15:10lot of opportunities to restear in that
  531. 15:11moment. And so what this looks like is
  532. 15:13basically where are we going? What does
  533. 15:15the final solution look like? And it
  534. 15:17has, you know, the current state, the
  535. 15:18desired end state. It has the patterns
  536. 15:20to follow. How many of you have ever
  537. 15:21like sent a coding agent and it like
  538. 15:23found the wrong way to do a thing in
  539. 15:25your code base and it followed the bad
  540. 15:26patterns? Yes?
  541. 15:28Right. This is your chance to go read
  542. 15:30all the patterns it found that it thinks
  543. 15:31are relevant and be like, "Nope, that's
  544. 15:33not how we do atomic SQL updates. That's
  545. 15:35some engineer that doesn't work here
  546. 15:36anymore and it's crazy and everyone
  547. 15:37hates it. Go find the way we do it over
  548. 15:38there."
  549. 15:39Um it'll keep track of resolved design
  550. 15:41decisions that we've made. It will ask
  551. 15:43open questions. This is sort of like
  552. 15:45taking Claude code plan mode and the ask
  553. 15:47user question tool and just brain
  554. 15:49dumping it all to the single document
  555. 15:50that you can interact with is like
  556. 15:52moldable and flexible.
  557. 15:54Um Matt Pocock has this idea, he calls
  558. 15:55it the design concept and it's this idea
  559. 15:57of like the thing that is locked up in
  560. 15:59this context window that is the shared
  561. 16:02understanding between you and the agent
  562. 16:04of what's being built and how.
  563. 16:06Uh so we put it into an underlying
  564. 16:07markdown artifact.
  565. 16:09Um and so we now have human agent
  566. 16:11alignment. And the idea here is like
  567. 16:12you're forcing the agent to brain dump
  568. 16:14out all the things it found, all the
  569. 16:15things it wants to do, all the things it
  570. 16:17thinks you want, and ask you questions
  571. 16:19about things it doesn't know. So you can
  572. 16:20do brain surgery on the agent before you
  573. 16:22proceed downstream. And it's all about
  574. 16:24do not outsource the thinking. You want
  575. 16:26to give the agent every single
  576. 16:27opportunity to show you what it's wrong
  577. 16:29about before you go write 2,000 lines of
  578. 16:31code.
  579. 16:33So uh 200 lines instead of a thousand, a
  580. 16:35little bit more leverage. We also get
  581. 16:37better leverage from the outline. So if
  582. 16:39design is like, "Where are we going?"
  583. 16:41the structure outline is, "How do we get
  584. 16:43there?" Or if you're an engineer who is
  585. 16:45miserable cuz of sitting in meetings all
  586. 16:46day, there's the like architecture
  587. 16:48review and then there's the sprint
  588. 16:50planning meeting. What are we going to
  589. 16:51build and then how do we break it down
  590. 16:53into tasks?
  591. 16:54And so we take our design and we take it
  592. 16:56to ticket in the research and we build
  593. 16:57up a new context window and we create
  594. 16:59the structure outline.
  595. 17:01And this is basically a high-level
  596. 17:02overview of the phases, not the exact
  597. 17:04code we're going to write, but just kind
  598. 17:06of what it's going to look like, what
  599. 17:07order we're going to do the changes in,
  600. 17:09and how we're going to test it along the
  601. 17:10way. Now, I don't actually test in
  602. 17:12between every phase everything I'm
  603. 17:13building, but if it's sensitive or if
  604. 17:15it's hard or if it's complex, I want to
  605. 17:17be able to catch it before it goes and
  606. 17:19writes all the code. I want to make sure
  607. 17:21each two, three, 400 line block is
  608. 17:23correct. Um and these docs mean lighter
  609. 17:25reviews. Instead of reviewing the plan,
  610. 17:27this is two things for the same feature.
  611. 17:28Plan eight pages, structure outline
  612. 17:30two-ish pages, much shorter. Um
  613. 17:33I like to think of this Has anyone ever
  614. 17:34written a C header file, a .h file?
  615. 17:37Yeah, okay. So, if the plan is the
  616. 17:39implementation, the outline is the C
  617. 17:41header files. Just here's the signatures
  618. 17:42and the new types that we're changing.
  619. 17:44Enough again for you to see what the
  620. 17:46agent is thinking and correct it if it's
  621. 17:48wrong.
  622. 17:49Um and the reason why we do this is
  623. 17:51despite like every single model and
  624. 17:53trying to prompt this out and eval the
  625. 17:54hell out of this, we cannot get models
  626. 17:56to stop writing horizontal plans. Or it
  627. 17:58like this is the best way to fix their
  628. 18:01need to write horizontal plans. And when
  629. 18:02I say horizontal plans, I basically mean
  630. 18:05you start with Models love to like we're
  631. 18:07going to do all the database and then
  632. 18:08we're going to do all the services and
  633. 18:09then we're going to do all the API and
  634. 18:11then we're going to do all the front end
  635. 18:11and before you know it, you're on the
  636. 18:13other side of 1,200 lines of code and
  637. 18:15it's not working.
  638. 18:17And now you have to go figure out which
  639. 18:18part is broken because there was no
  640. 18:19nothing really to test along the way,
  641. 18:21whether the model is verifying it or
  642. 18:23whether you the human are jumping in and
  643. 18:24checking it's correct. And so what we've
  644. 18:27seen work really, really well across
  645. 18:28orgs of all sizes
  646. 18:30um is what I call vertical plans. This
  647. 18:32is how I build when I'm like before AI,
  648. 18:34I would like make a mock API endpoint
  649. 18:36and then get it working in the front end
  650. 18:37and then wire that and then mock out the
  651. 18:39services layer and then do the database
  652. 18:41migration and then put everything
  653. 18:43together.
  654. 18:44And so, even though it's the same amount
  655. 18:46of code, you have these like checkpoints
  656. 18:48where you can see if it's working and if
  657. 18:49it's not, you can pause and fix it
  658. 18:51before you go try to do the rest of it.
  659. 18:54So, these are just markdown docs too.
  660. 18:55Like you can and should ask for more
  661. 18:56detail. They start high level, but like
  662. 18:58here's an example of like I don't think
  663. 18:59you're going to get this right. Tell me
  664. 19:01what you're thinking and then like
  665. 19:02dumped out the types and the signatures.
  666. 19:04Um and then getting better leverage from
  667. 19:06the plan itself, I mean, again, like
  668. 19:08usual, like we've been doing, we just
  669. 19:09take that artifact, we build it up with
  670. 19:11all the previous artifacts, and then we
  671. 19:13can go build the plan.
  672. 19:14Um and this is the same if you use
  673. 19:16create plan, it's the exact same
  674. 19:17template, exact same setup, exact same
  675. 19:18prompt. But this is a tactical doc for
  676. 19:20the agent. We've already done enough
  677. 19:22aligning that like I'm just going to
  678. 19:24spot-check this, and then we save the
  679. 19:25deep review for the actual code. And so
  680. 19:28if you've used any of the RPI plans,
  681. 19:30they look like this. It's the model
  682. 19:31saying, "Hey, here's all the changes I'm
  683. 19:32going to make."
  684. 19:33Um
  685. 19:34the most important part of this leverage
  686. 19:36is not just about you and the agent
  687. 19:37though. Like human agent alignment is
  688. 19:39important, and knowing what the agent's
  689. 19:40going to do and correcting that is is
  690. 19:42good, but it's also, you know, if you're
  691. 19:44working with a team of engineers, we've
  692. 19:46found a lot of value from taking these
  693. 19:48design discussions, these structure
  694. 19:50outlines, and review. I said don't
  695. 19:51review the plans, but these shorter docs
  696. 19:54are really, really good. Uh
  697. 19:56I I am not the code owner of most of our
  698. 19:59code at Human Layer. My uh co-founder
  699. 20:00is, and I send him my design discussions
  700. 20:03on purpose. We don't have a required
  701. 20:05step, but I want to I want to know that
  702. 20:07when we get to code review, it's just
  703. 20:09going to be like, "Yep, that's That's
  704. 20:10what I wanted. That's it. That's it." So
  705. 20:11any any of my bad decisions are headed
  706. 20:14off on a 200-line doc before I've gone
  707. 20:16and written the code and gotten it
  708. 20:17working and I'm attached to it. And so
  709. 20:19this is really, really powerful. Um
  710. 20:22before AI, we would basically the way
  711. 20:23Another way to think about it is like
  712. 20:24time savings. You would say, "Okay, it's
  713. 20:26a 2-day feature. I got to do all this
  714. 20:28stuff. The coding's probably 2 to 4
  715. 20:29hours."
  716. 20:30If you just pick up Claude code and use
  717. 20:32it to ship for you, you do get some
  718. 20:33speed up because now the coding takes 20
  719. 20:35minutes. It's still a 2-day feature cuz
  720. 20:37I still have to like align with my team
  721. 20:39on what we're going to do. I still have
  722. 20:40to get a code review and fix stuff.
  723. 20:42Maybe I'm working across repos that I
  724. 20:43don't personally own, and then we still
  725. 20:45have to verify and test it.
  726. 20:47But if you use AI to help you with your
  727. 20:49planning and alignment, then you also
  728. 20:51save time there, and I think you get
  729. 20:54much better alignment. Um and so your
  730. 20:56code review and rework is also much
  731. 20:58shorter because you already know what's
  732. 20:59coming. The team that's reviewing it
  733. 21:00already kind of like had their chance to
  734. 21:02restear you. And really good teams do
  735. 21:03this. It's They have a meeting that's
  736. 21:05called architecture review where we
  737. 21:06decide, you know, what's our technical
  738. 21:07design doc on how we're going to build
  739. 21:09this.
  740. 21:10So, um as far as testing and verifying,
  741. 21:12sorry, I don't have a good answer for
  742. 21:13you. It's a whole other talk. If you
  743. 21:15went to Drew's talk downstairs, go find
  744. 21:17Drew Brignac after this. He will tell
  745. 21:18you all about testing and verifying.
  746. 21:20Um let's put this all together.
  747. 21:22So, we have these five stages of
  748. 21:24research and planning.
  749. 21:25Um the process is basically questions,
  750. 21:27research, design, structure outline,
  751. 21:30plan, work tree, implement, finally the
  752. 21:32pull request.
  753. 21:34Uh that didn't make a very good acronym
  754. 21:35though, so we just picked the ones we
  755. 21:36liked and uh we're calling this crispy.
  756. 21:39Uh
  757. 21:40So, RPI to crispy, that's the There you
  758. 21:42go. Um what's next and what did I not
  759. 21:45have time to talk about today?
  760. 21:46Um three steps is already a lot for some
  761. 21:48people to learn and now there are seven.
  762. 21:50I thought we were supposed to make this
  763. 21:51easier for teams to learn this and adopt
  764. 21:52it. We can talk about how we're like
  765. 21:54thinking about that. Um the idea of how
  766. 21:56do you measure the impact of doing this
  767. 21:58um in engineering teams? I think it's
  768. 22:00like we've been trying to measure
  769. 22:01developer productivity for 50 years and
  770. 22:03we still don't know how to do it very
  771. 22:04well.
  772. 22:05Um and then it's like if you're a
  773. 22:07central kind of platform team rolling
  774. 22:09out changes to everybody in your org,
  775. 22:11how do you make these prompts better?
  776. 22:13How do you make this engineering system
  777. 22:15better? I mean,
  778. 22:16um we're just talking about like, "Oh,
  779. 22:17every team has a skill now and we want
  780. 22:19to consolidate and make that shared and
  781. 22:20let people benefit from each other's
  782. 22:22learnings." How do you make that stuff
  783. 22:24better without like breaking somebody's
  784. 22:26workflow or regressing it for some some
  785. 22:28team?
  786. 22:29Uh if you want to help us, if you're in
  787. 22:32San Francisco and you're working on
  788. 22:33critical systems and you want to like
  789. 22:35figure out how to get coding agents to
  790. 22:36do more, uh let's chat. We're also
  791. 22:39hiring. Um send us a note either way
  792. 22:41[email protected].
  793. 22:43Uh we're building a IDE that
  794. 22:45orchestrates this stuff for you. Uh you
  795. 22:47don't need this to get this value out of
  796. 22:49this, but this is the kind of stuff
  797. 22:50we're working on. Uh if you want to hang
  798. 22:52out, I'm doing a sandbox research
  799. 22:54hackathon on Saturday. We're going to
  800. 22:56just get together with a bunch of cool
  801. 22:57builders, test all of the sandbox
  802. 22:59providers together,
  803. 23:00uh and see which one's the best, and
  804. 23:02then share our learnings. I'll be also
  805. 23:04be at the Daytona Compute Conference.
  806. 23:06And if you feel like coming to Miami,
  807. 23:08AI Engineer Miami is going to be really
  808. 23:09fun. We'll be giving
  809. 23:11the updated version of this talk with
  810. 23:13more stuff that I didn't have time to
  811. 23:14get to today.
  812. 23:15Thank you so much to all of you for your
  813. 23:17energy, to Demetrius and the entire
  814. 23:19organizing squad.
  815. 23:21Good luck.
  816. 23:22>> Questions. Who's got a question for Dex?
  817. 23:25That was super fast. I like it. I was
  818. 23:29very doubtful that you were going to get
  819. 23:30through it, but I like it.
  820. 23:32All right. I'm I'm curious about reading
  821. 23:33the code. Like it's not scalable, right?
  822. 23:36Like are we Are you going to be saying
  823. 23:37the same thing in 6 months?
  824. 23:39>> I mean, 6 months ago I said not to read
  825. 23:41it.
  826. 23:41Anyway, so I think everyone who is
  827. 23:43saying don't read the code now is going
  828. 23:44to be in 6 months being like, yeah, we
  829. 23:46had to throw that out. There's something
  830. 23:47There's something in the middle, right?
  831. 23:49We're binary searching through the space
  832. 23:51of how much of the code should you read.
  833. 23:55I think yeah, the idea is if you still
  834. 23:57read the code, you can still get 2 to 3x
  835. 23:59speed it up, and that's actually better
  836. 24:01business outcomes than
  837. 24:05than going 10x faster and shipping a
  838. 24:06bunch of slop and hoping that, you know,
  839. 24:08GPT-7 will fix it for you.
  840. 24:12>> Yeah, hit thanks, Dex.
  841. 24:14Awesome talk. Curious your thoughts on
  842. 24:18like the software factory. I think it's
  843. 24:20like strong DM that's saying the
  844. 24:22opposite, which is like never have a
  845. 24:25human read either side of it, and I
  846. 24:27think that pushes us further into evals
  847. 24:31and stuff like that. So, what is your
  848. 24:34thinking on that?
  849. 24:35>> Yeah, there is a whole class of like
  850. 24:37there's a whole rabbit hole you can go
  851. 24:38down with like formal verification and
  852. 24:40TLA+ or I talked to a guy who's building
  853. 24:43a new TLA+ that is TLA++. That is like,
  854. 24:45okay, what if we don't read the code?
  855. 24:47How can we actually like formally verify
  856. 24:49everything that's working?
  857. 24:50I think there's a lot more to be built,
  858. 24:53and I think there's a lot of people
  859. 24:55right now who need to ship like code to
  860. 24:57production systems faster. So like maybe
  861. 25:00someday, but like I used to cite Shawn
  862. 25:01Grose's talk where he was like it's just
  863. 25:03the spec. Just write the document that
  864. 25:05explains the desired behavior and you
  865. 25:07treat the code like it's assembly and
  866. 25:08you never read it anymore. Um
  867. 25:11I do not endorse that. Let's put it that
  868. 25:14way.
  869. 25:17>> We got one more. All right. Last one.
  870. 25:20>> Uh I know you mentioned one of the
  871. 25:22slides about the like context window and
  872. 25:25uh the the dumb zone, right? Uh I know
  873. 25:27you researched that like heavily a few
  874. 25:30like Have you Have you like gone back to
  875. 25:32look at that again to see how true that
  876. 25:33still is after certain like context
  877. 25:36window especially with like all the auto
  878. 25:38compaction they have now and other
  879. 25:40methods for that.
  880. 25:41>> I mean I think like
  881. 25:44for if you were have been using AI
  882. 25:47coding agents for 6 to 9 months and you
  883. 25:50use them for 60 hours a week like the
  884. 25:52dumb zone is not a useful concept to
  885. 25:54you. I will regularly go up to 60. I
  886. 25:56will regularly like aggressively keep it
  887. 25:58below 30. It depends on the complexity
  888. 26:00of your task, the amount of instructions
  889. 26:03versus information. So like your mileage
  890. 26:06may vary. If you are using coding agents
  891. 26:08for the first time, this is what I This
  892. 26:09is what we teach people is like if you
  893. 26:11don't know what to do and you haven't
  894. 26:12developed that intuition, then like
  895. 26:14shoot to keep it under 40 and if you get
  896. 26:15up to 60 like think about wrapping it up
  897. 26:17and like you can keep iterating on the
  898. 26:19same doc. That's what's also nice about
  899. 26:20these is like we don't use the built-in
  900. 26:22compaction because everything that
  901. 26:23matters is going into static assets. And
  902. 26:26so you can always resume from where you
  903. 26:27left off without having to worry about
  904. 26:28the quality of an auto compact or manual
  905. 26:30compact.
  906. 26:32>> Brilliant. Dex.
  907. 26:35Well done, dude. Thank you. Let's give
  908. 26:37it up for him, huh?
  909. 26:39Yes.

About this transcript

This page contains the full transcript of Everything We Got Wrong About Research-Plan-Implement - Dexter Horthy by AAIF Live, generated from the public captions YouTube serves with the video. The transcript has 6,100 words across 909 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.