Everything We Got Wrong About Research-Plan-Implement - Dexter Horthy — Transcript
Full transcript
- 0:04All right, before I bring up Dex, I got
- 0:07to say I met James and James, yeah, can
- 0:11you stand up real fast? James told me
- 0:13this morning that he is going We can't
- 0:15see the QR code. You got to hit it up.
- 0:17He's going to be having dinner tonight
- 0:19and everybody's invited. So,
- 0:22if you want to go,
- 0:24just scan that QR code real fast and go
- 0:27hang out with James. That's awesome. And
- 0:30that's what we're going for here. That
- 0:31is great.
- 0:32Yeah, you rock. Yeah, that's I I like
- 0:34it. I like it, too. Okay, so I'm going
- 0:37to bring up Dex.
- 0:38Earlier somebody said we need to have a
- 0:40mustache competition. I don't think
- 0:42that's going to happen, but I will say
- 0:44that he gave us 200 slides that he's
- 0:47about to present.
- 0:48>> Woah, woah, 100 158.
- 0:50>> 158. We're going to keep him honest on
- 0:53the timing, all right? So, let it start
- 0:55now.
- 0:56>> Amazing. Let's do it. What's up,
- 0:58everybody?
- 1:02Uh
- 1:02I am Dex. Uh this is a talk with a very
- 1:05long title that I'm not going to read
- 1:06because we're on the clock now.
- 1:08Uh I have been talking about coding
- 1:09agents for quite a while, basically
- 1:11since like August. Um we did a long talk
- 1:13in November. Um there's this methodology
- 1:16that we've been talking about a lot uh
- 1:17called research plan implement. Um
- 1:20lot of upvotes on Hacker News. There's
- 1:21probably 10,000 people who have gone to
- 1:23our open source and grabbed our prompts
- 1:25and are using them internally from small
- 1:26startups up to the enterprise.
- 1:28Um it all started with this guy. Um this
- 1:30guy
- 1:31Has anyone seen this talk?
- 1:33Yes? Okay, so Igor went in and he said
- 1:35like, "Okay, cool. We're using a lot of
- 1:36tokens. We're spending a lot of money to
- 1:37get AI developer productivity." But what
- 1:39he found was that it actually tends to
- 1:41lead to a lot of rework. Like you are
- 1:42shipping 50% more, but half of that is
- 1:45just cleaning up the slop from last
- 1:47week. And the other thing they found,
- 1:48and these are last year's numbers, so
- 1:50this is not account for Opus 4.5. So, I
- 1:52would I would inflate this a little bit,
- 1:53but like it's great for low complexity
- 1:56greenfield tasks, not great for
- 1:59high complexity brownfield tasks. Um,
- 2:02and so I could give you a talk about RPI
- 2:04and why it's great. Uh, but that would
- 2:06be boring. And there's other talks. So
- 2:08if you haven't seen them, go watch them.
- 2:09They'll give more context, but uh
- 2:12I'm going to tell you everything we got
- 2:13wrong about RPI today. Um, we thought we
- 2:15had this AI thing thing figured out. I
- 2:18uh am am uh humble enough to admit when
- 2:20I was wrong.
- 2:21Uh, so we got a couple things wrong.
- 2:23One thing that's very relevant if you've
- 2:24been on Twitter today is uh I don't
- 2:26think it's okay to not read the code.
- 2:28Uh,
- 2:29I also don't think you should read
- 2:30really long plan files. Uh, well those
- 2:32two are related.
- 2:34Uh, and no, Claude should not be allowed
- 2:36to have If you If you're writing
- 2:37production code that is used by users
- 2:39and you're going to get paged at 3:00
- 2:40a.m. if it's broken, uh
- 2:42no slop. This is the year, 2026, no more
- 2:45slop. So we're all on this journey.
- 2:46We're all figuring this out. We all are
- 2:48wrong all the time. We did get a couple
- 2:50things right. Uh, there is no magic
- 2:51prompt. Uh, do not outsource the
- 2:53thinking. You the engineer are an
- 2:55important part of this process, and seek
- 2:57leverage. There's a lot of code being
- 2:59written. Find ways to make sure it's
- 3:01correct without having to read all of it
- 3:02and re-steer after the fact. Um
- 3:05So lots of people I'm sure have heard of
- 3:06research plan implement. Has anyone
- 3:08actually run this Claude command,
- 3:10research code base?
- 3:11Okay, cool. Leave your hand up if you've
- 3:13run it like this. Tell me how this
- 3:15system works.
- 3:16Has anyone run it like this of like,
- 3:18"Hey, I want to build this thing. Go do
- 3:18the research."
- 3:20Okay.
- 3:21Um, or maybe go fetch a ticket or
- 3:22something. What about the create plan uh
- 3:25prompt? Okay, couple hands. Uh,
- 3:28how many of you run it like that? You
- 3:29know, "Hey, we got to go build this
- 3:30thing." Yes?
- 3:32Has anyone run it like this?
- 3:33Work back and forth with me starting
- 3:35with your open questions and outline
- 3:36before writing the plan.
- 3:37Okay, some of you found out about the
- 3:39magic words. A lot of people didn't
- 3:41though. Uh, we'll get into why that's a
- 3:42problem. Um, so since October we've
- 3:45basically worked with thousands of
- 3:46engineers from tiny startups all the way
- 3:48up to Fortune 500s.
- 3:49Um, and we would find over and over
- 3:51again we would give these tools to an
- 3:52expert uh and they would get great
- 3:54results. They would go sit and talk to
- 3:55Claude for 70 hours a week and they
- 3:56would start shipping like crazy.
- 3:58And then they would go give it to their
- 3:59team
- 4:00and the results were not always so good.
- 4:02And so people weren't getting good
- 4:04results. And so we got in the trenches
- 4:05with our users and we went to go figure
- 4:07out what was going wrong. And the first
- 4:09thing that was going wrong was people
- 4:10were not getting good research.
- 4:12So we talked about this in November.
- 4:14This is
- 4:15one of the only slides I'm ever using,
- 4:16but you would pick a zone of your code
- 4:17base. You would say, "Oh, we're going to
- 4:18build something over here." And then you
- 4:20would launch it in coding agent session
- 4:21to go send these sub agents through
- 4:23these deep vertical slices through the
- 4:24code base for just the context
- 4:26compressed context about what is the
- 4:28thing we're about to go build.
- 4:30Right? And we said, "Keep things
- 4:32objective. Discourage opinions. Don't
- 4:35actually put any implementation details
- 4:36in there. You just want to compress the
- 4:38truth. What is true about how the code
- 4:40works today?"
- 4:42And a skilled engineer was really good
- 4:43at taking, "Okay, here's my ticket. Let
- 4:45me write some questions that will cause
- 4:47the model to go touch all the parts of
- 4:48the code base that matter." So if it
- 4:50was, you know, add a new endpoint to
- 4:51reticulate splines across tenants, we
- 4:54would say something like, "Okay, tell me
- 4:55how endpoint endpoints work and trace
- 4:57the logic flow for everything that
- 4:58touches splines and go find the workers
- 5:00that do all the reticulation."
- 5:02So if this is your ticket, a lot of
- 5:04people would run it like this. They
- 5:06would just say, "Hey, research code
- 5:07base. Here's what I'm building."
- 5:08And the problem is that good research is
- 5:10all facts, but if you tell the model
- 5:11what you're building, then you get
- 5:12opinions. And we don't We'll get into
- 5:14why the model shouldn't have opinions
- 5:15later. Comes back to this thing that
- 5:17Jake from Netflix came up with, which is
- 5:19do not outsource the thinking.
- 5:21The other thing that wasn't working is
- 5:23people were getting not great plans.
- 5:25And basically there were these steps
- 5:27built into this planning prompt
- 5:29that was this single giant like
- 5:31monolithic thing with 85 or more
- 5:33instructions. And it had these steps in
- 5:36it of like, "Cool, present design
- 5:37options to the user, get feedback on the
- 5:39structure before you actually go write
- 5:41the plan." And so a good planning
- 5:43session would look something like, you
- 5:45know, you have your Claude system tools
- 5:47in your prompt and then you say, "Hey,
- 5:49create plan."
- 5:50Loads the skill, looks at your ticket,
- 5:52loads your research doc, launch a bunch
- 5:54of sub agents to go find a bunch of
- 5:55things that are true about the code
- 5:56base, just confirm some stuff that
- 5:58wasn't maybe in the research. This is
- 5:59all one big context window, by the way.
- 6:00Usually, I'll use these columns to mean
- 6:02separate context windows, but today this
- 6:04is all one session. I just slides are
- 6:06sideways, so I had to put them on next
- 6:08to each other.
- 6:09Um but the agent would come and ask
- 6:11questions. Say, "Okay, here's our
- 6:12options for question one." User would
- 6:13pick an option. User would pick an
- 6:15option. And then eventually, it would
- 6:16say, "Cool, here's the order we're going
- 6:17to do the things. Um what do you think?"
- 6:19And the user could say, "Well, we need
- 6:20to add a testing step up front and I
- 6:21want to swap phases three and four."
- 6:23Assistant would give the new outline of
- 6:25the phases.
- 6:26Then the user would approve it. And only
- 6:28then
- 6:29would we write our plan file.
- 6:30Um
- 6:32complex process of aligning with the
- 6:33user on what was what was going to be
- 6:35built.
- 6:36Um but for about 50% of people, maybe
- 6:38more, if you didn't prompt it with this
- 6:40work back and forth with me or Opus was
- 6:42just feeling dumb for that that
- 6:44particular hour of the day, um
- 6:47it would just take the stuff and it
- 6:48would just immediately go and write the
- 6:50plan out. And so, you would get this and
- 6:52they'd be like, "Cool, I wrote the
- 6:53plan." Didn't ask me any questions, made
- 6:55all the decisions for me. Yikes.
- 6:57So, we give the tools to people and some
- 7:00people got good results and some people
- 7:01didn't. And we dug in and we were like,
- 7:02"What's the difference?"
- 7:04Uh
- 7:05and people would literally say this to
- 7:06me. They'd be like, "Well, you have to
- 7:07say the magic words." And I found myself
- 7:09in workshops full of enterprise
- 7:10engineers saying, "Well, guys, guys,
- 7:12guys, guys, yeah, here's the software,
- 7:13but don't forget to say the magic
- 7:14words." It was um quite frankly, it was
- 7:16embarrassing.
- 7:17But if you said this, work back and
- 7:18forth with me starting with your open
- 7:19questions and outline before writing the
- 7:21plan, then the agent would actually ask
- 7:22you the questions. And this isn't the
- 7:24user's fault. If you built a tool that
- 7:26requires hours and hours of training and
- 7:28reps to get like good results from, go
- 7:31fix the tool. And so, I'll talk about
- 7:33how we did that. Um but why these steps
- 7:35were getting skipped, the one of the big
- 7:36takeaways I'll give you today is like
- 7:38you have an instruction budget. Uh
- 7:41my co-founder Kyle is somewhere over
- 7:42here. He wrote this really good blog
- 7:44post in December or November, I guess
- 7:46technically, um that basically cited
- 7:47this archive paper, which again, this is
- 7:49from last year, so the number is
- 7:51probably a little bit higher now, but
- 7:52that frontier LLMs could only follow
- 7:54about 150 to 200 instructions with like
- 7:57good consistency. Anything more than
- 7:58that and it's kind of half attending to
- 8:00all of them and you're rolling the dice.
- 8:02So, if you have a prompt with 85
- 8:04instructions and your Claude MD and your
- 8:06system prompt and your tools and your
- 8:08MCP, um yeah, you're not likely to get
- 8:12full adherence to the workflow. So, more
- 8:14on how we fix this later.
- 8:16The other thing that I think, um really
- 8:18wasn't working for people was like we
- 8:20advocated for reading the plans that
- 8:21were output. This is me on stage in
- 8:24November telling people, you have to
- 8:25read the plan, otherwise it won't work.
- 8:28Um some people even would PR their plans
- 8:30and code review them together. But a
- 8:32thousand line plan tends to be about a
- 8:34thousand lines of code within 10% or so,
- 8:37and plans can have surprises. So, you
- 8:38would go and you would review the plan
- 8:40and then you would go right to code and
- 8:42it would be different. And so, you're
- 8:43telling you're asking one of your
- 8:44co-workers like, okay, you go spend an
- 8:46hour reading this and tell me what's
- 8:47wrong with it, and then you would go
- 8:48implement it and it would be different.
- 8:49They'd have to go read the code again
- 8:50and see what the surprises were and what
- 8:52changed. Um and so, this isn't leverage.
- 8:55Leverage is about like do less work to
- 8:57get more output. So, the new advice, uh
- 9:01don't read the plans.
- 9:02Please, read the code.
- 9:04Uh just cuz it's it's the same amount of
- 9:06work and like look for leverage
- 9:07elsewhere and I'll talk about how we
- 9:08found better leverage. Um and you may
- 9:10say, "Hey Dex, in August you said don't
- 9:12read the code. You said that the plans
- 9:14are enough. Just don't just just go just
- 9:15ship and let Claude do its thing."
- 9:17I was wrong. I am humble enough to admit
- 9:20when I was wrong. Uh this is actually a
- 9:21very big conversation right now. Please,
- 9:24please read the code. We tried not
- 9:25reading the code for like 6 months.
- 9:27Uh it did not end well. We had to rip
- 9:28out and replace large parts of that
- 9:30system. Um
- 9:31and you may say, "Hey Dex, but other
- 9:33people don't read the code." Beats,
- 9:34300,000 lines and counting. Uh no one's
- 9:37read that code, allegedly.
- 9:39Uh open claw, Pete's like, "Okay, you
- 9:40know, I know the structure and the
- 9:42pieces and how they fit together, but I
- 9:43don't read every line of every PR."
- 9:46Um these are OSS projects. They don't
- 9:48charge money.
- 9:49Nobody gets paged at 3:00 a.m. if it's
- 9:51broken, and no one gets fined millions
- 9:53of dollars if it's done wrong. I will
- 9:55also say though,
- 9:56these are OSS. They are very, very cool
- 9:58projects. I am humbled, deeply humbled
- 10:01by the accomplishments of the
- 10:02maintainers, and the stakes are still
- 10:04high. Like, if you break open claw, a
- 10:06lot of people are going to be upset. But
- 10:08they are different than if you were,
- 10:09say, working in a regulated industry
- 10:11shipping production SAS code.
- 10:13Um so, if you have people who depend on
- 10:14your code,
- 10:16please, I'm begging you, please read it.
- 10:19Please read it. We have a profession to
- 10:20uphold. 2026 is supposed to be the year
- 10:23of no more slop. Uh literally everyone
- 10:26is talking about the difference between
- 10:27slop and craft.
- 10:29Uh this is why I'm a little mid on agent
- 10:31swarms and the whole gas town thing
- 10:33because you still need to be able to
- 10:35ensure quality, and like going 10 times
- 10:37faster doesn't matter if you're going to
- 10:38throw it all away in 6 months. So, shoot
- 10:41for 2 to 3x. That's actually another
- 10:42talk of like how you measure this and
- 10:44how you actually get there and maintain
- 10:45like a near human level of quality. Um
- 10:48but I'll talk about the goals and like
- 10:49what you should think about if you want
- 10:50to get there is you should have high
- 10:51leverage planning.
- 10:53You should not outsource the thinking.
- 10:55Read and own the code.
- 10:56And ideally we will avoid uh
- 10:59magic words.
- 11:00So, uh
- 11:01we got better research, we got better
- 11:02plans, we got better leverage. I'm going
- 11:04to talk about each of those um as far as
- 11:06like, in general what we in like it's
- 11:08specifically what we did, and also some
- 11:10general concepts as you're building
- 11:11workflows and systems around coding
- 11:13agents, what you can do.
- 11:14So, we talked about a skilled This is
- 11:15the least exciting one, but talked about
- 11:17how a skilled engineer could detangle
- 11:18the ticket to the questions to the
- 11:20research,
- 11:21uh and then the research would be very
- 11:22objective.
- 11:23Um
- 11:24basically we just hide the ticket from
- 11:26the context window that's doing
- 11:27research, and we do it
- 11:28deterministically. So, basically you
- 11:30have one context window to generate
- 11:31questions, and then a fresh context
- 11:33window with no knowledge of what we're
- 11:34building to go make your research doc.
- 11:37Um this is pretty trivial. If you're
- 11:38familiar with the concept of query
- 11:39planning,
- 11:40um
- 11:41it's
- 11:41similar in concept but for, you know,
- 11:43LLMs reading through codebases.
- 11:46Um so, I've been hacking on agents for a
- 11:48while and before we did the coding agent
- 11:49stuff, I wrote this paper called 12
- 11:50factor agents, which was uh allegedly
- 11:53the first time anyone was like talking a
- 11:54lot about context engineering. Uh
- 11:57there's two ways to read context
- 11:59engineering and most people jumped in.
- 12:01Is anyone building like rag pipelines?
- 12:02Is it Raise your hand if you built a rag
- 12:04pipeline.
- 12:05Okay, some people are feeling uh not
- 12:07like not raising their hands today. Um
- 12:10But it's like, okay, put more
- 12:11information in, the model can't make
- 12:13sense of it. I actually think the more
- 12:15interesting read of context engineering
- 12:17is like better instructions and simpler
- 12:19tasks and smaller context windows. Of
- 12:21course, we all know Jeff now. I don't
- 12:22have to introduce him anymore. He used
- 12:23to have to I used to have to tell people
- 12:25who Jeff was when I was talking. Um we
- 12:28talked about this like context window
- 12:29thing as the idea of the dumb zone,
- 12:31which is, you know, you have about
- 12:33168,000 tokens and 200,000 but some of
- 12:36them are reserved for output. You have
- 12:38various things that they're for and
- 12:39around like 40% on average depending on
- 12:41what you're doing and how much of your
- 12:42context is user messages versus files
- 12:44and all of this stuff, you hit this
- 12:46point where you have degrading results.
- 12:48And obviously sometimes you can get
- 12:49still get good enough for you results at
- 12:5160% but the less of the context window
- 12:54you use, the better results you will
- 12:55get. Um our friends at Databricks were
- 12:57just talking about you have too many
- 12:58MCPs. The whole context window is full
- 13:00of instructions about how to use a bunch
- 13:01of tools that you don't care about and
- 13:03then by the time you're writing code,
- 13:04the model's like not good at following
- 13:05your instructions.
- 13:06So, you're not just giving the model too
- 13:07much information, you're also probably
- 13:10giving it too many instructions.
- 13:12And so, the idea of what we're doing was
- 13:14this thing like makes a lot of sense,
- 13:16use prompts for control flow. This is a
- 13:18customer support example but, you know,
- 13:19if it's a complaint, go do this. If it's
- 13:21product feedback, go do this. If it's a
- 13:23billing issue, go do this.
- 13:25Um
- 13:25And what you could do instead is you can
- 13:27instead of using prompts for control
- 13:28flow,
- 13:29you can kind of classify the input and
- 13:31then feed it to a series of smaller,
- 13:33more focused prompts where there are far
- 13:35fewer instructions and far fewer actions
- 13:37to choose from. I'm sure many of us have
- 13:38already done things like this to improve
- 13:40the performance of pipelines. Um so this
- 13:42was a single mega prompt with 85
- 13:44instructions.
- 13:45Um and if you did it right, you would go
- 13:46through all these different steps. All
- 13:48these different phases were part of
- 13:49that. And if any of the instructions
- 13:50didn't get followed, you would skip the
- 13:52things that made this really high
- 13:53leverage.
- 13:54Um so we split it across several
- 13:55prompts.
- 13:56And so like before it was research,
- 13:57plan, implement, now it's questions,
- 13:59research, design, structure, plan, work
- 14:00tree, implement, PR. We're not actually
- 14:01not going to have time to talk about the
- 14:03implement side of the thing today. But
- 14:05um
- 14:06we split up the planning into a design
- 14:07discussion, an outline, and a plan.
- 14:10And before it was 85 instructions, now
- 14:13they're all less than 40, which is
- 14:14really exciting. And I think some of
- 14:15them could actually be even smaller.
- 14:17We're still iterating on them. The
- 14:18lesson is don't use prompts for control
- 14:20flow if you can use control flow for
- 14:21control flow. Like the if statement is
- 14:23really, really powerful and LLMs are
- 14:25really good at classifying things. This
- 14:26is not just true for coding agents. This
- 14:28is any AI LLM-based system you're
- 14:29building.
- 14:30Um and it's really funny cuz we were
- 14:32writing all this stuff and we got on
- 14:33stage and we said like full fat agents
- 14:35don't work. Don't just call tools in a
- 14:37loop, do context engineering and build
- 14:38workflows and graphs and micro agents.
- 14:40We told everybody don't do this. And
- 14:42then we turned around in August and
- 14:43we're like, "Oh,
- 14:44all right, but this Claude code thing is
- 14:46pretty good." And we turned around and
- 14:47we wrote this giant monolithic prompt.
- 14:49So we figured it was time to actually go
- 14:50drink our own Kool-Aid.
- 14:52Um
- 14:53mind your instruction budget.
- 14:56How do we get better leverage?
- 14:57So we split things up to get better
- 14:58instruction following, right?
- 15:00These three different phases. But we
- 15:02also got more leverage. I'm going to
- 15:03talk about why. Because even if the plan
- 15:05is a thousand lines and the code is a
- 15:06thousand lines, your design discussion
- 15:08might only be 200 lines. And you get a
- 15:10lot of opportunities to restear in that
- 15:11moment. And so what this looks like is
- 15:13basically where are we going? What does
- 15:15the final solution look like? And it
- 15:17has, you know, the current state, the
- 15:18desired end state. It has the patterns
- 15:20to follow. How many of you have ever
- 15:21like sent a coding agent and it like
- 15:23found the wrong way to do a thing in
- 15:25your code base and it followed the bad
- 15:26patterns? Yes?
- 15:28Right. This is your chance to go read
- 15:30all the patterns it found that it thinks
- 15:31are relevant and be like, "Nope, that's
- 15:33not how we do atomic SQL updates. That's
- 15:35some engineer that doesn't work here
- 15:36anymore and it's crazy and everyone
- 15:37hates it. Go find the way we do it over
- 15:38there."
- 15:39Um it'll keep track of resolved design
- 15:41decisions that we've made. It will ask
- 15:43open questions. This is sort of like
- 15:45taking Claude code plan mode and the ask
- 15:47user question tool and just brain
- 15:49dumping it all to the single document
- 15:50that you can interact with is like
- 15:52moldable and flexible.
- 15:54Um Matt Pocock has this idea, he calls
- 15:55it the design concept and it's this idea
- 15:57of like the thing that is locked up in
- 15:59this context window that is the shared
- 16:02understanding between you and the agent
- 16:04of what's being built and how.
- 16:06Uh so we put it into an underlying
- 16:07markdown artifact.
- 16:09Um and so we now have human agent
- 16:11alignment. And the idea here is like
- 16:12you're forcing the agent to brain dump
- 16:14out all the things it found, all the
- 16:15things it wants to do, all the things it
- 16:17thinks you want, and ask you questions
- 16:19about things it doesn't know. So you can
- 16:20do brain surgery on the agent before you
- 16:22proceed downstream. And it's all about
- 16:24do not outsource the thinking. You want
- 16:26to give the agent every single
- 16:27opportunity to show you what it's wrong
- 16:29about before you go write 2,000 lines of
- 16:31code.
- 16:33So uh 200 lines instead of a thousand, a
- 16:35little bit more leverage. We also get
- 16:37better leverage from the outline. So if
- 16:39design is like, "Where are we going?"
- 16:41the structure outline is, "How do we get
- 16:43there?" Or if you're an engineer who is
- 16:45miserable cuz of sitting in meetings all
- 16:46day, there's the like architecture
- 16:48review and then there's the sprint
- 16:50planning meeting. What are we going to
- 16:51build and then how do we break it down
- 16:53into tasks?
- 16:54And so we take our design and we take it
- 16:56to ticket in the research and we build
- 16:57up a new context window and we create
- 16:59the structure outline.
- 17:01And this is basically a high-level
- 17:02overview of the phases, not the exact
- 17:04code we're going to write, but just kind
- 17:06of what it's going to look like, what
- 17:07order we're going to do the changes in,
- 17:09and how we're going to test it along the
- 17:10way. Now, I don't actually test in
- 17:12between every phase everything I'm
- 17:13building, but if it's sensitive or if
- 17:15it's hard or if it's complex, I want to
- 17:17be able to catch it before it goes and
- 17:19writes all the code. I want to make sure
- 17:21each two, three, 400 line block is
- 17:23correct. Um and these docs mean lighter
- 17:25reviews. Instead of reviewing the plan,
- 17:27this is two things for the same feature.
- 17:28Plan eight pages, structure outline
- 17:30two-ish pages, much shorter. Um
- 17:33I like to think of this Has anyone ever
- 17:34written a C header file, a .h file?
- 17:37Yeah, okay. So, if the plan is the
- 17:39implementation, the outline is the C
- 17:41header files. Just here's the signatures
- 17:42and the new types that we're changing.
- 17:44Enough again for you to see what the
- 17:46agent is thinking and correct it if it's
- 17:48wrong.
- 17:49Um and the reason why we do this is
- 17:51despite like every single model and
- 17:53trying to prompt this out and eval the
- 17:54hell out of this, we cannot get models
- 17:56to stop writing horizontal plans. Or it
- 17:58like this is the best way to fix their
- 18:01need to write horizontal plans. And when
- 18:02I say horizontal plans, I basically mean
- 18:05you start with Models love to like we're
- 18:07going to do all the database and then
- 18:08we're going to do all the services and
- 18:09then we're going to do all the API and
- 18:11then we're going to do all the front end
- 18:11and before you know it, you're on the
- 18:13other side of 1,200 lines of code and
- 18:15it's not working.
- 18:17And now you have to go figure out which
- 18:18part is broken because there was no
- 18:19nothing really to test along the way,
- 18:21whether the model is verifying it or
- 18:23whether you the human are jumping in and
- 18:24checking it's correct. And so what we've
- 18:27seen work really, really well across
- 18:28orgs of all sizes
- 18:30um is what I call vertical plans. This
- 18:32is how I build when I'm like before AI,
- 18:34I would like make a mock API endpoint
- 18:36and then get it working in the front end
- 18:37and then wire that and then mock out the
- 18:39services layer and then do the database
- 18:41migration and then put everything
- 18:43together.
- 18:44And so, even though it's the same amount
- 18:46of code, you have these like checkpoints
- 18:48where you can see if it's working and if
- 18:49it's not, you can pause and fix it
- 18:51before you go try to do the rest of it.
- 18:54So, these are just markdown docs too.
- 18:55Like you can and should ask for more
- 18:56detail. They start high level, but like
- 18:58here's an example of like I don't think
- 18:59you're going to get this right. Tell me
- 19:01what you're thinking and then like
- 19:02dumped out the types and the signatures.
- 19:04Um and then getting better leverage from
- 19:06the plan itself, I mean, again, like
- 19:08usual, like we've been doing, we just
- 19:09take that artifact, we build it up with
- 19:11all the previous artifacts, and then we
- 19:13can go build the plan.
- 19:14Um and this is the same if you use
- 19:16create plan, it's the exact same
- 19:17template, exact same setup, exact same
- 19:18prompt. But this is a tactical doc for
- 19:20the agent. We've already done enough
- 19:22aligning that like I'm just going to
- 19:24spot-check this, and then we save the
- 19:25deep review for the actual code. And so
- 19:28if you've used any of the RPI plans,
- 19:30they look like this. It's the model
- 19:31saying, "Hey, here's all the changes I'm
- 19:32going to make."
- 19:33Um
- 19:34the most important part of this leverage
- 19:36is not just about you and the agent
- 19:37though. Like human agent alignment is
- 19:39important, and knowing what the agent's
- 19:40going to do and correcting that is is
- 19:42good, but it's also, you know, if you're
- 19:44working with a team of engineers, we've
- 19:46found a lot of value from taking these
- 19:48design discussions, these structure
- 19:50outlines, and review. I said don't
- 19:51review the plans, but these shorter docs
- 19:54are really, really good. Uh
- 19:56I I am not the code owner of most of our
- 19:59code at Human Layer. My uh co-founder
- 20:00is, and I send him my design discussions
- 20:03on purpose. We don't have a required
- 20:05step, but I want to I want to know that
- 20:07when we get to code review, it's just
- 20:09going to be like, "Yep, that's That's
- 20:10what I wanted. That's it. That's it." So
- 20:11any any of my bad decisions are headed
- 20:14off on a 200-line doc before I've gone
- 20:16and written the code and gotten it
- 20:17working and I'm attached to it. And so
- 20:19this is really, really powerful. Um
- 20:22before AI, we would basically the way
- 20:23Another way to think about it is like
- 20:24time savings. You would say, "Okay, it's
- 20:26a 2-day feature. I got to do all this
- 20:28stuff. The coding's probably 2 to 4
- 20:29hours."
- 20:30If you just pick up Claude code and use
- 20:32it to ship for you, you do get some
- 20:33speed up because now the coding takes 20
- 20:35minutes. It's still a 2-day feature cuz
- 20:37I still have to like align with my team
- 20:39on what we're going to do. I still have
- 20:40to get a code review and fix stuff.
- 20:42Maybe I'm working across repos that I
- 20:43don't personally own, and then we still
- 20:45have to verify and test it.
- 20:47But if you use AI to help you with your
- 20:49planning and alignment, then you also
- 20:51save time there, and I think you get
- 20:54much better alignment. Um and so your
- 20:56code review and rework is also much
- 20:58shorter because you already know what's
- 20:59coming. The team that's reviewing it
- 21:00already kind of like had their chance to
- 21:02restear you. And really good teams do
- 21:03this. It's They have a meeting that's
- 21:05called architecture review where we
- 21:06decide, you know, what's our technical
- 21:07design doc on how we're going to build
- 21:09this.
- 21:10So, um as far as testing and verifying,
- 21:12sorry, I don't have a good answer for
- 21:13you. It's a whole other talk. If you
- 21:15went to Drew's talk downstairs, go find
- 21:17Drew Brignac after this. He will tell
- 21:18you all about testing and verifying.
- 21:20Um let's put this all together.
- 21:22So, we have these five stages of
- 21:24research and planning.
- 21:25Um the process is basically questions,
- 21:27research, design, structure outline,
- 21:30plan, work tree, implement, finally the
- 21:32pull request.
- 21:34Uh that didn't make a very good acronym
- 21:35though, so we just picked the ones we
- 21:36liked and uh we're calling this crispy.
- 21:39Uh
- 21:40So, RPI to crispy, that's the There you
- 21:42go. Um what's next and what did I not
- 21:45have time to talk about today?
- 21:46Um three steps is already a lot for some
- 21:48people to learn and now there are seven.
- 21:50I thought we were supposed to make this
- 21:51easier for teams to learn this and adopt
- 21:52it. We can talk about how we're like
- 21:54thinking about that. Um the idea of how
- 21:56do you measure the impact of doing this
- 21:58um in engineering teams? I think it's
- 22:00like we've been trying to measure
- 22:01developer productivity for 50 years and
- 22:03we still don't know how to do it very
- 22:04well.
- 22:05Um and then it's like if you're a
- 22:07central kind of platform team rolling
- 22:09out changes to everybody in your org,
- 22:11how do you make these prompts better?
- 22:13How do you make this engineering system
- 22:15better? I mean,
- 22:16um we're just talking about like, "Oh,
- 22:17every team has a skill now and we want
- 22:19to consolidate and make that shared and
- 22:20let people benefit from each other's
- 22:22learnings." How do you make that stuff
- 22:24better without like breaking somebody's
- 22:26workflow or regressing it for some some
- 22:28team?
- 22:29Uh if you want to help us, if you're in
- 22:32San Francisco and you're working on
- 22:33critical systems and you want to like
- 22:35figure out how to get coding agents to
- 22:36do more, uh let's chat. We're also
- 22:39hiring. Um send us a note either way
- 22:41[email protected].
- 22:43Uh we're building a IDE that
- 22:45orchestrates this stuff for you. Uh you
- 22:47don't need this to get this value out of
- 22:49this, but this is the kind of stuff
- 22:50we're working on. Uh if you want to hang
- 22:52out, I'm doing a sandbox research
- 22:54hackathon on Saturday. We're going to
- 22:56just get together with a bunch of cool
- 22:57builders, test all of the sandbox
- 22:59providers together,
- 23:00uh and see which one's the best, and
- 23:02then share our learnings. I'll be also
- 23:04be at the Daytona Compute Conference.
- 23:06And if you feel like coming to Miami,
- 23:08AI Engineer Miami is going to be really
- 23:09fun. We'll be giving
- 23:11the updated version of this talk with
- 23:13more stuff that I didn't have time to
- 23:14get to today.
- 23:15Thank you so much to all of you for your
- 23:17energy, to Demetrius and the entire
- 23:19organizing squad.
- 23:21Good luck.
- 23:22>> Questions. Who's got a question for Dex?
- 23:25That was super fast. I like it. I was
- 23:29very doubtful that you were going to get
- 23:30through it, but I like it.
- 23:32All right. I'm I'm curious about reading
- 23:33the code. Like it's not scalable, right?
- 23:36Like are we Are you going to be saying
- 23:37the same thing in 6 months?
- 23:39>> I mean, 6 months ago I said not to read
- 23:41it.
- 23:41Anyway, so I think everyone who is
- 23:43saying don't read the code now is going
- 23:44to be in 6 months being like, yeah, we
- 23:46had to throw that out. There's something
- 23:47There's something in the middle, right?
- 23:49We're binary searching through the space
- 23:51of how much of the code should you read.
- 23:55I think yeah, the idea is if you still
- 23:57read the code, you can still get 2 to 3x
- 23:59speed it up, and that's actually better
- 24:01business outcomes than
- 24:05than going 10x faster and shipping a
- 24:06bunch of slop and hoping that, you know,
- 24:08GPT-7 will fix it for you.
- 24:12>> Yeah, hit thanks, Dex.
- 24:14Awesome talk. Curious your thoughts on
- 24:18like the software factory. I think it's
- 24:20like strong DM that's saying the
- 24:22opposite, which is like never have a
- 24:25human read either side of it, and I
- 24:27think that pushes us further into evals
- 24:31and stuff like that. So, what is your
- 24:34thinking on that?
- 24:35>> Yeah, there is a whole class of like
- 24:37there's a whole rabbit hole you can go
- 24:38down with like formal verification and
- 24:40TLA+ or I talked to a guy who's building
- 24:43a new TLA+ that is TLA++. That is like,
- 24:45okay, what if we don't read the code?
- 24:47How can we actually like formally verify
- 24:49everything that's working?
- 24:50I think there's a lot more to be built,
- 24:53and I think there's a lot of people
- 24:55right now who need to ship like code to
- 24:57production systems faster. So like maybe
- 25:00someday, but like I used to cite Shawn
- 25:01Grose's talk where he was like it's just
- 25:03the spec. Just write the document that
- 25:05explains the desired behavior and you
- 25:07treat the code like it's assembly and
- 25:08you never read it anymore. Um
- 25:11I do not endorse that. Let's put it that
- 25:14way.
- 25:17>> We got one more. All right. Last one.
- 25:20>> Uh I know you mentioned one of the
- 25:22slides about the like context window and
- 25:25uh the the dumb zone, right? Uh I know
- 25:27you researched that like heavily a few
- 25:30like Have you Have you like gone back to
- 25:32look at that again to see how true that
- 25:33still is after certain like context
- 25:36window especially with like all the auto
- 25:38compaction they have now and other
- 25:40methods for that.
- 25:41>> I mean I think like
- 25:44for if you were have been using AI
- 25:47coding agents for 6 to 9 months and you
- 25:50use them for 60 hours a week like the
- 25:52dumb zone is not a useful concept to
- 25:54you. I will regularly go up to 60. I
- 25:56will regularly like aggressively keep it
- 25:58below 30. It depends on the complexity
- 26:00of your task, the amount of instructions
- 26:03versus information. So like your mileage
- 26:06may vary. If you are using coding agents
- 26:08for the first time, this is what I This
- 26:09is what we teach people is like if you
- 26:11don't know what to do and you haven't
- 26:12developed that intuition, then like
- 26:14shoot to keep it under 40 and if you get
- 26:15up to 60 like think about wrapping it up
- 26:17and like you can keep iterating on the
- 26:19same doc. That's what's also nice about
- 26:20these is like we don't use the built-in
- 26:22compaction because everything that
- 26:23matters is going into static assets. And
- 26:26so you can always resume from where you
- 26:27left off without having to worry about
- 26:28the quality of an auto compact or manual
- 26:30compact.
- 26:32>> Brilliant. Dex.
- 26:35Well done, dude. Thank you. Let's give
- 26:37it up for him, huh?
- 26:39Yes.
About this transcript
This page contains the full transcript of Everything We Got Wrong About Research-Plan-Implement - Dexter Horthy by AAIF Live, generated from the public captions YouTube serves with the video. The transcript has 6,100 words across 909 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.