No Vibes Allowed: Solving Hard Problems in Complex Codebases – Dex Horthy, HumanLayer — Transcript
Full transcript
- 0:13[music]
- 0:20Hi everybody. How y'all doing?
- 0:24>> It's exciting. I'm Dex. Uh, as they did
- 0:26in the great intro, I've been hacking on
- 0:27agents for a while. Um, our talk 12
- 0:30factor agents at AI engineer in June was
- 0:32one of the top talks of all time. Uh, I
- 0:35think top eight or something. One of the
- 0:36best ones from from AI engineer in June.
- 0:38May or may not have said something about
- 0:39context engineering. Um, why am I here
- 0:42today? What am I here to talk about? Um,
- 0:44I want to talk about one of my favorite
- 0:45talks from AI engineer in June. And I
- 0:47know we all got the update from Eigor
- 0:48yesterday, but they wouldn't let me
- 0:50change my slides. So, this is going to
- 0:51be about what Eigor talked about in
- 0:53June. uh basically that they surveyed a
- 0:55100,000 developers across all company
- 0:57sizes and they found that most of the
- 0:59time you use AI for software engineering
- 1:01you're doing a lot of rework a lot of
- 1:02codebase churn uh and it doesn't really
- 1:05work well for complex tasks brownfield
- 1:07code bases um and you can see in the
- 1:10chart basically you are shipping a lot
- 1:11more but a lot of it is just reworking
- 1:13the slop that you shipped last week so
- 1:16uh and then the other side right was
- 1:18that uh if you're doing green field
- 1:20little versel dashboard something like
- 1:22this then it's going to work great. Uh
- 1:25if you're going to go in a 10-year-old
- 1:27Java at codebase, maybe not so much. And
- 1:29this matched my experience personally
- 1:31and talking to a lot of founders and
- 1:32great engineers, too much slop uh tech
- 1:35debt factory. It's just it's not going
- 1:36to work from our codebase. Like maybe
- 1:37someday when the models get better, but
- 1:40that's what context engineering is all
- 1:41about. How can we get the most out of
- 1:43today's models? How do we manage our
- 1:45context window? So we talked about this
- 1:47in August. Um I have to confess
- 1:49something. The first time I used cloud
- 1:51code, I was not impressed. It was like,
- 1:53okay, this is a little bit better. I get
- 1:54it. I like the UX. Um, but since then,
- 1:57we as a team figured something out. Um,
- 2:00that we were actually able to get, you
- 2:01know, 2 to 3x more throughput. And we
- 2:03were shipping so much that we had no
- 2:05choice but to change the way we
- 2:07collaborated. We rewired everything
- 2:09about how we build software. Uh, it was
- 2:11a team of three. It took eight weeks. It
- 2:13was really freaking hard. Uh, but now
- 2:15that we solved it, we're we're never
- 2:16going back. This is the whole no slop
- 2:18thing. I think I think we got somewhere
- 2:20with this went super viral on HackerNews
- 2:22in September. Uh we have thousands of
- 2:24folks who have gone on to GitHub and
- 2:25grabbed our you know research plan
- 2:26implement prompt system. Um so the goals
- 2:29here which we kind of backed our way
- 2:31into we need AI that can work well in
- 2:34brownfield code bases that can solve
- 2:36complex problems. No slop, right? No
- 2:39more slop. Uh and we had to maintain
- 2:42mental alignment. I'll talk a little bit
- 2:43more about what that means in a minute.
- 2:44And of course we want to spend with
- 2:46everything we want to spend as many
- 2:47tokens as possible. what we can offload
- 2:48meaningfully to the AI is really really
- 2:51important. Um, super high leverage. So,
- 2:53this is advanced context engineering for
- 2:55coding agents. Um, I'll start with kind
- 2:57of like framing this. The most naive way
- 3:00to use a coding agent is to ask it for
- 3:02something and then tell it why it's
- 3:03wrong and resteere it and ask and ask
- 3:05and ask until you run out of context or
- 3:07you give up or you cry. Um, we can be a
- 3:10little bit smarter about this. Most
- 3:11people discover this pretty early on in
- 3:13their AI like exploration. uh is that it
- 3:17might be better if you start a
- 3:18conversation and you're off track that
- 3:22uh you just start a new context window.
- 3:24You say, "Okay, we went down that path.
- 3:25Let's start again. Same prompt, same
- 3:26task, but this time we're going to go
- 3:28down this path and like don't go over
- 3:29there cuz that doesn't work." So, uh how
- 3:32do you know when it's time to start
- 3:34over?
- 3:35If you see this,
- 3:39it's probably time to start over, right?
- 3:41This is what Claude says when you tell
- 3:43it it's screwing up.
- 3:46Um, so we can be even smarter about
- 3:47this. We can do what I call intentional
- 3:49compaction. Um, and this is basically
- 3:51whether you're on track or not, you can
- 3:53take uh your existing context window and
- 3:56ask the agent to compress it down into a
- 3:58markdown file. You can review this, you
- 4:00can tag it, and then when the new agent
- 4:01starts, it gets straight to work instead
- 4:03of having to do all that searching and
- 4:04codebase understanding and getting
- 4:06caught up. Um, what goes into
- 4:08compaction? Well, the question is like
- 4:10what takes up space in your context
- 4:12window. So, um, it's looking for files,
- 4:15it's understanding code flow, it's
- 4:17editing files, it's test and build
- 4:19output. And if you have one of those
- 4:20MCPs that's dumping JSON and a bunch of
- 4:22UU ids into your context window, you
- 4:25know, God help you. Uh, so what should
- 4:28we compact? I'll get more specifics
- 4:29here, but this is a really good
- 4:30compaction. This is exactly what we're
- 4:32working on. The exact files and line
- 4:34numbers that matter to the problem that
- 4:35we're solving. Um, why are we so
- 4:38obsessed with context? Because LMS are
- 4:41actually got roasted on YouTube for this
- 4:42one. And they're not pure functions cuz
- 4:43they're nondeterministic, but they are
- 4:45stateless. And the only way to get
- 4:46better better performance out of an LLM
- 4:49is to put better tokens in and then you
- 4:51get better tokens out. And so every turn
- 4:53of the loop when Claude is picking the
- 4:54next tool or any coding agent is picking
- 4:56the next and there could be hundreds of
- 4:57right next steps and hundreds of wrong
- 4:59next steps. But the only thing that
- 5:01influences what comes out next is what
- 5:03is in the conversation so far. So we're
- 5:05going to optimize this context window
- 5:07for correctness, completeness, size, and
- 5:10a little bit of trajectory. And the
- 5:11trajectory one is interesting because a
- 5:13lot of people say, "Well, I I told the
- 5:14agent to do something and it did
- 5:16[clears throat] something wrong. So, I
- 5:17corrected it and I yelled at it and then
- 5:19it did something wrong again and then I
- 5:20yelled at it." And then the LM is
- 5:21looking at this conversation says,
- 5:23"Okay, cool. I did something wrong. The
- 5:24human yelled at me and then I did
- 5:25something wrong and the human yelled at
- 5:26me." So, the next most likely conver
- 5:27token in this conversation is I better
- 5:30do something wrong so the human can yell
- 5:31at me again. So, mind be mindful of your
- 5:34trajectory. If you were going to invert
- 5:36this, the worst thing you can have is
- 5:37incorrect information, then missing
- 5:39information, and then just too much
- 5:41noise. Um, if you like equations,
- 5:43there's a dumb equation if you want to
- 5:45think about it this way. Um, Jeff
- 5:48Huntley uh did a lot of research on
- 5:49coding agents. Uh, he put it really
- 5:51well. Just the more you use the context
- 5:53window, the worse outcomes you'll get.
- 5:55This leads to a concept I'm in a very
- 5:56very academic concept called the dumb
- 5:58zone. So, you have your context window.
- 6:01You have 168,000 tokens roughly. Some
- 6:03are reserved for output and compaction.
- 6:05This varies by model. Um, but we'll use
- 6:07cloud code as an example here. Around
- 6:09the 40% line is where you're going to
- 6:10start to see some diminishing returns
- 6:12depending on your task. Um, if you have
- 6:15too many MCPs in your coding agent, you
- 6:17are doing all your work in the dumb zone
- 6:18and you're never going to get good
- 6:20results. People talked about this. I'm
- 6:22not going to talk about that one. Your
- 6:23mileage may vary. 40% is like it depends
- 6:24on how complex the task is, but this is
- 6:26kind of a good guideline. Um so back to
- 6:29compaction or as I will call it from now
- 6:31on cleverly avoiding the dumb zone. Um
- 6:35we can do sub agents. Um if you have a
- 6:37front-end sub aent and a backend sub
- 6:39aent and a QA sub aent and a data data
- 6:41scientist sub aent
- 6:43please stop. Sub aents are not for
- 6:46anthropomorphizing roles. They are for
- 6:47controlling context. And so what you can
- 6:49do is if you want to go find how
- 6:51something works in a large codebase um
- 6:53you can steer the coding agent to do
- 6:55this if it supports sub agents or you
- 6:56can build your own sub agent system. But
- 6:58basically you say hey go find how this
- 7:00works and it can fork out a new context
- 7:02window that is going to go do all that
- 7:04reading and searching and finding and
- 7:06reading entire files and understanding
- 7:08the codebase and then just return a
- 7:11really really succinct message back up
- 7:13to the parent agent of just like hey the
- 7:15file you want is here. parent agent can
- 7:18read that one file and get straight to
- 7:20work. And so this is really powerful. If
- 7:22you wield these correctly, you can get
- 7:24good responses like this and then you
- 7:26can manage your context really, really
- 7:27well. Um, what works even better than
- 7:29sub agents or like a layer on top of sub
- 7:31aents is a workflow I call frequent
- 7:33intentional compaction. We're going to
- 7:35talk about research plan implement in a
- 7:37minute, but like the point is you're
- 7:38constantly st keeping your context
- 7:40window small. You're building your
- 7:42entire workflow around context
- 7:44management. So comes in three phases.
- 7:46research, plan, implement. Um, and we're
- 7:49going to try to stay in the smart zone
- 7:50the whole time. So, the research is all
- 7:52about understanding how the system
- 7:53works, finding the right files, staying
- 7:55objective. Here's a prompt you can use
- 7:57to do research. Here's the output of um,
- 7:59a research prompt. These are all open
- 8:01source. You can go grab them and play
- 8:02with them yourself. Um, planning, you're
- 8:05going to outline the exact steps. You're
- 8:06going to include file names and line
- 8:07snippets. You're going to be very
- 8:08explicit about how we're going to test
- 8:09things after every change. Here's a good
- 8:11planning prompt. Here's one of our
- 8:13plans. It's got actual code snippets in
- 8:14it. Um, and then we're gonna implement.
- 8:16And if you read one of these plans, you
- 8:18can see very easily how the dumbest
- 8:20model in the world is probably not going
- 8:21to screw this up. Um, so we just go
- 8:23through and we run the plan and we keep
- 8:25the context low. As a planning prompt,
- 8:27like I said, it's the least exciting
- 8:28part of the process. Um, I wanted to put
- 8:30this into practice. So, working for us,
- 8:32uh, I do a podcast with my buddy uh,
- 8:34Vibv, who's the CEO of a company called
- 8:35Boundary ML. Uh, and I said, "Hey, I'm
- 8:38going to try to oneshot a fix to your
- 8:39300,000line Rust codebase for a
- 8:41programming language."
- 8:43Um, and the whole episode goes in, it's
- 8:45like an hour and a half. Uh, I'm not
- 8:46going to talk through it right now, but
- 8:47we built a bunch of research and then we
- 8:48threw them out because they were bad.
- 8:49And then we made a plan and we made a
- 8:50plan without research and with research
- 8:52and compared all the results. It's a fun
- 8:53time. Uh, by that was Monday night. By
- 8:56Tuesday morning, we were on the show and
- 8:57the CTO had like seen the PR and like
- 9:00didn't realize I was doing it as a bit
- 9:01for a podcast and basically was like,
- 9:03"Yeah, this looks good. We'll get in the
- 9:04next release." He I think he was a
- 9:06little confused. Um, here's the the
- 9:08plan. But anyways, uh, yeah, confirmed
- 9:11works in brownfield code bases and no
- 9:14slop. But I wanted to see if we could
- 9:16solve complex problems. So, Vib was
- 9:18still a little skeptical. I sat down, we
- 9:20sat down for like 7 hours on a Saturday
- 9:21and we shipped 35,000 lines of code to
- 9:24BAML. One of the PRs got merged like a
- 9:26week later. I will say some of this is
- 9:28codegen. You know, you update your
- 9:29behavior. All the golden files update
- 9:31and stuff, but we shipped a lot of code
- 9:32that day. Um, he estimates it was about
- 9:341 to 2 weeks and 7 hours. And uh, so
- 9:37cool. We can solve complex problems.
- 9:40There are limits to this. I sat down
- 9:41with my buddy Blake. We tried to remove
- 9:43Hadoop dependencies from Parket Java. If
- 9:46you know what Paret Java is, I'm sorry
- 9:49uh for whatever happened to you to get
- 9:51you to this point in your career. Uh it
- 9:53did not go well. Uh here's the plans,
- 9:55here's the research. Uh at a certain
- 9:57point, we threw everything out and we
- 9:58actually went back to the whiteboard. We
- 10:00had to actually once we had learned
- 10:01where were the where all the foot guns
- 10:03were, we we went back to okay, how is
- 10:05this actually going to fit together? Um,
- 10:07and this brings me to a really
- 10:08interesting point that Jake's going to
- 10:09talk about later. Uh, do not outsource
- 10:12the thinking. AI cannot replace
- 10:14thinking. It can only amplify the
- 10:16thinking you have done or the lack of
- 10:17thinking you have done. So people ask,
- 10:20so Dex, this is specri development,
- 10:22right? No, specri development is broken.
- 10:28Not the idea, but the phrase. Um,
- 10:32it's not well defined. This is Brietta
- 10:34from Thought Works. Um, and a lot of
- 10:36people just say spec and they mean a
- 10:37more detailed prompt. Does anyone
- 10:40remember this picture? Does anyone know
- 10:41what this is from? All right, that's a
- 10:43deep cut. Uh, there will never be a year
- 10:45of agents because of semantic diffusion.
- 10:47Martin Fowler said this in 2006. We come
- 10:49up with a good term with a good
- 10:51definition and then everybody gets
- 10:53excited and everybody starts meaning it
- 10:55to mean a hundred things to a 100
- 10:56different people and it becomes useless.
- 10:59We had an agent is a person, an agent is
- 11:01a micros service. An agent is a chatbot.
- 11:03An agent is a workflow. And thank you,
- 11:05Simon. We're back to the beginning. An
- 11:07agent is just tools in a loop. Um, this
- 11:10is happening to spec driven dev. I used
- 11:12to have Sean's uh slide in the beginning
- 11:14of this talk, but it caused a bunch of
- 11:15people to focus on the wrong things. His
- 11:17thing of like, forget the code. It's
- 11:18like assembly now and you just focus on
- 11:20the markdown. Very cool idea, but people
- 11:23say Spectrum Dev is writing a better
- 11:24prompt, a product requirements document.
- 11:26Sometimes it's using like verifiable
- 11:28feedback loops and back pressure. Maybe
- 11:30it is treating the code like assembly
- 11:32like Sean taught us. Um, but a lot of
- 11:34people is just using a bunch of markdown
- 11:35files while you're coding. Or my
- 11:37favorite, I just stumbled upon this last
- 11:39week. Uh, a spec is documentation for an
- 11:42open source library. So it's gone. It's
- 11:44as specri dev is overhyped. It's useless
- 11:47now. It's semantically diffused.
- 11:50Um, so I want to talk about like four
- 11:52things that actually work today. The
- 11:54tactical and practical steps that we
- 11:55found working internally and with a
- 11:57bunch of users. Um, we do the research,
- 11:59we figure out how the system works. Um,
- 12:02remember Momento? This is the best the
- 12:04best movie on context engineering, as
- 12:06Peter says it. Guy wakes up, he has no
- 12:08memory. He has to like read his own
- 12:10tattoos to figure out who he is and what
- 12:12he's up to. If you don't onboard your
- 12:14agents, they will make stuff up. And so,
- 12:17if this is your team, this is very
- 12:18simplified for most of you. Most of you
- 12:20have much bigger orgs than this. But
- 12:21let's say you want to do some work over
- 12:22here. Um, one thing you could do is you
- 12:25could put onboarding into every repo.
- 12:27You put a bunch of context. Here's the
- 12:28repo. Here's how it works. This is a
- 12:30compression of all the context in the
- 12:32codebase that the agent can see ahead of
- 12:34time before actually getting to work.
- 12:36This is challenging because
- 12:38sometimes it gets too long. As your
- 12:40codebase gets really big, you either
- 12:41have to make this longer or you have to
- 12:43leave information out. And so as you uh
- 12:47are reading through this, you're going
- 12:49to read the context of this big 5
- 12:50million line monor repo and you're going
- 12:52to use all the smart zone just to learn
- 12:54how it works. And you're not going to be
- 12:55able to do any good tool calling in the
- 12:56dumb zone. So that's uh you can
- 13:01you can shard this down the stack. You
- 13:02can do they're just talking about
- 13:03progressive disclosure. You could split
- 13:04this up, right? You could just put a
- 13:06file in the root of every repo and then
- 13:08like at every level you have like
- 13:10additional context based on if you're
- 13:12working here, this is what you need to
- 13:14know. Uh we don't document the files
- 13:16themselves cuz they're the source of
- 13:17truth. But then as your agent is
- 13:19working, you know, you pull in the root
- 13:20context and then you pull in the
- 13:21subcontext. We won't talk about any
- 13:22specific like you could use cloudd for
- 13:24this, you can use hooks for this,
- 13:25whatever it is. Um, but then you still
- 13:27have plenty of room in the smart zone
- 13:28because you're only pulling in what you
- 13:29need to know. Um, the problem with this
- 13:32is that it gets out of date. And so
- 13:34every time you ship a new feature, you
- 13:36need to kind of like cache and validate
- 13:38and rebuild large parts of this internal
- 13:40documentation. And you could use a lot
- 13:43of AI and make it part of your process
- 13:44to update this. Um, but I want to ask a
- 13:48question between the actual code, the
- 13:49function names, the comments, and the
- 13:50documentation. Does anyone want to guess
- 13:52what is on the y-axis of this chart?
- 13:57slop
- 13:57>> slop. It's actually the amount of lies
- 14:00you can find in any one part of your
- 14:02codebase.
- 14:04Um, so you could make a part of your
- 14:05process to update this, but you probably
- 14:07shouldn't cuz you probably won't. What
- 14:08we prefer is on demand compressed
- 14:10context. So if I'm building a feature
- 14:12that relates to SCM providers and Jira
- 14:14and Linear, um, I would just give it a
- 14:16little bit of steering. I would say,
- 14:17hey, we're going over in like this like
- 14:19part of the codebase over here. Um, and
- 14:22a good research uh prompt or or slash
- 14:24command might take you or skill even uh
- 14:27launch a bunch of sub aents to take
- 14:29these vertical slices through the
- 14:30codebase and then build up a research
- 14:32document that is just a snapshot of the
- 14:34actually true based on the code itself
- 14:37parts of the codebase that matter. We
- 14:39are compressing truth. Um, planning is
- 14:42leverage. Planning is about compression
- 14:44of intent. Um, and in plan we're going
- 14:46to outline the exact steps. We take our
- 14:49research and our PRD or our bug ticket
- 14:50or our whatever it is and we create a
- 14:52plan and we create a plan file. So we're
- 14:54compacting again. And I want to pause to
- 14:56talk about mental alignment. Um does
- 14:58anyone know what code review is for?
- 15:03>> Mental alignment. Mental alignment is it
- 15:06is about finding making sure things are
- 15:07correct and stuff but the most important
- 15:09thing is how do we keep everybody on the
- 15:10team on the same page about how the
- 15:12codebase is changing and why. And I can
- 15:14read a thousand lines of Golang every
- 15:16week. Uh sorry I can't read a thousand.
- 15:18It's hard. I can do it. I don't want to.
- 15:20Um, and as our team grows, I all the
- 15:22code gets reviewed. We don't not read
- 15:24the code, but I, as you know, a
- 15:25technical leader in the in on the team,
- 15:27I can read the plans and I can keep up
- 15:29to date and I can that's enough. I can
- 15:31catch some problems early and I maintain
- 15:33understanding of how the system is
- 15:34evolving. Um, Mitchell had this really
- 15:36good post about how he's been putting
- 15:37his AMP threads on his pull requests so
- 15:39that you can see not just, hey, here's a
- 15:41wall of green text in GitHub, but here's
- 15:43the exact steps, here's the prompts, and
- 15:44hey, I ran the build at the end and it
- 15:46passed. This takes the reviewer on a
- 15:48journey in a way that a GitHub PR just
- 15:50can't. And as you're shipping more and
- 15:52more in two to three times as much code,
- 15:54it's really on you to find ways to keep
- 15:57your team on the same page and show them
- 15:59here's the steps I did and here's how we
- 16:00tested it manually. Um, your goal is
- 16:03leverage. So you want high confidence
- 16:04that the model will actually do the
- 16:05right thing. I can't read this plan and
- 16:07know what actually is going to happen
- 16:09and what code changes are going to
- 16:10happen. So we've over time iterated
- 16:12towards our plans include actual code
- 16:15snippets of what's going to change. So
- 16:17your goal is leverage. You want
- 16:18compression of intent and you want
- 16:20reliable execution. Um and so I don't
- 16:22know I have a physics background. We
- 16:23like to draw lines through the center of
- 16:26peaks and curves. Uh as your plans get
- 16:28longer, reliability goes up, readability
- 16:30goes down. There's a sweet spot for you
- 16:32and your team and your codebase. you
- 16:34should try to find it because when we
- 16:35review the research and the plans, if
- 16:37they're good, then we can get mental
- 16:39alignment. Um, don't outsource the
- 16:42thinking. I've said this before, this is
- 16:44not magic. There is no perfect prompt.
- 16:46You still will not work if you do not
- 16:49read the plan. So, we built our entire
- 16:51process around you, the builder, are in
- 16:53back and forth with the agent reading
- 16:55the plans as they're created. And then
- 16:57if you need peer review, you can send it
- 16:58to someone and say, "Hey, does this plan
- 16:59look right? Is this the right approach?
- 17:00Is this the right order to look at these
- 17:02things?" Um Jake again wrote a really
- 17:04good blog post about like the thing that
- 17:05makes research plan implementing
- 17:06valuable is you the human in the loop
- 17:09making sure it's correct. So if you take
- 17:12one thing away from this talk it should
- 17:14be that a bad line of code is a bad line
- 17:16of code and a bad part of a plan is
- 17:20could be a hundred bad lines of code and
- 17:22a bad line of research like a
- 17:24misunderstanding of how the system works
- 17:26and where things are your whole thing is
- 17:28going to be hosed. You're going to be
- 17:29telling sending the model off in the
- 17:30wrong direction. And so when we're
- 17:32working internally and with users, we're
- 17:34constantly trying to move human effort
- 17:36and focus to the highest leverage parts
- 17:38of this pipeline. Um, don't outsource
- 17:40the thinking. Watch out for tools that
- 17:42just spew out a bunch of markdown files
- 17:44just to make you feel good. I'm not
- 17:46going to name names here. Uh, sometimes
- 17:48this is overkill. And the way I like to
- 17:49think about this is like, yeah, you
- 17:51don't always need a full research plan
- 17:53implement. Sometimes you need more,
- 17:55sometimes you need less. If you're
- 17:56changing the color of a button, just
- 17:58talk to the agent and tell it what to
- 17:59do. Um, if you're doing like a simple
- 18:02plan and as a small feature, if you're
- 18:04doing medium features across multiple
- 18:06repos, then do one research, then build
- 18:08a plan. Basically, the hardest problem
- 18:10you can solve, the ceiling goes up the
- 18:12more of this context engineering
- 18:14compaction you're willing to do. Um, and
- 18:16so if you're in the top right corner,
- 18:18you're probably going to have to do
- 18:19more. A lot of people ask me, "How do I
- 18:21know how much context engineering to
- 18:22use?" It takes reps. You will get it
- 18:25wrong. You have to get it wrong over and
- 18:26over and over again. Sometimes you'll go
- 18:27too big. Sometimes you go too small.
- 18:29Pick one tool and get some reps. I
- 18:32recommend against minmaxing across cloud
- 18:34and codeex and all these different
- 18:35tools. Um, so I'm not a big acronym guy.
- 18:39Uh, we said specri dev was broken. Uh,
- 18:42research plan and implement I don't
- 18:44think will be the steps. The important
- 18:45part is compaction and context
- 18:46engineering and staying in the smart
- 18:47zone. But people are calling this RPI
- 18:50and there's nothing I can do about it.
- 18:52So, uh, just be wary. There is no
- 18:54perfect prompt. There is no silver
- 18:55bullet. Um, if you really want a hypy
- 18:58word, you can call this harness harness
- 19:00engineering, which is part of context
- 19:01engineering, and it's how you integrate
- 19:02with the integration points on codeex,
- 19:04claude, cursor, whatever. How you
- 19:06customize your codebase. Um, so what's
- 19:09next? I think the coding agent stuff is
- 19:12actually going to be commoditized.
- 19:13People are going to learn how to do this
- 19:14and get better at it. And the hard part
- 19:16is going to be how do you adapt your
- 19:17team and your workflow and the SDLC to
- 19:20work in a world where 99% of your code
- 19:22is shipped by AI. Uh, and if you can't
- 19:25figure this out, you're hosed because
- 19:26there's kind of a rift growing where
- 19:28like staff engineers don't adopt AI
- 19:29because it doesn't make them that much
- 19:30faster. And then junior mid-levels
- 19:32engineers use a lot because it fills in
- 19:34skill gaps and then it also produces
- 19:36some slop. And then the senior engineers
- 19:37hate it more and more every week because
- 19:39they're cleaning up slop that was
- 19:40shipped by cursor the week before. Uh,
- 19:43this is not AI's fault. This is not the
- 19:44mid-level engineers fault. Like if
- 19:46cultural change is really hard and it
- 19:48needs to come from the top if it's going
- 19:49to work. So if you're a technical leader
- 19:51at your company, pick one tool and get
- 19:53some reps. If you want to help, we are
- 19:56hiring. We're building an Aentic IDE to
- 19:58help teams of all sizes speedrun the
- 20:00journey to 99% AI generated code. Uh if
- 20:04we'd love to we'd love to talk if you
- 20:05want to work with us. Uh go go hit our
- 20:07website, send us an email, come find me
- 20:08in the hallway. Uh thank you all so much
- 20:10for your energy.
- 20:14[music]
- 20:16Heat.
- 20:30[music]
About this transcript
This page contains the full transcript of No Vibes Allowed: Solving Hard Problems in Complex Codebases – Dex Horthy, HumanLayer by AI Engineer, generated from the public captions YouTube serves with the video. The transcript has 4,538 words across 624 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.