Birgitta Böckeler - State of Play: AI Coding Assistants - AI Native DevCon June 2026 — Transcript
Full transcript
- 0:00Birgitta Buckler is a is a distinguished
- 0:02engineer at ThoughtWorks and um
- 0:05I've been chatting to Birgitta for a
- 0:07little while and actually uh was really
- 0:09impressed by a very a post that went
- 0:12viral on Martin Fowler's uh site that
- 0:15Birgitta wrote uh talking about specs
- 0:18and talking about uh the the comparison
- 0:21at the time between Spec Kit uh Tessel
- 0:24and uh Kiro, I think it was. And uh
- 0:27okay, Tessel's moved on a little bit
- 0:29since then, but it's still amazing to
- 0:30see how many people go to go to that uh
- 0:33site. Uh Birgitta's an amazing person. I
- 0:35very much encourage you to to follow
- 0:36her. She's she's very thoughtful in with
- 0:38so many uh posts and blogs that she
- 0:40writes. And
- 0:42a a really wonderful way to finish off
- 0:44uh this conference with a very visionary
- 0:47uh session whereby we Birgitta's going
- 0:48to look at the last 12 months where
- 0:50what's been changing, where we are
- 0:52today, and what we can look forward to.
- 0:53So, uh please give a very a a Native Dev
- 0:56warm welcome to Birgitta Buckler.
- 0:59>> [applause]
- 1:06>> Okay, yeah, thanks Simon. Thanks Simon
- 1:08and Patrick for inviting me. Yeah, so
- 1:10I'm a distinguished engineer at
- 1:12ThoughtWorks and what that means for me
- 1:13specifically is that uh 3 years ago I
- 1:16got a full-time role to just be immersed
- 1:19in the space of AI coding or in general
- 1:21using AI on software teams to help my
- 1:23colleagues, to help our clients like I
- 1:25don't stay on top of it. So, I talk a
- 1:27lot to
- 1:28uh our teams, our clients, and so on.
- 1:30And then I write about it, for example,
- 1:32on my colleague Martin Fowler's website.
- 1:34Um and so that's kind of like what what
- 1:36all of this is um based on. And yeah,
- 1:39it's kind of like a tough task of
- 1:41wrapping up after everybody heard about
- 1:43this topic for 2 days. Uh so, I'm going
- 1:45to try and help you see the forest for
- 1:48the trees or the multiple forests for
- 1:49all of the the trees. So, I'll kind of
- 1:52do like a recap. I'll start with the
- 1:54recap slide. Simon was briefly confused
- 1:56and thought maybe my slide setup was
- 1:58wrong, but I'll it's kind of like recap
- 2:00a lot of what the stuff that you heard
- 2:01also over the last two days, but also
- 2:04kind of what happened in the last 12
- 2:05months like advancements as well as
- 2:07things that are maybe not going so well
- 2:09or that are kind of like the all the
- 2:11second order consequences and
- 2:13implications that we're experiencing
- 2:14right now. So that when you get back to
- 2:16work tomorrow and your colleague who
- 2:19maybe isn't as immersed in the space ask
- 2:21you, "So,
- 2:22what should I know, right?" I hope I can
- 2:24help you answer that question.
- 2:26And I'll start with yeah, the reason why
- 2:29all of this is happening which is
- 2:31the model so kind of first and most
- 2:33obvious.
- 2:34There wasn't even that much talk about
- 2:36models here at the conference, I think
- 2:38which is not surprising and totally
- 2:40okay, I think because for me this is not
- 2:42really the most exciting part to be
- 2:43honest. I'm much more interested in
- 2:45everything that's now happening around
- 2:47it, the ecosystem, all of the
- 2:48integrations, and so on. And I mean
- 2:51obviously if we talk about the last 12
- 2:52months, there was the Opus 4.5 moment
- 2:55kind of last year that made a lot of
- 2:59people kind of like come back to this
- 3:00that hadn't maybe tried AI coding for a
- 3:02while.
- 3:04Um
- 3:05and
- 3:06yeah, so that was maybe the biggest
- 3:08event in models that happened like I
- 3:09mean almost every week there's a new
- 3:11model, but I usually don't even follow
- 3:13it that much because like I said it's
- 3:15kind of interesting to to see all the
- 3:17stuff around it. So if we think about
- 3:20models like what are the core things
- 3:22kind of as users who use them for coding
- 3:25that we need to know or learn almost
- 3:27like a kind of like learning map, right?
- 3:29So the first thing I always try to get
- 3:31out of the way is that they are not
- 3:33magic, right? They are very, very
- 3:35impressive and very, very useful math,
- 3:37but unfortunately even like a lot of
- 3:40technologists, a lot of our peers kind
- 3:41of like it's very easy to fall into that
- 3:43trap, right? Like of
- 3:45thinking of them as as more than that,
- 3:47right? So I kind of like this
- 3:48visualization to remind ourselves that,
- 3:51you know, even though we don't really
- 3:52know what's like why it works and what's
- 3:54happening, it's still like very
- 3:56impressive math, right? So, first of
- 3:58all, they're not magic. Um I mention
- 4:01here as the second point that their
- 4:02statelessness, because that's also
- 4:04something that I often notice that
- 4:06people haven't quite grasped, right? So,
- 4:08the the model doesn't have a session,
- 4:10right? So, the longer our conversation
- 4:12with them gets, the longer our session
- 4:13with them gets, every single time the
- 4:16our our agent, our harness basically
- 4:18sends the whole history of the
- 4:20conversation, right? Maybe not quite,
- 4:22there's caching and all kinds of clever
- 4:23ways that uh different tools try to
- 4:26optimize that, but they are stateless,
- 4:28right? So, that is a a factor that
- 4:30happens like the longer our
- 4:32conversations get with them. So, I think
- 4:33that's a really important core thing
- 4:35that um
- 4:37uh yeah, people need to understand.
- 4:41Uh the third thing we need to know about
- 4:43is, of course, the size of the context
- 4:46window and um in relationship, but also
- 4:48in relationship to what that means for
- 4:49attention, right? So, even though
- 4:51technically the context windows have
- 4:52gotten a lot bigger, um it comes with a
- 4:55trade-off on like how well the models
- 4:57are able to keep attention on all of the
- 4:59many instructions and all of the context
- 5:01that we're trying to feed them now. So,
- 5:03there's something there to be understood
- 5:05by everybody who uses this about what
- 5:07that trade-off is.
- 5:08Um and then um finally this uh this and
- 5:12that's maybe the the biggest area. So,
- 5:15those first things are kind of like you
- 5:16could you can learn them in a formal
- 5:17training and kind of like understand the
- 5:19basics, right? But this last one is a
- 5:21lot more about using the models and
- 5:23figuring this out, right? Which model do
- 5:25we use for which task, right? So,
- 5:28there's I mean, these are just like a
- 5:29few illustrative examples, like uh we we
- 5:32have autocomplete, like there's still
- 5:34people who use that a lot, uh let's say,
- 5:37or let's say you want to just change a
- 5:38few specific files and you have very
- 5:40clear instructions and a very clear idea
- 5:41of what you want to do, or you have a
- 5:43larger and more complex change that
- 5:45needs a bunch of code research before
- 5:47the model actually does it or the agent
- 5:49actually does that. Or you have like
- 5:51tasks like planning, debugging,
- 5:53designing that maybe have a lot more
- 5:55like things like asking you questions
- 5:57and a lot more reasoning involved. So,
- 5:59um different tasks will have different
- 6:02levels of reasoning that are useful, uh
- 6:05will need different uh sizes of context,
- 6:08and will have different levels of need
- 6:10for tool calling. And I actually think
- 6:12that in this area here, we don't even
- 6:15have to know that much about the details
- 6:17of all of the different features of the
- 6:19models, but it's a lot more about us
- 6:21reflecting on these types of tasks,
- 6:23right? Which is something that we've
- 6:24done in the profession for a for a long
- 6:26time, right? Like all of those like
- 6:28philosophical discussions about what
- 6:29does complexity mean, right? Like that
- 6:31we have in in estimations or stuff like
- 6:34that, right? So, um you have to think
- 6:36about like how many files might this
- 6:37involve, what's the blast radius, how
- 6:40much uncertainty do I still have, all of
- 6:42those things. So, I think yeah, like I
- 6:44said, I think it's um much more
- 6:46important for us to to work on this like
- 6:48reflection on the types of tasks that we
- 6:50have and then match them to different uh
- 6:53levels of power that the the models
- 6:55have.
- 6:57And you if you've been bathing in
- 6:59powerful models for the last few months,
- 7:01then it's actually a nice exercise to
- 7:03remember what we're already taking for
- 7:05granted these days when we try to run a
- 7:07smaller model on our developer laptop,
- 7:10right? And try to see what it can do and
- 7:12what it cannot do. And
- 7:16so here's like an example. This is just
- 7:18to show you kind of the speed, right?
- 7:20The speed has actually come a really
- 7:21long way. So, this is Gwen 3.6 running
- 7:24on my Apple M3 with 48 GB of RAM. And
- 7:28it's actually not even that much slower
- 7:30than like some of the stuff that happens
- 7:31in Claude code uh with with Sonnet or
- 7:34Opus, right? This is uh this is open
- 7:36code, um by the way. So, the speed has
- 7:39actually come quite a bit of a way. Tool
- 7:41calling is still kind of even with this
- 7:43like quite powerful model that I could
- 7:46hardly even run on a 64 GB RAM MacBook.
- 7:50So, it it's it crashed after my my first
- 7:52attempts to use it. So, even that model
- 7:54was still struggling a bit with tool
- 7:56calling, which is really crucial for our
- 7:57agentic workflows, right?
- 8:00But, it also has gotten a lot better
- 8:02than it than it was like, I don't know,
- 8:046 months ago or so. Um and complexity of
- 8:07all the instructions, I mean, uh you
- 8:09know, many of you might have seen this
- 8:11when you when you try smaller models or
- 8:14I remember when I first used Gemini for
- 8:16for coding like a year ago or so, I kept
- 8:18I kept getting these as well. So, um
- 8:21that is still happening, but at the same
- 8:23time, like I was for example recently
- 8:25with uh Gemma 4, which is like a model
- 8:27that everybody's talking about right now
- 8:29in terms of like a smaller model that is
- 8:31quite capable at coding comparatively,
- 8:34you know, to uh to other small models.
- 8:36And so, I used this recently. Again,
- 8:38this was the type of task where I knew
- 8:39exactly what needed to be done, but I
- 8:41couldn't be bothered to type it out
- 8:43myself, right? So, I gave it like one
- 8:46small paragraph of instructions, and it
- 8:47was actually really good at doing that.
- 8:49I went after that back and forth a
- 8:50little bit, had it refactor it a little
- 8:52bit because it was quite uh
- 8:54cyclomatically complex, and then I had
- 8:56my little utility script. So, for this,
- 8:58I could actually use it on my [snorts]
- 9:01uh on my M3, right? So again, it's a lot
- 9:03about like knowing the tasks and then
- 9:04kind of uh mapping that to the level of
- 9:07power you want to use.
- 9:09So, moving on from the model, then the
- 9:11next thing that we have around the model
- 9:13is what we what uh we've now kind of
- 9:15come to calling the coding harness,
- 9:17right? Or
- 9:18uh you know, also sometimes more
- 9:20colloquially still like the coding
- 9:21agent, right? So, um that's the thing
- 9:24that's kind of helping us leverage the
- 9:26model for our coding tasks, and it has
- 9:29things under the hood like a system
- 9:31prompt or all kinds of like like other
- 9:32prompts that we usually can't even see
- 9:34when unless it's open source, right? I
- 9:36mean, we recently got a glimpse into the
- 9:38cloud code once, but
- 9:40a lot of them are a lot of the big ones
- 9:42that we actually use a lot are closed
- 9:44source, right?
- 9:46It comes according to Harness, it comes
- 9:48with the tool integrations, all of the
- 9:49standard stuff that we'll definitely
- 9:51need, right?
- 9:52Changing files, reading files, code
- 9:54search is a big one, right? So, we get
- 9:56like a type of code search out of the
- 9:58box which each with each Harness that we
- 10:00pick. It has like all kinds of
- 10:03orchestration, like most of them has
- 10:05have sub agents now, for example, and
- 10:07also decide when to spawn off certain
- 10:09sub agents or they kind of decide like
- 10:13how many tool calls at once they pass on
- 10:15to the model and all of those types of
- 10:17things. Maybe there's some caching
- 10:19involved in some of them. They have a
- 10:22user interface, of course, right? Like
- 10:24some of them have a terminal-based user
- 10:26interface, some of them are like in VS
- 10:28Code or or other more graphical user
- 10:31interfaces. And they also all come with
- 10:34different levels of extensibility and
- 10:35observability. So, extensibility,
- 10:38famously, the pie coding agent is very
- 10:40popular right now as one that has kind
- 10:42of brought that more to our attention of
- 10:44having a coding agent that we can
- 10:45actually like that is malleable, that we
- 10:47can change, right? And observability is
- 10:50also a space where there's a lot
- 10:52happening right now in terms of uh you
- 10:55know, having traces of what the agent is
- 10:57doing,
- 10:59how can we use that to, for example,
- 11:01analyze more like how we can improve
- 11:03how we're using the agent
- 11:06getting some visibility into, for
- 11:09example, I've done some stuff about like
- 11:10visualizing for myself during session
- 11:12which files is it is it reading and
- 11:14which files is it writing, so I could
- 11:16get an idea of like blast radius of
- 11:18stuff. So, I think there's still a lot
- 11:20of potential here for us to also make
- 11:22these things part of the the review
- 11:24cycle.
- 11:26They're kind of like first rumblings
- 11:28about the bloat, right? We always have
- 11:30the cycle in software, right? We we get
- 11:32like a new tool, like let's say spring,
- 11:34right? And it's like super lightweight
- 11:37and we're excited because it's much more
- 11:38lightweight than the bloated thing that
- 11:40we had before, and then it takes like a
- 11:42number of years, right? And then we kind
- 11:44of feel like that's also too big, right?
- 11:46With Claude code, it hasn't even taken a
- 11:48year and we feel like it's maybe a bit
- 11:50much, right?
- 11:51>> [laughter]
- 11:52>> Um so here's like a comparison of like
- 11:54when you start out with a session in pie
- 11:56versus open AI codex versus Claude code,
- 11:58all of the stuff that's already in
- 12:00there, right? Um [snorts]
- 12:03but yeah, so we need to as when we come
- 12:05back to like what what do we need to
- 12:06know, right? What do we need to
- 12:08understand as as engineers using these
- 12:10tools, we need to kind of like
- 12:11understand these features and how they
- 12:13distinguish between the agents, right?
- 12:16So just to give like an example, um
- 12:18I remember when Claude code came out and
- 12:20it became super popular, um I saw a lot
- 12:24of conflation of the interface with the
- 12:27reason why Claude code was good, right?
- 12:30So I saw that for a while that people
- 12:31said, "Oh yeah, terminal based, that's
- 12:33the way to go because it's really
- 12:34powerful." But actually when you have
- 12:36for example one of my favorite harnesses
- 12:38next to Claude code is cursor is also
- 12:40really really good under the hood with
- 12:42all of those things that you see in the
- 12:43circle below there. So it's not
- 12:45necessarily because of the terminal,
- 12:46right? So this is just like an example
- 12:48of why it's uh it's important to kind of
- 12:51distinguish um
- 12:54you know, kind of know as the engineer
- 12:56what what's actually happening in the
- 12:58tools.
- 13:01So yeah, we have to understand their
- 13:02footprint, understand their features so
- 13:04that we can use these features
- 13:06effectively, right? And so most like one
- 13:09of the big things that we want to do is
- 13:11we want to understand how to regulate
- 13:13the context, how to tune the context
- 13:14that we give to the agent with the use
- 13:16of the features that the coding harness
- 13:18provides us, right?
- 13:20So about 12 months ago, the main way
- 13:22that we did that was rules files or
- 13:24instruction files, right? So we would
- 13:26have agents MD or Claude MD and maybe
- 13:29write down the typical pitfalls. And
- 13:31we're still doing that, but these days,
- 13:33of course, there's so many different
- 13:35like features in the coding harnesses
- 13:37that help us do that in a more
- 13:38sophisticated way, right? So, there's
- 13:40skills, of course.
- 13:42MCP servers also existed about a year
- 13:45ago as well. There's now sub agents,
- 13:47there's extensions, plugins, hooks, all
- 13:49of those things. And it's becoming a bit
- 13:51overwhelming and confusing, right? But
- 13:54it's kind of like the storming phase of
- 13:57this of this technology. And this is
- 13:59also how to use these features is, I
- 14:02would say, was probably at least like
- 14:0350% of the talks at this at this
- 14:06conference as well, unsurprisingly.
- 14:08So, it's kind of context engineering for
- 14:11coding agents. And that means we're
- 14:12expanding the harness. We're
- 14:14using the features of the harness to to
- 14:17help us do this for our specific code
- 14:19base, for our
- 14:21use case, right?
- 14:24Um
- 14:26Yeah. So, we're expanding that harness.
- 14:29Um so, I'm I'm calling it coder harness
- 14:31here. I have to say like I'm a little
- 14:33bit unsure. So, this is term that
- 14:36that has now gotten traction since
- 14:38February, maybe, of harness engineering,
- 14:40right? Which is basically this, right?
- 14:43Expanding this coding harness. But in in
- 14:45other areas,
- 14:47you know, people also use harness
- 14:48engineering to talk about like how do
- 14:50you make the coding harness itself
- 14:51better, right? So, it's still like a
- 14:52little bit clunky term, I would say. I
- 14:54wish we had we had a better one. I also
- 14:56jumped on the bandwagon and wrote an
- 14:58article about harness engineering.
- 15:00Um yeah, it would be great. May- maybe
- 15:02somebody comes up with an even better
- 15:04word. Like the folks at Tesla here at
- 15:06the conference, they've basically
- 15:07they've also had this kind of like
- 15:09trifecta that I'm presenting here,
- 15:10right? Like of the model, the harness,
- 15:12and then the context, right? So, I think
- 15:14it's like it's reasonable to think of
- 15:17harness engineering as context
- 15:19engineering for coding agents, right?
- 15:21So, that's just like to get the
- 15:23terminology out of the way a little bit.
- 15:26Um
- 15:28Yeah. So let's look at like uh how this
- 15:31like area of harness engineering,
- 15:32context engineering for coding agents,
- 15:34like um a mental model, like how I think
- 15:37about this, like beyond the features,
- 15:39right? Yes, this is about skills, this
- 15:41is about MCP servers, but conceptually,
- 15:43like what are we actually doing there
- 15:44when we use these features?
- 15:46So one thing, and that's the most common
- 15:49thing right now, is this uh
- 15:51way of putting conventions, product
- 15:53context, workflow, the prompts, like
- 15:56basically markdown files into our
- 15:59uh code base somehow, right? Or like
- 16:00making them accessible through skills
- 16:02and so on. And in those markdown files
- 16:04is actually lots of different things
- 16:05going on, right? We have some normative
- 16:07stuff, like coding conventions, we have
- 16:09some informative stuff, like yeah,
- 16:12product context, what are we actually
- 16:13doing here? Maybe reference
- 16:15documentation,
- 16:17uh and then we also have instructions,
- 16:18right? Like uh always help me build in
- 16:21the following workflow, or always write
- 16:23a failing test first, stuff like that.
- 16:25So there's actually lots of different
- 16:26things going on in something that looks
- 16:28just like a bunch of uh text at first.
- 16:31And then also some of them we just have
- 16:32directly in the workspace, and others
- 16:34are maybe more dynamically loaded from
- 16:37other data sources, right? And um
- 16:40so these are all for me like kind of
- 16:41feed forward, so we're trying to
- 16:43anticipate what the agent might do
- 16:45wrong, and we're also trying to
- 16:47anticipate, of course, what we want it
- 16:49to do. And we're feeding it all of this
- 16:51information, these instructions, these
- 16:52norms, so that hopefully in its initial
- 16:54generation of code it's already doing
- 16:56perfectly, right? But as we know, that's
- 16:58not um always happening. So we're
- 17:01starting with these guides, but then we
- 17:03also want to give it feedback, right? So
- 17:05um
- 17:06ideally, uh so that we can trigger
- 17:08immediately a self-correction loop
- 17:10before we even look at the code, so that
- 17:12we don't have to like have all those
- 17:14low-hanging fruits still in there. So um
- 17:16the most common way that people do that
- 17:19right now is with like code review
- 17:20agents, right? But, there's also all of
- 17:22these other tools that we have in our
- 17:24toolbox from before AI, like static code
- 17:27analysis. And then, of course, we can
- 17:29also an agent usually has access to the
- 17:31logs, so it can start the application,
- 17:33see what what logs come out of it. Uh,
- 17:35many people give an agent access to the
- 17:37browser, so it can look at like
- 17:39something when it has changed a web
- 17:41component or something like that.
- 17:43Um,
- 17:44and there's actually, um,
- 17:46uh, a difference kind of between these.
- 17:48So, like a review agent is an LLM
- 17:50judging the work of another LLM, right?
- 17:52So, it's kind of inferential. It's
- 17:54running on the GPU. But, we have a bunch
- 17:56of tools as well that are, uh,
- 17:58computational, as I decided to call them
- 18:01here. So, kind of things that run on the
- 18:02CPU, right? Like, the static code
- 18:04analysis is the best example, I think,
- 18:06to to think about this.
- 18:08Um,
- 18:09yeah, and we have the same distinction
- 18:11on the feed forward on the guide side.
- 18:12So, we we can, uh, we
- 18:15um, can also think about computational
- 18:16guides on that side. And the best
- 18:18example, uh, for me there is code mods.
- 18:21Uh, Ian from Meta also just mentioned
- 18:23those. Um,
- 18:24which is, for example, tools like open
- 18:26rewrite that are really good at doing,
- 18:29uh, version upgrades and my migrations
- 18:31of, uh,
- 18:32of, um,
- 18:33frameworks. I don't know if you remember
- 18:35like quite a while ago Amazon had a
- 18:37really big headline about saving 400 or
- 18:40500 developer years or something for
- 18:42Java upgrades. That was under the hood,
- 18:45actually, mostly code mods being made
- 18:47available to AI. So, that combination is
- 18:49really powerful, right? So, all of these
- 18:51things, uh, or maybe providing a
- 18:53different type of code search that that
- 18:55is more effective for your really large
- 18:56code base. All of those are ways, again,
- 18:59to increase the probability that AI does
- 19:01what you want in the first go.
- 19:04So, that's then the expanded harness.
- 19:06And then, as a human, as we've heard in
- 19:08a few talks, uh, here as well, as a
- 19:10human our job, in part becomes kind of
- 19:13steering this set of guides and sensors.
- 19:17So as Mitchell Hashimoto says in his
- 19:19blog post, it's the idea that anytime
- 19:21you find an agent makes a mistake, you
- 19:23take the time to engineer a solution
- 19:25such that the agent never makes that
- 19:27mistake again.
- 19:30And of course, AI can help us engineer
- 19:33those little solutions, right? Which is
- 19:34really useful.
- 19:36So just two quick examples to to bring
- 19:39that home. So um
- 19:41here's like something in the agents MD,
- 19:43do not use console log, you know, we
- 19:45have a custom logger that does
- 19:47structured logging and so on. It's in
- 19:48the following place, right? Instead, I
- 19:50could have like a linting rule um that I
- 19:53customize where I customize the message
- 19:56and point uh the agent at that log file
- 20:00through that, right? So especially when
- 20:02you have things that it doesn't even do
- 20:03wrong that often, it happens maybe once
- 20:05a week, it's much more effective to have
- 20:07this linting rule than to like stuff it
- 20:09into your context every single time the
- 20:11agent runs. Or here's an another
- 20:13example, let's say you have a back-end
- 20:14coding conventions a skill that talks
- 20:17about the back-end layers that the agent
- 20:19should respect uh in terms of like which
- 20:22uh modules are allowed to call which
- 20:24other modules. There are some tools in
- 20:26most language ecosystems that help you
- 20:28scan the different imports between
- 20:30files. And so again, you can like come
- 20:32up together with AI with some rules in
- 20:34those tools that help you uh already
- 20:37catch the low-hanging fruit of those
- 20:39modularity um
- 20:41uh violations.
- 20:44And then you should think about like how
- 20:46you where you put those sensors, when
- 20:48you run them, right? So uh kind of like
- 20:50strategically think about your path to
- 20:52production and think about when you want
- 20:54to run them. So do you want to
- 20:56run them in the coding session, right?
- 20:58Which is I think uh
- 21:00whenever that's possible in terms of
- 21:02like how cheap is it if how fast it is
- 21:03to run a sensor, I think you should run
- 21:06them like even before you commit, right?
- 21:09So, I have this box here about
- 21:10integration, right? So, it kind of like
- 21:12depends what that means for you. So,
- 21:14probably 80% of my commits in the last
- 21:1615 years have been put straight onto the
- 21:18main branch, which is probably not the
- 21:20case for most of you.
- 21:22Um so, integration could either be like
- 21:24for you to say, "Okay, I want to do all
- 21:26of those things before I even create a
- 21:27commit." Or it could be as part of the
- 21:29pull request uh process where you run
- 21:32some additional like uh inferential
- 21:34sensors or something like that, right?
- 21:36Then we have lots of stuff in our
- 21:38continuous integration pipeline already,
- 21:39right? You probably don't want any
- 21:41inferential sensors in there because you
- 21:44you don't want, you know, the the green
- 21:46or red state of your pipeline to depend
- 21:48on semantic interpretation of an LLM,
- 21:51right? But we have like lots of uh
- 21:53computational sensors in there. And then
- 21:56also what uh I've heard a lot of stories
- 21:58now about teams doing in ThoughtWorks
- 22:00and also lots of people writing about
- 22:02that. Also the um
- 22:04the team Ryan's team who had a um
- 22:07a presentation yesterday at OpenAI, they
- 22:09call it garbage collection, right? So,
- 22:11kind of like continuous drift detection
- 22:13for the technical debt that still
- 22:15accumulates. Um where you can probably
- 22:18put like a lot of inferential sensors,
- 22:20right? So, in the code base where I'm uh
- 22:21where I'm setting up all of these
- 22:23things, I have like a modularity review,
- 22:25like a dependency uh freshness review,
- 22:29uh security review that doesn't run
- 22:30every single time the pipeline runs, but
- 22:33I trigger it maybe like once a week just
- 22:34to see if there's something new that
- 22:36came up.
- 22:37Um and these like continuous drift
- 22:39detection things, they're probably also
- 22:41like require some kind of process on
- 22:42your team to deal with them, right?
- 22:44There's a lot of parallels to, for
- 22:46example, security vulnerabilities and
- 22:48how you deal with them, right? Like uh
- 22:50you know, they keep popping up again. Do
- 22:52you want to suppress them because you
- 22:53cannot fix them right now, but then you
- 22:55might forget about them. So, I suspect
- 22:58we'll have all of these challenges there
- 22:59with this type of stuff as well. Like
- 23:01how is the team going to going to deal
- 23:03with these? Again, you can maybe like
- 23:05have AIs, of course, I create
- 23:07you can have agents create pull requests
- 23:09for you, but still
- 23:11how do you deal with those, right?
- 23:14And then finally, I should also mention
- 23:16this of course also this way of having
- 23:18sensors like in production that you give
- 23:20AI access to, right? Especially when it
- 23:23comes to your architecture fitness,
- 23:24things like scalability, latency, all of
- 23:27those things. There were a few talks
- 23:28here at the conference as well about
- 23:30using observability data and
- 23:33you know, both to help you fix
- 23:34incidents, but also just to like monitor
- 23:37how you can make your
- 23:38your runtime better.
- 23:44Okay, so as users of coding agents, we
- 23:47have to know a few key model
- 23:49capabilities and like I said kind of
- 23:50like have good reflection on the types
- 23:52of tasks that we're that we're using and
- 23:55what complexity means for us in
- 23:56relationship to what models can do.
- 23:59We should understand the key features of
- 24:01harnesses and and how they differ and
- 24:03not just like look at them as like
- 24:05you know, these blobs and one is a
- 24:07terminal and one is an IDE.
- 24:09But the biggest area is really this like
- 24:11how do we use them to our advantage,
- 24:14right?
- 24:15Yeah, how to apply this tool to our
- 24:18domain, which is software engineering.
- 24:21So with all of these things like models
- 24:22getting much more capable, harnesses
- 24:24getting much more capable, we're kind of
- 24:26like getting more sophisticated in a
- 24:27figuring out how do we use this, how do
- 24:29we provide our context. So this kind of
- 24:32like race continues, right? We want more
- 24:33agent autonomy and we want less human
- 24:36supervision. And as part of that also
- 24:38something that has happened over the
- 24:39last 12 months is that it has become a
- 24:41lot easier to run agents with
- 24:44no supervision, right? So this
- 24:47screenshot is from like I probably took
- 24:49it last year in June, July or something.
- 24:51That was the first version of Codex and
- 24:54over time now also most of the the
- 24:56harness products, the coding agent
- 24:57products have now come come out with
- 25:00like platforms where you can always
- 25:02decide like do I want to run this coding
- 25:03session locally on my machine or do I
- 25:05want to do it in the cloud? And so it's
- 25:07become a lot easier
- 25:10to do this and to to try this out,
- 25:13right? Like whatever size you feel
- 25:14comfortable size and complexity of tasks
- 25:17you want to do this for like some people
- 25:19run it to actually like build full
- 25:21features and others maybe like just dip
- 25:23their toes in like cleaning up this
- 25:25feature toggle or like small clean up
- 25:27tasks, right?
- 25:29And also in the meantime we've kind of
- 25:31like taking this like to more extremes,
- 25:34right? Which is more in experimental
- 25:35stage right now I would say
- 25:37which is this idea of like swarms or
- 25:39really brute force just sending lots and
- 25:41lots of agents out there and also having
- 25:44the agents decide how many agents they
- 25:46need, right?
- 25:48Gastown got a lot of attention in I
- 25:51think it came out in January. There were
- 25:53these two like big or even more
- 25:55experiments from like cursor for example
- 25:57or anthropic to build C compiler in the
- 25:59browser.
- 26:00Cloud flow I think it has a different
- 26:02name now that was probably even earlier
- 26:04than Gastown last year. Kind of people
- 26:06were playing around with that. So that's
- 26:08kind of like taking it to the extreme
- 26:10and seeing how we can push the
- 26:11boundaries and actually have AI build
- 26:13much bigger things more autonomously.
- 26:17Um
- 26:20Yeah, so this is our fourth year into
- 26:22this.
- 26:23>> [laughter]
- 26:24>> Um so we start with auto complete and we
- 26:26had a bit more like integration into the
- 26:28IDEs, more context. Cloud 3.5 Sonnet was
- 26:32like an early model moment I think where
- 26:35I certainly from that point on just just
- 26:38always almost always use Cloud Sonnet
- 26:40because it just like
- 26:41felt so much better at coding than the
- 26:43other ones. Um
- 26:45So then
- 26:46we got these like at the time most
- 26:48people are I also my presentations still
- 26:51call it agentic coding modes, right? So
- 26:53that's when like the the like like
- 26:57cursor and so on they got like these
- 26:58modes where they could also run terminal
- 27:00commands which which had been a thing
- 27:02that was already out there in open
- 27:03source, but not as widely used and FCP.
- 27:06So that was only about 1 and 1/2 years
- 27:07ago.
- 27:09Then shortly after that the vibe coding
- 27:12term got coined by
- 27:14Andre Karpathy which then led to like a
- 27:17lot of attention of people discovering
- 27:19these new agentic modes and going oh,
- 27:21this is like there was like a wave of
- 27:23like people picking this up again and
- 27:24saying oh, this is has actually improved
- 27:26quite a bit.
- 27:28Then we got these kind of background
- 27:29agents that I just talked about, right?
- 27:31Like so Codex for example, you know,
- 27:33allowing you to run things unsupervised
- 27:35in the background. Cloud code is
- 27:39I think generally available probably
- 27:41about a year old. It's like it started a
- 27:44little bit earlier than that. The
- 27:45context engineering term also started
- 27:47gaining traction about a year ago. Then
- 27:50we had that Cloud Opus moment. We got
- 27:52skills. Open Claw is maybe also relevant
- 27:56relevant moment even though it's not
- 27:58quite directly about coding.
- 28:00And then yes, so like kind of beginning
- 28:02of this year we got this like next wave
- 28:04of people paying attention again and
- 28:06going oh, this has actually changed
- 28:08quite a bit, right? I can actually see
- 28:10those two yellow spots in our internal
- 28:12AI coding chat community in in
- 28:15ThoughtWorks. I can see the spikes in
- 28:18activity and then kind of like holding
- 28:20and then there's like another like I
- 28:22don't know
- 28:23250 messages a week or something median.
- 28:28Yeah, and then so yeah, the gas town and
- 28:29all of the swarms and all of that
- 28:31started those big experiments started
- 28:32happening beginning of this year then
- 28:34the harness engineering term is now like
- 28:36a buzzword. So and you know, I don't
- 28:39have any additional boxes here because I
- 28:40think right now it's like us all like
- 28:43processing, right? And actually using
- 28:44all of these things that are happening.
- 28:45So there's been like I think not like a
- 28:47big coinage of anything or like a new
- 28:50buzzword other than these things, but
- 28:51we're just like grappling with it and
- 28:53and using this stuff, right?
- 28:58So, it's come a long way, but the costs
- 29:01have as well, and I don't just mean the
- 29:02token costs.
- 29:04So, one is security, right? So, our it's
- 29:07it's both about about like secrets
- 29:09potentially leaking from our
- 29:11environments, from our machines, but
- 29:12also our ecosystem is under attack,
- 29:15right? So, um
- 29:17we have to like think even more than
- 29:19before about dependency management,
- 29:20about sandboxing, and and stuff like
- 29:22that.
- 29:23Um on the other hand, there was also a
- 29:25good talk by Joseph from GitHub here
- 29:27yesterday about like how we can use AI
- 29:29to improve our security, right? So,
- 29:31there's usually both sides to it.
- 29:33Stability, right? So, in the Dora
- 29:35report, um
- 29:37this was one of the big like, let's say,
- 29:39negative trend findings about stability
- 29:42actually getting worse according to
- 29:43their um data, but also there were a lot
- 29:46of talks here about like how we can use
- 29:48AI to help us improve our stability.
- 29:51Um then changeability is a big thing,
- 29:54right? So, like
- 29:55code quality defined as code that
- 29:58remains easy to change and remains where
- 30:01remains easy to change it with low risk,
- 30:03right? So, this was a change uh I
- 30:06recently introduced into a still
- 30:08relatively new codebase that was all
- 30:10created with AI, and I made a change
- 30:13that touched 41 files, but it shouldn't
- 30:15have. It wasn't like that big of a deal,
- 30:17and so it was like a clear smell that
- 30:19there was already like accumulated tech
- 30:21debt that was making changes more risky
- 30:23and more
- 30:24uh costly. So, but then also here the
- 30:27question is like how far can we push
- 30:29this more and improve this more with
- 30:31guides, with sensors, with static code
- 30:33analysis, and so on.
- 30:35Token cost is the most obvious cost
- 30:37thing, of course. In the beginning of
- 30:382024, I I at a keynote presentation
- 30:41where the speaker said, "Generating 100
- 30:43lines of code only costs about 12 cents.
- 30:46Can you imagine?"
- 30:48>> [laughter]
- 30:48>> And he was comparing that to developer
- 30:50salaries. I mean, regardless of, you
- 30:52know, lines of code is not
- 30:54a measure of value, right? Let's just
- 30:56get that out of the way as well. But of
- 30:58course now we have like there was some
- 31:01numbers or some quotes in the Pragmatic
- 31:02Engineer newsletter recently where
- 31:04somebody for example, there were
- 31:05multiple of these quotes, but somebody
- 31:07said, "Some developers are now spending
- 31:09$500 a day." Which if we take that
- 31:12analogy of a developer salary is over
- 31:14$100,000 a year salary, which is a
- 31:16pretty decent salary in the even in the
- 31:18richest countries in the world, right?
- 31:21Um
- 31:22Another type of cost is like cognitive
- 31:24load and burnout, right? Like
- 31:27who would have thought, right? It
- 31:28doesn't actually make us make our lives
- 31:30more relaxed on the contrary, right?
- 31:32Like some people are working more even
- 31:34though they they already create more
- 31:36output. And Steve Yegge had this analogy
- 31:38with the energy vampire from What We Do
- 31:41in the Shadows. I don't know if any of
- 31:42you have seen that
- 31:44that show, but he basically yeah, he's a
- 31:46vampire that doesn't suck blood that but
- 31:48that sucks
- 31:49energy. So there's lots of stories about
- 31:51people saying, "Oh, I can only do this
- 31:52like 3 hours in a row and then I have to
- 31:55like take a nap."
- 31:57Um then we have the review crisis of
- 31:59course, right? We have like higher
- 32:02coding throughput, but then can we if we
- 32:05can code faster, can we review faster,
- 32:07test faster, ship faster, review faster?
- 32:10We've definitely so far the state is no.
- 32:12We cannot review faster, right?
- 32:13Everybody's complaining about this about
- 32:16this pain. Uh but it's not just coding.
- 32:19Everybody can create more with AI now,
- 32:21right? So I was talking to a colleague
- 32:23the other day who was telling me a story
- 32:24about an organization
- 32:27on this side of like if you can code
- 32:28faster, can you fill the backlog faster?
- 32:31So here, you know, there was like
- 32:33another
- 32:34another kind of like thing popping up
- 32:36before the coding where the the product
- 32:39managers were actually like churning out
- 32:41lots and lots of prototypes and lots and
- 32:43lots of ideas, right? So now there was
- 32:45this weird like bottleneck between this
- 32:47pile of prototypes and the pile of code
- 32:50because they couldn't get it like to
- 32:51sync up again and nobody really could
- 32:53figure out how to converge on what they
- 32:55actually wanted to build cuz they were
- 32:57doing this kind of like in these two
- 33:00silos. So that's maybe like something
- 33:03like an open question that we're already
- 33:05seeing a little bit but are we are we
- 33:07heading in general towards like a flow
- 33:09crisis, right? And whenever I want to
- 33:11understand flow better, I turn to my
- 33:13colleague James Lewis who like has done
- 33:15a lot of very interesting presentations
- 33:17about flow that you know it's easy to
- 33:20find on on YouTube. Um but here's one
- 33:23where he's talking about congestion
- 33:25collapse
- 33:27and so he has this
- 33:29this prediction apparently has a bet
- 33:31open with Gene Kim for a crate of beer
- 33:34that this will happen and that this will
- 33:36become like a big topic of conversation
- 33:38that we're just like overloading and
- 33:40also overloading in these different
- 33:41silos and that everything will just
- 33:43become super slow at some point and just
- 33:45collapse. So this comes you know from
- 33:48thing theory of constraints and and
- 33:50stuff like that and here he's he's
- 33:51quoting the Don Reinertsen book about
- 33:54the principles of product development
- 33:55flow
- 33:56as well.
- 33:59So then humans are seen as the
- 34:00bottleneck, right? That's what what
- 34:02multiple people also quoted or said here
- 34:06at the conference.
- 34:10So
- 34:11it remains this big question of like
- 34:13trust in the code and how much
- 34:14supervision do we want to have and how
- 34:17much review.
- 34:20And of course it depends, right? So the
- 34:22autonomy of AI coding agents is here.
- 34:25It's just unevenly distributed, right?
- 34:28Lots of you are probably already using
- 34:29it for some things, right? But I think
- 34:32it will never be that we can use AI
- 34:35coding agents for any type of task in
- 34:37any situation, right? It's just always
- 34:39it depends on the situation. And um
- 34:43what does it depend on, right? So, for
- 34:45me, the way I think about it right now
- 34:47is as this risk assessment uh out of
- 34:49probability, impact, and detectability,
- 34:52right? Which is very typical kind of
- 34:54components of risk assessment in all
- 34:55kinds of areas. So, first I think about
- 34:57the probability that AI gets something
- 34:59wrong or gets something right. And
- 35:01that's all about me knowing the things
- 35:03that I talked about before, my AI tool,
- 35:06like what my context is, did I even give
- 35:08the agent a chance to do it right? Um
- 35:10and it's also about me reflecting on my
- 35:12confidence in my requirements, right?
- 35:14Like how [snorts] certain am I even that
- 35:16I even know what to do? Um then I think
- 35:18about the impact if AI gets something
- 35:20wrong. So, that's all about the use case
- 35:22criticality, of course. So, is this like
- 35:25a super critical business flow that uh
- 35:28you know, will wake me up at 2:00 a.m.
- 35:29on Saturday because I'm on call? Or is
- 35:32this something a lot less uh
- 35:34a lot less important? Then maybe like
- 35:36I'm a little bit more loose with uh how
- 35:39I review. And the third thing is I
- 35:40reflect on detectability that AI got
- 35:43something wrong. Will I notice, right?
- 35:45And um
- 35:46by the way, all of this starts with
- 35:48knowing what right and wrong means,
- 35:49right? I should So, and and often it's
- 35:52like is it appropriate, right?
- 35:54Uh so, you have to know your uh feedback
- 35:56loops, basically. And then based on
- 35:58these things, I decide which workflow do
- 36:00I use, how much review do I do, and how
- 36:03long let do I let it go without
- 36:04supervision. For example, if I don't
- 36:07even know myself yet quite what I need,
- 36:09then I won't let it go off for like half
- 36:11an hour and then realize that it was all
- 36:13for nothing, right?
- 36:16And yeah, so you have to kind of be this
- 36:17tall to ride the roller coaster. You
- 36:19have to be this tall to reduce
- 36:20supervision, right? So, you can think
- 36:22about uh you know, feedback loop in your
- 36:24team, in your organization. And you can
- 36:26also increase the probabilities by
- 36:29improving your context engineering, your
- 36:31harness engineering,
- 36:33um and also
- 36:35refactoring and modernizing and you
- 36:37know, all of those things because
- 36:40AI can also deal with a well-factored
- 36:42code base much better
- 36:45than with a messy code base.
- 36:48So, we're tempted to move from in the
- 36:49loop to on the loop to out of the loop.
- 36:52There's I I feel this every day being
- 36:54drawn to like, "Ugh, I don't want to
- 36:56look at this. I don't want to look at
- 36:57the code anymore." It's like
- 36:59Yeah, there's all of these forces that
- 37:01pull us there, but we're also starting
- 37:03to actually feel the costs. It's not
- 37:04just speculation anymore. It's not just
- 37:07like dooming kind of predictions. We're
- 37:09actually feeling the cost of tokens,
- 37:11risks, and cognitive X, right? Cognitive
- 37:14load, cognitive debt, right? This idea
- 37:18of like we don't even understand anymore
- 37:19how our code base is structured.
- 37:22I sometimes think of like cognitive
- 37:23deferral as well. It feels like we keep
- 37:25like deferring the review to other
- 37:28people or like deferring processing what
- 37:31actually happened. And recently there
- 37:34was like a new cognitive X term coined
- 37:37that
- 37:38I I saw because Adi Osmany wrote about
- 37:40it, which is cognitive surrender, right?
- 37:43So, in this paper they put AI into the
- 37:46context of the system one system two
- 37:50thinking fast and slow. I think probably
- 37:52lots of you have heard about that book.
- 37:53If not, then look it up. It's like
- 37:55really great. And so, they are they're
- 37:57talking about a mode where we're
- 37:58basically displacing system two like our
- 38:01actually like active thinking with AI
- 38:04and they call it cognitive surrender.
- 38:06And apart from like this paper and the
- 38:08cognitive part of it, this term like
- 38:11surrender has just been stuck in my head
- 38:12like ever since I heard this. And I feel
- 38:15like we're There's like so many things
- 38:17where we we're in danger of just like
- 38:19surrendering right now, not just
- 38:21cognitively.
- 38:22Um and so I think we have to be careful
- 38:25about what we are surrendering, right?
- 38:26And like really think about be mindful
- 38:29of
- 38:30where that's worth doing. So I'll just
- 38:32give like some examples, right? Um
- 38:35So it's like ah it's too much to reason
- 38:37about this big change. It'll be fine,
- 38:39right?
- 38:40I do it myself myself. I'll just do it
- 38:42myself quickly instead of teaching
- 38:43somebody, right? Everybody's talking
- 38:45about oh how will juniors learn? How
- 38:47will we do this? But I don't see that
- 38:49much like
- 38:50active action, right? Because
- 38:52everybody's just in a tunnel of you know
- 38:54I see experienced people building tools
- 38:56for themselves to use AI better, but
- 38:58like why are we not taking more
- 39:00initiative to think about how this is
- 39:03sustainable for people who don't have
- 39:04all of this experience already?
- 39:07Or this like ah I'll just use a big
- 39:08model, you know. I can't be bothered to
- 39:11then like try retry it again, so let's
- 39:13just use the most expensive biggest
- 39:14model.
- 39:15Or this surrendering to like ah I don't
- 39:17want to solve all these problems. Models
- 39:19will get better. Tokens will be cheaper
- 39:20again. We just have to wait it out,
- 39:23right? Um we're working more, we're
- 39:26producing more, but we're still getting
- 39:28the same compensation.
- 39:30Is that also a type of surrender?
- 39:32Um or this like ah I don't have time to
- 39:34find better approaches, which is by the
- 39:36way not necessarily the individual's
- 39:38fault. It's also like what's happening
- 39:40around them and incentives and
- 39:41pressures, right? Uh sandboxing is too
- 39:44tedious. Surely nothing will go wrong,
- 39:46right?
- 39:48Or I I can't work on this code base
- 39:50without AI anymore, right? That's the
- 39:51the original cognitive surrender
- 39:53cognitive debt kind of definition.
- 39:56Um this was like Hannah yesterday in a
- 39:58talk and she was kind of talking about
- 39:59what she's keeping, what she's trashing,
- 40:02and what she's trying. And I think
- 40:03that's similar in terms of like we have
- 40:05to think about what we're surrendering
- 40:07and what we should be uh what we should
- 40:09be keeping. So if you are a person of
- 40:11influence in your engineering
- 40:12organization, are you creating an
- 40:14environment that leads to surrender,
- 40:16right? That makes me people feel like
- 40:18they just have to crank out the PRs and
- 40:20don't have time to like actually figure
- 40:22out how to improve the environment, how
- 40:24to improve the context engineering, and
- 40:26so on. Um if you're if you maybe feel a
- 40:28bit like powerless or you feel like you
- 40:30can't influence it and there's all these
- 40:32pressures around you, still like try to
- 40:34think about your sphere of influence and
- 40:35the small things you can do. Like what
- 40:37do you really have to surrender, right?
- 40:39Um so there's lots of options between
- 40:42not using AI at all and like just total
- 40:45surrender and hope and hoping the models
- 40:46will fix it all in the future.
- 40:49And again, like if you feel kind of like
- 40:51powerless and like this is like washing
- 40:52over you, we can also work together and
- 40:54collaborate on these things, right? So
- 40:56if you feel like you're not a good
- 40:57communicator about this, maybe look for
- 41:00somebody in your team or organization
- 41:01who's good at that and share your data,
- 41:03share your observations with them,
- 41:04right? So like we can collectively do a
- 41:07lot about this as well.
- 41:09In terms of skills, we need to know our
- 41:10toolbox, past and present. We also need
- 41:12to rediscover some things.
- 41:14Uh we need critical thinking, risk
- 41:16assessment, and some form of like
- 41:18patience, right? So we need to evaluate
- 41:21that productivity maybe doesn't just
- 41:22mean like typing. And for leadership,
- 41:25you know, this is still horizon two,
- 41:27right? This is not horizon one yet. So
- 41:29maybe we don't know the ROI yet. Maybe
- 41:32we just have to like give people some
- 41:33time to um
- 41:35to create a good setup so that we can
- 41:37continue to safely and quickly deliver
- 41:41software to users in a sustainable way.
- 41:45Thank you.
- 41:47>> [applause]
- 41:49[music]
About this transcript
This page contains the full transcript of Birgitta Böckeler - State of Play: AI Coding Assistants - AI Native DevCon June 2026 by AI Native Dev, generated from the public captions YouTube serves with the video. The transcript has 8,038 words across 1,203 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.