YouTube2Text

Birgitta Böckeler - State of Play: AI Coding Assistants - AI Native DevCon June 2026 — Transcript

by AI Native Dev · 8,038 words · 1,203 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Birgitta Buckler is a is a distinguished
  2. 0:02engineer at ThoughtWorks and um
  3. 0:05I've been chatting to Birgitta for a
  4. 0:07little while and actually uh was really
  5. 0:09impressed by a very a post that went
  6. 0:12viral on Martin Fowler's uh site that
  7. 0:15Birgitta wrote uh talking about specs
  8. 0:18and talking about uh the the comparison
  9. 0:21at the time between Spec Kit uh Tessel
  10. 0:24and uh Kiro, I think it was. And uh
  11. 0:27okay, Tessel's moved on a little bit
  12. 0:29since then, but it's still amazing to
  13. 0:30see how many people go to go to that uh
  14. 0:33site. Uh Birgitta's an amazing person. I
  15. 0:35very much encourage you to to follow
  16. 0:36her. She's she's very thoughtful in with
  17. 0:38so many uh posts and blogs that she
  18. 0:40writes. And
  19. 0:42a a really wonderful way to finish off
  20. 0:44uh this conference with a very visionary
  21. 0:47uh session whereby we Birgitta's going
  22. 0:48to look at the last 12 months where
  23. 0:50what's been changing, where we are
  24. 0:52today, and what we can look forward to.
  25. 0:53So, uh please give a very a a Native Dev
  26. 0:56warm welcome to Birgitta Buckler.
  27. 0:59>> [applause]
  28. 1:06>> Okay, yeah, thanks Simon. Thanks Simon
  29. 1:08and Patrick for inviting me. Yeah, so
  30. 1:10I'm a distinguished engineer at
  31. 1:12ThoughtWorks and what that means for me
  32. 1:13specifically is that uh 3 years ago I
  33. 1:16got a full-time role to just be immersed
  34. 1:19in the space of AI coding or in general
  35. 1:21using AI on software teams to help my
  36. 1:23colleagues, to help our clients like I
  37. 1:25don't stay on top of it. So, I talk a
  38. 1:27lot to
  39. 1:28uh our teams, our clients, and so on.
  40. 1:30And then I write about it, for example,
  41. 1:32on my colleague Martin Fowler's website.
  42. 1:34Um and so that's kind of like what what
  43. 1:36all of this is um based on. And yeah,
  44. 1:39it's kind of like a tough task of
  45. 1:41wrapping up after everybody heard about
  46. 1:43this topic for 2 days. Uh so, I'm going
  47. 1:45to try and help you see the forest for
  48. 1:48the trees or the multiple forests for
  49. 1:49all of the the trees. So, I'll kind of
  50. 1:52do like a recap. I'll start with the
  51. 1:54recap slide. Simon was briefly confused
  52. 1:56and thought maybe my slide setup was
  53. 1:58wrong, but I'll it's kind of like recap
  54. 2:00a lot of what the stuff that you heard
  55. 2:01also over the last two days, but also
  56. 2:04kind of what happened in the last 12
  57. 2:05months like advancements as well as
  58. 2:07things that are maybe not going so well
  59. 2:09or that are kind of like the all the
  60. 2:11second order consequences and
  61. 2:13implications that we're experiencing
  62. 2:14right now. So that when you get back to
  63. 2:16work tomorrow and your colleague who
  64. 2:19maybe isn't as immersed in the space ask
  65. 2:21you, "So,
  66. 2:22what should I know, right?" I hope I can
  67. 2:24help you answer that question.
  68. 2:26And I'll start with yeah, the reason why
  69. 2:29all of this is happening which is
  70. 2:31the model so kind of first and most
  71. 2:33obvious.
  72. 2:34There wasn't even that much talk about
  73. 2:36models here at the conference, I think
  74. 2:38which is not surprising and totally
  75. 2:40okay, I think because for me this is not
  76. 2:42really the most exciting part to be
  77. 2:43honest. I'm much more interested in
  78. 2:45everything that's now happening around
  79. 2:47it, the ecosystem, all of the
  80. 2:48integrations, and so on. And I mean
  81. 2:51obviously if we talk about the last 12
  82. 2:52months, there was the Opus 4.5 moment
  83. 2:55kind of last year that made a lot of
  84. 2:59people kind of like come back to this
  85. 3:00that hadn't maybe tried AI coding for a
  86. 3:02while.
  87. 3:04Um
  88. 3:05and
  89. 3:06yeah, so that was maybe the biggest
  90. 3:08event in models that happened like I
  91. 3:09mean almost every week there's a new
  92. 3:11model, but I usually don't even follow
  93. 3:13it that much because like I said it's
  94. 3:15kind of interesting to to see all the
  95. 3:17stuff around it. So if we think about
  96. 3:20models like what are the core things
  97. 3:22kind of as users who use them for coding
  98. 3:25that we need to know or learn almost
  99. 3:27like a kind of like learning map, right?
  100. 3:29So the first thing I always try to get
  101. 3:31out of the way is that they are not
  102. 3:33magic, right? They are very, very
  103. 3:35impressive and very, very useful math,
  104. 3:37but unfortunately even like a lot of
  105. 3:40technologists, a lot of our peers kind
  106. 3:41of like it's very easy to fall into that
  107. 3:43trap, right? Like of
  108. 3:45thinking of them as as more than that,
  109. 3:47right? So I kind of like this
  110. 3:48visualization to remind ourselves that,
  111. 3:51you know, even though we don't really
  112. 3:52know what's like why it works and what's
  113. 3:54happening, it's still like very
  114. 3:56impressive math, right? So, first of
  115. 3:58all, they're not magic. Um I mention
  116. 4:01here as the second point that their
  117. 4:02statelessness, because that's also
  118. 4:04something that I often notice that
  119. 4:06people haven't quite grasped, right? So,
  120. 4:08the the model doesn't have a session,
  121. 4:10right? So, the longer our conversation
  122. 4:12with them gets, the longer our session
  123. 4:13with them gets, every single time the
  124. 4:16our our agent, our harness basically
  125. 4:18sends the whole history of the
  126. 4:20conversation, right? Maybe not quite,
  127. 4:22there's caching and all kinds of clever
  128. 4:23ways that uh different tools try to
  129. 4:26optimize that, but they are stateless,
  130. 4:28right? So, that is a a factor that
  131. 4:30happens like the longer our
  132. 4:32conversations get with them. So, I think
  133. 4:33that's a really important core thing
  134. 4:35that um
  135. 4:37uh yeah, people need to understand.
  136. 4:41Uh the third thing we need to know about
  137. 4:43is, of course, the size of the context
  138. 4:46window and um in relationship, but also
  139. 4:48in relationship to what that means for
  140. 4:49attention, right? So, even though
  141. 4:51technically the context windows have
  142. 4:52gotten a lot bigger, um it comes with a
  143. 4:55trade-off on like how well the models
  144. 4:57are able to keep attention on all of the
  145. 4:59many instructions and all of the context
  146. 5:01that we're trying to feed them now. So,
  147. 5:03there's something there to be understood
  148. 5:05by everybody who uses this about what
  149. 5:07that trade-off is.
  150. 5:08Um and then um finally this uh this and
  151. 5:12that's maybe the the biggest area. So,
  152. 5:15those first things are kind of like you
  153. 5:16could you can learn them in a formal
  154. 5:17training and kind of like understand the
  155. 5:19basics, right? But this last one is a
  156. 5:21lot more about using the models and
  157. 5:23figuring this out, right? Which model do
  158. 5:25we use for which task, right? So,
  159. 5:28there's I mean, these are just like a
  160. 5:29few illustrative examples, like uh we we
  161. 5:32have autocomplete, like there's still
  162. 5:34people who use that a lot, uh let's say,
  163. 5:37or let's say you want to just change a
  164. 5:38few specific files and you have very
  165. 5:40clear instructions and a very clear idea
  166. 5:41of what you want to do, or you have a
  167. 5:43larger and more complex change that
  168. 5:45needs a bunch of code research before
  169. 5:47the model actually does it or the agent
  170. 5:49actually does that. Or you have like
  171. 5:51tasks like planning, debugging,
  172. 5:53designing that maybe have a lot more
  173. 5:55like things like asking you questions
  174. 5:57and a lot more reasoning involved. So,
  175. 5:59um different tasks will have different
  176. 6:02levels of reasoning that are useful, uh
  177. 6:05will need different uh sizes of context,
  178. 6:08and will have different levels of need
  179. 6:10for tool calling. And I actually think
  180. 6:12that in this area here, we don't even
  181. 6:15have to know that much about the details
  182. 6:17of all of the different features of the
  183. 6:19models, but it's a lot more about us
  184. 6:21reflecting on these types of tasks,
  185. 6:23right? Which is something that we've
  186. 6:24done in the profession for a for a long
  187. 6:26time, right? Like all of those like
  188. 6:28philosophical discussions about what
  189. 6:29does complexity mean, right? Like that
  190. 6:31we have in in estimations or stuff like
  191. 6:34that, right? So, um you have to think
  192. 6:36about like how many files might this
  193. 6:37involve, what's the blast radius, how
  194. 6:40much uncertainty do I still have, all of
  195. 6:42those things. So, I think yeah, like I
  196. 6:44said, I think it's um much more
  197. 6:46important for us to to work on this like
  198. 6:48reflection on the types of tasks that we
  199. 6:50have and then match them to different uh
  200. 6:53levels of power that the the models
  201. 6:55have.
  202. 6:57And you if you've been bathing in
  203. 6:59powerful models for the last few months,
  204. 7:01then it's actually a nice exercise to
  205. 7:03remember what we're already taking for
  206. 7:05granted these days when we try to run a
  207. 7:07smaller model on our developer laptop,
  208. 7:10right? And try to see what it can do and
  209. 7:12what it cannot do. And
  210. 7:16so here's like an example. This is just
  211. 7:18to show you kind of the speed, right?
  212. 7:20The speed has actually come a really
  213. 7:21long way. So, this is Gwen 3.6 running
  214. 7:24on my Apple M3 with 48 GB of RAM. And
  215. 7:28it's actually not even that much slower
  216. 7:30than like some of the stuff that happens
  217. 7:31in Claude code uh with with Sonnet or
  218. 7:34Opus, right? This is uh this is open
  219. 7:36code, um by the way. So, the speed has
  220. 7:39actually come quite a bit of a way. Tool
  221. 7:41calling is still kind of even with this
  222. 7:43like quite powerful model that I could
  223. 7:46hardly even run on a 64 GB RAM MacBook.
  224. 7:50So, it it's it crashed after my my first
  225. 7:52attempts to use it. So, even that model
  226. 7:54was still struggling a bit with tool
  227. 7:56calling, which is really crucial for our
  228. 7:57agentic workflows, right?
  229. 8:00But, it also has gotten a lot better
  230. 8:02than it than it was like, I don't know,
  231. 8:046 months ago or so. Um and complexity of
  232. 8:07all the instructions, I mean, uh you
  233. 8:09know, many of you might have seen this
  234. 8:11when you when you try smaller models or
  235. 8:14I remember when I first used Gemini for
  236. 8:16for coding like a year ago or so, I kept
  237. 8:18I kept getting these as well. So, um
  238. 8:21that is still happening, but at the same
  239. 8:23time, like I was for example recently
  240. 8:25with uh Gemma 4, which is like a model
  241. 8:27that everybody's talking about right now
  242. 8:29in terms of like a smaller model that is
  243. 8:31quite capable at coding comparatively,
  244. 8:34you know, to uh to other small models.
  245. 8:36And so, I used this recently. Again,
  246. 8:38this was the type of task where I knew
  247. 8:39exactly what needed to be done, but I
  248. 8:41couldn't be bothered to type it out
  249. 8:43myself, right? So, I gave it like one
  250. 8:46small paragraph of instructions, and it
  251. 8:47was actually really good at doing that.
  252. 8:49I went after that back and forth a
  253. 8:50little bit, had it refactor it a little
  254. 8:52bit because it was quite uh
  255. 8:54cyclomatically complex, and then I had
  256. 8:56my little utility script. So, for this,
  257. 8:58I could actually use it on my [snorts]
  258. 9:01uh on my M3, right? So again, it's a lot
  259. 9:03about like knowing the tasks and then
  260. 9:04kind of uh mapping that to the level of
  261. 9:07power you want to use.
  262. 9:09So, moving on from the model, then the
  263. 9:11next thing that we have around the model
  264. 9:13is what we what uh we've now kind of
  265. 9:15come to calling the coding harness,
  266. 9:17right? Or
  267. 9:18uh you know, also sometimes more
  268. 9:20colloquially still like the coding
  269. 9:21agent, right? So, um that's the thing
  270. 9:24that's kind of helping us leverage the
  271. 9:26model for our coding tasks, and it has
  272. 9:29things under the hood like a system
  273. 9:31prompt or all kinds of like like other
  274. 9:32prompts that we usually can't even see
  275. 9:34when unless it's open source, right? I
  276. 9:36mean, we recently got a glimpse into the
  277. 9:38cloud code once, but
  278. 9:40a lot of them are a lot of the big ones
  279. 9:42that we actually use a lot are closed
  280. 9:44source, right?
  281. 9:46It comes according to Harness, it comes
  282. 9:48with the tool integrations, all of the
  283. 9:49standard stuff that we'll definitely
  284. 9:51need, right?
  285. 9:52Changing files, reading files, code
  286. 9:54search is a big one, right? So, we get
  287. 9:56like a type of code search out of the
  288. 9:58box which each with each Harness that we
  289. 10:00pick. It has like all kinds of
  290. 10:03orchestration, like most of them has
  291. 10:05have sub agents now, for example, and
  292. 10:07also decide when to spawn off certain
  293. 10:09sub agents or they kind of decide like
  294. 10:13how many tool calls at once they pass on
  295. 10:15to the model and all of those types of
  296. 10:17things. Maybe there's some caching
  297. 10:19involved in some of them. They have a
  298. 10:22user interface, of course, right? Like
  299. 10:24some of them have a terminal-based user
  300. 10:26interface, some of them are like in VS
  301. 10:28Code or or other more graphical user
  302. 10:31interfaces. And they also all come with
  303. 10:34different levels of extensibility and
  304. 10:35observability. So, extensibility,
  305. 10:38famously, the pie coding agent is very
  306. 10:40popular right now as one that has kind
  307. 10:42of brought that more to our attention of
  308. 10:44having a coding agent that we can
  309. 10:45actually like that is malleable, that we
  310. 10:47can change, right? And observability is
  311. 10:50also a space where there's a lot
  312. 10:52happening right now in terms of uh you
  313. 10:55know, having traces of what the agent is
  314. 10:57doing,
  315. 10:59how can we use that to, for example,
  316. 11:01analyze more like how we can improve
  317. 11:03how we're using the agent
  318. 11:06getting some visibility into, for
  319. 11:09example, I've done some stuff about like
  320. 11:10visualizing for myself during session
  321. 11:12which files is it is it reading and
  322. 11:14which files is it writing, so I could
  323. 11:16get an idea of like blast radius of
  324. 11:18stuff. So, I think there's still a lot
  325. 11:20of potential here for us to also make
  326. 11:22these things part of the the review
  327. 11:24cycle.
  328. 11:26They're kind of like first rumblings
  329. 11:28about the bloat, right? We always have
  330. 11:30the cycle in software, right? We we get
  331. 11:32like a new tool, like let's say spring,
  332. 11:34right? And it's like super lightweight
  333. 11:37and we're excited because it's much more
  334. 11:38lightweight than the bloated thing that
  335. 11:40we had before, and then it takes like a
  336. 11:42number of years, right? And then we kind
  337. 11:44of feel like that's also too big, right?
  338. 11:46With Claude code, it hasn't even taken a
  339. 11:48year and we feel like it's maybe a bit
  340. 11:50much, right?
  341. 11:51>> [laughter]
  342. 11:52>> Um so here's like a comparison of like
  343. 11:54when you start out with a session in pie
  344. 11:56versus open AI codex versus Claude code,
  345. 11:58all of the stuff that's already in
  346. 12:00there, right? Um [snorts]
  347. 12:03but yeah, so we need to as when we come
  348. 12:05back to like what what do we need to
  349. 12:06know, right? What do we need to
  350. 12:08understand as as engineers using these
  351. 12:10tools, we need to kind of like
  352. 12:11understand these features and how they
  353. 12:13distinguish between the agents, right?
  354. 12:16So just to give like an example, um
  355. 12:18I remember when Claude code came out and
  356. 12:20it became super popular, um I saw a lot
  357. 12:24of conflation of the interface with the
  358. 12:27reason why Claude code was good, right?
  359. 12:30So I saw that for a while that people
  360. 12:31said, "Oh yeah, terminal based, that's
  361. 12:33the way to go because it's really
  362. 12:34powerful." But actually when you have
  363. 12:36for example one of my favorite harnesses
  364. 12:38next to Claude code is cursor is also
  365. 12:40really really good under the hood with
  366. 12:42all of those things that you see in the
  367. 12:43circle below there. So it's not
  368. 12:45necessarily because of the terminal,
  369. 12:46right? So this is just like an example
  370. 12:48of why it's uh it's important to kind of
  371. 12:51distinguish um
  372. 12:54you know, kind of know as the engineer
  373. 12:56what what's actually happening in the
  374. 12:58tools.
  375. 13:01So yeah, we have to understand their
  376. 13:02footprint, understand their features so
  377. 13:04that we can use these features
  378. 13:06effectively, right? And so most like one
  379. 13:09of the big things that we want to do is
  380. 13:11we want to understand how to regulate
  381. 13:13the context, how to tune the context
  382. 13:14that we give to the agent with the use
  383. 13:16of the features that the coding harness
  384. 13:18provides us, right?
  385. 13:20So about 12 months ago, the main way
  386. 13:22that we did that was rules files or
  387. 13:24instruction files, right? So we would
  388. 13:26have agents MD or Claude MD and maybe
  389. 13:29write down the typical pitfalls. And
  390. 13:31we're still doing that, but these days,
  391. 13:33of course, there's so many different
  392. 13:35like features in the coding harnesses
  393. 13:37that help us do that in a more
  394. 13:38sophisticated way, right? So, there's
  395. 13:40skills, of course.
  396. 13:42MCP servers also existed about a year
  397. 13:45ago as well. There's now sub agents,
  398. 13:47there's extensions, plugins, hooks, all
  399. 13:49of those things. And it's becoming a bit
  400. 13:51overwhelming and confusing, right? But
  401. 13:54it's kind of like the storming phase of
  402. 13:57this of this technology. And this is
  403. 13:59also how to use these features is, I
  404. 14:02would say, was probably at least like
  405. 14:0350% of the talks at this at this
  406. 14:06conference as well, unsurprisingly.
  407. 14:08So, it's kind of context engineering for
  408. 14:11coding agents. And that means we're
  409. 14:12expanding the harness. We're
  410. 14:14using the features of the harness to to
  411. 14:17help us do this for our specific code
  412. 14:19base, for our
  413. 14:21use case, right?
  414. 14:24Um
  415. 14:26Yeah. So, we're expanding that harness.
  416. 14:29Um so, I'm I'm calling it coder harness
  417. 14:31here. I have to say like I'm a little
  418. 14:33bit unsure. So, this is term that
  419. 14:36that has now gotten traction since
  420. 14:38February, maybe, of harness engineering,
  421. 14:40right? Which is basically this, right?
  422. 14:43Expanding this coding harness. But in in
  423. 14:45other areas,
  424. 14:47you know, people also use harness
  425. 14:48engineering to talk about like how do
  426. 14:50you make the coding harness itself
  427. 14:51better, right? So, it's still like a
  428. 14:52little bit clunky term, I would say. I
  429. 14:54wish we had we had a better one. I also
  430. 14:56jumped on the bandwagon and wrote an
  431. 14:58article about harness engineering.
  432. 15:00Um yeah, it would be great. May- maybe
  433. 15:02somebody comes up with an even better
  434. 15:04word. Like the folks at Tesla here at
  435. 15:06the conference, they've basically
  436. 15:07they've also had this kind of like
  437. 15:09trifecta that I'm presenting here,
  438. 15:10right? Like of the model, the harness,
  439. 15:12and then the context, right? So, I think
  440. 15:14it's like it's reasonable to think of
  441. 15:17harness engineering as context
  442. 15:19engineering for coding agents, right?
  443. 15:21So, that's just like to get the
  444. 15:23terminology out of the way a little bit.
  445. 15:26Um
  446. 15:28Yeah. So let's look at like uh how this
  447. 15:31like area of harness engineering,
  448. 15:32context engineering for coding agents,
  449. 15:34like um a mental model, like how I think
  450. 15:37about this, like beyond the features,
  451. 15:39right? Yes, this is about skills, this
  452. 15:41is about MCP servers, but conceptually,
  453. 15:43like what are we actually doing there
  454. 15:44when we use these features?
  455. 15:46So one thing, and that's the most common
  456. 15:49thing right now, is this uh
  457. 15:51way of putting conventions, product
  458. 15:53context, workflow, the prompts, like
  459. 15:56basically markdown files into our
  460. 15:59uh code base somehow, right? Or like
  461. 16:00making them accessible through skills
  462. 16:02and so on. And in those markdown files
  463. 16:04is actually lots of different things
  464. 16:05going on, right? We have some normative
  465. 16:07stuff, like coding conventions, we have
  466. 16:09some informative stuff, like yeah,
  467. 16:12product context, what are we actually
  468. 16:13doing here? Maybe reference
  469. 16:15documentation,
  470. 16:17uh and then we also have instructions,
  471. 16:18right? Like uh always help me build in
  472. 16:21the following workflow, or always write
  473. 16:23a failing test first, stuff like that.
  474. 16:25So there's actually lots of different
  475. 16:26things going on in something that looks
  476. 16:28just like a bunch of uh text at first.
  477. 16:31And then also some of them we just have
  478. 16:32directly in the workspace, and others
  479. 16:34are maybe more dynamically loaded from
  480. 16:37other data sources, right? And um
  481. 16:40so these are all for me like kind of
  482. 16:41feed forward, so we're trying to
  483. 16:43anticipate what the agent might do
  484. 16:45wrong, and we're also trying to
  485. 16:47anticipate, of course, what we want it
  486. 16:49to do. And we're feeding it all of this
  487. 16:51information, these instructions, these
  488. 16:52norms, so that hopefully in its initial
  489. 16:54generation of code it's already doing
  490. 16:56perfectly, right? But as we know, that's
  491. 16:58not um always happening. So we're
  492. 17:01starting with these guides, but then we
  493. 17:03also want to give it feedback, right? So
  494. 17:05um
  495. 17:06ideally, uh so that we can trigger
  496. 17:08immediately a self-correction loop
  497. 17:10before we even look at the code, so that
  498. 17:12we don't have to like have all those
  499. 17:14low-hanging fruits still in there. So um
  500. 17:16the most common way that people do that
  501. 17:19right now is with like code review
  502. 17:20agents, right? But, there's also all of
  503. 17:22these other tools that we have in our
  504. 17:24toolbox from before AI, like static code
  505. 17:27analysis. And then, of course, we can
  506. 17:29also an agent usually has access to the
  507. 17:31logs, so it can start the application,
  508. 17:33see what what logs come out of it. Uh,
  509. 17:35many people give an agent access to the
  510. 17:37browser, so it can look at like
  511. 17:39something when it has changed a web
  512. 17:41component or something like that.
  513. 17:43Um,
  514. 17:44and there's actually, um,
  515. 17:46uh, a difference kind of between these.
  516. 17:48So, like a review agent is an LLM
  517. 17:50judging the work of another LLM, right?
  518. 17:52So, it's kind of inferential. It's
  519. 17:54running on the GPU. But, we have a bunch
  520. 17:56of tools as well that are, uh,
  521. 17:58computational, as I decided to call them
  522. 18:01here. So, kind of things that run on the
  523. 18:02CPU, right? Like, the static code
  524. 18:04analysis is the best example, I think,
  525. 18:06to to think about this.
  526. 18:08Um,
  527. 18:09yeah, and we have the same distinction
  528. 18:11on the feed forward on the guide side.
  529. 18:12So, we we can, uh, we
  530. 18:15um, can also think about computational
  531. 18:16guides on that side. And the best
  532. 18:18example, uh, for me there is code mods.
  533. 18:21Uh, Ian from Meta also just mentioned
  534. 18:23those. Um,
  535. 18:24which is, for example, tools like open
  536. 18:26rewrite that are really good at doing,
  537. 18:29uh, version upgrades and my migrations
  538. 18:31of, uh,
  539. 18:32of, um,
  540. 18:33frameworks. I don't know if you remember
  541. 18:35like quite a while ago Amazon had a
  542. 18:37really big headline about saving 400 or
  543. 18:40500 developer years or something for
  544. 18:42Java upgrades. That was under the hood,
  545. 18:45actually, mostly code mods being made
  546. 18:47available to AI. So, that combination is
  547. 18:49really powerful, right? So, all of these
  548. 18:51things, uh, or maybe providing a
  549. 18:53different type of code search that that
  550. 18:55is more effective for your really large
  551. 18:56code base. All of those are ways, again,
  552. 18:59to increase the probability that AI does
  553. 19:01what you want in the first go.
  554. 19:04So, that's then the expanded harness.
  555. 19:06And then, as a human, as we've heard in
  556. 19:08a few talks, uh, here as well, as a
  557. 19:10human our job, in part becomes kind of
  558. 19:13steering this set of guides and sensors.
  559. 19:17So as Mitchell Hashimoto says in his
  560. 19:19blog post, it's the idea that anytime
  561. 19:21you find an agent makes a mistake, you
  562. 19:23take the time to engineer a solution
  563. 19:25such that the agent never makes that
  564. 19:27mistake again.
  565. 19:30And of course, AI can help us engineer
  566. 19:33those little solutions, right? Which is
  567. 19:34really useful.
  568. 19:36So just two quick examples to to bring
  569. 19:39that home. So um
  570. 19:41here's like something in the agents MD,
  571. 19:43do not use console log, you know, we
  572. 19:45have a custom logger that does
  573. 19:47structured logging and so on. It's in
  574. 19:48the following place, right? Instead, I
  575. 19:50could have like a linting rule um that I
  576. 19:53customize where I customize the message
  577. 19:56and point uh the agent at that log file
  578. 20:00through that, right? So especially when
  579. 20:02you have things that it doesn't even do
  580. 20:03wrong that often, it happens maybe once
  581. 20:05a week, it's much more effective to have
  582. 20:07this linting rule than to like stuff it
  583. 20:09into your context every single time the
  584. 20:11agent runs. Or here's an another
  585. 20:13example, let's say you have a back-end
  586. 20:14coding conventions a skill that talks
  587. 20:17about the back-end layers that the agent
  588. 20:19should respect uh in terms of like which
  589. 20:22uh modules are allowed to call which
  590. 20:24other modules. There are some tools in
  591. 20:26most language ecosystems that help you
  592. 20:28scan the different imports between
  593. 20:30files. And so again, you can like come
  594. 20:32up together with AI with some rules in
  595. 20:34those tools that help you uh already
  596. 20:37catch the low-hanging fruit of those
  597. 20:39modularity um
  598. 20:41uh violations.
  599. 20:44And then you should think about like how
  600. 20:46you where you put those sensors, when
  601. 20:48you run them, right? So uh kind of like
  602. 20:50strategically think about your path to
  603. 20:52production and think about when you want
  604. 20:54to run them. So do you want to
  605. 20:56run them in the coding session, right?
  606. 20:58Which is I think uh
  607. 21:00whenever that's possible in terms of
  608. 21:02like how cheap is it if how fast it is
  609. 21:03to run a sensor, I think you should run
  610. 21:06them like even before you commit, right?
  611. 21:09So, I have this box here about
  612. 21:10integration, right? So, it kind of like
  613. 21:12depends what that means for you. So,
  614. 21:14probably 80% of my commits in the last
  615. 21:1615 years have been put straight onto the
  616. 21:18main branch, which is probably not the
  617. 21:20case for most of you.
  618. 21:22Um so, integration could either be like
  619. 21:24for you to say, "Okay, I want to do all
  620. 21:26of those things before I even create a
  621. 21:27commit." Or it could be as part of the
  622. 21:29pull request uh process where you run
  623. 21:32some additional like uh inferential
  624. 21:34sensors or something like that, right?
  625. 21:36Then we have lots of stuff in our
  626. 21:38continuous integration pipeline already,
  627. 21:39right? You probably don't want any
  628. 21:41inferential sensors in there because you
  629. 21:44you don't want, you know, the the green
  630. 21:46or red state of your pipeline to depend
  631. 21:48on semantic interpretation of an LLM,
  632. 21:51right? But we have like lots of uh
  633. 21:53computational sensors in there. And then
  634. 21:56also what uh I've heard a lot of stories
  635. 21:58now about teams doing in ThoughtWorks
  636. 22:00and also lots of people writing about
  637. 22:02that. Also the um
  638. 22:04the team Ryan's team who had a um
  639. 22:07a presentation yesterday at OpenAI, they
  640. 22:09call it garbage collection, right? So,
  641. 22:11kind of like continuous drift detection
  642. 22:13for the technical debt that still
  643. 22:15accumulates. Um where you can probably
  644. 22:18put like a lot of inferential sensors,
  645. 22:20right? So, in the code base where I'm uh
  646. 22:21where I'm setting up all of these
  647. 22:23things, I have like a modularity review,
  648. 22:25like a dependency uh freshness review,
  649. 22:29uh security review that doesn't run
  650. 22:30every single time the pipeline runs, but
  651. 22:33I trigger it maybe like once a week just
  652. 22:34to see if there's something new that
  653. 22:36came up.
  654. 22:37Um and these like continuous drift
  655. 22:39detection things, they're probably also
  656. 22:41like require some kind of process on
  657. 22:42your team to deal with them, right?
  658. 22:44There's a lot of parallels to, for
  659. 22:46example, security vulnerabilities and
  660. 22:48how you deal with them, right? Like uh
  661. 22:50you know, they keep popping up again. Do
  662. 22:52you want to suppress them because you
  663. 22:53cannot fix them right now, but then you
  664. 22:55might forget about them. So, I suspect
  665. 22:58we'll have all of these challenges there
  666. 22:59with this type of stuff as well. Like
  667. 23:01how is the team going to going to deal
  668. 23:03with these? Again, you can maybe like
  669. 23:05have AIs, of course, I create
  670. 23:07you can have agents create pull requests
  671. 23:09for you, but still
  672. 23:11how do you deal with those, right?
  673. 23:14And then finally, I should also mention
  674. 23:16this of course also this way of having
  675. 23:18sensors like in production that you give
  676. 23:20AI access to, right? Especially when it
  677. 23:23comes to your architecture fitness,
  678. 23:24things like scalability, latency, all of
  679. 23:27those things. There were a few talks
  680. 23:28here at the conference as well about
  681. 23:30using observability data and
  682. 23:33you know, both to help you fix
  683. 23:34incidents, but also just to like monitor
  684. 23:37how you can make your
  685. 23:38your runtime better.
  686. 23:44Okay, so as users of coding agents, we
  687. 23:47have to know a few key model
  688. 23:49capabilities and like I said kind of
  689. 23:50like have good reflection on the types
  690. 23:52of tasks that we're that we're using and
  691. 23:55what complexity means for us in
  692. 23:56relationship to what models can do.
  693. 23:59We should understand the key features of
  694. 24:01harnesses and and how they differ and
  695. 24:03not just like look at them as like
  696. 24:05you know, these blobs and one is a
  697. 24:07terminal and one is an IDE.
  698. 24:09But the biggest area is really this like
  699. 24:11how do we use them to our advantage,
  700. 24:14right?
  701. 24:15Yeah, how to apply this tool to our
  702. 24:18domain, which is software engineering.
  703. 24:21So with all of these things like models
  704. 24:22getting much more capable, harnesses
  705. 24:24getting much more capable, we're kind of
  706. 24:26like getting more sophisticated in a
  707. 24:27figuring out how do we use this, how do
  708. 24:29we provide our context. So this kind of
  709. 24:32like race continues, right? We want more
  710. 24:33agent autonomy and we want less human
  711. 24:36supervision. And as part of that also
  712. 24:38something that has happened over the
  713. 24:39last 12 months is that it has become a
  714. 24:41lot easier to run agents with
  715. 24:44no supervision, right? So this
  716. 24:47screenshot is from like I probably took
  717. 24:49it last year in June, July or something.
  718. 24:51That was the first version of Codex and
  719. 24:54over time now also most of the the
  720. 24:56harness products, the coding agent
  721. 24:57products have now come come out with
  722. 25:00like platforms where you can always
  723. 25:02decide like do I want to run this coding
  724. 25:03session locally on my machine or do I
  725. 25:05want to do it in the cloud? And so it's
  726. 25:07become a lot easier
  727. 25:10to do this and to to try this out,
  728. 25:13right? Like whatever size you feel
  729. 25:14comfortable size and complexity of tasks
  730. 25:17you want to do this for like some people
  731. 25:19run it to actually like build full
  732. 25:21features and others maybe like just dip
  733. 25:23their toes in like cleaning up this
  734. 25:25feature toggle or like small clean up
  735. 25:27tasks, right?
  736. 25:29And also in the meantime we've kind of
  737. 25:31like taking this like to more extremes,
  738. 25:34right? Which is more in experimental
  739. 25:35stage right now I would say
  740. 25:37which is this idea of like swarms or
  741. 25:39really brute force just sending lots and
  742. 25:41lots of agents out there and also having
  743. 25:44the agents decide how many agents they
  744. 25:46need, right?
  745. 25:48Gastown got a lot of attention in I
  746. 25:51think it came out in January. There were
  747. 25:53these two like big or even more
  748. 25:55experiments from like cursor for example
  749. 25:57or anthropic to build C compiler in the
  750. 25:59browser.
  751. 26:00Cloud flow I think it has a different
  752. 26:02name now that was probably even earlier
  753. 26:04than Gastown last year. Kind of people
  754. 26:06were playing around with that. So that's
  755. 26:08kind of like taking it to the extreme
  756. 26:10and seeing how we can push the
  757. 26:11boundaries and actually have AI build
  758. 26:13much bigger things more autonomously.
  759. 26:17Um
  760. 26:20Yeah, so this is our fourth year into
  761. 26:22this.
  762. 26:23>> [laughter]
  763. 26:24>> Um so we start with auto complete and we
  764. 26:26had a bit more like integration into the
  765. 26:28IDEs, more context. Cloud 3.5 Sonnet was
  766. 26:32like an early model moment I think where
  767. 26:35I certainly from that point on just just
  768. 26:38always almost always use Cloud Sonnet
  769. 26:40because it just like
  770. 26:41felt so much better at coding than the
  771. 26:43other ones. Um
  772. 26:45So then
  773. 26:46we got these like at the time most
  774. 26:48people are I also my presentations still
  775. 26:51call it agentic coding modes, right? So
  776. 26:53that's when like the the like like
  777. 26:57cursor and so on they got like these
  778. 26:58modes where they could also run terminal
  779. 27:00commands which which had been a thing
  780. 27:02that was already out there in open
  781. 27:03source, but not as widely used and FCP.
  782. 27:06So that was only about 1 and 1/2 years
  783. 27:07ago.
  784. 27:09Then shortly after that the vibe coding
  785. 27:12term got coined by
  786. 27:14Andre Karpathy which then led to like a
  787. 27:17lot of attention of people discovering
  788. 27:19these new agentic modes and going oh,
  789. 27:21this is like there was like a wave of
  790. 27:23like people picking this up again and
  791. 27:24saying oh, this is has actually improved
  792. 27:26quite a bit.
  793. 27:28Then we got these kind of background
  794. 27:29agents that I just talked about, right?
  795. 27:31Like so Codex for example, you know,
  796. 27:33allowing you to run things unsupervised
  797. 27:35in the background. Cloud code is
  798. 27:39I think generally available probably
  799. 27:41about a year old. It's like it started a
  800. 27:44little bit earlier than that. The
  801. 27:45context engineering term also started
  802. 27:47gaining traction about a year ago. Then
  803. 27:50we had that Cloud Opus moment. We got
  804. 27:52skills. Open Claw is maybe also relevant
  805. 27:56relevant moment even though it's not
  806. 27:58quite directly about coding.
  807. 28:00And then yes, so like kind of beginning
  808. 28:02of this year we got this like next wave
  809. 28:04of people paying attention again and
  810. 28:06going oh, this has actually changed
  811. 28:08quite a bit, right? I can actually see
  812. 28:10those two yellow spots in our internal
  813. 28:12AI coding chat community in in
  814. 28:15ThoughtWorks. I can see the spikes in
  815. 28:18activity and then kind of like holding
  816. 28:20and then there's like another like I
  817. 28:22don't know
  818. 28:23250 messages a week or something median.
  819. 28:28Yeah, and then so yeah, the gas town and
  820. 28:29all of the swarms and all of that
  821. 28:31started those big experiments started
  822. 28:32happening beginning of this year then
  823. 28:34the harness engineering term is now like
  824. 28:36a buzzword. So and you know, I don't
  825. 28:39have any additional boxes here because I
  826. 28:40think right now it's like us all like
  827. 28:43processing, right? And actually using
  828. 28:44all of these things that are happening.
  829. 28:45So there's been like I think not like a
  830. 28:47big coinage of anything or like a new
  831. 28:50buzzword other than these things, but
  832. 28:51we're just like grappling with it and
  833. 28:53and using this stuff, right?
  834. 28:58So, it's come a long way, but the costs
  835. 29:01have as well, and I don't just mean the
  836. 29:02token costs.
  837. 29:04So, one is security, right? So, our it's
  838. 29:07it's both about about like secrets
  839. 29:09potentially leaking from our
  840. 29:11environments, from our machines, but
  841. 29:12also our ecosystem is under attack,
  842. 29:15right? So, um
  843. 29:17we have to like think even more than
  844. 29:19before about dependency management,
  845. 29:20about sandboxing, and and stuff like
  846. 29:22that.
  847. 29:23Um on the other hand, there was also a
  848. 29:25good talk by Joseph from GitHub here
  849. 29:27yesterday about like how we can use AI
  850. 29:29to improve our security, right? So,
  851. 29:31there's usually both sides to it.
  852. 29:33Stability, right? So, in the Dora
  853. 29:35report, um
  854. 29:37this was one of the big like, let's say,
  855. 29:39negative trend findings about stability
  856. 29:42actually getting worse according to
  857. 29:43their um data, but also there were a lot
  858. 29:46of talks here about like how we can use
  859. 29:48AI to help us improve our stability.
  860. 29:51Um then changeability is a big thing,
  861. 29:54right? So, like
  862. 29:55code quality defined as code that
  863. 29:58remains easy to change and remains where
  864. 30:01remains easy to change it with low risk,
  865. 30:03right? So, this was a change uh I
  866. 30:06recently introduced into a still
  867. 30:08relatively new codebase that was all
  868. 30:10created with AI, and I made a change
  869. 30:13that touched 41 files, but it shouldn't
  870. 30:15have. It wasn't like that big of a deal,
  871. 30:17and so it was like a clear smell that
  872. 30:19there was already like accumulated tech
  873. 30:21debt that was making changes more risky
  874. 30:23and more
  875. 30:24uh costly. So, but then also here the
  876. 30:27question is like how far can we push
  877. 30:29this more and improve this more with
  878. 30:31guides, with sensors, with static code
  879. 30:33analysis, and so on.
  880. 30:35Token cost is the most obvious cost
  881. 30:37thing, of course. In the beginning of
  882. 30:382024, I I at a keynote presentation
  883. 30:41where the speaker said, "Generating 100
  884. 30:43lines of code only costs about 12 cents.
  885. 30:46Can you imagine?"
  886. 30:48>> [laughter]
  887. 30:48>> And he was comparing that to developer
  888. 30:50salaries. I mean, regardless of, you
  889. 30:52know, lines of code is not
  890. 30:54a measure of value, right? Let's just
  891. 30:56get that out of the way as well. But of
  892. 30:58course now we have like there was some
  893. 31:01numbers or some quotes in the Pragmatic
  894. 31:02Engineer newsletter recently where
  895. 31:04somebody for example, there were
  896. 31:05multiple of these quotes, but somebody
  897. 31:07said, "Some developers are now spending
  898. 31:09$500 a day." Which if we take that
  899. 31:12analogy of a developer salary is over
  900. 31:14$100,000 a year salary, which is a
  901. 31:16pretty decent salary in the even in the
  902. 31:18richest countries in the world, right?
  903. 31:21Um
  904. 31:22Another type of cost is like cognitive
  905. 31:24load and burnout, right? Like
  906. 31:27who would have thought, right? It
  907. 31:28doesn't actually make us make our lives
  908. 31:30more relaxed on the contrary, right?
  909. 31:32Like some people are working more even
  910. 31:34though they they already create more
  911. 31:36output. And Steve Yegge had this analogy
  912. 31:38with the energy vampire from What We Do
  913. 31:41in the Shadows. I don't know if any of
  914. 31:42you have seen that
  915. 31:44that show, but he basically yeah, he's a
  916. 31:46vampire that doesn't suck blood that but
  917. 31:48that sucks
  918. 31:49energy. So there's lots of stories about
  919. 31:51people saying, "Oh, I can only do this
  920. 31:52like 3 hours in a row and then I have to
  921. 31:55like take a nap."
  922. 31:57Um then we have the review crisis of
  923. 31:59course, right? We have like higher
  924. 32:02coding throughput, but then can we if we
  925. 32:05can code faster, can we review faster,
  926. 32:07test faster, ship faster, review faster?
  927. 32:10We've definitely so far the state is no.
  928. 32:12We cannot review faster, right?
  929. 32:13Everybody's complaining about this about
  930. 32:16this pain. Uh but it's not just coding.
  931. 32:19Everybody can create more with AI now,
  932. 32:21right? So I was talking to a colleague
  933. 32:23the other day who was telling me a story
  934. 32:24about an organization
  935. 32:27on this side of like if you can code
  936. 32:28faster, can you fill the backlog faster?
  937. 32:31So here, you know, there was like
  938. 32:33another
  939. 32:34another kind of like thing popping up
  940. 32:36before the coding where the the product
  941. 32:39managers were actually like churning out
  942. 32:41lots and lots of prototypes and lots and
  943. 32:43lots of ideas, right? So now there was
  944. 32:45this weird like bottleneck between this
  945. 32:47pile of prototypes and the pile of code
  946. 32:50because they couldn't get it like to
  947. 32:51sync up again and nobody really could
  948. 32:53figure out how to converge on what they
  949. 32:55actually wanted to build cuz they were
  950. 32:57doing this kind of like in these two
  951. 33:00silos. So that's maybe like something
  952. 33:03like an open question that we're already
  953. 33:05seeing a little bit but are we are we
  954. 33:07heading in general towards like a flow
  955. 33:09crisis, right? And whenever I want to
  956. 33:11understand flow better, I turn to my
  957. 33:13colleague James Lewis who like has done
  958. 33:15a lot of very interesting presentations
  959. 33:17about flow that you know it's easy to
  960. 33:20find on on YouTube. Um but here's one
  961. 33:23where he's talking about congestion
  962. 33:25collapse
  963. 33:27and so he has this
  964. 33:29this prediction apparently has a bet
  965. 33:31open with Gene Kim for a crate of beer
  966. 33:34that this will happen and that this will
  967. 33:36become like a big topic of conversation
  968. 33:38that we're just like overloading and
  969. 33:40also overloading in these different
  970. 33:41silos and that everything will just
  971. 33:43become super slow at some point and just
  972. 33:45collapse. So this comes you know from
  973. 33:48thing theory of constraints and and
  974. 33:50stuff like that and here he's he's
  975. 33:51quoting the Don Reinertsen book about
  976. 33:54the principles of product development
  977. 33:55flow
  978. 33:56as well.
  979. 33:59So then humans are seen as the
  980. 34:00bottleneck, right? That's what what
  981. 34:02multiple people also quoted or said here
  982. 34:06at the conference.
  983. 34:10So
  984. 34:11it remains this big question of like
  985. 34:13trust in the code and how much
  986. 34:14supervision do we want to have and how
  987. 34:17much review.
  988. 34:20And of course it depends, right? So the
  989. 34:22autonomy of AI coding agents is here.
  990. 34:25It's just unevenly distributed, right?
  991. 34:28Lots of you are probably already using
  992. 34:29it for some things, right? But I think
  993. 34:32it will never be that we can use AI
  994. 34:35coding agents for any type of task in
  995. 34:37any situation, right? It's just always
  996. 34:39it depends on the situation. And um
  997. 34:43what does it depend on, right? So, for
  998. 34:45me, the way I think about it right now
  999. 34:47is as this risk assessment uh out of
  1000. 34:49probability, impact, and detectability,
  1001. 34:52right? Which is very typical kind of
  1002. 34:54components of risk assessment in all
  1003. 34:55kinds of areas. So, first I think about
  1004. 34:57the probability that AI gets something
  1005. 34:59wrong or gets something right. And
  1006. 35:01that's all about me knowing the things
  1007. 35:03that I talked about before, my AI tool,
  1008. 35:06like what my context is, did I even give
  1009. 35:08the agent a chance to do it right? Um
  1010. 35:10and it's also about me reflecting on my
  1011. 35:12confidence in my requirements, right?
  1012. 35:14Like how [snorts] certain am I even that
  1013. 35:16I even know what to do? Um then I think
  1014. 35:18about the impact if AI gets something
  1015. 35:20wrong. So, that's all about the use case
  1016. 35:22criticality, of course. So, is this like
  1017. 35:25a super critical business flow that uh
  1018. 35:28you know, will wake me up at 2:00 a.m.
  1019. 35:29on Saturday because I'm on call? Or is
  1020. 35:32this something a lot less uh
  1021. 35:34a lot less important? Then maybe like
  1022. 35:36I'm a little bit more loose with uh how
  1023. 35:39I review. And the third thing is I
  1024. 35:40reflect on detectability that AI got
  1025. 35:43something wrong. Will I notice, right?
  1026. 35:45And um
  1027. 35:46by the way, all of this starts with
  1028. 35:48knowing what right and wrong means,
  1029. 35:49right? I should So, and and often it's
  1030. 35:52like is it appropriate, right?
  1031. 35:54Uh so, you have to know your uh feedback
  1032. 35:56loops, basically. And then based on
  1033. 35:58these things, I decide which workflow do
  1034. 36:00I use, how much review do I do, and how
  1035. 36:03long let do I let it go without
  1036. 36:04supervision. For example, if I don't
  1037. 36:07even know myself yet quite what I need,
  1038. 36:09then I won't let it go off for like half
  1039. 36:11an hour and then realize that it was all
  1040. 36:13for nothing, right?
  1041. 36:16And yeah, so you have to kind of be this
  1042. 36:17tall to ride the roller coaster. You
  1043. 36:19have to be this tall to reduce
  1044. 36:20supervision, right? So, you can think
  1045. 36:22about uh you know, feedback loop in your
  1046. 36:24team, in your organization. And you can
  1047. 36:26also increase the probabilities by
  1048. 36:29improving your context engineering, your
  1049. 36:31harness engineering,
  1050. 36:33um and also
  1051. 36:35refactoring and modernizing and you
  1052. 36:37know, all of those things because
  1053. 36:40AI can also deal with a well-factored
  1054. 36:42code base much better
  1055. 36:45than with a messy code base.
  1056. 36:48So, we're tempted to move from in the
  1057. 36:49loop to on the loop to out of the loop.
  1058. 36:52There's I I feel this every day being
  1059. 36:54drawn to like, "Ugh, I don't want to
  1060. 36:56look at this. I don't want to look at
  1061. 36:57the code anymore." It's like
  1062. 36:59Yeah, there's all of these forces that
  1063. 37:01pull us there, but we're also starting
  1064. 37:03to actually feel the costs. It's not
  1065. 37:04just speculation anymore. It's not just
  1066. 37:07like dooming kind of predictions. We're
  1067. 37:09actually feeling the cost of tokens,
  1068. 37:11risks, and cognitive X, right? Cognitive
  1069. 37:14load, cognitive debt, right? This idea
  1070. 37:18of like we don't even understand anymore
  1071. 37:19how our code base is structured.
  1072. 37:22I sometimes think of like cognitive
  1073. 37:23deferral as well. It feels like we keep
  1074. 37:25like deferring the review to other
  1075. 37:28people or like deferring processing what
  1076. 37:31actually happened. And recently there
  1077. 37:34was like a new cognitive X term coined
  1078. 37:37that
  1079. 37:38I I saw because Adi Osmany wrote about
  1080. 37:40it, which is cognitive surrender, right?
  1081. 37:43So, in this paper they put AI into the
  1082. 37:46context of the system one system two
  1083. 37:50thinking fast and slow. I think probably
  1084. 37:52lots of you have heard about that book.
  1085. 37:53If not, then look it up. It's like
  1086. 37:55really great. And so, they are they're
  1087. 37:57talking about a mode where we're
  1088. 37:58basically displacing system two like our
  1089. 38:01actually like active thinking with AI
  1090. 38:04and they call it cognitive surrender.
  1091. 38:06And apart from like this paper and the
  1092. 38:08cognitive part of it, this term like
  1093. 38:11surrender has just been stuck in my head
  1094. 38:12like ever since I heard this. And I feel
  1095. 38:15like we're There's like so many things
  1096. 38:17where we we're in danger of just like
  1097. 38:19surrendering right now, not just
  1098. 38:21cognitively.
  1099. 38:22Um and so I think we have to be careful
  1100. 38:25about what we are surrendering, right?
  1101. 38:26And like really think about be mindful
  1102. 38:29of
  1103. 38:30where that's worth doing. So I'll just
  1104. 38:32give like some examples, right? Um
  1105. 38:35So it's like ah it's too much to reason
  1106. 38:37about this big change. It'll be fine,
  1107. 38:39right?
  1108. 38:40I do it myself myself. I'll just do it
  1109. 38:42myself quickly instead of teaching
  1110. 38:43somebody, right? Everybody's talking
  1111. 38:45about oh how will juniors learn? How
  1112. 38:47will we do this? But I don't see that
  1113. 38:49much like
  1114. 38:50active action, right? Because
  1115. 38:52everybody's just in a tunnel of you know
  1116. 38:54I see experienced people building tools
  1117. 38:56for themselves to use AI better, but
  1118. 38:58like why are we not taking more
  1119. 39:00initiative to think about how this is
  1120. 39:03sustainable for people who don't have
  1121. 39:04all of this experience already?
  1122. 39:07Or this like ah I'll just use a big
  1123. 39:08model, you know. I can't be bothered to
  1124. 39:11then like try retry it again, so let's
  1125. 39:13just use the most expensive biggest
  1126. 39:14model.
  1127. 39:15Or this surrendering to like ah I don't
  1128. 39:17want to solve all these problems. Models
  1129. 39:19will get better. Tokens will be cheaper
  1130. 39:20again. We just have to wait it out,
  1131. 39:23right? Um we're working more, we're
  1132. 39:26producing more, but we're still getting
  1133. 39:28the same compensation.
  1134. 39:30Is that also a type of surrender?
  1135. 39:32Um or this like ah I don't have time to
  1136. 39:34find better approaches, which is by the
  1137. 39:36way not necessarily the individual's
  1138. 39:38fault. It's also like what's happening
  1139. 39:40around them and incentives and
  1140. 39:41pressures, right? Uh sandboxing is too
  1141. 39:44tedious. Surely nothing will go wrong,
  1142. 39:46right?
  1143. 39:48Or I I can't work on this code base
  1144. 39:50without AI anymore, right? That's the
  1145. 39:51the original cognitive surrender
  1146. 39:53cognitive debt kind of definition.
  1147. 39:56Um this was like Hannah yesterday in a
  1148. 39:58talk and she was kind of talking about
  1149. 39:59what she's keeping, what she's trashing,
  1150. 40:02and what she's trying. And I think
  1151. 40:03that's similar in terms of like we have
  1152. 40:05to think about what we're surrendering
  1153. 40:07and what we should be uh what we should
  1154. 40:09be keeping. So if you are a person of
  1155. 40:11influence in your engineering
  1156. 40:12organization, are you creating an
  1157. 40:14environment that leads to surrender,
  1158. 40:16right? That makes me people feel like
  1159. 40:18they just have to crank out the PRs and
  1160. 40:20don't have time to like actually figure
  1161. 40:22out how to improve the environment, how
  1162. 40:24to improve the context engineering, and
  1163. 40:26so on. Um if you're if you maybe feel a
  1164. 40:28bit like powerless or you feel like you
  1165. 40:30can't influence it and there's all these
  1166. 40:32pressures around you, still like try to
  1167. 40:34think about your sphere of influence and
  1168. 40:35the small things you can do. Like what
  1169. 40:37do you really have to surrender, right?
  1170. 40:39Um so there's lots of options between
  1171. 40:42not using AI at all and like just total
  1172. 40:45surrender and hope and hoping the models
  1173. 40:46will fix it all in the future.
  1174. 40:49And again, like if you feel kind of like
  1175. 40:51powerless and like this is like washing
  1176. 40:52over you, we can also work together and
  1177. 40:54collaborate on these things, right? So
  1178. 40:56if you feel like you're not a good
  1179. 40:57communicator about this, maybe look for
  1180. 41:00somebody in your team or organization
  1181. 41:01who's good at that and share your data,
  1182. 41:03share your observations with them,
  1183. 41:04right? So like we can collectively do a
  1184. 41:07lot about this as well.
  1185. 41:09In terms of skills, we need to know our
  1186. 41:10toolbox, past and present. We also need
  1187. 41:12to rediscover some things.
  1188. 41:14Uh we need critical thinking, risk
  1189. 41:16assessment, and some form of like
  1190. 41:18patience, right? So we need to evaluate
  1191. 41:21that productivity maybe doesn't just
  1192. 41:22mean like typing. And for leadership,
  1193. 41:25you know, this is still horizon two,
  1194. 41:27right? This is not horizon one yet. So
  1195. 41:29maybe we don't know the ROI yet. Maybe
  1196. 41:32we just have to like give people some
  1197. 41:33time to um
  1198. 41:35to create a good setup so that we can
  1199. 41:37continue to safely and quickly deliver
  1200. 41:41software to users in a sustainable way.
  1201. 41:45Thank you.
  1202. 41:47>> [applause]
  1203. 41:49[music]

About this transcript

This page contains the full transcript of Birgitta Böckeler - State of Play: AI Coding Assistants - AI Native DevCon June 2026 by AI Native Dev, generated from the public captions YouTube serves with the video. The transcript has 8,038 words across 1,203 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.