YouTube2Text

Guy Podjarny - Spec Driven Dev From Single Player to Multiplayer to Ecosystem | DevCon Fall 2025 — Transcript

by AI Native Dev · 6,288 words · 876 segments · language en · Watch on YouTube

Full transcript

  1. 0:09Now, our next speaker is the uh is the
  2. 0:12uh founder and CEO of a little company
  3. 0:15called Tessle. Uh 18-month old company.
  4. 0:18Um GIO previously is he's a serial
  5. 0:21founder. He he he founded Blaze, which
  6. 0:23was acquired. He founded Sneak, a
  7. 0:25massively successful developer security
  8. 0:28uh company and organization that
  9. 0:29actually when we think about it, a lot
  10. 0:31of the challenges that that had in and
  11. 0:33around moving the the thought process,
  12. 0:36moving the moving the mindset of people
  13. 0:38to a developer security path is actually
  14. 0:41very similar to how we're thinking today
  15. 0:43about A&E development. And now we're
  16. 0:46going to talk about how AI can be done
  17. 0:48collaboratively. So please give it up
  18. 0:50for our first keynote speaker, Guy
  19. 0:51Pjani.
  20. 0:53>> [applause]
  21. 0:56[music]
  22. 1:03>> Hello everyone. Uh, welcome to the first
  23. 1:06inerson and native DevCon. Uh, love in
  24. 1:08person events, you know, like I think
  25. 1:10you you you convey things virtually. I
  26. 1:11consume things virtually, but nothing
  27. 1:13beats, you know, talking to people and
  28. 1:14engaging with them. And I hope you're
  29. 1:15all feeling that way. So, thanks for
  30. 1:17shipping in, especially in the rain here
  31. 1:19in New York. and uh uh and I hope you
  32. 1:21you enjoy the event and I look forward
  33. 1:23to meeting and learning from many of
  34. 1:24you. Uh I am Gapari or I am the
  35. 1:27founder and CEO of TESL. Uh and I'm here
  36. 1:30to talk about specdriven development's
  37. 1:32growth uh from kind of single player to
  38. 1:35really how do you use it in a team and
  39. 1:37how do you think about ecosystems? Um
  40. 1:39let me set the stage first. So hopefully
  41. 1:43we're all aligned that AI is
  42. 1:44transforming software development. It's
  43. 1:46not a little hiccup. It's not a little
  44. 1:47change. It's something that is
  45. 1:48substantial that is happening and we've
  46. 1:51seen a bunch of evolution already. We've
  47. 1:53had two kind of waves of what I think of
  48. 1:55as AI augmented software development uh
  49. 1:58with tools that were pioneered by
  50. 2:00copilot and cursor that autocomplete
  51. 2:02things better you know that that helped
  52. 2:04a lot that build in chat and allow you
  53. 2:06to talk to your code and understand
  54. 2:08things that cursor pioneered and then we
  55. 2:10uh in this last sort of especially 6
  56. 2:12months maybe 12 we've seen the rise of
  57. 2:14more true AI native which I think of
  58. 2:17them as more AI native dev tools because
  59. 2:19they're more about delegation they're
  60. 2:21more uh ask the agent to do a thing for
  61. 2:24you and then interact with it. So, it's
  62. 2:26a it's a much more substantial change
  63. 2:28for us and there's a lot to learn. And
  64. 2:31while all of these tools are are
  65. 2:32valuable in their own right, they've all
  66. 2:34kind of consolidated towards this
  67. 2:36agentic behavior. Uh and I think it's
  68. 2:38pretty clear that that is the future.
  69. 2:40So, I'll focus on where is that taking
  70. 2:43us? You know, how do we make this
  71. 2:44agentic development work?
  72. 2:47So, the reality today is that agents are
  73. 2:50are super powerful, uh, but they're
  74. 2:52massively unreliable, right? They're
  75. 2:53sort of sometimes spooky, sometimes
  76. 2:54kooky. You know, they're they're uh
  77. 2:56they're they're amazing and they do
  78. 2:58these things that astound you, but then
  79. 2:59they just make these foolish mistakes.
  80. 3:01Um, they're also very untrustworthy.
  81. 3:03They'll tell you something is done when
  82. 3:05it isn't really. They will break things
  83. 3:06along the way as they, you know, you ask
  84. 3:08them to build a feature, they'll burn
  85. 3:09down the house to to get it done. Uh,
  86. 3:11and so they're very powerful, uh, but
  87. 3:14they're kind of hard to use. Um, and
  88. 3:17this creates this capability reliability
  89. 3:19gap, you know, and you see studies like
  90. 3:21this one from meter where they they
  91. 3:25assess, you know, how how much do
  92. 3:27developers think AI tools will make them
  93. 3:30better. And you see this is the bar at
  94. 3:32the left and people developers, well, I
  95. 3:34think it's going to make me faster. I'll
  96. 3:35complete something faster. And then they
  97. 3:36observe how long it actually takes them
  98. 3:38to do it. And it turns out it actually
  99. 3:40was slower. There's a lot of caveats, a
  100. 3:42lot of flaws in this study and in any of
  101. 3:44the others. It's a very complex area,
  102. 3:46but I think everybody relates to the
  103. 3:48fact that it's not easy to turn this
  104. 3:51power into improved productivity into
  105. 3:54something that actually works.
  106. 3:57And and so what do we do about it? Well,
  107. 4:00what we can't do about it is say, you
  108. 4:02know, bury your head in the sand like
  109. 4:04let's not use agents is tempting. And,
  110. 4:06you know, a good number of people,
  111. 4:08probably none in this room are uh are
  112. 4:10saying that. Uh but you know it's
  113. 4:13clearly shortsighted. I mean this is a
  114. 4:15technology that's too powerful that uh
  115. 4:18to to kind of ignore or think it would
  116. 4:20fade away. It's here. It's going to
  117. 4:22shape our development. And we need to
  118. 4:24think about how do we embrace it? Not
  119. 4:26how do we ignore it? And so the real
  120. 4:28question is how do we help agents? How
  121. 4:30do we help agents succeed and close this
  122. 4:33capability reliability gap?
  123. 4:36Let's look a little bit at a at a
  124. 4:38history of attempting to do so. You
  125. 4:39know, this is uh our quest for silver
  126. 4:41bullets that would make LLMs work uh the
  127. 4:44way we want, right? And a bit of a loose
  128. 4:46timeline. We start off with, you know,
  129. 4:48this is us as the industry thinking
  130. 4:50about fine-tuning. Hey, we're going to
  131. 4:51tell the LLM, here are like tens,
  132. 4:53hundreds, maybe thousands of examples of
  133. 4:56how I like to code, how you code in my
  134. 4:58environment on it, and the LLM's weights
  135. 5:00will magically adapt uh and work like
  136. 5:03that. And it turns out that this works
  137. 5:05very well for brand new skills, for
  138. 5:07cases in which you're teaching the LLM
  139. 5:09to do something that it didn't know
  140. 5:10before. But if it already has opinions,
  141. 5:12if it already knows how to code, it's
  142. 5:14very hard to sway it. It's kind of like
  143. 5:15people, you know, we if if can't teach
  144. 5:17an old dog new tricks, right? It's hard
  145. 5:19to sway someone will behave. And so
  146. 5:22fine-tuning is useful, but it's not a
  147. 5:24silver bullet. And in fact, in many
  148. 5:25cases, uh it is it is hard to to really
  149. 5:28have it shine.
  150. 5:31The second kind of wave of technology of
  151. 5:33this will solve our problems was was
  152. 5:35rag. It's like I will take my prompt and
  153. 5:38I will put it into this magical kind of
  154. 5:40embedding algorithms and vector
  155. 5:41databases will pull out the right
  156. 5:43context that I need everything precisely
  157. 5:45that I need to know to be able to
  158. 5:47address this problem. And again rag very
  159. 5:49powerful when you have your own keywords
  160. 5:51and terminology and things like that but
  161. 5:54uh context gathering is more complicated
  162. 5:56than that. uh there's there's a lot of
  163. 5:58subtlety often time things that are not
  164. 5:59in your prompt they're not in that
  165. 6:01initial path and so it remained a useful
  166. 6:03tool but was deemed not uh not a silver
  167. 6:06bullet
  168. 6:08um the next kind of this will save us
  169. 6:10all was the world of big context you
  170. 6:11know what we don't need to worry about
  171. 6:12all this stuff you know we've got like
  172. 6:14million token 2 million token context
  173. 6:16windows we'll just tell it everything uh
  174. 6:19and it will figure it out it'll just
  175. 6:20sort of know what to do and it turns out
  176. 6:22you know like I often get uh the like
  177. 6:25through my career I've gotten the often
  178. 6:26times the guidance that says if you have
  179. 6:2810 smart things to say and you say all
  180. 6:30of them each of them will get less
  181. 6:32attention than if you said three of them
  182. 6:34uh and and I still struggle to not say
  183. 6:37everything I think but uh it but it's
  184. 6:39true you know if you say fewer things
  185. 6:41each of them gets more attention and
  186. 6:42NLMs are the same they if you if you
  187. 6:45give it more information in its context
  188. 6:46each of these bits would get less
  189. 6:48attention it would be less focused and
  190. 6:50we'll show a little bit of that in a sec
  191. 6:53and so we get into modern world uh and
  192. 6:55we have these two additional tools. One
  193. 6:57is agentic search. You know what? Let's
  194. 7:00delegate the actual problem of context
  195. 7:02creation to the agent. Let's give it
  196. 7:04tools to find the information. That's
  197. 7:06what we do as humans. We go, we Google,
  198. 7:08we look at files, we uh we do that work.
  199. 7:10And that is useful, but again, it's
  200. 7:12useful when it's constrained. If you
  201. 7:15tell someone just just go find, you
  202. 7:17know, the information about anything,
  203. 7:19it'll take time to find it. It'll take
  204. 7:22uh you might find the wrong information.
  205. 7:24uh and you know it's a in it's both
  206. 7:26inefficient and unbounded and you can
  207. 7:28get down rabbit holes that really kind
  208. 7:30of make things worse and so it's a very
  209. 7:32very powerful tool. We'll refer to it
  210. 7:34back again but again it didn't just
  211. 7:35solve the problem because there's
  212. 7:37another piece that we need on top of
  213. 7:38that and that is context engineering and
  214. 7:40that's the sort of the state of affairs
  215. 7:42today which is you actually have to give
  216. 7:46up on the magic solutions and think
  217. 7:48about what do you want to tell the
  218. 7:50agent? what information does the agent
  219. 7:52need to succeed? And I love uh this is
  220. 7:55uh Toby, the uh founder and CEO of
  221. 7:57Shopify. He's a very serial guy, super
  222. 8:00smart guy, and he was on the podcast uh
  223. 8:02uh a couple of months ago. And I like
  224. 8:04his quotes here, which is I like the
  225. 8:05term context engineering because I think
  226. 8:07the fundamental skill of using AI well
  227. 8:10is to be able to, and this is the key
  228. 8:11word, a state of problem with enough
  229. 8:14context in such a way that without any
  230. 8:16additional pieces of information, the
  231. 8:18task is plausibly solvable. I love that
  232. 8:20it's this notion of like let it lean
  233. 8:23into intelligence but don't assume mind
  234. 8:25readading you know like tell it what it
  235. 8:27needs to know so that it can succeed.
  236. 8:31So context engineering in my mind and I
  237. 8:34might be biased is basically the same as
  238. 8:36specs. You know this is about we just
  239. 8:38had this long conversation about you
  240. 8:39know should we call something spec or
  241. 8:41context. uh it doesn't really matter
  242. 8:42like for me it is about getting out of
  243. 8:45your head and out of your knowledge
  244. 8:46getting the the nuggets of knowledge
  245. 8:48that you need to uh give the agent to
  246. 8:50succeed and so I'm going to use the term
  247. 8:52context and specs interchangeably
  248. 8:54throughout this talk that's the way I
  249. 8:55see the world
  250. 8:58so I'm going to go into that sort of
  251. 8:59single multi-ecosystem player but before
  252. 9:02that I want to remind us that you can't
  253. 9:04optimize what you can't measure this is
  254. 9:06a known phrase in the world of software
  255. 9:08development but really mostly you hear
  256. 9:10it about ops because in ops we are
  257. 9:13familiar with these systems that
  258. 9:14sometimes behave well and sometimes
  259. 9:16don't right the the servers sometimes
  260. 9:17work sometimes don't sometimes they
  261. 9:19crash and so we have to measure them so
  262. 9:20we know what's happening so we can
  263. 9:21optimize it now we have the same
  264. 9:23statistical behavior in our development
  265. 9:26you know we have these statistical
  266. 9:27creatures agents they build stuff
  267. 9:29sometimes they work sometimes they don't
  268. 9:31and that means two things one we need to
  269. 9:33leave the notion of agents can
  270. 9:35absolutely do this or cannot absolutely
  271. 9:38and embrace a statistical number around
  272. 9:40their success uccess rate and two is we
  273. 9:42have to measure that we have to think
  274. 9:43about how do we evaluate how do we
  275. 9:44assess success and I'll try to give some
  276. 9:46data to back some of my statements in
  277. 9:47this talk
  278. 9:50so now we'll actually get to the thing
  279. 9:51that is on the title which is let's talk
  280. 9:53about uh single player how do you do
  281. 9:55context engineering in a single player
  282. 9:57context so um just to demonstrate a
  283. 10:01little bit the the base this is a a um a
  284. 10:04little sample to-do application with a
  285. 10:06Jurassic theme that I worked very hard
  286. 10:08on uh to just sort of demonstrate
  287. 10:10And the base context that an agent has
  288. 10:13is your project. Here I'm asking the
  289. 10:16agent to add an edit button to each of
  290. 10:18the to-do items in my list. And you'll
  291. 10:20see the agent automatically go off and
  292. 10:22read files in the project. It doesn't
  293. 10:24read all of them. It's already doing a
  294. 10:26gentic search to load uh some of that
  295. 10:29content. Uh and then it would load that
  296. 10:32code, understand that code, and uh
  297. 10:36sorry, I'll just sort of pause here. uh
  298. 10:38and um uh it it can load loads that code
  299. 10:41and reads that into context. So, I
  300. 10:43didn't need to tell it what the to-do
  301. 10:44item is. I didn't need to tell it where
  302. 10:46to put that button. It had enough
  303. 10:48information in the code and this is the
  304. 10:50base context that's important. However,
  305. 10:53you'll notice for instance that it
  306. 10:54created this blue edit button and that
  307. 10:56doesn't fit my beautiful Jurassic theme.
  308. 10:58Uh and it turns out that in the code
  309. 11:01some information is not there. It
  310. 11:02doesn't know that I want to stick to my
  311. 11:05theme colors. And so let's tell it to do
  312. 11:08that. So we can add explicit context
  313. 11:11that doesn't rely on agentic search like
  314. 11:14an agent MD file that says always stick
  315. 11:17to the theme colors used British
  316. 11:19spelling and it will go off and I will
  317. 11:22ask the same action you know I will say
  318. 11:23add an edit button and it will go off.
  319. 11:25it'll do the same type of like dynamic
  320. 11:27loading and uh maybe I'll sort of spare
  321. 11:30us a moment here uh as we get to the end
  322. 11:33you will see that it indeed created uh
  323. 11:35the button with my beautiful theme
  324. 11:37colors uh and and so this is this is 101
  325. 11:40of context engineering at at the base
  326. 11:42your project's context your project's
  327. 11:44code is the agent's context how well
  328. 11:47documented it is whether it is confusing
  329. 11:48and has references to old data
  330. 11:50structures that you have uh how um uh
  331. 11:53reasonable are your file names all of
  332. 11:55Those things help the agent load the
  333. 11:57right files, but then there's going to
  334. 11:59be some information that's not in the
  335. 12:00code and you put in your explicit agents
  336. 12:02MD, claude MD, cursor rules, local
  337. 12:05context. You probably know much of this
  338. 12:07already. The challenge is this is nice
  339. 12:10when you want to tell it stick to the
  340. 12:12theme colors, but typically you have a
  341. 12:14lot more to say. You there's just a lot
  342. 12:16that you want to tell it uh about how to
  343. 12:18work in your environment, in your style.
  344. 12:21And as we talked about before, if you
  345. 12:24tell it 10 things, it's going to get,
  346. 12:26you know, give each of them less
  347. 12:27attention than if you told it three. Uh,
  348. 12:29and so this is uh a problem. Before I
  349. 12:32talk about how to solve it, let's maybe
  350. 12:34back that claim that this is a problem
  351. 12:36with some data. So to do that, we took a
  352. 12:40little uh game, the team took a little
  353. 12:42sort of dodge the blocks game uh that
  354. 12:44that was vivecoded and asked uh the
  355. 12:47asked an agent in this case Claude code.
  356. 12:49Actually we tried three agents say add a
  357. 12:51login to this like add sessionbacked
  358. 12:53authentication so this game will be
  359. 12:55behind the login. The other thing we did
  360. 12:56is we created again with agents uh an
  361. 12:59evaluation scorecard like how good is
  362. 13:01the security of this thing that you've
  363. 13:04just added just like a benchmark that
  364. 13:06are standards today for good rubrics and
  365. 13:08then we ran that in three different
  366. 13:10modes. One is we ran the agents without
  367. 13:13any context and we scored them. The
  368. 13:15second is we put into the agents MD or
  369. 13:17the equivalent uh the OASP the open web
  370. 13:19application security project uh guidance
  371. 13:22around authentication about 3 kilobytes
  372. 13:24of data and the second is the longer 20
  373. 13:27kilobytes of full security guidance that
  374. 13:29includes the 3 kilobytes there's no word
  375. 13:31that didn't exist in the second one and
  376. 13:34indeed as we ran this we saw and I'm
  377. 13:37just sort of you know uh this is no
  378. 13:39surprise that without any guidance the
  379. 13:43agent scored a 65 85% score and with the
  380. 13:47guidance with the precise guidance it
  381. 13:48got 85% score and if we told it more
  382. 13:51things it got a lower score it got 81.
  383. 13:54So this was good and it's already
  384. 13:56validation. We need to know that even
  385. 13:57though we expect this behavior now we
  386. 13:59have a bit more confidence and there are
  387. 14:01many examples of this but also we find
  388. 14:03information that is maybe a bit
  389. 14:04unexpected. For instance we ran this on
  390. 14:08three different agents. We ran it on
  391. 14:10claude on codeex and on cursor and what
  392. 14:13you can see is they all fared about the
  393. 14:16same mediocre results without any
  394. 14:18guidance. There are many pitfalls that
  395. 14:19you can have from a security perspective
  396. 14:21when you add a login page, but there's a
  397. 14:24big difference in how they handle the
  398. 14:25deeper context. Claude shined when you
  399. 14:27gave it precise instructions, but
  400. 14:28dropped a fair bit uh when you gave it
  401. 14:30the longer instruction. Codeex didn't do
  402. 14:32as well in total, but dealt with the
  403. 14:34longer context a lot more. And it's
  404. 14:37interesting because it's not obvious to
  405. 14:39people that not all agents listen the
  406. 14:42same. You we told them the same words.
  407. 14:44Some of these actually use the same
  408. 14:45model behind them, but agents process
  409. 14:48and make many decisions. And you know,
  410. 14:51granted, there are 10 runs behind each
  411. 14:52of these columns. So maybe there's also
  412. 14:54not enough data for statistical
  413. 14:55significance, but the results are
  414. 14:57substantially different. And you have to
  415. 14:58think not just about what is the
  416. 14:59question you're asking and the
  417. 15:00saturation, but also who are you asking
  418. 15:02it of?
  419. 15:05So, so we have this challenge. We want
  420. 15:07to tell agents a lot of things and they
  421. 15:10uh you know, we can't because it's going
  422. 15:12to deteriorate their performance. what
  423. 15:13do we do? So, it's a broader problem and
  424. 15:16there are a lot there's a lot to say
  425. 15:17about it, but one common pattern is the
  426. 15:19separation between rules and knowledge.
  427. 15:22Rules would be a small amount of data
  428. 15:24that you're providing it that you're
  429. 15:25kind of shoving down their throat.
  430. 15:27You're saying this will be in your
  431. 15:28agent. going to see this every single
  432. 15:30time and they have to be succinct and
  433. 15:32what they should do is they should
  434. 15:34include links and references to guide
  435. 15:36agentic search around where can it fetch
  436. 15:38additional information and that is the
  437. 15:40knowledge and I'll give an example from
  438. 15:42Tesla's content. So hopefully many
  439. 15:45people on this uh uh in this event uh
  440. 15:48here or virtually know that Tesla has a
  441. 15:50spec registry and we it's a a dependency
  442. 15:54system for knowledge. you can publish
  443. 15:56context into it in this form we call
  444. 15:58tiles and then you can consume that down
  445. 16:00to your uh to your consumers right to
  446. 16:03your agents uh and it will adapt it to
  447. 16:06the local agent and you can do it for
  448. 16:07your context but we also pre-populated
  449. 16:10the registry with over 10,000 specs or
  450. 16:14context uh that uh that that helps
  451. 16:17agents use open source libraries better.
  452. 16:19We've analyzed those libraries, we've
  453. 16:21iterated, we've evaluated and we created
  454. 16:23those. So let's let's look a little bit
  455. 16:25at what is the structure of those. Um in
  456. 16:29first of all we define a knowledge base.
  457. 16:31So well what is the knowledge that you
  458. 16:33want? So you do that in this tessle.json
  459. 16:35file. In our example there are
  460. 16:36variations of it we'll touch uh to say
  461. 16:39this is the information I want. In this
  462. 16:40case I have a bunch of libraries. I want
  463. 16:42to know how to use them. And then you'll
  464. 16:44see that each one of those is fragmented
  465. 16:46into pieces. So the agent doesn't need
  466. 16:49to suck it all up right away. Especially
  467. 16:51for some of these libraries, there's a
  468. 16:52lot to say. You don't want to hog the
  469. 16:54context window. So there's a bunch of
  470. 16:56different files.
  471. 16:58The second thing we would do is we will
  472. 17:00put a little bit of knowledge or a
  473. 17:02little bit of rules in the agents MD
  474. 17:04file and we'll say here's a link to your
  475. 17:06knowledge. Uh we actually it's a bit
  476. 17:08longer over here but we also give a
  477. 17:10little bit of instruction of uh what is
  478. 17:12the structure of the knowledge you might
  479. 17:13find to the agent but it's a very small
  480. 17:16amount of context that is in the roles
  481. 17:17and always loaded. And then if you click
  482. 17:20that knowledgemd file, you will get a
  483. 17:22markdown file that the agent can easily
  484. 17:24find out which libraries do you have a
  485. 17:26little bit of information about them and
  486. 17:28where can it learn more and then down
  487. 17:30the rabbit hole it can learn more. It
  488. 17:32has an index file it explains about in
  489. 17:34this case next.js and then it can in
  490. 17:38that file find links to additional
  491. 17:39information. So this is like a human
  492. 17:41clicking links on a website
  493. 17:43[clears throat] right you have a summary
  494. 17:44thing you're expanding a section or
  495. 17:46you're clicking through and it's very
  496. 17:47effective. This is a good way to give
  497. 17:50the agent uh a way to find the right
  498. 17:53information. Lean into its agentic
  499. 17:55search capabilities, but don't ask it to
  500. 17:58sort of browse the web and just figure
  501. 17:59it out.
  502. 18:02So hopefully at this point you're
  503. 18:03saying, well, show me some data. I
  504. 18:06realized as I was building it that uh
  505. 18:07this is a Jaring Maguire reference and
  506. 18:09not everybody knows that. It might be
  507. 18:11like a my age is starting to defer of
  508. 18:13it. Like it does say show me the money
  509. 18:15typically. Um so okay can we see some
  510. 18:18data right I make this claim that this
  511. 18:20is a good way and a good effective way
  512. 18:21to to make the agents work. Let's see
  513. 18:23some data. I'll show two examples of
  514. 18:25this. First is uh thanks to the great
  515. 18:28work that uh the team at Versell has
  516. 18:31done on Nex.js. So Nex.js JS is a very
  517. 18:33well uh wellused very popular JavaScript
  518. 18:36framework that is maintained and built
  519. 18:39today by the Verscell team and they
  520. 18:41published as a thought leader in the AI
  521. 18:43space uh a benchmark about that tries to
  522. 18:46measure how well can agents use Nex.js
  523. 18:49and they created about 50 different test
  524. 18:51cases uh and uh they ran it through
  525. 18:54different agents to see how well they
  526. 18:56can address those. It's interesting, by
  527. 18:59the way, immediately to see that even
  528. 19:01though Nex.js is so popular and has so
  529. 19:04much content, it's like the sweetheart
  530. 19:05of the training data, right? It's well
  531. 19:07doumented. It's an amazingly like clear
  532. 19:10set of data. The top agents, the ones
  533. 19:13that did the best, still only got a 42%
  534. 19:15score in this benchmark. And so that's
  535. 19:17already interesting and a bit sobering.
  536. 19:19But it goes to show, okay, can we
  537. 19:21improve this with context engineering?
  538. 19:22and in general that if you measure this
  539. 19:25now I bet you all the agents are now
  540. 19:27working on optimizing it and so our
  541. 19:30question was okay we've done this can we
  542. 19:32optimize this with context can we test
  543. 19:35our uh uh our our knowledge like these
  544. 19:38tiles that we create do they actually
  545. 19:39make it better and we ran this eval
  546. 19:42because we already had the eval so we
  547. 19:43just had to compare the two test cases
  548. 19:45and we're happy to see that yes yes yes
  549. 19:47we can with the knowledge the success
  550. 19:50rate like for us it was about 40% % in
  551. 19:53our uh uh test environment. Success rate
  552. 19:55of using SJS on this benchmark, it
  553. 19:58jumped up to 92% if we got that info.
  554. 20:00This is if we got you know the very
  555. 20:02thoroughly created piece of information.
  556. 20:04But we also find for instance that a
  557. 20:06smaller prompt midway that just guides
  558. 20:08the agent to look at the things that you
  559. 20:10care about. Uh in this case, some
  560. 20:13linting and some build errors actually
  561. 20:15also bumped up the process
  562. 20:16substantially. So sometimes it's
  563. 20:17knowledge and sometimes it's steering
  564. 20:20around what is it to pay attention to.
  565. 20:24This is great for next.js
  566. 20:25[clears throat] but this was a very
  567. 20:26thorough benchmark set that required a
  568. 20:29lot of attention. We can't do that for
  569. 20:31everything. But what you can do when you
  570. 20:32evaluate your content or when you have
  571. 20:34something that is uh you don't have the
  572. 20:36sort of the bandwidth to create a deep
  573. 20:38benchmark for is that you can use LLMs
  574. 20:40to create evaluation data. And so we
  575. 20:43took 270 random libraries and we asked
  576. 20:46the LLM to create test exercises for
  577. 20:49them, a scorecard to evaluate how well
  578. 20:52they did on that test and then we ran it
  579. 20:54through agents. The data I have here is
  580. 20:56for cloud code. We saw similar data for
  581. 20:58cursor and then eventually you ask an
  582. 21:00LLM to judge the result with that
  583. 21:02rubric. And indeed when we ran that we
  584. 21:04saw that on average across those 270 we
  585. 21:08bumped up the success rate from about
  586. 21:1060% to about 81% uh in open source in
  587. 21:14agents ability to consume open source
  588. 21:15across these 270 libraries and we did
  589. 21:17that with a little bit of a shortening
  590. 21:19of the time. So the agents were a lot
  591. 21:21more successful in a little bit less
  592. 21:22time a little bit faster which was
  593. 21:24really encouraging. So this was great
  594. 21:26and this is again the validation part.
  595. 21:28We also had you always find these bits
  596. 21:30of information that you don't know how
  597. 21:31to expect. uh some interesting
  598. 21:33information about the age of a library.
  599. 21:35And so this graph shows with the same
  600. 21:37test the success rate of the agent. This
  601. 21:40is cloud code across uh correlated to
  602. 21:44the uh age of the libraries. The
  603. 21:46libraries at the left are the oldest.
  604. 21:47The libraries on the right are the
  605. 21:49newest that this is release date just as
  606. 21:51a simple measure. And it's interesting
  607. 21:53to see how the libraries that are 5 to
  608. 21:5510 years old, those are the ones that
  609. 21:57the agents do the best with. And we've
  610. 21:59seen this with other agents as well. Uh,
  611. 22:01and the distance is substantial. And
  612. 22:03that makes sense because very old
  613. 22:05libraries, they might have they might
  614. 22:07have confusing information on the web.
  615. 22:10Uh, maybe they don't have as much web
  616. 22:11representation if they're truly old. And
  617. 22:13the very new libraries often times they
  618. 22:15haven't accumulated enough information
  619. 22:16on the web and some of their data is
  620. 22:18actually after the training data. And
  621. 22:20our theory was that with uh with the
  622. 22:22tiles, we actually can flatten that
  623. 22:24line, right? We because we gave them the
  624. 22:26information, they have it, so they
  625. 22:27should use it. And indeed, we've seen
  626. 22:28that we elevate the result more broadly.
  627. 22:31And this was useful because it matches
  628. 22:32our kind of anecdotal experience that um
  629. 22:36using this type of knowledge, using the
  630. 22:37right context, sometimes just helps and
  631. 22:39sometimes it turns it from it doesn't
  632. 22:41work to it works.
  633. 22:45Um I'll move on to multiplayer, but just
  634. 22:48one thing to say is that in context
  635. 22:49engineering is actually a lot more than
  636. 22:50what I'm showing right now. And there
  637. 22:52are many many different tools and one
  638. 22:54notable one is the ability to use a sub
  639. 22:56agent in claude and now a few other
  640. 22:58agents to be able to give it its context
  641. 23:01and focus it. So there's a lot more to
  642. 23:02do here and it's a worthy domain to
  643. 23:04invest in.
  644. 23:07So with that let's talk from single
  645. 23:09player talk about multiplayer. In
  646. 23:10multiplayer, the question is in your
  647. 23:13team, in your organization, what context
  648. 23:16do you want to provide and when? What
  649. 23:19are the ways in which you can share
  650. 23:20context across the team? I'll talk about
  651. 23:23three versions of this. The first and
  652. 23:26kind of easiest most sort of uh direct
  653. 23:29way to share context is to put stuff in
  654. 23:31the repo. As we said, the project itself
  655. 23:34is the base context for an agent. And so
  656. 23:37if you check in an agent MD file,
  657. 23:39anybody that uses that repository would
  658. 23:40load it. It's the repo or the path
  659. 23:42within the repo. This is great and you
  660. 23:44should do it and it is a perfect fit for
  661. 23:47things that are repo specific knowledge.
  662. 23:49Right? This is documentation about how
  663. 23:50this application behaves or things like
  664. 23:52that. Awesome. Perfect place to do it.
  665. 23:54Highly recommended. Uh where it slightly
  666. 23:57falters is uh one is cross repository
  667. 24:00information. your be best practices,
  668. 24:02shared libraries and all those feels a
  669. 24:04bit odd to check those into a
  670. 24:06repository. Would you check them into
  671. 24:08multiple repositories? You might want a
  672. 24:09different solution. And the second is uh
  673. 24:12and I'm picking a little bit on backlog
  674. 24:13MD over here. You know, it's a great
  675. 24:15framework. We had a great workshop
  676. 24:16yesterday about it. Um and you can see
  677. 24:18though that you know those who adopt
  678. 24:20agents and lean into them, what you find
  679. 24:22in the repos is you find many duplicate
  680. 24:23files. You find your sort of Gemini MD
  681. 24:25and cloud MD and agent MD and cursor
  682. 24:27rules and you'll find skills. You'll
  683. 24:28find GitHub stuff.
  684. 24:30and and that's not awesome. Uh and it's
  685. 24:33not that you have this many files that
  686. 24:34is the problem, but rather that they are
  687. 24:36duplicative. There's often content that
  688. 24:38is very similar within each one of them.
  689. 24:40So that's a challenge.
  690. 24:43The second tool that you have to
  691. 24:45accumulate and to provide shared context
  692. 24:48amongst your team is to give agents
  693. 24:50tools to gather context. uh most common
  694. 24:53is of course browsing the web but also
  695. 24:55things that are more focused like hey
  696. 24:57use this to uh uh this MCP tool to be
  697. 25:00able to go and fetch uh code from GitHub
  698. 25:03repositories and load them. So this is
  699. 25:05I'd say you know first of all super
  700. 25:07useful for discovery and exploration
  701. 25:09like you know you you don't know which
  702. 25:10library you're going to use clearly you
  703. 25:12don't have the context for that library
  704. 25:13as you're exploring you should have
  705. 25:15dynamic tools to go and fetch that. It's
  706. 25:16also great for dynamic data like uh uh
  707. 25:19production data about how your system
  708. 25:21ran or you know something that is
  709. 25:22constantly changing. So for those it is
  710. 25:24amazing. Uh it is not great when there's
  711. 25:28a high potential to get the wrong
  712. 25:29answer. Versioning is a is a strong
  713. 25:32example of that. If you're using a
  714. 25:33version that's not the latest and your
  715. 25:35agent went to the latest GitHub repo,
  716. 25:37it's going to get the wrong information
  717. 25:38to operate with. And for complex topics
  718. 25:42where there's subtleties, you know,
  719. 25:43authentication related topics or um just
  720. 25:46sort of complex environments,
  721. 25:49it's also highly inefficient for
  722. 25:52repeatedly needed info. So let's say you
  723. 25:55can go off and find the relevant commit
  724. 25:57at the relevant repo and read the code.
  725. 25:59If you need this information seven times
  726. 26:01a day, you know, five days a week, why
  727. 26:05like why wouldn't you sort of gather
  728. 26:06that context one time, store it,
  729. 26:08evaluate it and then optimize for that?
  730. 26:11Which gets me to the third mode which is
  731. 26:13today kind of the cutting edge and it's
  732. 26:16the notion of curated context or curated
  733. 26:18knowledge. What you would have seen is
  734. 26:20over the last couple of months you've
  735. 26:22seen a lot of progress from all the big
  736. 26:24agents uh uh providing reusable context
  737. 26:28frameworks. So they have you have you
  738. 26:30saw claude skills you saw cursor team
  739. 26:32rules you saw GitHub copilot spaces uh
  740. 26:35and you know these are these are kind of
  741. 26:37acknowledgements from these agent
  742. 26:39companies to say well we don't want the
  743. 26:40agent to reinvent it every time we
  744. 26:42actually want to allow you to share
  745. 26:44things and to share practices and that's
  746. 26:46great I highly recommend these platforms
  747. 26:48you should learn about your environment
  748. 26:49for it they're great for cross
  749. 26:51repository knowledge uh they're great
  750. 26:53when you're using often times you start
  751. 26:54seeing people use uh maybe like a you
  752. 26:57have a security code review bot that
  753. 26:59runs in your review surrounding, but
  754. 27:01that same context, you also want it when
  755. 27:03you're building locally and maybe when
  756. 27:04you're assessing an incident with
  757. 27:06another agent. So, it's great to have
  758. 27:08not just crossdevelopment repo, but also
  759. 27:10cross use case knowledge. Uh, and
  760. 27:12because they're built natively, they
  761. 27:14tend to be kind of nice and elegant in
  762. 27:15terms of how they interact, you know,
  763. 27:16clawed with its skills, cursor with its
  764. 27:18rules, etc. Uh, they're a bit of an
  765. 27:20overkill for single repo information.
  766. 27:22Maybe that's not too bad. But really the
  767. 27:24biggest challenge with them is that they
  768. 27:26are single agent. Um and that kind of
  769. 27:29begs the question of do we think long
  770. 27:31term that knowledge and context should
  771. 27:34be managed per agent, right? Do you
  772. 27:36think in your organization, in your
  773. 27:38practice, do you expect that you would
  774. 27:39use one agent for your org over time?
  775. 27:43And I think the answer to that is no.
  776. 27:45Like I think just like any technologies,
  777. 27:47technologies will be better at one thing
  778. 27:49versus the other. They will move. they
  779. 27:51will change. Companies are big. People
  780. 27:53will have preferences
  781. 27:55and that knowledge as a whole should be
  782. 27:57something that is a core competency of
  783. 27:59the knowledge that you have and it
  784. 28:01should be adapted per agent but it
  785. 28:02should be your asset that is broader. So
  786. 28:05that's the approach we have at Tessle.
  787. 28:06You've sort of seen a bit of this
  788. 28:07already. We have the Tessle JSON and
  789. 28:09when you install this knowledge then it
  790. 28:12will it will adapt it to the relevant
  791. 28:15agent. It will store it as a cursor role
  792. 28:16or it would store it as a as an agent MD
  793. 28:18or whatever that is. Uh and then also
  794. 28:20once we download that knowledge, the
  795. 28:22knowledge is there as files and so you
  796. 28:24can use it in whatever way that you
  797. 28:26want. We do provide tools to make it
  798. 28:29easier and better to work with it. We
  799. 28:30want it to be seamless, but you're not
  800. 28:32locked in to you need to use those tools
  801. 28:34if you want to use different tools to
  802. 28:36consume it. So we think this is
  803. 28:37important when we think about uh uh kind
  804. 28:40of broader long-term context management
  805. 28:42and development.
  806. 28:44And in fact, we're doubling down on
  807. 28:46this. And just as a slight plug here,
  808. 28:49you know, we are expanding the registry
  809. 28:51into a full-blown TESL uh agent
  810. 28:54enablement platform uh that will help
  811. 28:56you gather and and this type of
  812. 28:58knowledge like the packages you've seen,
  813. 29:00create and run evaluations for them,
  814. 29:02distribute them so that every repository
  815. 29:04gets its knowledge uh and then help you
  816. 29:06optimize that over time. Uh so this is a
  817. 29:08commercial product and if you're
  818. 29:10interested in being uh being part of it,
  819. 29:12uh come come talk to us, see us or email
  820. 29:15us at contact.io. fail if you're online.
  821. 29:19So that's a good segue for just the
  822. 29:21closing points around ecosystem context.
  823. 29:24And you know when I just sort of sang
  824. 29:27the sort of the uh the praises of why
  825. 29:29ecosystemwide context from the registry
  826. 29:32is helpful and you should use it and
  827. 29:33it'll help you today. But in the long
  828. 29:35run you do have to wonder who should
  829. 29:38create this type of open source context.
  830. 29:40know we we come along to someone else's
  831. 29:42library and we create these agent docs
  832. 29:44for them and we track them and that you
  833. 29:46know we do it today because it's helpful
  834. 29:48today and we want to help but long-term
  835. 29:50is that right
  836. 29:52and I think the answer is no like I
  837. 29:54think over time kind of open source
  838. 29:56context is really an aspect of agent
  839. 29:59experience and it is the creator's
  840. 30:01responsibility if you are an open source
  841. 30:03maintainer if you are a vendor it's your
  842. 30:06responsibility to have a great developer
  843. 30:07experience it's also your experience
  844. 30:10your responsibility to have a great
  845. 30:12agent experience. You know, it is it is
  846. 30:14in your control. It is something that
  847. 30:15you can choose where to invest and
  848. 30:18context is your best tool for that.
  849. 30:20Context is both your documentation and
  850. 30:21your UX. You know, you need to think
  851. 30:23about what is it that you say there. How
  852. 30:25do you distribute that to your users?
  853. 30:28You know, do you want this indeed to be
  854. 30:29single agent or multi- aent? And eval
  855. 30:32are your tests. They're your moment of
  856. 30:34defining what is correct and what is not
  857. 30:36correct. And so, you should invest in
  858. 30:38those. You should define them. And you
  859. 30:39should ask your users for feedback on
  860. 30:41when they work and when they don't work.
  861. 30:44And you're it's your responsibility, but
  862. 30:46you're not alone. You know, you have us
  863. 30:48to help with our technology with our
  864. 30:49platform on it. You know, feel free to
  865. 30:51talk to us. You have the amazing AI
  866. 30:53native dev uh community that you should
  867. 30:55tap and they will help you. And
  868. 30:57together, I think we want to create this
  869. 30:59successful methodologies and actual
  870. 31:02context that will help us all consume
  871. 31:04software together better. Uh and if we
  872. 31:06achieve all of that then we'll help
  873. 31:08we'll successfully put everything in the
  874. 31:10right context. Thank you.
  875. 31:16[music]
  876. 31:34>> [music]

About this transcript

This page contains the full transcript of Guy Podjarny - Spec Driven Dev From Single Player to Multiplayer to Ecosystem | DevCon Fall 2025 by AI Native Dev, generated from the public captions YouTube serves with the video. The transcript has 6,288 words across 876 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.