Spec-Driven Development: Agentic Coding at FAANG Scale and Quality — Al Harris, Amazon Kiro — Transcript
Full transcript
- 0:13[music]
- 0:20For those of you who haven't heard of
- 0:22us, Kira is an agentic ID.
- 0:24[clears throat] Um, we launched
- 0:25generally available this most recent
- 0:27Monday, I think the 17th, but we
- 0:29launched public preview on, uh, in July,
- 0:32>> uh, I think July 14th. So, out there for
- 0:35a few months getting customer feedback,
- 0:37um, all that good stuff. We're going to
- 0:38talk a little bit about using Spectriven
- 0:40development to sharpen your AI toolbox.
- 0:42I did a show of hands. About a quarter
- 0:43of the people here familiar with
- 0:44Spectrum and Dev. My name is Al Harris.
- 0:46Um, principal engineer at Amazon. I've
- 0:48been working on Curo for the last. Uh,
- 0:50and we're a very small team. We were
- 0:52basically three or four people sitting
- 0:54in a closet doing what we thought we
- 0:56could do to improve um the software
- 0:58development life cycle for customers. So
- 1:01we were ch we were charged with building
- 1:03a development tool that's that answered
- 1:05um that improved the experience for
- 1:07spectrum and development. We were
- 1:09theoretically funded out of the org that
- 1:11supported things like QDV but we were
- 1:13purposefully a very different product
- 1:14suite from the QE system to just take a
- 1:17different take on these things. So we
- 1:19wanted to work on scaling, you know,
- 1:20helping you scale AI dev to more complex
- 1:22problems. Uh improve the amount of
- 1:24control you have over AI agents and
- 1:26improve the code quality and maintain uh
- 1:28reliability, I should say, of what you
- 1:30got out the other end of the pipe. Now
- 1:32we're back to new content. Um so our
- 1:35solution was specri. We took a look at
- 1:37some existing stuff out there and said,
- 1:38"Hey, vibe coding is great, but vibe
- 1:40coding relies a lot on me as the
- 1:42operator getting things right. That is
- 1:44me giving guardrails to the system. And
- 1:45that is me uh putting the agent through
- 1:48a uh kind of a strict workflow. We
- 1:50wanted Spectri driven dev to sort of
- 1:52represent the holistic SDLC because
- 1:54we've got you know 25 30 years of
- 1:56industry experience um building uh
- 1:59software building it well and building
- 2:01it with different practices right we've
- 2:03gone through waterfall at XP um we have
- 2:06all these different ways that we
- 2:07represent what a system should do and we
- 2:09want to effectively respect what came
- 2:11before.
- 2:12So uh this animation looked a lot
- 2:15better. It was initially just the left
- 2:17diamond but I the idea was hey you know
- 2:19you basically are iterating on an idea.
- 2:21I think like half of software
- 2:23development is discovery requirements.
- 2:25Um and that discovery doesn't just
- 2:26happen by sitting there and thinking
- 2:28about what what should the system do?
- 2:29What can the system do? We we realized
- 2:32though kind of working on this that the
- 2:33best way to make these systems work is
- 2:35to actually synthesize the output and be
- 2:37able to feed that back really quickly.
- 2:38things like your input requirements um
- 2:41to actually do the design and feedback
- 2:43you know realize oh actually if we do
- 2:45this there's a side effect here we
- 2:46didn't consider we need to feed that
- 2:47back to the input requirements and so
- 2:50this compression of the SDLC evolved to
- 2:52bring structure into the software
- 2:54development flow we wanted to take um
- 2:58the artifacts that you generate as part
- 2:59of a design that's the requirements that
- 3:01maybe a product manager or developer
- 3:03writes that's going to be the acceptance
- 3:04criteria what does success look like at
- 3:07the end of this and then we want to the
- 3:08design artifacts that you might review
- 3:10with your dev team, you might review
- 3:11with you know stakeholders and say this
- 3:13is what we're going to go build and
- 3:14implement the thing and we want to make
- 3:16sure that you can do this all in some
- 3:17tight inner loop. Um and ult that was
- 3:20initially what spectriven dev was
- 3:23um what spectriven development in hero
- 3:27is today or at least was before it went
- 3:29g was uh you give us a prompt and we
- 3:32will take that and turn it into a set of
- 3:34clear requirements with acceptance
- 3:35criteria. We represent these acceptance
- 3:37criteria in the EARS format. EARS stands
- 3:39for the easy approach to requirement
- 3:41syntax. Um, and this lets you really
- 3:44easily uh it's effectively a structured
- 3:46natural language representation of what
- 3:48we you want the system to do. Now, for
- 3:51the first four and a half months this
- 3:52product existed, the ears format looked
- 3:54like kind of an interest decision we
- 3:56made, but just that sort of interesting.
- 3:58Um and with our launch, our general
- 4:00availability launch on Monday, we have
- 4:02finally started to roll out some of the
- 4:03side effects of which is property based
- 4:06testing. Um so now your ears
- 4:08requirements can be translated directly
- 4:10into properties of the system which are
- 4:12effectively invariants that you want to
- 4:13deliver. Um, for those of you who have
- 4:16or like have not I guess done property
- 4:19based testing in the past using
- 4:20something like I think it's a hypothesis
- 4:23in Python or fast check and node um
- 4:27closures spec library is another
- 4:29example. These are uh approaches to
- 4:32testing your software system where
- 4:34you're effectively trying to produce a
- 4:35single uh test case that that falsifies
- 4:38the invariant that you want to prove.
- 4:40And if you can find any uh contraositive
- 4:44then you can say this requirement is not
- 4:45met. If you cannot you have some high
- 4:47degree of confidence where the word high
- 4:50there is doing a little bit of heavy
- 4:51lifting because it depends on how well
- 4:52you write your tests but you can say
- 4:56with a high degree of confidence that
- 4:57the system does exactly what you're
- 5:00saying it does. Um yeah, so a property
- 5:05we we'll get a little bit more into
- 5:07property based testing and PBTs a little
- 5:08later, but this is the first step of
- 5:11many we're taking to actually take these
- 5:13structured natural language requirements
- 5:15and then tie this with a throughine all
- 5:17the way to the finished code and say if
- 5:19your code if the properties of the code
- 5:22meet the initial requirements, we have a
- 5:25high degree of confidence that you have
- 5:27re uh reliably shipped the the software
- 5:29you expected to ship.
- 5:31So with spectriven dev, we take your
- 5:34prompt, we turn it into requirements, we
- 5:36pull a design out of that, we define
- 5:38properties of the system and then we
- 5:40build a task list and we go and you can
- 5:42run your task list. Effectively the spec
- 5:45then becomes the natural language
- 5:46representation of your system. It has
- 5:48constraints, it has concerns um around
- 5:52functional requirements, non-functional
- 5:53requirements and it's this set of
- 5:55artifacts uh that you're delivering. So
- 5:57I don't think I have the slide in this
- 5:59deck, but ultimately the way I look at
- 6:00spec is that it is one a set of
- 6:02artifacts that represent sort of the
- 6:04state of your system at a point in time
- 6:05t. It is two a structured workflow that
- 6:08we push you through to reliably deliver
- 6:10high-quality software and that is the
- 6:12requirements design um and execution
- 6:15phases. And then three it is a set of
- 6:18tools and and um systems on top of that
- 6:20that help us deliver reproducible
- 6:22results where one example of that is
- 6:24property based testing. Another example
- 6:26of that which is a little less obvious
- 6:28but we can talk about later is going to
- 6:29be um I don't even know what to call it
- 6:33uh requirements verification. So we scan
- 6:35your requirements for over ambiguity. We
- 6:37scan your requirements for um invalid
- 6:41constraints eg uh you have conflicting
- 6:44requirements and we help you resolve
- 6:46those ambiguities using sort of classic
- 6:48uh automated reasoning techniques. Um
- 6:51and I could talk a little bit more about
- 6:52sort of the the features of Kira. I
- 6:55think that's maybe less interesting for
- 6:56this talk because we want to talk about
- 6:57spectrum and dev. We have all the stuff
- 6:59you would expect though. We have
- 7:01steering which is sort of memory and
- 7:02sort of cursor rules. We have MCP
- 7:05integration. We have you know image yada
- 7:08yada. Um so we have ways to and we have
- 7:10software hooks. Um so let's talk a
- 7:13little bit about sharpening your tool
- 7:15chain. And I'm going to take a break
- 7:16really quick here. Uh just pause for a
- 7:18moment for folks in the room who had
- 7:20maybe tried downloading Curo um or
- 7:24something else and just say are there
- 7:25any questions right now before we dive
- 7:27into how to actually use spec to achieve
- 7:29a goal?
- 7:32No questions. It could be a good sign.
- 7:35Could mean I'm not uh talking about
- 7:37anything that's particularly
- 7:37interesting. So um I actually want to
- 7:40like talk in some concrete detail here.
- 7:43Uh this is a talk I gave a few months
- 7:45ago on how to use MCPS in Kira. And so
- 7:48one of the challenges that people who
- 7:49had tested out Kira had that might be a
- 7:52little easier to see was that they
- 7:56um they felt that the flow we were
- 7:58pushing them through was a little bit
- 8:00too structured like you don't have
- 8:02access to external data, you don't have
- 8:04access to the to all these other things
- 8:05you want. And so one thing that we said
- 8:07on our journey here towardsing your um
- 8:12oh you know what this out of order
- 8:13here's my nice AI generated image. So
- 8:16you can use MCP. Everybody here I assume
- 8:18is familiar with MCP at this point. But
- 8:20uh Curo integrates MCP the same way all
- 8:23the other tools do. Uh but what I think
- 8:26people don't do enough is use their MCPs
- 8:28when they're building their specs. And
- 8:30so you can use your MCP servers in any
- 8:33phase of the specdriven development
- 8:34workflow. That's going to be
- 8:36requirements generation, design, um, and
- 8:38implementation. Um, and you can use,
- 8:42we'll go through an example of each. So,
- 8:45first of all, to set up a spec in Kuro
- 8:47is fairly straightforward. We have the
- 8:48Kuro panel here, which there's a little
- 8:51ghosty um, and then you can go down to
- 8:54your MCP servers and click the plus
- 8:56button. You can also just my favorite
- 8:57way to do it is to ask Kirro to add an
- 9:00MCP uh and then give it some some
- 9:03information on where it is and it can go
- 9:05figure it out usually from there or you
- 9:07just give it the JSON blob and it'll
- 9:08figure it out. Once you have your MCP
- 9:10added, you'll see it in the control
- 9:11panel down here and you can enable it,
- 9:13disable it, allow list tools, disable
- 9:15tools, etc. So you can manage context
- 9:17that way. Worth noting changing MCP and
- 9:20changing tools in general is a caching
- 9:22operation. So if you're very deep into a
- 9:24long session, maybe don't tweak your MCP
- 9:26config because it will slow you down
- 9:28dramatically. But let's talk about um
- 9:31MCP inspect generation. So something I
- 9:34the Curo team uses a um for reasons I
- 9:38don't know, but it's our task tracker of
- 9:41choice. Uh but so one thing I want to do
- 9:43is uh maybe go and say I don't want to
- 9:46write the requirements for a spec from
- 9:47scratch. My product team has already
- 9:49done some thinking. We've iterated in a
- 9:50sauna to kind of break a project down.
- 9:52This is not always how things work, but
- 9:54sometimes how things work. So in this
- 9:55case, I have I have a task in a sauna.
- 9:58Oh no, I did the wrong thing.
- 10:02That's what I get for zooming. So I have
- 10:05this task in in a sauna that says add
- 10:07the view model and controller to this
- 10:08API. In this case, this was a demo app
- 10:11that I can figure in a few minutes. And
- 10:13we even had like it's kind of peeking
- 10:16under here, but we had some details
- 10:17about what we wanted to have happen. Now
- 10:19I can go into Kira and just say start
- 10:21executing task XYZ URL from ASA and Kira
- 10:25is going to recognize this is an Asana
- 10:27URL. I had the ASAN MCP installed. It
- 10:29goes and pulls down all the metadata
- 10:31there. Um da da da. So it's going to
- 10:33break out and from there start um
- 10:36start determining what to work on. Um
- 10:43oh it's funny these titles are
- 10:44backwards.
- 10:46basically create a spec for my open
- 10:48asauna tasks. Again, go pull from a
- 10:50sauna all the tasks and then for each
- 10:52one generate um requirements based on
- 10:55those tasks. So I think I had like six
- 10:56tasks assigned to me. One is do user
- 10:59management, do some sort of um
- 11:04uh property management da da da it
- 11:06pulled them in generated the
- 11:07requirements and then in this case title
- 11:10is wrong apologies start executing task.
- 11:13this is I want to go and do the code
- 11:14synthesis for this um and I will take a
- 11:17quick break here to talk about how you
- 11:20can do this in practice. So for those of
- 11:21you who are you know following along in
- 11:23room uh feel free to fire up your curo
- 11:26open a project and then picking a an MCP
- 11:29server. I'll share a few repos here
- 11:31really quick that you can play around
- 11:32with.
- 11:35So
- 11:37I have an MCP server implemented.
- 11:41I have
- 11:48this lofty views which I think
- 11:50implements the asauna. Um and then these
- 11:53should all be public. Let me just double
- 11:55check.
- 11:57Yeah. Okay. So for example, if you
- 11:59wanted to extend my I have a Nobel Prize
- 12:01MCP which curls perhaps unsurprisingly
- 12:05there is a Nobel Prize API.
- 12:07>> [clears throat]
- 12:07>> Um, so you can use UVX to install it or
- 12:09you can get clone this Al Harris at
- 12:11Nobelmcp.
- 12:13Uh, this is just one example. Another
- 12:15one here is if you want to play around
- 12:16with the sample that's in the video. Um,
- 12:18I have Al Harris atlofty Views. Um, I'll
- 12:21leave these both sort of up on the
- 12:23screen for a few moments for folks who
- 12:25do want to copy the uh the URLs.
- 12:31But while that is happening,
- 12:34oh no, let's put you on the same window.
- 12:41[clears throat]
- 12:48So what I'll demo quick is the usage of
- 12:51an MCP to make like spec generation much
- 12:54easier or more reliable. So here I have
- 12:58let's see Got
- 13:01a lot of MCPs. Which ones do I actually
- 13:04want to use?
- 13:11Let's use the GitHub MCP.
- 13:15Oh, no.
- 13:18Ignore me.
- 13:23That's better. Okay. Well, I have the
- 13:24fetch MCP. So in this case I could for
- 13:27example come in here and say hey I've
- 13:30generated a bunch of tasks lofty views
- 13:33app. This is basically a very simple
- 13:34CRUD web app. Um but I want Kira to
- 13:41uh use the fetch MCP to pull examples
- 13:44from similar products that exist on the
- 13:47internet. You could also use you know
- 13:48Brave search or Tavlet search MCP
- 13:50servers but in this case I'll just use
- 13:52fetch because I've got it enabled. Um,
- 13:54so let's say,
- 13:57oh actually we can run the web server
- 13:59and use fetch. That's a good example.
- 14:14[clears throat]
- 14:20This is one example of you can at any
- 14:22point in the workflow generating a spec
- 14:24go through and um you know use your MCP
- 14:27servers to get things working. No, this
- 14:31is what I get for not using a project in
- 14:32a while.
- 14:36We'll cancel that. We can actually do
- 14:38something a little more interesting
- 14:39which is a separate project I've been
- 14:41working on. Um, so I've been working on
- 14:43a an agent core agent and that might be
- 14:47I I know the project works, which is the
- 14:49reason I'll fire it up here. Should I
- 14:51call it?
- 15:04Well, maybe we'll do live demos at the
- 15:05end.
- 15:07So that's sort of like the most basic
- 15:09thing you can do with Kira is just use
- 15:11MCP servers, but any tool uses MCP
- 15:13servers. I actually don't think that's
- 15:15particularly interesting. So let's say
- 15:17in sort of this process of trying to
- 15:19sharpen our our spec dev toolkit, we've
- 15:21finished up with the 200 grit. We've
- 15:23added some capabilities with MCP. It's
- 15:25useful, but it's not going to be a
- 15:26gamecher for us. I want to come in here
- 15:28and actually get up to the 400 grit.
- 15:30Let's get start to get a really good
- 15:31polish on this thing. I want to
- 15:33customize the artifacts produced because
- 15:35you've got this task list, you've got
- 15:37this requirements list and I don't agree
- 15:38with what you put in there, Al. Um, you
- 15:41could say that a lot of people do and I
- 15:43that's a a great starting point. So,
- 15:46here's something I heard earlier in the
- 15:48week at um, you know, earlier in the
- 15:50conference is that people like to do
- 15:51things like use wireframes in their
- 15:53mocks. Um, use wireframe mocks because
- 15:55in your specs are natural language,
- 15:58you're using specs as a control surface
- 16:00to explain what you want the system to
- 16:01do. Uh therefore I want to be able to
- 16:03actually put UI mocks in here. So the
- 16:06trivial case is that I just come in here
- 16:07and say Kuro's asked me here does does
- 16:10the design look good? Are you happy? And
- 16:12I said this looks great but could you
- 16:13include wireframe diagrams and ask you
- 16:15for the screens we're going to build
- 16:17here. I'm adding this is again from that
- 16:20lofty views thing. I'm adding a user
- 16:21management UI but I want to actually see
- 16:24what we're sort of proposing building
- 16:25not just the architecture of the thing.
- 16:27So your cure is going to sit here and
- 16:28churn for a few seconds, but you can add
- 16:30whatever you want to any of these
- 16:31artifacts because they're natural
- 16:32language. So they're structured, which
- 16:34means we want some re um some sort of
- 16:38reproducibility in what they look like,
- 16:40but ultimately what they look like
- 16:41doesn't matter because we've got the the
- 16:43any machine here, the agent sitting that
- 16:45can help translate it to what it needs
- 16:46to be. So Kira's churning away here.
- 16:49It's thinking thinking and then it's
- 16:51going to spit out these uh text wrapped
- 16:54asy diagrams. I'll fix the wrapping here
- 16:56in a second in the video, but ultimately
- 16:58like
- 17:01you know it does whatever you want. So
- 17:03if you want additional data in your
- 17:06requirements, you can do that. If you
- 17:08want additional data in the design like
- 17:10this, uh you can easily add that. Here
- 17:13we've got sort of these wireframes in
- 17:14ASKI that help me sort of rationalize
- 17:16what we're actually about to ship. Um,
- 17:18and then I can again continue to chat
- 17:20and say actually in the design I don't
- 17:22want um, you know, maybe I don't want
- 17:25this add user button to be up at the top
- 17:26the entire time in which case I could
- 17:28chat with it to make that change easily
- 17:30and now we're on the same page up front
- 17:32instead of later during implementation
- 17:34time. So we've again sort of left
- 17:35shifted some of the concerns. Um, so
- 17:38that's one example. You know, I want to
- 17:39add UI mocks to the design of a system.
- 17:42Another example though could be this.
- 17:43Um, oh, this is a just a quick snapshot
- 17:46of the end state there where now my
- 17:48design does have these UI mocks.
- 17:51Um, but another example that I actually
- 17:53like a little bit more is this uh
- 17:55including test cases in the definition
- 17:57and tasks. So today the tasks that cure
- 18:00will give you will be kind of the bullet
- 18:02points of the requirements and the
- 18:03acceptance criteria you need to hit. But
- 18:06I want to know that at the end state of
- 18:08this task being executed, we have a
- 18:10really crisp understanding that it is
- 18:12correct. It's not just like done because
- 18:14the a anybody who's used an agent can
- 18:16probably testify that um the LMS are
- 18:18very good at saying I'm done. I'm happy.
- 18:20I'm sure you're happy. I'm just going to
- 18:22be complete. Oh, yeah. The tests don't
- 18:23pass but they're annoying. I tried three
- 18:26times them to work. I'm just going to
- 18:27move on. Um no, I don't want that. I
- 18:30want to actually know that things are
- 18:31working. So, in this case, I've asked
- 18:32Hero to um include explicit unit test
- 18:35cases that are going to be covered. So
- 18:37my task here for example in create
- 18:38creating this agent core memory checkp
- 18:40pointer is going to have all the test
- 18:42cases that need to pass before it's
- 18:43complete and then I can use things like
- 18:45agent hooks to ensure those are correct.
- 18:47We'll run this uh sample a little later
- 18:48in the talk. Um this is the thing I'm
- 18:51ready to little demo.
- 18:53Uh yeah, so this is another example
- 18:55where you can again you're you're
- 18:57working on your toolbench. You're sort
- 18:58of you have all these capabilities and
- 19:00primitives at your control and you can
- 19:03tweak the process to work for you, not
- 19:05just the process that I think is the
- 19:07best one. And then sort of last but not
- 19:09least, the 800 grit. At this point,
- 19:11we're getting a final polish on the
- 19:13tool. Uh we might be stropping necks,
- 19:15but we want to, you know, you can
- 19:17iterate on your artifacts, but you can
- 19:18also iterate on the actual process that
- 19:21runs. So, one thing you might have, and
- 19:24I do this a lot, is I'll I'll be
- 19:25chatting with Kira, and I say, "Hey, I
- 19:28want to um in this case, I want to add
- 19:31memory to my agent in agent core. Um,
- 19:35let's dump conversations to an S3 file
- 19:37at the end of every execution." Cur is
- 19:39going to say, "That's great. I know how
- 19:40to do that. I'm going to research
- 19:41exactly how to do that thing. I will
- 19:43achieve this goal for you." But
- 19:45ultimately what I've done is actually
- 19:47introduce a bias up front which is I'm
- 19:49steering the whole agent using S3 as
- 19:51this storage solution just because maybe
- 19:53I'm familiar with it but it's probably
- 19:55not the best way to go about it. So then
- 19:57after it had synthesized the design and
- 19:59all the tasks and all this stuff I came
- 20:01back and said well like we don't need to
- 20:03stick to this rigid spectriven dev
- 20:04workflow that I've that has been defined
- 20:06by Kirao. I can ask for alternatives
- 20:08like is this the idiomatic way to
- 20:09achieve session persistence? I don't
- 20:12know maybe there's a better way. Maybe
- 20:14if we're talking AWS services, it's not
- 20:16S3, it's Dynamo or yada yada. Uh Kira's
- 20:19going to come in here and say, you know,
- 20:21good question. Uh da da da. Let me
- 20:23research. It's going to go through call
- 20:25a bunch of MCP tools that I've given it
- 20:27access to. This kind of ties back to
- 20:29that you should be using MCP. And then
- 20:31it comes back with this recommendation
- 20:32that I didn't know was a feature, which
- 20:34is Asian core memory. Um it says it's
- 20:37more idiomatic and future proof that
- 20:39maybe is TBD and should be checked a
- 20:41little closer. Um, but [snorts] uh or
- 20:44you could use S3, which is the thing you
- 20:46recommend. Now, actually, I I bet
- 20:48there's far more than two options here.
- 20:50So, you could probably keep asking the
- 20:51agent, are there other options, yada
- 20:53yada, and it would go and continue to
- 20:54investigate, but you should not lock
- 20:56yourself into the rigid flow that is
- 20:58sort of the starting point here. Um,
- 21:01yeah. So, that that's actually I think
- 21:03it for my deck. Um what I will talk
- 21:06about
- 21:08is let's just run through that sample I
- 21:10just had up there which is that um
- 21:15so
- 21:17basically let me delete delete it and
- 21:20I'll just do a live demo of sort of
- 21:21specs in Curo and how we can fine-tune
- 21:24things a little bit. So this project is
- 21:27a Node.js app. It is a um it's a CDK.
- 21:32Again, I'm not trying to sell
- 21:34[clears throat] more AWS. This is just
- 21:36the technologies I'm familiar with, so I
- 21:38can move a lot more quickly. So, I
- 21:40wanted to know a little bit about agent
- 21:41core, which is a new AWS offering. And
- 21:43as somebody building an agent, I should
- 21:44probably be familiar with it. So, and
- 21:46I'm not familiar enough with it. So,
- 21:48I've got we've got some other people
- 21:50here who know a lot about it. So, put my
- 21:52hand up a little bit and you know, you
- 21:53caught me. So, I set up a CDK stack,
- 21:56which is just um you know, IA technology
- 21:59to deploy software. I'm familiar with it
- 22:01and I love it. Uh, so I have a stack
- 22:03here that lets me deploy whatever an
- 22:06agent core runtime is. I don't know. I
- 22:08asked Kira to do it. We vibe coded this
- 22:09part. So we vibe coded the general
- 22:11structure. We got an agent. We got IA
- 22:13set up. I then vibe code added commit
- 22:16lint. I added husky. A few things like
- 22:18this that I like for my own TypeScript
- 22:19projects. Um, prettier and eslint I
- 22:22think. So we have a basic product here
- 22:24or like a basic project here that I know
- 22:26I can deploy to my personal AWS account.
- 22:29Um, now I'm going to come in here and
- 22:32oh, and then importantly, this is super
- 22:34important because I don't know how the
- 22:35hell agent core works. And I could go
- 22:38read the docs, but the docs are long and
- 22:39they're complicated and I'm really just
- 22:41trying to build out a PC to to like
- 22:43learn about it myself. So, I added two
- 22:47MCP servers. Oh, no, maybe I didn't. Let
- 22:51me check. Oh, okay. Yes, sorry. Buried
- 22:55down here at the bottom. So this is my
- 22:57Kira MCP config. I added one important
- 23:00MCP server here which is the AWS
- 23:02documentation one. There's other ways to
- 23:04get documentation. You can use things
- 23:06like um Tessle level 7 but in this case
- 23:09this is vended by AWS. So I have some
- 23:11confidence that it might be correct. So
- 23:13I used this to help the agent have
- 23:17knowledge about sort of what
- 23:18technologies exist. And I think I used
- 23:19fetch quite a bit as well. So these are
- 23:21the two sets of um
- 23:24these are the two step sets of uh MCP
- 23:27servers I provided the system. That's
- 23:29great. Move on.
- 23:32Confirm. So
- 23:34and I'll just rerun this from scratch.
- 23:37So what I had done yesterday evening or
- 23:39maybe the evening before was I sat down
- 23:42and I have this system sort of basically
- 23:46working and now I want to start doing
- 23:47specri development. So, I want to add
- 23:49this uh session ID concept and then I
- 23:52want to read conversation to an S3 file
- 23:54blah blah blah. This is the whole sort
- 23:56of bias thing I showed you earlier.
- 23:58We're going to fire that off through
- 23:59Curo. It's going to start running uh
- 24:01chugging away and then it's going to,
- 24:04you know, see if the spec exists. Uh,
- 24:06okay, the folder does exist. It's
- 24:08probably going to realize there's no
- 24:09files there and start working away. But,
- 24:12um, from here I'll sort of live demo.
- 24:14It's going to read through require. It's
- 24:16going to read through existing docs.
- 24:17It's going to read through existing
- 24:18files, gather the context it needs.
- 24:20Sure, in a way. Um,
- 24:23but in a moment once it generates sort
- 24:25of the initial requirements and design,
- 24:27I am going to challenge it to use its
- 24:28own, you know, MCQ servers. I want you
- 24:31to go and do some research on the best
- 24:32way to do this and provide me some
- 24:33proposals. Um, and this is why I was
- 24:36hoping to get the clip on mic working
- 24:38because I've got to set this down for a
- 24:39moment.
- 25:05Okay. So, you [clears throat] know, I
- 25:07don't know if this is the best way to do
- 25:08this. Um, go read docs, go use fetch. D.
- 25:11It's going to keep kind of churning away
- 25:13here and then come back to me after it's
- 25:15probably got a few ideas and proposed
- 25:17it. But, um, this is an example of me
- 25:20just using additional capabilities. uh
- 25:22use fetch, use the docs MCP, use
- 25:25whatever you can to get the best
- 25:27information and don't take at face value
- 25:28the things that I said. These are
- 25:30usually things we have to prompt pretty
- 25:31hard to get the agent to do, but if
- 25:33you're doing it in real time, it works
- 25:35fairly well. Um, again, the agent, all
- 25:38of these agents are going to be very
- 25:39easy to please. So, you know, just cuz I
- 25:42said something in the stupid docs, it
- 25:43may or may not actually be the most
- 25:45important thing from the agents
- 25:46perspective down the road. So, you know,
- 25:49okay, so it's done a little bit of
- 25:50research. It understands the lang graph
- 25:52which is the agent framework we're using
- 25:54already has this knowledge of
- 25:55persistence
- 25:57um da da da and actually in this case it
- 26:01didn't find it did not use the mcp for
- 26:03uh agent core docs who didn't find that
- 26:05agent core has this knowledge of
- 26:06persistence um so maybe you like let's
- 26:10assume I don't I still don't know that
- 26:11exists because I didn't dry run this a
- 26:13few days ago um we might have to find
- 26:15that later the design phase so first
- 26:18thing it's going to do is kind of
- 26:19iterate over all my requirements
- 26:20requirements here. Um, you know, it's
- 26:23changed the requirements based on what
- 26:24it now knows about Langraph and how it
- 26:26can natively integrate with the uh
- 26:28checkpointing, but it's still really
- 26:30crisply bound to this like S3 decision
- 26:32that I made implicitly in the ask. Um,
- 26:35so that is just something to be aware
- 26:36of. Anything you put in the prompt is
- 26:39effectively rounding the agent. Um, for
- 26:42better or for worse. I see it's still
- 26:44iterating. So, yeah, comes through says,
- 26:47does this look good? We changed duh. I'm
- 26:49going to say looks great. Let's go to
- 26:50the design phase. So now Curo is going
- 26:52to take my requirements and take me into
- 26:53the design phase of this project. I can
- 26:55make this
- 26:58so things are a little bit bigger.
- 27:00But
- 27:02um here's an example of what I meant by
- 27:04these ears requirements. So the user
- 27:07story here is as a dev I want to
- 27:08implement a custom S3based checkpoint so
- 27:10the agent can use Langraph's native
- 27:12persistence mechanism with S3. Great.
- 27:14That sounds reasonable to me as a person
- 27:16you know sort of co-authoring these
- 27:18requirements.
- 27:19This here, this sort of when then shall
- 27:22syntax. This is the years format and the
- 27:25structured natural language is really
- 27:26important for us to pass this through
- 27:27non LLM based models and give you more
- 27:30deterministic results when we parse out
- 27:32your requirements because ultimately our
- 27:33goal is to actually use the LM for as
- 27:35little not as little as possible but
- 27:36less and less over time. We want to use
- 27:38classic automated reasoning techniques
- 27:40to give you high quality results not
- 27:42just you know whatever the latest model
- 27:44is going to tell you. Um, so here's gone
- 27:48through spits out a design doc. Let's
- 27:50actually just look at this in markdown.
- 27:53This sure you got a server da da checkpo
- 27:57pointer ghost s3 that makes sense pseudo
- 28:00code again in a real scenario. Maybe I
- 28:02read this a little bit more closely
- 28:05and what's actually this is the new
- 28:07thing we shipped in um on the 17th is
- 28:10that now cur is going to go through and
- 28:11do this formalizing requirements for
- 28:13correctness properties. Um and so right
- 28:16now what the system is doing is it's
- 28:18taking a look at those requirements you
- 28:19generated uh the requirements we agreed
- 28:21upon with the system earlier. These look
- 28:23good. I agree with them. yada yada. It's
- 28:25taking a look at the design and it's
- 28:27extracting correctness properties about
- 28:28the system that we want to run property
- 28:30based testing for down the road. This is
- 28:32something that may or may not matter for
- 28:33you in the prototyping phase but should
- 28:35matter for you significantly when you're
- 28:37going to production. because these
- 28:38properties are correct and these
- 28:40properties are all met. The system
- 28:42aligns one to one with the input
- 28:44requirements you provided. Um yeah, so
- 28:48while this is chugging away, any
- 28:49questions yet? Any folks kind of curious
- 28:53about this?
- 28:55>> Um yeah,
- 28:57>> we're here and then there.
- 28:58>> Um what would you say is the main
- 29:01difference between
- 29:03that has?
- 29:05Uh I haven't used the planning mode in a
- 29:07couple of weeks. So it's I'm things move
- 29:09so fast it's a little wild. Um but I
- 29:11think ultimately uh what we would say is
- 29:14that Kuro's spectrum and dev is not just
- 29:18LLM driven but it is actually driven by
- 29:20like a structured system. Um and so
- 29:22planning mode I'm not sure if there's
- 29:24actually like a workflow behind it that
- 29:25takes you through things but um yeah
- 29:28this is our take on it for sure.
- 29:31>> I'm not familiar enough to give like a
- 29:32more concrete example unfortunately.
- 29:34similar I mean it doesn't give you like
- 29:36this I think that this document is cool
- 29:39is bringing you the school but uh what
- 29:42Cer does is to basically create you a
- 29:44plan that's
- 29:46>> just an execution plan okay
- 29:48>> oh I see so I think that the fundamental
- 29:51difference there uh does that plan get
- 29:54committed anywhere or is it just
- 29:56ephemeral
- 29:57>> uh it's kind of
- 29:59>> okay so what I want over time is not is
- 30:02not just how we make the changes we care
- 30:05about but it is actually the
- 30:06documentation and specification about
- 30:07what the system does. Um so the
- 30:09long-term goal I have is that as Kira we
- 30:12were able to do sort of a birectional
- 30:13sync that is as you continue to work
- 30:16with Kira you're not just acrewing these
- 30:19sort of task lists uh and so I'm just
- 30:22going to say go for it to go to the
- 30:23tasks um but we're not just acrewing
- 30:25task list but actually if I come back
- 30:27and let's say change the requirements
- 30:29down the road we will mutate a previous
- 30:31spec. So I'm looking at really just a
- 30:33diff of requirements which as you go
- 30:36through the green field process you're
- 30:37going to produce a lot of green in your
- 30:38PRs which is maybe not the best because
- 30:41I'm just reviewing three new huge
- 30:42markdown files but on the next time or
- 30:46the subsequent times that I go and open
- 30:47that doc up I want to be seeing oh
- 30:50you've actually you know you've relaxed
- 30:52this previous requirement you've added a
- 30:53requirement that actually has this
- 30:55implication on the design doc um that is
- 30:57the process the curo team internally
- 30:59uses to talk about changes to the curo
- 31:01So we review our design docs have in
- 31:04general been uh replaced by spec
- 31:08reviews. So we will you know somebody
- 31:10will take a spec from markdown they'll
- 31:13blast it into our wiki basically using
- 31:14an MCP tool we use internally and then
- 31:17we'll review that thing and comment on
- 31:18it in sort of a design session as
- 31:20opposed to you know I wrote this
- 31:22markdown file or a wiki from scratch. Um
- 31:25so it becomes sort of if uh well it's
- 31:28actually not like an ADR because it's
- 31:29not point in time. It is like this
- 31:31living documentation about the system.
- 31:34Um but yeah thanks for the question.
- 31:37There's one over here.
- 31:39>> Um
- 31:41this may be more a spectrum development
- 31:43question but are there like like is
- 31:45there like a template for a set of files
- 31:48that you fill out? Like right now you're
- 31:50in the design.md.
- 31:52>> Are there like
- 31:53>> is this is the designd the spec and it's
- 31:56a single doc or are there
- 31:59>> oh great question. So the yeah the
- 32:01question was um are there and correct me
- 32:04if I'm wrong here but question is are
- 32:05there a set of templates that are used
- 32:07for the system and is the question
- 32:09you're driving at can you change the
- 32:11templates or is just are there okay so
- 32:14the yeah question is are there a set of
- 32:15templates um there are implicitly in our
- 32:18system prompts for how we take care of
- 32:20your specs so you'll see here at the top
- 32:22navbar here right now we're really rigid
- 32:24about this requirement design task list
- 32:26phase but we know that doesn't work for
- 32:28everybody for example if you're starting
- 32:30we get this feedback from a lot of
- 32:31internal Amazonians actually that I want
- 32:33to start with a I have an idea for a
- 32:34technical design and I don't necessarily
- 32:36know what the requirements are yet but I
- 32:38know I want to make maybe design is even
- 32:39the wrong word I want to start with a
- 32:41technical note like I want to refac this
- 32:44comes up a lot for refactoring actually
- 32:47um so I want to refactor this to no
- 32:50longer have a dependency on
- 32:52um here's a good example here we use a
- 32:54ton of mutxes around the system to make
- 32:56sure that we're locking appropriately
- 32:57when the agent is taking certain actions
- 32:59because we don't want different agents
- 33:00to step on each other's toes. But maybe
- 33:02I want to challenge the requirements of
- 33:04the system so I can remove one of these
- 33:05mutexes uh or semaphors I should say. Um
- 33:09so I might start with something like a
- 33:12technical note and then from there sort
- 33:14of extract the the requirements that I
- 33:16want to share with the team and say hey
- 33:17you know I had to kind of play with it
- 33:18for a little while to understand what I
- 33:20wanted to build but I still want to
- 33:21generate all these rich artifacts. So
- 33:24today it's this structured workflow.
- 33:25We're playing a lot around with making
- 33:26that a little bit more flexible. But the
- 33:28the structure is important because the
- 33:30structure lets us build reproducible
- 33:31tooling that is not just an L. So I
- 33:35think that that's an important
- 33:36distinction we make is that our agent is
- 33:37not just an LLM with a workflow on top
- 33:40of it. The backend may or may not be an
- 33:42LLM or it may or may not be other
- 33:44neurosymbolic reasoning tools under the
- 33:46hood. Um, and so we we try to keep that
- 33:49distinction a little bit clear, uh, that
- 33:51you're not just talking to like Sonnet
- 33:53or Gemini or whatever. You're talking to
- 33:55sort of an amalgam of systems based on
- 33:56what type of task you're executing at
- 33:58any point in time. Um, although when
- 34:00you're chatting, you are talking to just
- 34:02an LLM.
- 34:04Um, but yeah, so we have a template for
- 34:05the requirements. We have a template for
- 34:07this design doc because there's sections
- 34:08that we think are important to cover. Um
- 34:11and again like if you disagree and
- 34:13you're like I don't care about the
- 34:15testing strategy section just ask the do
- 34:18it and similarly the task list has is
- 34:21structured because we have sort of UI
- 34:23elements that are built on top of it as
- 34:24well like task management and um do we
- 34:28have [snorts] we'll get there when we do
- 34:30some property based testing but um
- 34:32there's some additional UI we'll add for
- 34:35things like optional you can have
- 34:36optional tasks and stuff like that and
- 34:38so we we need the structure there for
- 34:40our uh taskless LSP to work for example.
- 34:44Um yeah, thank you for the question.
- 34:46Anything else before we truck on?
- 34:50Cool. Uh I may need somebody to remind
- 34:52me what we were doing. Oh, that's right.
- 34:55So, we went through and we synthesized
- 34:58the spec for adding memory and some
- 35:00amount of persistence to my agent. By
- 35:02the way, I didn't introduce you to this
- 35:04project. This project is called Gramps.
- 35:06It is uh it is an agent that I'm
- 35:09deploying to agent core to learn about
- 35:10it. I mentioned that. But what I didn't
- 35:12tell you is that is it is uh a dad joke
- 35:16generator.
- 35:17A very expensive one since we're
- 35:19powering it via LLMs, but effectively
- 35:22you're a dad joke generator. Jokes
- 35:25should be clean. They should be based on
- 35:26puns, you know, obviously bon bonus
- 35:28points if they're slightly corny but
- 35:30endearing. Um yada yada. So we're
- 35:32deploying this to the back end. So, the
- 35:34reason I want memory is because every
- 35:35time I ask the dad joke generator for a
- 35:37joke, it gives me the same damn joke and
- 35:40that's just super boring and my kids are
- 35:41not going to be excited about that. So,
- 35:43I want memory so that as I come back for
- 35:44the same session, I get different jokes
- 35:46over and over again. Um, that's the
- 35:49context on the project. So, we've come
- 35:51through here and we actually said we
- 35:52generated this thing, we did the task
- 35:54list. I said, "Hey, is this the
- 35:56idiomatic way to do it?" But what I know
- 35:58is that we didn't actually uh we're not
- 36:01using Agent Core's memory feature, which
- 36:03is probably a big oops. Um, and so, you
- 36:06know, quick show hands. Do we want to
- 36:07make the mistake and go all the way to
- 36:08synthesis and deployment, or should we
- 36:10fix it now?
- 36:11>> Who wants to fix it now because we know
- 36:12better?
- 36:14>> No, I want to make the mistake. Let's
- 36:16keep on trucking. I I had three yeses in
- 36:18a room full of nothing. So, we're going
- 36:19to make the mistake and then come back
- 36:21and fix it later. So, uh, let's say run
- 36:26all tasks
- 36:29in order.
- 36:33Uh, the reason I mention in order, which
- 36:35seems very specific, is because this is
- 36:36a preview build of Kira. Um, and so
- 36:39somebody just added to the system prompt
- 36:40I should only do one task at a time. And
- 36:42I found that if I say run all tasks, it
- 36:44thinks I somehow mean do them all in
- 36:45parallel. So, we'll that'll be fixed
- 36:47before these changes get out to
- 36:49production. So Kira's going to keep kind
- 36:52of going through here and chewing away
- 36:53on the system in the back end. Um, it
- 36:55has steering docs that explain how to do
- 36:57its job. It has, which I guess I should
- 37:00show you guys. Steering again is like
- 37:01memory. So I have some steering how to
- 37:04do commits. Uh, you know, how I like to
- 37:06have commits, but also steering on
- 37:07things like how do you actually deploy
- 37:09this thing? Um, how do you deal with
- 37:11agent core? And then how do you run the
- 37:13commands that are necessary for you to
- 37:14deploy this to my local dev account. Um,
- 37:17and then those are mostly just an
- 37:19example again of sharpening your tools
- 37:21like uh I went through this kind of
- 37:22painful process of figuring out oh you
- 37:25know you have to use this parameter on
- 37:26the CDK
- 37:28the CDK command you have to use this lag
- 37:31otherwise it doesn't work correctly and
- 37:32so once I go through that pain of
- 37:34learning I just say kira write what you
- 37:36learned into a steering doc and it will
- 37:37usually do a very good job of
- 37:38summarizing um and so it generated
- 37:40automatically this Asian core langraph
- 37:42workflow MD file um yeah so I mean it's
- 37:46just going to kind of go away here and
- 37:48truck truck on and do its job and we can
- 37:51watch it in the background. But in the
- 37:52interim, um I think at this point we're
- 37:54at a pretty flexible spot. Uh so for
- 37:56folks who want feel free to use Kira,
- 37:59try out Spectriven Dev on your own. I'm
- 38:01going to keep just kind of running this
- 38:02in the background and taking questions
- 38:04and comments. But that's kind of it for
- 38:06the scheduled part of today.
- 38:09>> Yep.
- 38:10>> How does Carol work for like existing
- 38:12large code bases or this?
- 38:14>> Yeah.
- 38:16>> Yeah. question was how does cure work
- 38:17for large and existing code bases
- 38:19basically the brownfield use case uh and
- 38:21the answer is it depends on what you're
- 38:23trying to do um for spec driven dev you
- 38:25can ask cure to do research into what
- 38:26already exists so when you start a new
- 38:28spec it will usually start by reading
- 38:29through the the working tree um but the
- 38:33agent is generally starting from a a
- 38:35scratch perspective right it needs to
- 38:36understand the system um in practice
- 38:39what that means is that you're going to
- 38:40end up with a bunch of things like if
- 38:42your system already had good separation
- 38:44of concerns uh your the components in
- 38:47your system are highly cohesive and
- 38:49they're sort of highly coherent and
- 38:51highly cohesive, it's going to have a
- 38:53great job, right? It's going to be able
- 38:54to say this is the module that does this
- 38:56thing. I don't need to keep 18 things in
- 38:58my context to do my job and it's going
- 39:00to do well. Um if you let's just take an
- 39:04example that's off the top of my head.
- 39:06if you were trying to launch an IDE very
- 39:08quickly uh leading up to an AWS launch
- 39:11and you um you know took a lot of tech
- 39:13debt along the way that you need to
- 39:14unwind and you know nobody here would do
- 39:17that I'm sure but um in case you did
- 39:19that like me then your agent might
- 39:21actually have a much harder time
- 39:23traversing the codebase in the same way
- 39:24that a dev would right so uh from just
- 39:28kind of that perspective the more
- 39:30reliable things like your test suite are
- 39:31and the more understandable things like
- 39:33module separation and sort of
- 39:35decomposition of concerns are the better
- 39:37the agent will do. Um and versus true of
- 39:41course. Now for things like uh
- 39:44understanding the code base, this is a
- 39:46bad example because this is a very small
- 39:48code base, but uh we do have things like
- 39:52you know code search and workspace. Um
- 39:56uh I don't know what to call these
- 39:58context providers. Um, so you can come
- 40:00in here and just say I want to do code.
- 40:03Uh, what is it?
- 40:06I might have turned this off actually.
- 40:08Oh, I did turn it off because the code
- 40:10base isn't big enough. We'll do things
- 40:11like indexing in the background so the
- 40:13agent like you can do semantic search
- 40:15over what you've got um if you're just
- 40:17chatting. But in general, uh, Cur should
- 40:20go in and do sort of background search
- 40:22to figure out how to do its job. like as
- 40:24the codebase scales up, it's going to be
- 40:26less do probably less well overall. But
- 40:28that's one thing we're working on as a
- 40:30team.
- 40:31Did that answer your question or did I
- 40:33kind of glance off the side a bit?
- 40:35>> Yeah, I think I got it.
- 40:36>> Okay, cool.
- 40:38>> Anybody else?
- 40:44>> Uh, how long are you willing to wait for
- 40:46indexing to complete?
- 40:48>> [laughter]
- 40:48>> Uh so one example I have is that the
- 40:52code OSS um if it's not supremely
- 40:55obvious by looking at it cur is a code
- 40:56OSS fork just like you know cursor winds
- 40:59surf um
- 41:01one of the challenges we've had is the
- 41:03code OSS codebase is very large fairly
- 41:05large there's other big ones out there
- 41:07but that's kind of my large code base
- 41:09because I'm not forced get to work in it
- 41:12fairly frequently um and so there
- 41:14there's definitely some perceived
- 41:17slowdown when you're dealing with
- 41:18something large like that, especially
- 41:20when you talk about codebased indexing.
- 41:21It's a very active area of work for us
- 41:23though. So, we're trying to do things
- 41:24like um either remove indexing from the
- 41:27critical path so that you're not waiting
- 41:29there on some kind of slowed down render
- 41:32thread because indexing is running. Um
- 41:34but in practice, there should not be. I
- 41:37mean, again, the agent may practically
- 41:39do less well, but we're going to be
- 41:41talking in a couple weeks at reinvent
- 41:42about how some of the temple features in
- 41:45Curo were built via spec in a codebase
- 41:47we did not understand particularly well
- 41:48because we're just not VS code devs. Um,
- 41:52and Curo did a fine job of it. But
- 41:54again, that's a testament to the fact
- 41:56that codebase is reasonably well um
- 41:58structured
- 42:00>> and like if you've taken the time to
- 42:02understand how it works, it's very
- 42:04understandable. If you have not, it will
- 42:05might be a little bit opaque to stare
- 42:07at.
- 42:09>> Yeah.
- 42:10>> Uh in terms of indexing, is it like just
- 42:13just putting um um as much information
- 42:16from the code base into context or it
- 42:18just
- 42:19>> is there a way to like create some kind
- 42:21of like vector database of all the
- 42:25code base and then like query it? I just
- 42:30>> Yes. Um [clears throat] so the question
- 42:33was what do you mean by indexing? Um
- 42:36because indexing can mean a bunch of
- 42:37different things and what I mean is that
- 42:39um the agent is actually not provided
- 42:42the
- 42:43>> I'm going to keep the agent context as
- 42:44small as possible. We use the uh the
- 42:46index for most like secondary effects
- 42:48things like if you're doing a uh a code
- 42:52search or if I do something like search
- 42:53for um pound uh what the file in here
- 42:58http server like we use it more for
- 43:01these types of UI um than giving it to
- 43:03the agent because the agent does this is
- 43:05sort of anecdotal and based on our
- 43:07benchmarks does better when given less
- 43:10context but given the tools to
- 43:11understand where to go find things. Um,
- 43:13something we've heard a lot about is
- 43:14sort of incremental disclosure here at
- 43:16this conference. And that's again, we
- 43:18don't want to load too much at the
- 43:19beginning of the context and
- 43:20conversation with the agent. We want the
- 43:22agent to self-discover the right context
- 43:23for the task. Yeah.
- 43:26>> Thank you.
- 43:28>> Yeah.
- 43:28>> You guys managing session length like is
- 43:30there any kind of compression or
- 43:32pruning?
- 43:36>> Yeah. So, um, question was how do we
- 43:38manage session length? We have no
- 43:40incremental pruning today or incremental
- 43:42summary. Um you basically just accrete
- 43:44context until you hit your limit which I
- 43:46think right now I'm on auto which has
- 43:49like a 200k token limit um similar to
- 43:53the sonnetss. Um uh so we don't have a
- 43:57very sophisticated algorithm here yet.
- 43:58We've looked at a few things but our
- 44:00number one concern actually is um prompt
- 44:03caching hit rate. And so in a normal use
- 44:06case, I can achieve something like 90
- 44:0895% cash token usage here on per turn,
- 44:11which means that my interactions are
- 44:12very fast. And that's or they're much
- 44:14faster than the alternative, which is
- 44:16I'm sending 160k tokens to to bedrock
- 44:19cold. Um, so that's one of the reasons
- 44:22we've actually not done much
- 44:23experimentation with incremental
- 44:24summary. Um, our summarization feature
- 44:27exists. When you hit the cap, it's not
- 44:30great. It's something we're trying to uh
- 44:31ship an improved version very very
- 44:33shortly. Um eg in the next couple of
- 44:36weeks which should be faster. Today it's
- 44:38like a one-off operation that can take
- 44:41up to 30 or 45 seconds which is a
- 44:42horrendous experience. We're hoping to
- 44:44fix that here and make it sort of a
- 44:46real-time experience.
- 44:48The follow
- 44:50>> managing stapleness between sessions
- 44:52then is that how why you're relying on a
- 44:54stereopated
- 44:56spectrum.
- 44:59[gasps]
- 45:01>> So sort of um
- 45:05that is not the only reason I mean the
- 45:06spect the spectrum of dev is less to do
- 45:08with performance and more to do with
- 45:10reproducibility and accuracy of the
- 45:11agent. Um because if we can give you the
- 45:16right result,
- 45:17the the the way I and I think that we
- 45:20talk about it internally as this team is
- 45:22if I spend 10 seconds giving a prompt to
- 45:24the agent and then it goes off and it
- 45:26gets it wrong, it's like it's kind of no
- 45:28skin off my back, right? I burned
- 45:30however many tokens and you know,
- 45:32[clears throat] a couple cents of credit
- 45:33usage with whoever my LM provider is,
- 45:36but I spent 10 seconds generating a
- 45:38prompt. If I spend five to 10 minutes
- 45:40with the system producing a detailed
- 45:43design doc or let's just say even a
- 45:45detailed set of requirements I wanted to
- 45:46do a fairly good job. If I spend an hour
- 45:50generating a design doc reviewing it
- 45:52with my team and then synthesizing from
- 45:54that I wanted to get it right. So the
- 45:56goal necessarily is not just latency but
- 45:58actually accuracy when we talk about
- 46:00that. No, it's a both and. You need to
- 46:01do both. But um spec comes more from a
- 46:04uh the goal to have um highly
- 46:07reproducible output.
- 46:12I'm going to go over here first and then
- 46:13you
- 46:14>> Yeah. How did each of these task agents
- 46:16pass context to each other? And then are
- 46:18you only supposed to run this this
- 46:20parent task? Because it just finished
- 46:22all like 3.1 3.2 3.3 but then it still
- 46:26thought that 3.1 wasn't done and ran
- 46:28that in 3.2. too.
- 46:30>> Oh, did it?
- 46:31>> Yeah. Well, no, mine right.
- 46:32>> Oh, okay. Yeah. Yeah. Um, so
- 46:36if you
- 46:39the uh the question is if you're in the
- 46:41UI and you're like running tasks and I
- 46:43can just kind of pull up my task list
- 46:44here. Um, so if I just hit start, start
- 46:47start each of these is going to be a new
- 46:49session which means the context is
- 46:51completely unique. Um, personally I like
- 46:53to just if I can if I've got the context
- 46:55base to afford it, I just say do all the
- 46:56tasks because I find that more
- 46:59understandable and I think I actually
- 47:00get better performance. But by default,
- 47:02each task will be a new session that has
- 47:04no shared context with the previous
- 47:05ones. So the session is effectively just
- 47:08seated with your specification and then
- 47:10like here you're working on a spec that
- 47:12does all this stuff block of text um and
- 47:16you are doing this task da da da don't
- 47:18do any other tasks just do this. Um, so
- 47:21that sounds like a bug. Um,
- 47:22>> they ever spin up sub agents for certain
- 47:25things.
- 47:25>> We don't have sub agents yet in Caro,
- 47:27some we're working on.
- 47:28>> Yeah. Yeah. Because I mean, ideally,
- 47:30right, if we click on task three and
- 47:33I've got 31, 32, 33, and they're
- 47:35separated, there's no good reason I
- 47:36couldn't have different systems working
- 47:37on them. Yeah.
- 47:40>> Uh, right here,
- 47:42>> we do have in the Curo CLI custom agents
- 47:45that you can also run off.
- 47:47>> Yeah. Curli is a concept of custom
- 47:49agents. um which can be run sort of as a
- 47:52task um and it's something we're playing
- 47:53with right now in Curo Desktop um and I
- 47:55think you had another one
- 47:57>> yeah I'm sorry if I missed this but in
- 47:59the spec folder
- 48:01um as you do more and more of these
- 48:03tasks over time
- 48:05>> y is it just all in one design
- 48:08requirements tasks your whole project is
- 48:11defined there or did it group by
- 48:13>> that's a good question um yeah so I will
- 48:16have many I will have uh the question
- 48:18was as you do more you generate let's
- 48:21say more specs over time. Are you sort
- 48:23of just creating one massive spec and
- 48:26no? Uh let me open a different project.
- 48:45>> [clears throat]
- 48:49>> So this is for example the curo
- 48:51extension which is like a 1p extension
- 48:53inside the curo IDE. This is where the
- 48:55agent itself lives. And so we have
- 48:57pruned some specs but there are specs in
- 48:58here that we can talk through or I can
- 49:00just kind of demo. Um
- 49:04so these are the way I think about it is
- 49:06that the spec sort of represents a
- 49:07feature or a problem area in the in the
- 49:09project. And so for example, I can blast
- 49:12this a little larger. So for example, we
- 49:15have um like some of these are just
- 49:19tests. We've done things like oh could
- 49:20we have a prompt registry? Could we have
- 49:22a prompt registry file loader? They may
- 49:24or may not make it all the way to
- 49:25production. Um I want telemetry on the
- 49:27chat UI. So these are just like somebody
- 49:29will go off and spend maybe represents a
- 49:32few days of work for an SD. Um, agents
- 49:35MD support is a good one where we just,
- 49:37you know, I sort of said research what
- 49:39agents MD is and build it in the way you
- 49:41build steering in like support in the
- 49:42same way. This spec is fairly unlikely
- 49:44for us to come back and revisit in the
- 49:46future. So I may actually just delete
- 49:47it. Um, which is what we've done with
- 49:49some of the older ones. But a good
- 49:50example of one that we might come back
- 49:51to is our message history sanitizer. So,
- 49:54one thing we've had issues with or we
- 49:56had issues with early in the the
- 49:58development of Kira is that we would
- 49:59send these sort of invalid um sequences
- 50:02of messages because let's say the
- 50:04anthropic API required tools to be in
- 50:08the same order they were invoked and the
- 50:09responses but the system wasn't doing
- 50:11that. So we built this whole sanitizer
- 50:12system that has a bunch of requirements
- 50:14around um
- 50:17let's see very specifically
- 50:20yeah when conversation is validated the
- 50:23system shall verify that each user input
- 50:24is either non-MPT content or tool
- 50:26responses. So we had things where like
- 50:28empty strings would get passed in but
- 50:30there was a tool response. This is a
- 50:32good example where we've come in over
- 50:33time and actually just added maybe not
- 50:35to the requirements but to the to the
- 50:37acceptance criteria of the requirements
- 50:38as new validation rules are uncovered.
- 50:42>> Yeah.
- 50:42>> So how do you handle like that? So for
- 50:45example you have like
- 50:47>> telemetry up there y feature that needs
- 50:50telemetry is it going to go back and
- 50:52update that spec too or you're just
- 50:53>> it should. Yeah. So, if you usually
- 50:55you'll see and let me just ask uh
- 50:58a new chat here.
- 51:05No, that's a terrible idea.
- 51:21So here I've asked I've made a inspect
- 51:24mode I've made some requests to um add
- 51:26UI telemetry to the thing I'll help you
- 51:28add it let me first check if there's any
- 51:30relevant runbooks then explore the
- 51:31codebase and sand the implementation it
- 51:34might go do a little bit of research
- 51:35here and then flip of a coin again it's
- 51:38an LLM so it may or may not discover the
- 51:41existing uh spec but ideally it will
- 51:44after doing its research say there
- 51:46exists a spec already for things like UI
- 51:49telemetry, I'm going to go and amend
- 51:50that one. Um, and if it doesn't in this
- 51:52case, like I would come in and just ask
- 51:54it to um as sort of the operator of the
- 51:57system. But over time, again, we want
- 51:58that to be easier for you as a user to
- 52:00not have to think about so much.
- 52:04We can watch it while it chugs along.
- 52:10>> Is there anything reconfigured in Kira
- 52:13that makes it better to work with AWS?
- 52:16trans.
- 52:18>> No, not really. Um,
- 52:21>> was that a question?
- 52:23>> Oh, question was, uh, is there anything
- 52:25in Kira that that's preconfigured to
- 52:27make it work better with AWS? No. Um, we
- 52:30are sort of purposefully we're in we are
- 52:33brought to you by AWS, which so you
- 52:35know, uh, Andy Jasse and Jeffy B pay my
- 52:38check, but um, we're not like an AWS
- 52:41product that's deeply deeply integrated
- 52:43with the rest of the AWS ecosystem. Now
- 52:45that said, I still answer emails when
- 52:47somebody says, "Why is this other thing
- 52:48we built with AWS not working with
- 52:50Curo?" Yay. But um similarly like if
- 52:53you're building on GC or Azure, whatever
- 52:56um or you're running some on-rem system,
- 52:58the product should work just as well for
- 53:00you. That's our goal.
- 53:02>> Good a good answer potentially is the
- 53:04AWS documentation MCP server.
- 53:07>> Yes.
- 53:08>> So there are MCP servers that you can
- 53:10add into any of these things that will
- 53:12make better.
- 53:14Yeah, that's a good point. So, like in
- 53:16this case, I actually had to add the AWS
- 53:20MCP documentation here. We could of
- 53:22course have natively bundled this, but I
- 53:23don't want to ship this to customers who
- 53:24don't need it.
- 53:26Yeah, because again, AWS is not the only
- 53:29docs that we might care about. Um, by
- 53:31the way, coming back to your question,
- 53:32so it did find the existing spec for
- 53:34telemetry. It read it, it read different
- 53:36sections of it, and now it's actually
- 53:38making amendments to it. So, we can
- 53:40follow the diff as it shows up here. So,
- 53:41it's added new uh requirements. um to
- 53:44the pre-existing specs. So, this is
- 53:46effectively another case where we're
- 53:48mutating the system as opposed to just
- 53:50adding this sort of never- ending spiel
- 53:52of specs.
- 53:53>> I guess what I'm wondering is like how
- 53:56did it know or decide where to put the
- 53:59spec, you know, if you break down your
- 54:02project into these different categories?
- 54:04>> Y
- 54:04>> I would imagine like crossover.
- 54:07>> Yeah. I mean, it's that that's sort of
- 54:09like software development in a nutshell
- 54:12though, right? like how do you actually
- 54:13define the seams between different parts
- 54:14of your system different concerns the
- 54:16product
- 54:17>> right but if you want to like build
- 54:18something like I have a task and it's
- 54:20going to cost
- 54:21>> require changing like three or four
- 54:22things
- 54:23>> y
- 54:23>> it's going to change three or four specs
- 54:25and then run tasks across three or four
- 54:27>> oh yeah yeah no it should not do that it
- 54:29would probably so again I don't have a
- 54:31good example off hand that we can do for
- 54:33that but um my my perspective would be
- 54:36that if you're working on something that
- 54:37is a crossf functional uh by the way the
- 54:39question was um if I'm working on
- 54:41something that let's say I have a spec
- 54:43for security requirements and I have a
- 54:46spec for API design uh like the API
- 54:50shapes and I have a spec for
- 54:52logging and I am changing something in
- 54:55the API public interface that is a
- 54:58securityf facing concern because we're
- 55:01redacting logging PII um I think that's
- 55:04maybe a semi-tangible use case uh that
- 55:07we can all imagine coming down from our
- 55:09governance teams um I want to
- 55:13I would imagine that you either pick one
- 55:15of those to load the requirements into
- 55:17or you create sort of a cross functional
- 55:19spec, but that would come down to I
- 55:20think you as a as an operator making
- 55:22that decision in much the same way that
- 55:24if I how you actually implement it might
- 55:27be you you would not necessarily
- 55:30implement my PII API redaction module.
- 55:33It's a standalone thing. It's going to
- 55:34be a crosscutting theme across your
- 55:36codebase, I'd imagine. And it's also a
- 55:39good example. There's like multi group
- 55:40workspace [laughter] came out when it
- 55:42went to G on Monday and now you can like
- 55:44drag different. So like in your example
- 55:46you just went through with like APIs and
- 55:48off and like even the front ending you
- 55:51can bring in those projects if you have
- 55:53them separately and then still work.
- 55:58>> Yeah. Thanks bro.
- 56:04the mental model the spec generates the
- 56:07code after that like what code you can
- 56:09specify how does that work
- 56:12>> yeah so um we have now synthesized
- 56:16effectively the spec so we we sat down
- 56:19we defined the requirements design and
- 56:21task list I've had Kira now go through
- 56:23and run all the tasks in this spec so it
- 56:26ran them one at a time it basically
- 56:27worked on small bite-sized pieces of
- 56:28work uh chunk by chunk and then uh now
- 56:33this is done So what we've actually
- 56:35produced is not just like the completed
- 56:37spec, but it went here into my agent and
- 56:40it did a few things in the CDK repo
- 56:43because it's doing persistence to S3.
- 56:45I'm sure it added a bucket. Yep. Some
- 56:47new bucket encryption and yada yada. It
- 56:49then went in to the agent, added the S3
- 56:52checkpoint saver. It looks like it, you
- 56:55know, created a checkpointer. It adds
- 56:57this to the graph and it kind of passes
- 57:00this all the way through the system. And
- 57:02the S3 checkpointer here I'm sure has
- 57:03some knowledge of how to write the
- 57:05checkpoints to and from S3. So like we
- 57:07have gone not just for defining the
- 57:09system but we've now um produced it end
- 57:11to end or we've uh delivered it end to
- 57:13end including property tests I believe.
- 57:16Um yeah.
- 57:20>> Oh, I have a answer to an earlier
- 57:23question related to like um some
- 57:25specific AWS related features like that
- 57:28makes it easier to work with. The Curo
- 57:30CLI comes with the use AWS tool which
- 57:33helps with the CLI.
- 57:36>> Yeah. Yep. So, uh, what Rob's pointing
- 57:38out is the Curo CLI, which we just
- 57:39rebranded, um, this week, has a use AWS
- 57:42tool, which is basically a wrapper over
- 57:43the AWS SDK, um, to make some of those
- 57:46things easy. Uh, but again,
- 57:49BYO use GCP tool as an FCP server if you
- 57:52were so inclined, if that's your uh,
- 57:55tool of choice. And I believe, don't
- 57:57quote me on this, um, because the CLI is
- 58:00kind of new to my new to me I should
- 58:02say. Um, but I believe you can turn off
- 58:05tools in the CLI as well. Let me know if
- 58:07that's not right, Rob.
- 58:09>> Yeah. So, that's like you're actually
- 58:11not strict. Uh, in the desktop product
- 58:14today, you can't control the tools, the
- 58:16native tools built in, but in CLI, you
- 58:18can.
- 58:21>> Um, so I I intuitively get the benefits
- 58:24of having a spec. Have you done any work
- 58:26to empirically see like how a project or
- 58:31a problem would have worked with or
- 58:32without?
- 58:34>> Yeah. Um we do have benchmarks uh
- 58:36covering the data off hand. Um I think
- 58:39part of that's in our blogs. So if you
- 58:41go to the cure.deblog
- 58:43or it's on the site, we we talk really
- 58:45crisply about some of the lift things
- 58:46like property based testing give to task
- 58:48accuracy.
- 58:50Science team's always working on that
- 58:52stuff.
- 58:53a blog about specs. I'm curious about
- 59:00>> Yeah. Distinguish engineer for
- 59:02databases. Yeah.
- 59:03>> His blog post really steps it up. I
- 59:06don't think it has the D specific that
- 59:07you are asking for, but I think it will
- 59:09be useful.
- 59:10>> Yeah. Yeah.
- 59:14>> How does it work? I understand the
- 59:16feature side of it, but how does it work
- 59:18in a nonfunctional site like agency
- 59:21dealing with, you know, a little bit
- 59:23more harder problems?
- 59:25>> Well, yeah. I mean, that is ultimately
- 59:26the goal here, right? Is we're saying
- 59:28you're making a slightly larger
- 59:29investment up front, but we believe that
- 59:30the uh the structure we're bringing is
- 59:33going to help you get increase the
- 59:35accuracy of your uh result. So, um,
- 59:38while we've got a team of people who are
- 59:40basically working on making spec better,
- 59:41my job when I fly back to Seattle is to
- 59:43make cur as a whole much faster. Um,
- 59:46one, execution time and like kind of
- 59:48like laggginess in the UI, but two, how
- 59:50do we get tokens through the system
- 59:51faster? How do we get responses to you
- 59:53faster so that like you're not syncing
- 59:55as much cost into KO to use a spec?
- 59:58>> Yeah. Yeah. I'm not talking about the KO
- 1:00:00tool itself, the code generated from the
- 1:00:03spec.
- 1:00:04>> Oh. Oh, yeah. Okay. Yeah. you mean like
- 1:00:05the non-functional requirements of the
- 1:00:07generated code? So, uh that's going to
- 1:00:09come down to I think what you're
- 1:00:10specifically trying to do. So, you could
- 1:00:13add uh one of the slides I had here was
- 1:00:16talking a little bit about how to tweak
- 1:00:18the process and tweak the artifacts for
- 1:00:19your use cases. Um again, you could very
- 1:00:23easily add something like I want
- 1:00:24non-functional requirements for speed
- 1:00:26and runtime and things like lock
- 1:00:27contention to be considered in the
- 1:00:29design phase. Um yeah, [clears throat]
- 1:00:31something you could certainly add. So
- 1:00:33you could generate a code in Rust or or
- 1:00:36Java.
- 1:00:37>> Yeah, totally. Yeah.
- 1:00:39>> And it will vary in the functional
- 1:00:43depending on what language you
- 1:00:44generated.
- 1:00:45>> I mean it would it would have to like
- 1:00:46yeah there's no other way I think to
- 1:00:48approach it. Um again I'm just I'm
- 1:00:50familiar with node so I'm doing
- 1:00:51everything here in node but you can use
- 1:00:52this with any language. I think
- 1:00:54technically we say we support Java,
- 1:00:56Python, JavaScript, um, and
- 1:00:59Jesus, JavaScript, TypeScript, Java, and
- 1:01:03Rust. But in practice, there's no reason
- 1:01:05that this doesn't work with any
- 1:01:07language. I mean, it's just an LLM. The
- 1:01:09there's nothing language specific or
- 1:01:10framework specific in the system. And
- 1:01:12for those of you um, so there was a
- 1:01:14conference earlier this week hosted by
- 1:01:16Tessle, which are doing sort of specs
- 1:01:18for knowledge base. um as long as you've
- 1:01:20got the right grounding docks in there
- 1:01:22and this is sort of uh their argument is
- 1:01:25that it should not matter what you're
- 1:01:27building like that's all just informed
- 1:01:29by the the context you're building for
- 1:01:30your system.
- 1:01:32>> This is also a really good point for
- 1:01:34steering. So steering you can get the
- 1:01:36agent to develop code in the way you
- 1:01:38want. Like being a developer is all
- 1:01:39about making trade-offs and the problem
- 1:01:41with your out of the box is it's like so
- 1:01:43polite because it's trying to be
- 1:01:44everything to everyone. U and especially
- 1:01:46like with latency and cost and other
- 1:01:48things like that, just tell it in
- 1:01:50steering what you want it to prioritize
- 1:01:53and then that will influence any code
- 1:01:54that gets generated.
- 1:01:56>> Yep.
- 1:01:56>> Even like how it designs based on that
- 1:01:58as well. So if there's something that's
- 1:01:59very specific to your use case or your
- 1:02:01industry or whatever, just shove it in
- 1:02:02that steering file and then
- 1:02:05>> Yeah, that's exactly right. So, for
- 1:02:07example, I I will have Kira generate um
- 1:02:11commits for me. And one of the things I
- 1:02:13care I personally care about is that I
- 1:02:15can track commits I generate versus
- 1:02:17commits that Kira generates being the
- 1:02:18ones that come from the system. And so
- 1:02:20my steering dock while short includes
- 1:02:22things like very specifically my
- 1:02:25requirement for Curo is
- 1:02:28just use the UI
- 1:02:32um
- 1:02:34attributed to the co-author of Kuro
- 1:02:36agent um which is trivial but also I
- 1:02:40want it to happen every time. So in this
- 1:02:41case it just generated a commit
- 1:02:42co-authored by Kirao agent D. So that's
- 1:02:46an example of like you could add
- 1:02:47whatever you want in there, not just
- 1:02:49something related to get commits, but
- 1:02:51you could do code style, you could do um
- 1:02:53uh you know code style, code coverage.
- 1:02:56Uh whenever you add a spec or you're
- 1:02:57adding a new module, make sure that you
- 1:02:59annotate it with coverage minimums that
- 1:03:01are 90% because that's the thing I care
- 1:03:02about. Um [clears throat] you can kind
- 1:03:05of put anything you want up in there.
- 1:03:07The good news is it looks like what we
- 1:03:09built works. Um, Cur is very happy with
- 1:03:11itself at least and it looks like all
- 1:03:13tests passed. But um, yeah, so we'll we
- 1:03:16can deploy this to the back end and see
- 1:03:17how things work.
- 1:03:21We're uh technically just about time.
- 1:03:23So, you know, if anybody has any other
- 1:03:24questions, I'm going to stick around
- 1:03:26here for a while. But uh, thank you all
- 1:03:28for joining, listening, and uh, learning
- 1:03:30a little bit more about Spectrum and
- 1:03:32Dev. [music]
- 1:03:41>> [music]
- 1:03:47[music]
- 1:03:48>> Heat.
About this transcript
This page contains the full transcript of Spec-Driven Development: Agentic Coding at FAANG Scale and Quality — Al Harris, Amazon Kiro by AI Engineer, generated from the public captions YouTube serves with the video. The transcript has 12,123 words across 1,725 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.