[한영자막] Cursor 핵심 개발자 Lauren Tan: AI 에이전트를 실전에서 제대로 신뢰하는 법 (xAI GrokBot 워크숍) — Transcript
Full transcript
- 0:00Everyone, I am Lauren Lauren Tan. I
- 0:03guess not many people know my last name.
- 0:05Uh, I am Potato on Twitter. Uh, potato
- 0:10with spelled with an E. Um, and I have
- 0:13been at Cursor for about 5 months. Uh,
- 0:18previously I was at Meta where I worked
- 0:20on the React team. Uh, specifically
- 0:22working on the React compiler, uh, which
- 0:25was a whole lot of fun. Uh I'm still on
- 0:27the on the core team and and uh
- 0:29contributing to open source here and
- 0:30there. Uh so that that's really nice
- 0:33that they still let me do that. Uh and
- 0:36before Meta, I was at Netflix uh where I
- 0:38was uh both a tech lead and uh I
- 0:43transitioned to be be an engineering
- 0:45manager um for about two years.
- 0:49So I've had a I've have had a lot of
- 0:51experience going between engineering
- 0:54management and being an individual
- 0:56contributor. Uh and I think something
- 0:58I've noticed actually which is quite
- 1:00interesting is that there are so many
- 1:02parallels with you know management
- 1:04skills and how to like manage agents. Uh
- 1:07and that's actually a big part about
- 1:09what I wanted to chat with you and
- 1:11everybody else about today. Um but yeah
- 1:14that's that's me. Uh I do have some like
- 1:17very light slides but uh it's not going
- 1:20to be uh just rambling. So let me just
- 1:23share my screen
- 1:25and hope that I don't leak anything. Uh
- 1:33oh no I need to allow permissions.
- 1:36>> No worries. Oh take your time. There's
- 1:38always tech tech trouble. Uh give me one
- 1:41second to rejoin.
- 1:42>> Yeah go for it.
- 1:46I see many of you already know Lauren
- 1:48from from the looks of the chat here.
- 1:50Um, so yeah, it's exciting to to get a
- 1:53chance to chat with her uh and and go
- 1:55through some of her her recent work. As
- 1:57you guys heard, you know, a lot of
- 1:58recent experience from from Netflix to
- 2:01to Meta and then now over at Cursor. Uh,
- 2:04we're going to chat a little bit about
- 2:05Grockbot as well. So, uh, that'll be
- 2:08exciting. I don't know if you guys saw
- 2:09that was a recent release. I think
- 2:10literally maybe yesterday or the day
- 2:12before uh from the cursor team which is
- 2:14kind of like um let's call it like
- 2:16agents for everyone. You can go check it
- 2:18out if you want and learn a little bit
- 2:19more about the product but um but yeah
- 2:22we'll we'll explore that a little bit
- 2:23today as well.
- 2:26All righty. Welcome back.
- 2:40And Lauren, you're just on mute there if
- 2:42uh you want to hop off of mute if you're
- 2:43chatting.
- 2:44>> Yeah, sorry.
- 2:45>> No worries.
- 2:47>> It's 2026 and I still don't know how to
- 2:49use Zoom.
- 2:50>> That's [laughter] all good.
- 2:51>> Uh okay. So, I assume you can see my
- 2:53screen.
- 2:54>> Yes. Yeah, we're good. Um, so yeah,
- 2:56today, yeah, I think I think the big
- 2:58theme for me as I've been using agents
- 3:01to write code, and I'm sure a lot of you
- 3:03have had the same experience as well, is
- 3:06how do you trust it? You know,
- 3:07especially if you are an engineer that's
- 3:10been writing code for a very long time,
- 3:12you have a lot of opinions and lessons
- 3:15that you've learned about doing good
- 3:17engineering. And when you see agents
- 3:20just, you know, winging it and, you
- 3:22know, guessing, hallucinating,
- 3:24uh, you know, confidently stating that
- 3:27they found the smoking gun, uh, for the
- 3:29hundth time, uh, but it's actually not
- 3:32the real problem. You lose a lot of
- 3:34trust. And when you lose when you don't
- 3:36have much trust in your agents,
- 3:39I feel like you you really can't get the
- 3:41most out of them. And for me, the
- 3:43parallel is like with management. Uh so
- 3:46if I'm an a manager an engineering
- 3:48manager of a team and I have a bunch of
- 3:51you know I have a team of engineers uh
- 3:54on my team and I don't trust them then
- 3:57the mode of operation I'm going to be in
- 3:59is going to be like micromanagement
- 4:00right I'll have to spend a lot of time
- 4:03looking over my reports shoulders and
- 4:06checking that they're doing their work
- 4:08well you know that they're not shipping
- 4:09bugs to production
- 4:13and so I drew this chart cuz uh it's not
- 4:16it's not a very scientific chart but
- 4:18like this is how I imagine myself and my
- 4:22journey through using agents. So you
- 4:25know like fast forward or back forward
- 4:29uh or fast back uh fast backwards like a
- 4:33year or so when you know nobody was or
- 4:35not many people were using agents to
- 4:37code. Uh I think you
- 4:41uh you know get into this mode where you
- 4:43are
- 4:45in very heavily in the loop with one or
- 4:48several like a handful of agents and you
- 4:51find yourself just constantly fig uh you
- 4:54know trying to understand what your
- 4:55agents are doing uh and you're very very
- 4:57in loop. You're watching every single
- 4:59output. you are sitting there prompting
- 5:03um and you really can't parallelize
- 5:05beyond that because you don't again you
- 5:08don't have that trust right you can't go
- 5:09to a 100 agents uh like spawn 100 agents
- 5:13when you don't even trust the output of
- 5:14one agent
- 5:17so over the past 5 months I feel like
- 5:20I've really been able to uh like ascend
- 5:23this trust curve and now I'm at the
- 5:26point where uh I actually have this
- 5:29Sounds kind of scary to say this and it
- 5:31makes me sound like a slop artist, but I
- 5:33I promise I'm not, but I actually have
- 5:36my agents now um automerging PRs for me.
- 5:39Uh which is like a wild thing to say,
- 5:42but um like I woke up today and there
- 5:44were like 20 PRs landed and I just
- 5:47reviewed them on Maine like they were
- 5:49already landed and they were good. Uh so
- 5:52how did I get to that point is basically
- 5:55what I wanted to talk about today.
- 5:58Uh and again like yeah feel free to jump
- 6:00in if you have questions Colin. Um but
- 6:04uh oh yeah of course I got to show this
- 6:07this chart. Uh where uh
- 6:12no do not trust to someone requested to
- 6:16control my computer. Uh probably won't
- 6:18do that. Uh but yeah, so this chart I
- 6:22think I I'm I'm sharing this chart not
- 6:24to kind of like flex but to kind of show
- 6:27like the journey like so you can see
- 6:29like the curve like it sort of like
- 6:31inversely matches the contributions I've
- 6:34been able to land at cursor. So I joined
- 6:37five months ago and five months ago like
- 6:39I you know my first month I was like not
- 6:41very productive because I was ob you
- 6:43know I was learning the codebase didn't
- 6:45know what the heck was going on and as I
- 6:48got more confident in in my agents uh
- 6:51I've really been able to kind of ramp up
- 6:53my productivity uh and again like yeah
- 6:56like last month I shipped a thousand PRs
- 7:00which is ridiculous. Uh, and then this
- 7:02month we're only on the 12th. I'm
- 7:04already at like almost 800 PRs landed.
- 7:09Uh, so the velocity is definitely high
- 7:12and you you I'm sure a lot of you will
- 7:14definitely be questioning like how how
- 7:16much of this code is actually good. Um,
- 7:18and I think yeah, like that's definitely
- 7:21fair to question.
- 7:23Um, but uh, yeah, I think I think if you
- 7:28set up your agents well, you can
- 7:30definitely get to a very similar level.
- 7:34Um, and so I'm going to talk about how
- 7:35we do that.
- 7:38Uh so for me I think I'm curious like I
- 7:42guess call in your experience as well
- 7:44but uh for me I think the most important
- 7:48skill that you should have in your
- 7:51toolbox when you work with agents is
- 7:53verification.
- 7:55Uh and by verification I mean the
- 7:57ability for an agent to actually run the
- 8:00code uh or take CPU traces or heap
- 8:05snapshots or uh you know open an iOS
- 8:09simulator whatever you know however your
- 8:12application is exposed to your users it
- 8:15can do the same thing and uh run it for
- 8:19real and actually test and verify it
- 8:21don't work because that's the thing that
- 8:23really closes the loop. Uh it doesn't
- 8:26guarantee your agent writes good code.
- 8:29Uh but it allows them to at least write
- 8:31correct code. Uh which is a big a really
- 8:34big step forward for being able to trust
- 8:37your agents. Um
- 8:40I will I can share one example that we
- 8:43have uh within cursor. Uh
- 8:49oops
- 8:50where let me open this. Let's
- 9:00make this uh make me full screen. There
- 9:03you go.
- 9:05Uh so for the for cursors agent window
- 9:09uh so this is actually an interesting
- 9:11story but uh when I joined cursor 5
- 9:14months ago uh they're actually uh well I
- 9:18was supposed to join a different team. I
- 9:20was supposed to join like the cloud
- 9:21agents team. Uh but then since I have a
- 9:24lot of experience working on react and
- 9:26agents window is a react application. Uh
- 9:29I was
- 9:31uh I was asked to basically help out
- 9:34with uh the agent window work. Uh but um
- 9:40there wasn't really a lot of like skills
- 9:43to help me. So I just found myself like
- 9:45okay uh agents is going to launch in
- 9:47like a week right we have a really tight
- 9:49deadline and um there was uh you know I
- 9:54was just sitting there like okay I'm
- 9:55going to open up the performant the the
- 9:57chrome dev tools and just like take a
- 9:59trace look at it myself and try to make
- 10:01sense of this flame graph and keep in
- 10:03mind I was just like in my first week so
- 10:06I had no idea what I was looking at no
- 10:07idea what you know I mean I had some
- 10:09idea but you know the codebase was
- 10:11completely fresh to me Uh, and I
- 10:14realized like my agent had no idea
- 10:17either, you know, like I would take a
- 10:18screenshot of the tree, I download
- 10:19trades, I would send it to it, and it be
- 10:21like, "Yeah, it kind of looks like this,
- 10:23you know, uh, and it would like
- 10:25confidently state like it's this thing."
- 10:27And then I try to fix that. And turns
- 10:29out that's not the actual thing.
- 10:32So, this was very very slow process. And
- 10:34if you've ever done any like performance
- 10:36work yourself or you know just even
- 10:38development with an agent where you
- 10:40don't have a verification skill, you are
- 10:43the verifier, right? You you're the
- 10:45bottleneck. You you you tell your agent
- 10:47to do something and then it goes off and
- 10:49write some code. Then you open up your
- 10:51you know local dev build and then you
- 10:53start to say oh you know doesn't work.
- 10:54Then you got to copy paste screen uh you
- 10:56know screenshots or console errors or
- 10:59whatever. uh and then your agent like
- 11:01slowly kind of like uh you know works
- 11:04with that and then tries to understand
- 11:06it and um fix the thing but then you're
- 11:10constantly just in the loop and and
- 11:12being a bottleneck. So there's really no
- 11:13way to parallelize. So the control glass
- 11:16skill is like one of the first skills I
- 11:18built uh for cursor. Uh and glass by the
- 11:21way is the code name for agents window
- 11:24that we use internally but it's just
- 11:27cursor I guess. Um, and so this skill uh
- 11:31is I guess the the the code itself is
- 11:33not super interesting. Your agent can
- 11:35very easily make one for you. Uh where
- 11:38if if you're building an Electron app or
- 11:40a web app or even iOS uh applications,
- 11:45uh you can teach your agent how to use
- 11:47like the ChromeDev tools protocol or
- 11:50through uh Apple has some utilities as
- 11:53well for running the simulator and
- 11:55taking traces and controlling
- 11:56programmatic control.
- 11:58as well. Uh so that's really useful. Uh
- 12:03but one thing I want to talk about is uh
- 12:06the
- 12:08this thing
- 12:10uh where is the read me? Uh so this
- 12:14skill comes with this very unique
- 12:16feature called or not feature uh unique
- 12:19file called a feature map. And so the
- 12:23story then is like I built this skill
- 12:25and so now the agent was able to uh
- 12:28actually run the agent window and take
- 12:30traces and whatnot. Uh but it had no
- 12:33idea what what the agent window was. So
- 12:38um you know like someone would say like
- 12:40oh the the left sidebar is like laggy or
- 12:44something like that or you know the
- 12:46right side the the PR tab is not working
- 12:48and the agent would just be like kind of
- 12:50flailing around. it would spend a lot of
- 12:51time trying to like look up the code and
- 12:53you know where is this feature? How do I
- 12:55actually get to it on the UI which made
- 12:58it basically completely useless. Uh you
- 13:00know like we would I would run this
- 13:03skill locally and you know it would
- 13:05spawn a dev build uh but then it' just
- 13:08be churning like I just try to click
- 13:11here. It wouldn't know how to get to
- 13:13things um and it was just an awful
- 13:15experience. Uh so who is putting arrows
- 13:19on my screen? Um so uh yeah this this
- 13:24feature map has been really useful uh
- 13:26because it teaches the agent how to get
- 13:28to all of the features that you have. Um
- 13:31and in PAC the plugin that I I've made
- 13:35uh if you search for PAC cursor on
- 13:39Google you you'll find it. Uh, but there
- 13:41is a create verification skill in that
- 13:44plug-in where it actually helps you set
- 13:46up something like this for yourself. Um,
- 13:49including the feature map. So, it will
- 13:50actually explore the code and build up
- 13:52this initial feature map that tells your
- 13:54agent how to get to all of the different
- 13:57features that you have. Uh, and this is
- 13:59extremely powerful because now like you
- 14:02have these user reports that come in.
- 14:04Uh, you you can actually map even like a
- 14:07vague report or even a screenshot. So we
- 14:10have this uh internally at cursor where
- 14:14uh we have a slack channel where you
- 14:16know lots of people giving us feedback
- 14:18on the agents window and rockbot and
- 14:20whatnot. Uh, and often times the report
- 14:24is very bad, like very low quality, like
- 14:26someone will just put very often we get
- 14:29like a screenshot like and then someone
- 14:30just says question mark question mark
- 14:32question mark like what is this? And you
- 14:34know like without this your agent like I
- 14:37have no clue, right? But with a feature
- 14:39map like this, it has a lot more context
- 14:42and understanding of how to actually
- 14:45navigate, how to get to all the
- 14:47different features. Uh so like you know
- 14:49example like I guess like the sidebar
- 14:51like what is the sidebar uh you know
- 14:54like all the different sub features that
- 14:56are present in it um like from the user
- 14:59point of view here's where how do I get
- 15:01to it all the different keyboard
- 15:03shortcuts
- 15:05uh even like the the what do you call it
- 15:07the DOM elements or yeah like the
- 15:10attributes that you use for selecting
- 15:12things through the CDP uh are all there.
- 15:16So uh again yeah this is like really
- 15:19really powerful uh for for agents
- 15:23>> uh and pstack ships uh that create
- 15:26verification skill but also a maintain
- 15:28verification skill uh so you can keep
- 15:30this up to date.
- 15:32>> Cool. Yeah, I was just going to ask how
- 15:34you created that. So do you mind sharing
- 15:35a little bit more about um that that
- 15:37process in the context of Pstack and
- 15:39maybe just what Pstack is for the folks
- 15:40who aren't familiar?
- 15:42Yeah. So, P stack is pretty interesting
- 15:44because uh well, first of all, the name
- 15:47is kind of goofy. like the P the P and P
- 15:50stack is like potato potato snack
- 15:54because I um so uh uh there is a pretty
- 15:59uh famous person Gary Tan who is the CEO
- 16:02of Y Combinator and he's come up with
- 16:05this plugin called GStack uh Gary Stack
- 16:09and uh funnily enough we share the last
- 16:12name we have no relations uh but I
- 16:14thought it would be funny to kind of you
- 16:16know poke fun at Gary and make peace
- 16:18stack my version of of of of [laughter]
- 16:21his plugin. Uh but kind of tailor it to
- 16:24my own set of p engineering practices.
- 16:29Uh but I honestly actually never set out
- 16:31to build PAC. Uh it just started with a
- 16:33bunch of skills, right? Like I started
- 16:35with that control glass skill and then I
- 16:37started with another skill like called
- 16:39how which I also noticed through like
- 16:42observing agents. Um, so like you know
- 16:46in the early days of me, you know,
- 16:47trying to climb this ladder, I was like
- 16:49super in the loop and I was basically
- 16:51nitpicking my agents to an extreme
- 16:53degree. I was uh like I would tell it um
- 16:57you know this feature has stopped
- 16:59working here's a bug report like why
- 17:01isn't it working?
- 17:04And very often the agent would just like
- 17:06confidently state like oh it has to be
- 17:09this right it has to be this thing. And
- 17:11I noticed like when I looked at the
- 17:13actual tool calls, I noticed it wasn't
- 17:15actually reading the code that I thought
- 17:18should be affected. And that made me
- 17:20just extremely suspicious. And at that
- 17:22point, I was like, I'm not going to I
- 17:24can't trust any this agent anymore cuz
- 17:26it's just it's just completely
- 17:27hallucinating.
- 17:29And I think
- 17:31I think it's very easy to just, you
- 17:33know, like build up that distrust and
- 17:35not and kind of feel helpless like, you
- 17:38know, you don't know how to help your
- 17:40agents succeed. But like again, I think
- 17:43the the the management analogy is super
- 17:45helpful because like imagine if you were
- 17:46a manager of an engineering team and you
- 17:49had an engineer on your team who was a
- 17:52really good coder, no business context
- 17:54whatsoever. you know they they just you
- 17:56just hired them and they they onboarded
- 17:58you know like 5 seconds ago. Uh and so
- 18:02how do you actually teach that person to
- 18:04be effective?
- 18:05So how you do that is through a skill.
- 18:08uh a skill being just you know it's just
- 18:09markdown right but you know it encodes a
- 18:12lot of information instructions a lot of
- 18:16uh you can really draw out a lot of
- 18:19intelligence from an agent by well some
- 18:22people on Twitter call it like you know
- 18:23pull the agent to a different latent
- 18:26space which is kind of like a fancy way
- 18:28of just saying like since uh you LLMs
- 18:30are sort of like they predict the next
- 18:32token uh when you give it some high
- 18:35quality tokens uh to to begin with then
- 18:38you know it it can kind of pattern match
- 18:40on like a higher space that's you know
- 18:43smarter
- 18:45um so that's like a very interesting
- 18:48model there but yeah I built Pstack very
- 18:50very incrementally uh so uh started with
- 18:54just really observing how agents you
- 18:56know all the fail different failure
- 18:58modes of of that agents were having and
- 19:01every time I saw that I just okay I'm
- 19:02just going to make that a skill right
- 19:04like stop hallucinating actually go and
- 19:07search up, look up the code, use a lot
- 19:09of sign agents, uh, and yeah, stop
- 19:12guessing.
- 19:16>> Yeah, that makes sense. One, one kind of
- 19:18followup question here, uh, both from
- 19:19myself and from a bunch of people in the
- 19:20chat. So,
- 19:22>> I guess it's two two parts. So, one is
- 19:23like how do you maintain these skills?
- 19:25So, like the product changes over time.
- 19:27Obviously, there's a lot of people who
- 19:28are shipping against the codebase. So,
- 19:30how do these skills get maintained? Uh
- 19:32and then second to that is like how do
- 19:34you know when your verification is is
- 19:35good enough? Uh like and you know you
- 19:38can trust that the ver verification
- 19:40loops that you've built are going to I
- 19:42guess you trust the outputs uh when
- 19:44they're done.
- 19:47>> Uh yeah maybe I'll talk about um I think
- 19:50some are related. Maybe I'll start with
- 19:52this one first. So like how do I
- 19:56maintain these skills?
- 19:59So, um, if you're not familiar with this
- 20:01concept, an eval [clears throat] is
- 20:03essentially like a way to, uh, well, I
- 20:06the mental model I have is like it's
- 20:07like a unit test for an agent. Um, and,
- 20:11uh, you can actually make your own eval.
- 20:14You don't need like a special framework
- 20:15for them. You can build you can you can
- 20:18build one depending on like, you know,
- 20:20how scientific and how rigorous you want
- 20:22to be. Uh, my screen is red.
- 20:26>> Yeah, there's a little button. Um,
- 20:28sorry. Do you mind like disabling the
- 20:31drawing or something? I I can't see my
- 20:32screen.
- 20:33>> Yeah, sorry. If you guys could not draw
- 20:35on the screen, that'd be great. But, um,
- 20:36there's a little button in the
- 20:38>> Is that a troll?
- 20:39>> Yeah, the little drop down.
- 20:42>> How do I clear?
- 20:44>> Yeah, you got it. Perfect.
- 20:49>> Yeah. So, eval
- 20:52test your skills basically. And actually
- 20:55in Pstack we ship uh under potato mode
- 20:59there's a playbook if you search for it
- 21:00called eval playbook. Um and it's uh
- 21:05uh it's like not it's actually pretty
- 21:07pretty rigorous the way it's done. Uh
- 21:10but um essentially what I do is I spawn
- 21:13a lot of different sub agents. I have
- 21:15like my main coordinator agent uh come
- 21:18up with a rubric for uh what I want the
- 21:22skill to do. Um, and then it spawns all
- 21:25these sub aents and it it creates
- 21:28individual directories for them uh which
- 21:31are cleverly named to not let the sub
- 21:35agent know that it's being evaluated
- 21:37because uh agents can actually tell and
- 21:40when they do they change their behavior.
- 21:42Uh, but it does a bunch of stuff like
- 21:44that to um essentially yeah like test
- 21:49whether or not the skill I'm making or
- 21:51changing is actually doing what I think
- 21:53it does. Um, and one of the really nice
- 21:56things about cursor is that we are we we
- 21:58support so many different models. So you
- 22:01can actually eval your skill across all
- 22:03sorts of different models um and you
- 22:06know get a sense of how well it performs
- 22:08across that different matrix. um
- 22:12especially for the models that you use.
- 22:15Uh so I do this a lot. Every time I I
- 22:17modify a skill, I will run one of these
- 22:20uh like the ebal playbook uh and make
- 22:22sure that you know it's actually leading
- 22:25to a result I want. Uh but I will say
- 22:28like
- 22:29maintaining skills is actually pretty
- 22:32hard. Uh it requires I think a lot of
- 22:34taste and observation. So, you kind of
- 22:37need to be very good at being a backseat
- 22:40driver, you know what I mean? Like, if
- 22:42you do pair, if you ever done pair
- 22:44programming, for example, uh, and you
- 22:46watch a co-orker code and you just like
- 22:48you, you could probably do this better,
- 22:50you know, you could do, you know, like
- 22:51why did you not do this, right? You you
- 22:53ask a lot of questions to your coworker
- 22:55and it's kind of a similar thing here.
- 22:56You like you don't want to just be a
- 22:58passive observer of your agent. You want
- 23:00to be very in the driver seat in the
- 23:03initial stages when you're building up
- 23:04your own set of skills. uh you know
- 23:06obviously you can use something like PAC
- 23:08but if you're building your own set of
- 23:10skills it's very I think you know
- 23:13opening up the all the tool calls and
- 23:15like reading the code and reading all
- 23:17the uh the agent behavior and their
- 23:20thinking blocks is a really great way to
- 23:23see where they they fail right like what
- 23:26what you know where are they being done
- 23:29and then you can go and build a skill
- 23:30for that and then with verification how
- 23:33you trust it is it's I think it's also a
- 23:36very similar iteration loop uh where you
- 23:38know like I actually did the same
- 23:40process for verifying the verification
- 23:43skill where I actually get um so one
- 23:47thing that's interesting about eval is
- 23:49that you can sort of hill climb them
- 23:51meaning that uh your eval can produce a
- 23:54score right uh a score that you can get
- 23:57your coordinator to produce uh but also
- 24:00comp uh you can have a judge agent of a
- 24:02different model to uh kind cross
- 24:06reference and make sure that the first
- 24:08model is not being biased, right? The
- 24:10model that's judging all of the sub
- 24:12aents that are running the thing. Uh but
- 24:14you can also like hill climb. So meaning
- 24:16that you can you can use like /loop in
- 24:19cursor and you can say okay keep looping
- 24:22on this eval right until everything is
- 24:2510 out of 10 as an example. Uh and I did
- 24:28the same the basically the same approach
- 24:30with the control skill. And so I kind of
- 24:32it was very it was very hands-off
- 24:33actually. Uh so you know I uh I kind of
- 24:37built I built that skill that way like
- 24:39the CLI in that skill. Um and over time
- 24:43it's gotten really good. Uh but yeah it
- 24:46was definitely not super smooth at the
- 24:48beginning. It required a lot of
- 24:50iteration and I think there's an analogy
- 24:53here for me which is um well I make this
- 24:57analogy later in a different slide on my
- 24:59drawing here. Uh but I think of it like
- 25:03uh you know as a as a engineer now
- 25:06you're sort of more like you uh like
- 25:09maybe a manager or the analogy I like is
- 25:12like you're like a a chef in a
- 25:14restaurant. Uh you you're the head chef.
- 25:17Uh you're not cooking all the food
- 25:18yourself anymore. You have a team of
- 25:20cooks, right? You have line cooks, you
- 25:22have a sue chef, you have, you know, all
- 25:24these different stations.
- 25:26Um and it's your job to really design
- 25:28the environment. you know, you you're in
- 25:31charge of setting up the kitchen. You're
- 25:32in charge of, you know, like giving
- 25:35tasks to different people. So,
- 25:39um yeah, it's a very interesting way of
- 25:42working. Uh but yeah, that's that's how
- 25:45I've basically built uh these
- 25:47verification skills.
- 25:48>> Yeah, just just one followup there on
- 25:50like to go try to go one layer deeper.
- 25:52So, are you let's say we wanted to build
- 25:56um an eval or a skill for for something
- 25:59and we wanted to kind of get better on
- 26:01its own, which is is what I think you're
- 26:03suggesting. Uh are you doing that in
- 26:05like a work tree kind of isolated with
- 26:08like the sub aents and and then the
- 26:09reviewer agent and and all that? Is it
- 26:11happening like in some type of cloud
- 26:13hosted environment? Like what's the the
- 26:15more the practical steps? If I wanted to
- 26:17go do this uh and like set up a
- 26:19verification system for something, what
- 26:20would I what would I do or where would I
- 26:21start?
- 26:24Um I think that uh the best place to
- 26:27start is local because you can observe
- 26:31you can definitely observe what your
- 26:32agents are doing. So, uh, if you're
- 26:34building a verification skill for
- 26:36yourself, uh, I would definitely start
- 26:38local and just have your agent bring up
- 26:40the application, whether it's like a CLI
- 26:43or, uh, desktop app or whatever. And so,
- 26:46you can actually observe, right? You can
- 26:47see how the agent is interacting with
- 26:50the the application. You can see it, you
- 26:53know, how it calls like the different
- 26:56APIs that that allow it to interact with
- 26:58the uh the application.
- 27:02Um but uh for me personally uh I have
- 27:06basically been kind of all in mostly all
- 27:08in on cloud agents because they're
- 27:10extremely powerful. Uh and the really
- 27:13powerful thing about cursor is the the
- 27:15cloud agents actually where if you spend
- 27:18a little bit of time setting up your
- 27:19environment
- 27:21these control skills these verification
- 27:23skills pay a huge amount of dividends
- 27:26because it's not just something that
- 27:28makes you as a single engineer better.
- 27:31It actually levels up your whole team uh
- 27:33and even your whole company because uh
- 27:36you can actually start thinking about
- 27:38cloud agents. start thinking about
- 27:39automations that automatically do things
- 27:43like uh I I get I I kind of talk about
- 27:46this a bit later, but I'll just kind of
- 27:49get into it. Uh where where you know,
- 27:51for example, like I talk a lot about
- 27:53this agent we have called Benny, right?
- 27:55who who uh you know takes all of the bug
- 27:59reports that we get and it automatically
- 28:02goes off in the cloud, opens up a cloud
- 28:04uh it's you know its desktop. It runs
- 28:07cursor in its own computer and it uses
- 28:10the same control skills to interact with
- 28:13the application and try to reproduce the
- 28:15bug uh or the user report, right? And
- 28:17this is so so powerful because at once I
- 28:20can immediately I I get so much
- 28:22information from this automatically like
- 28:24here in this example you can see that uh
- 28:26the Benny actually reproduced the bug uh
- 28:30but it's already fixed on main. So it
- 28:33actually confirms that we fixed this
- 28:35problem already and all I need to do is
- 28:37just release another build of of cursor.
- 28:40Uh, so that's like huge information
- 28:42there that I didn't have to go off and
- 28:44sit with an agent, you know, and spend
- 28:46an hour trying to figure like is this
- 28:47fixed, is this not fixed. So you you you
- 28:50gain back so much time. Uh, but you
- 28:52know, everybody on my team benefits from
- 28:54this. Everybody at the company benefits
- 28:56from this. Uh, so definitely think that
- 29:00uh, you know, keeping these using cloud
- 29:03agents is super powerful. Uh, but yeah,
- 29:06it's like a journey. You have to trust
- 29:08it first, right? before you you get to
- 29:10this point. And that's it goes back to
- 29:12what I was saying here where, you know,
- 29:14it's very hard. It's almost impossible.
- 29:16And I would definitely encourage you not
- 29:18to try to jump from, you know, like if
- 29:21you're still in this zone, you don't
- 29:24want to jump to like I'm going to spawn
- 29:26a hundred of thousand or thousands of
- 29:28cloud agents right now because you're
- 29:30just going to waste a lot of tokens. Um,
- 29:32and it's going to be extremely
- 29:34expensive.
- 29:35>> Yeah. So just to kind of recap so far,
- 29:37basically the if we wanted to go on the
- 29:39journey that you've kind of gone on, it
- 29:40would be to start with verification,
- 29:43building some some skills and some some
- 29:46ways of determining that the agents are
- 29:47producing at least like correct code.
- 29:50Whether like you said whether it's good
- 29:51code or not is maybe a separate
- 29:52question, but like it's it's technically
- 29:53solving the problem by looking at you
- 29:56know stack traces, looking at you know
- 29:58the the actual behavior in the app and
- 29:59so on. Um, and then once we trust it
- 30:01locally, then we can start to think
- 30:03about scaling into the cloud and running
- 30:05more agents that are picking up signals,
- 30:07I guess, on their own, right? So whether
- 30:09that's like a bug report that comes in
- 30:10or something, they can go and pick it up
- 30:12and solve the problem and and give us
- 30:14back a PR. And then maybe the last step
- 30:16is like automerging the PRs, which uh is
- 30:19where you're at, maybe not where
- 30:20everyone is at.
- 30:21>> Um, and then and reviewing them on main,
- 30:23but um, is that is that about right?
- 30:26>> Yeah, exactly. I think yeah, that's why
- 30:27I drew this this uh this this curve,
- 30:29right? because that this this basically
- 30:31describes my journey of you know when I
- 30:34started barely could use a couple agents
- 30:36and I was just observing every single
- 30:38thing. I think there's really no
- 30:40shortcut for going from here to there
- 30:42because this is really about your
- 30:44personal level of trust in agents,
- 30:47right? Um obviously, you know, as a as
- 30:49engineer, you don't want to just slop
- 30:50code into production. So, how do you
- 30:53actually build up that trust? Takes um a
- 30:56lot of uh I guess taste and judgment. Um
- 30:59but uh you know, like I think plugins
- 31:02like Pstack definitely kind of help you
- 31:05uh get up to speed much quicker. Uh, and
- 31:08so I guess it's like if you trust me and
- 31:11you trust Pstack, then in by extension
- 31:14you can maybe trust your agents. But if
- 31:16you don't trust me, and I I definitely
- 31:18would not encourage people to blindly
- 31:20trust me. Uh, uh, you know, if you build
- 31:24up your own set of skills that you can
- 31:26obviously, you know, take a look at PA
- 31:28and kind of fork it, make it your own,
- 31:30improve the skills. Definitely encourage
- 31:32that. Uh but for me it's really all
- 31:35about it just keeps coming back to
- 31:37trust. You know every one of us here in
- 31:39this chat have a different standard for
- 31:41engineering. Uh and there are different
- 31:43things that are important for us in our
- 31:46codebase. And uh when you are able to
- 31:50encode all of that into skills and you
- 31:51can verify that your agent is actually
- 31:53doing them that allows you to really
- 31:55kind of ascend this curve and um uh you
- 32:00know start automating things. Uh there's
- 32:02another piece I wanted to talk about. Um
- 32:05if there's more
- 32:06>> Yeah, go for it. I'll I'll pick up more
- 32:08questions as I go. But
- 32:09>> yeah, I think there's a third part to
- 32:10this which I haven't talked about yet,
- 32:12which is kind of an interesting one
- 32:14which is like refactoring and rewriting
- 32:17like one of the uh I guess most
- 32:19controversial one of the most
- 32:21controversial topics in the industry I
- 32:23think is like should you rewrite your
- 32:25app or not? Um because I think engineers
- 32:30are very prone to this where especially
- 32:32when you join a company you come in and
- 32:34you see like the codebase and you're
- 32:35like oh man this is like who wrote
- 32:38this code you know it's terrible I want
- 32:40to rewrite the whole thing there is a
- 32:42very common inclination and I think a
- 32:44lot of you know before agents um and I
- 32:48guess arguably even now people will
- 32:50definitely discourage you from re
- 32:51rewriting stuff but I'm actually here to
- 32:54make a case for why you might want to
- 32:56consider But
- 32:58um because
- 33:00I think it really depends uh you know uh
- 33:03brownfield applications I think are
- 33:05actually in a pretty good spot
- 33:07especially if they're set up well
- 33:09already. Uh and like recently I've been
- 33:12talking to some people but uh you know I
- 33:14I was just observing I I just noticed
- 33:17this parallel which is that a lot of big
- 33:21tech company problems are now
- 33:23everybody's problems. Um, and the big
- 33:26tech company problem, you know, like
- 33:27when I was working at Meta, like we had
- 33:29this giant monora repo, we had like, I
- 33:31don't know, tens of thousands of
- 33:33engineers just, you know, like banging
- 33:35on their keyboards and and shipping code
- 33:38and
- 33:40a lot of really great engineers at Meta.
- 33:42Uh, but, uh, I'll say like, you know,
- 33:45you'll be surprised that the code
- 33:46quality is actually not that good.
- 33:48[laughter] Um, and so I often joke that
- 33:51like, you know, before AI sloth, we had
- 33:53human sloth. Um and so uh you know I
- 33:56think a lot of big tech infra like uh
- 33:58like what Meta has or Google you know
- 34:01you know really big tech companies are
- 34:03actually designed for that where you
- 34:05you're sort of like you're catering to
- 34:07the the you know like uh sounds that
- 34:10sounds so bad to say but like the the
- 34:11least capable engineer on your team
- 34:13right you build you build frameworks you
- 34:16build conventions you build guard rails
- 34:19you know you restrict credentials so
- 34:21that you know your intern doesn't wipe
- 34:22your production database
- 34:24Um
- 34:26there's uh you know if you have that
- 34:28level of infra already I think your
- 34:31agents can actually already do a very
- 34:33solid job right because they have the
- 34:35the guard rails are already in place for
- 34:38agents to not cause havoc or not cause
- 34:41too much havoc uh in your codebase um
- 34:44and you can always add more you know
- 34:46guard rails. Uh but I think like green
- 34:49field applications especially are you
- 34:50know like the brand new applications are
- 34:53like the biggest risk in my opinion. Uh
- 34:56and also the greatest opportunity
- 34:58because you know if you vibe code a
- 35:00project uh a prototype um like we did
- 35:04for Grockbot you know Grockbot was spun
- 35:05up very very very quickly. Um and if you
- 35:08if you haven't heard of of Grockbot it's
- 35:10like our a new application we just
- 35:12launched yesterday. Uh it's it's really
- 35:14cool. uh lets you orchestrate your
- 35:18create like individual agents that have
- 35:20their own identity and you can kind of
- 35:21orchestrate them. It's super cool.
- 35:23Definitely check it out. Um but yeah,
- 35:25that was it's like a very it was a very
- 35:26green field application like most
- 35:29prototypes are so like vibe coded very
- 35:32quickly. Humans were not reading the
- 35:34code at all. And uh I had this tweet
- 35:38recently uh where I said something about
- 35:41organic architecture. Um
- 35:45maybe I'll find it. Uh but the idea is
- 35:49that
- 35:50uh when you have a completely vibe coded
- 35:52application, you essentially have no
- 35:54guard rails whatsoever. So uh your
- 35:57agents
- 35:59when you give them a task, they will
- 36:01just solve it in whatever method is the
- 36:03most convenient. And over time you get
- 36:06into this uh situation where you have a
- 36:09codebase that is spiraling out of
- 36:11control because you don't understand it.
- 36:14Uh your agents understand it I guess in
- 36:16a way but like they've built something
- 36:18that is you know optimized for short for
- 36:21shortcuts. Uh and uh you know it will
- 36:24you will suffer you'll have a lot of of
- 36:27issues with that application.
- 36:30Uh so I think starting your codebase
- 36:33with uh like very strong constraints is
- 36:37very much needed uh because like when
- 36:40you have a codebase that you can trust,
- 36:43right? when you have guardrails that
- 36:44actually help you uh uh help your agents
- 36:48write good code, you can get into the
- 36:50you know like into this part of the
- 36:52curve where I I where like I I said you
- 36:55know I woke up today and I had like 20
- 36:57PRs merged u by my agents and that's
- 37:00because I invested a lot a lot of time
- 37:04uh over 600 PRs I I I calculated
- 37:06yesterday uh when I refactored all of
- 37:10Grockbot to this new architecture that
- 37:12I've been
- 37:14Um,
- 37:15and yeah, I've gotten to a point where I
- 37:19I don't really look I really don't look
- 37:20at the code anymore. And um, I say that
- 37:24not just, you know, to sell you tokens,
- 37:25but because I, you know, it it it took a
- 37:29lot of work to get to that point. I
- 37:30spent a lot of tokens to get the
- 37:32codebase to this point where I no longer
- 37:34have to look at it. Uh but I'm very
- 37:37excited because you know of the
- 37:39potential where you know it's not just
- 37:41this doesn't just benefit me it benefits
- 37:43everyone contributing to Grockot and it
- 37:46also empowers you know designers and
- 37:49product managers and you know pe uh even
- 37:52GTM people to add features to Grockbot
- 37:55and I don't have to worry you know I
- 37:56don't have to to wake up at night in in
- 37:59the middle of the night and worry like
- 38:00oh someone's just merged a perf
- 38:02regression right I have a ton of
- 38:05constraints and CI is like it's actually
- 38:07very annoying to write code in in graph
- 38:09web but like agents absorb all of that
- 38:11annoyance.
- 38:13Um but yeah I'm happy to talk about what
- 38:16exactly that is. Um
- 38:18>> yeah I think one question um
- 38:21>> before we get into the this part here is
- 38:23just around that element of like what
- 38:26your your your CI looks like or maybe
- 38:28some of the constraints and then also
- 38:29like the average PR size. I saw a
- 38:30question about that earlier just to give
- 38:32people you know kind of a a glance. It
- 38:35doesn't have to be like mathematically
- 38:36average, but just uh you know like what
- 38:38generally the size of the a PR is. Um if
- 38:41it's only a couple lines of code or you
- 38:43know um yeah
- 38:45>> um
- 38:47I think it depends. Uh let me
- 38:51I'm trying to do this in a way where I'm
- 38:53not going to like
- 38:54>> you. Yeah, you don't have to share the
- 38:55actual number like an actual average. I
- 38:57think this this is fine
- 38:58>> benchmark
- 38:59>> but like we have so okay this is not
- 39:02that interesting but uh well fun fact is
- 39:05that virtualization in grabbot and in uh
- 39:09cursor is actually powered by uh pretext
- 39:13uh which is a sort of new library that
- 39:16someone's built um that's really
- 39:19interesting you should you should check
- 39:20it out but that's not really that
- 39:22important uh I think the average PR size
- 39:25I actually don't know. I pro I don't
- 39:27know if I want to click on these. Uh I
- 39:29probably can, but I would say like they
- 39:32can range anywhere from a few hundred
- 39:34lines or 50 lines to like a thousand
- 39:37depending on what the thing is doing. Uh
- 39:40so like here I'm actually like deleting
- 39:42a bunch of files. So I expect that it's
- 39:44just like mostly deletion. Uh but yeah,
- 39:47it kind of varies.
- 39:49There's no like Yeah,
- 39:51>> there's no like hard cap or hard limit.
- 39:52Are there all like 50 line PR?
- 39:53>> There's no hard cap. Yeah, there's
- 39:54definitely no hard cap. But I I do
- 39:56encourage my agents to split up their
- 39:57work into multiple PRs. Uh I do that
- 40:01mostly because uh I like I like the idea
- 40:05of the I guess maybe this is much harder
- 40:07to do now as in the world of agents and
- 40:10you have like so many commits, but I
- 40:12like the idea that you know the git
- 40:13history is a very rich source of
- 40:15context. Uh, and I like the I like each
- 40:19PR to sort of atomically describe what
- 40:22that small piece of thing is doing,
- 40:25which also makes it easier for me to
- 40:26revert changes and like figure out, you
- 40:28know, oh, I shipped a bug and it's just
- 40:30it's here, right? It's not in this
- 40:3240,000 line PR where who knows what
- 40:36landed in there.
- 40:38Uh, but I don't have a hard cap on PR
- 40:41size.
- 40:42>> Cool. And then um yeah, also quick
- 40:45question on like CI. So again, you don't
- 40:46have to go into like uh the screen share
- 40:48of like your CI does, but just generally
- 40:51would you describe what the CI kind of
- 40:53looks like uh or how strict it is?
- 40:56>> Uh yeah, so
- 40:58uh well specifically for Grockbot. So
- 41:01Dune is the is the sort of cheeky code
- 41:04code name for the architecture that
- 41:06we've built for Grockbot. Um the CI
- 41:10looks pretty annoying because there's
- 41:13checks for everything. So like literally
- 41:16I have um uh well if you've written any
- 41:19React for example you know you know that
- 41:21one of the biggest foot guns in React is
- 41:22use effect. Uh so in
- 41:26uh Dune and in graphbot we've banned use
- 41:29effect. So Dune is just you can the the
- 41:32mental model of what Dune is uh you can
- 41:34kind of think of it as like Nex.js JS
- 41:36for uh electron apps and it's designed
- 41:39for agents to write uh and it's like
- 41:42custom for you know our agent powered
- 41:45applications. Um so the CI checks are
- 41:48very like specific to that like you know
- 41:50don't use use effect. It's it's it's
- 41:52banned like CI will fail uh and yell at
- 41:55you. We have like some of the more
- 41:57interesting ones that people might raise
- 41:59eyebrows is like I actually ban code
- 42:01comments as well uh which is very
- 42:04interesting. Uh, but I've noticed that
- 42:0899% of the time agents just write code
- 42:11comments that kind of describe some
- 42:13historical thing that is actually
- 42:15totally irrelevant to the code. Um, like
- 42:18it will often say like you know, oh
- 42:19Lauren said you should never do this and
- 42:21it's now in in a code comment like what
- 42:23like why that was I didn't say that as
- 42:26like a durable you know global rule. I
- 42:28just meant like your this PR sucks and
- 42:31you should change that part.
- 42:34agents don't really understand us that
- 42:36well surprisingly uh and or they kind of
- 42:39assume too much and they kind of do
- 42:41things in like very stupid ways. So like
- 42:44yeah we just ban everything everything
- 42:46you can imagine like the agents are bad
- 42:49at we ban. Uh so one example that we
- 42:53actually suffer a lot in the agents
- 42:55window is we have uh you know you if
- 42:58you've used the agents window you've
- 42:59definitely seen performance issues and
- 43:01you know we're constantly trying to fix
- 43:02them. Uh but it's like a it's a never-
- 43:06ending struggle because there's so many
- 43:08pull requests that get merged. Every any
- 43:10one of them could just regress
- 43:12performance or stability or reliability.
- 43:15Uh you know the agents window doesn't
- 43:16have this architecture yet. I plan to do
- 43:18bring this learning back there and kind
- 43:21of refactor everything there. Uh but uh
- 43:25it just regresses super often. uh
- 43:28because uh there's just one example is
- 43:30like we have very poor um isolation
- 43:34between processes. So like on you know
- 43:36on on Electron you have a renderer
- 43:38thread that renders your UI but you also
- 43:40have like a main thread that you can run
- 43:42other code that you know doesn't need to
- 43:44block the renderer.
- 43:46Um, but we do a poor job of separating
- 43:49those things and so often times you just
- 43:51accidentally have code that gets pulled
- 43:53into running on the renderer thread and
- 43:56then all of a sudden you're competing
- 43:57with the the renderer that you know that
- 44:00has a very if you want like 60 fps you
- 44:03have to every frame that gets drawn has
- 44:05to be done in 16 milliseconds. So very
- 44:08very small you know deadline per frame
- 44:11uh if you want you know a very smooth
- 44:13product. Uh and when you start building
- 44:15bringing in accidentally bringing in you
- 44:17know things that are like very
- 44:19computationally heavy or they have a lot
- 44:21of IO uh then you just get into like a
- 44:24lot of jank right your your FPS really
- 44:27drops you start uh you know losing
- 44:29frames you get long tasks that take more
- 44:32than 16 milliseconds and you just get
- 44:34this really choppy experience.
- 44:37So all of those patterns that we've
- 44:38learned basically building electron apps
- 44:40we've encoded into this framework and it
- 44:43becomes like a hard failure. So I
- 44:45literally in in grabbot we literally
- 44:47have a directory called electron main
- 44:49electron renderer and we have uh import
- 44:54uh CI guess where we actually check the
- 44:58dependency graph to make sure you're not
- 44:59accidentally importing code from one
- 45:02directory to another. Uh so that's
- 45:04enforced by CI um as well as bug bots uh
- 45:09which is our which cursors um like code
- 45:13review tool that runs on CI uh you know
- 45:15in our agents MD it's everywhere like so
- 45:18I I I I um I have this thing here where
- 45:22I I talk about like um you know like
- 45:25there are multiple layers I think for
- 45:27building a good codebase. Uh obviously
- 45:30the codebase is one where uh if you have
- 45:32an architecture like this where it's
- 45:34extremely strict uh you know the the the
- 45:38way to build features is very
- 45:39conventional that's like the strongest
- 45:42strongest level of enforcement because
- 45:44agents just love to copy existing
- 45:46patterns. So uh one example of this in
- 45:49rockbot is like we have this these
- 45:51concepts called like a feature and we
- 45:54have entry points and transcript cards
- 45:56like oh you know the cards that you see
- 45:57in the chat these are all like like
- 46:01nouns I guess in in the framework and so
- 46:03there's a very conventional way of
- 46:05creating them and so like a feature is
- 46:09all in in a single directory as an
- 46:10example and so all of the code that
- 46:13contributes to that feature lives in one
- 46:15directory so it's all collocated in one
- 46:17place makes it super easy. You know,
- 46:19agents don't have to like uh grap around
- 46:21and try to figure out like where all the
- 46:23things are. It just looks at the feature
- 46:25and like, oh, okay, I'm working on the
- 46:27onboarding feature in Grockbot. Uh I'm
- 46:31just going to work in this directory.
- 46:32And for 80% of the work, it's mostly
- 46:35just very encapsulated there. But, uh
- 46:38like it's like very it's like designed
- 46:41again for you know like the dumbest
- 46:43agent like you don't have to think,
- 46:45right? the the the one of the key
- 46:48principles I have for this framework is
- 46:50like the shortest the shortest path is
- 46:52the best path.
- 46:55So uh because that plays exactly to how
- 46:57agents love to write code is like they
- 46:59like to take shortcuts really you know
- 47:01they they'll find the quickest way to
- 47:04solve the problem. So why not make that
- 47:06the best way to solve the problem? Uh so
- 47:10I I I probably won't get into all the
- 47:12specific details. Um and uh the the this
- 47:16framework is really more of a collection
- 47:17of ideas and principles rather than
- 47:19something that will open source. Uh you
- 47:22can you can you know screenshot this I
- 47:23guess if you want and uh tell your agent
- 47:26to uh do some build build something like
- 47:29this for you too.
- 47:31Um yeah, but it's really all about the
- 47:34layers uh you know like the the codebase
- 47:36is one part with features uh and
- 47:38directories and you know import or
- 47:41blocking import dependencies uh that
- 47:44shouldn't be imported. Uh but and and it
- 47:47all enforces that and static analysis.
- 47:49So like uh there's CI checks, we have a
- 47:52lot of lints for bad patterns that we
- 47:55observe. Uh compiler diagnostics. Uh
- 47:59there's also rules and bugbot which are
- 48:02um I think like three four five are more
- 48:05soft right these two actually make make
- 48:09CI red right so that you know there's a
- 48:12hard constraint where the agent can't
- 48:13just write crappy code
- 48:16for rules and skills and buggbot your
- 48:20agents can still forget right you can
- 48:22still or it may not always consistently
- 48:25apply them so I like to layer them, but
- 48:30I don't I don't like to rely on them as
- 48:32the only source of enforcement because
- 48:35it's very very soft, right? And if you
- 48:37if you only have rules and bug bar and
- 48:39skills and a style guide for your code,
- 48:41you will it's only a matter of time
- 48:43before your codebase looks like complete
- 48:46trash. I'm sorry to say that, but uh I
- 48:49definitely recommend yeah like you know
- 48:50investing in you know things that can be
- 48:53hard and forced, right? And this is why
- 48:56you know maybe the choice of text stack
- 48:59that you use is also very important. Um
- 49:02like I think for example Rust is sort of
- 49:04making you know it's like getting super
- 49:06popular again. Uh because the compiler
- 49:10is so strict right the compiler enforces
- 49:12so many different things. you know,
- 49:14there's a borrow checker that you have
- 49:15to appease and if as long as you make
- 49:17sure your agents don't write unsafe code
- 49:20blocks, uh you can more or less feel
- 49:22somewhat confident that if the code
- 49:24compiles, it probably works and it's
- 49:26good. Uh but you see it gives you that
- 49:29level of trust and confidence that you
- 49:32as a human engineer no longer need to go
- 49:35and check it yourself. you know, you you
- 49:38rely on code and static analysis to
- 49:42actually make that uh a lot smoother.
- 49:46Um, and I I guess the worst part, the
- 49:49worst place to be in is if you are stuck
- 49:52in code review land where you actually
- 49:54enforce all of the constraints, the
- 49:56invariance in your codebase by literally
- 50:00the human person saying, you know,
- 50:02reading the code and like, okay, you
- 50:03should not do this, right?
- 50:06Every time you have to do that, you
- 50:07should consider that as a code smell,
- 50:08like a anti- pattern and you should say,
- 50:11"Okay, instead of me commenting on the
- 50:13PR, how do I turn this into a hard rule,
- 50:17right? How do I turn this into a lint
- 50:19rule? How do I turn this into a CI
- 50:21failure? Or how do I even categorically
- 50:23eliminate this problem uh entirely?" Uh
- 50:28I I can talk about another migration
- 50:30I've done, but I'll probably pause here.
- 50:32>> Sure.
- 50:33>> Yeah. I feel like that's that's where I
- 50:34am to be honest is is what you're
- 50:36describing right now which is that like
- 50:38I don't have all of these rules. So I
- 50:40have some things to go do after this
- 50:41session in terms of being able to scale
- 50:43my agents. I'm I'm definitely on like
- 50:45the uh you know maybe a couple of
- 50:47parallel ones locally stage. So like two
- 50:50to three locally and I'm sure most
- 50:51people here are on the same. So uh yeah
- 50:54I know we only have a couple minutes
- 50:55left. Lauren, was there anything else
- 50:57that you wanted to to highlight? I
- 50:58obviously there's lots of questions so I
- 51:00can grab more but I want to give you a
- 51:02few minutes uh if there's anything else
- 51:03you want to talk about
- 51:03>> I think I've been yapping for quite a
- 51:05lot so maybe let's just do questions.
- 51:08>> Okay cool. Uh one question that had uh a
- 51:10couple of uh came up a couple times was
- 51:12just around like token usage.
- 51:15>> So the the question is like is is what
- 51:17you're describing a realistic thing for
- 51:19people who are on you know uh a normal
- 51:23set of token usage they don't have you
- 51:25know basically unlimited tokens uh to
- 51:27work with.
- 51:29I think that's a really good point. I
- 51:30mean, like obviously, you know, I work
- 51:32at a AI lab where we have unlimited
- 51:34tokens. So, uh I definitely cannot
- 51:39say that, you know, this is something
- 51:41everyone should do in the exact same way
- 51:43that I did it. I think it's possible to
- 51:45get to this point without, you know,
- 51:47breaking the bank.
- 51:49But you know if you're like an
- 51:50engineering leader or you know you're
- 51:52you you have a startup that you lead um
- 51:55I think to me it's a question of ROI um
- 51:58and it's like uh yes you spend a lot of
- 52:03money on tokens in the upfront stage you
- 52:06know like refactoring a code base is
- 52:07going to take a lot of tokens uh adding
- 52:09all these things uh is going to take a
- 52:11bunch of tokens but if we're heading to
- 52:14a world where agents are writing all the
- 52:16code and you know You want to be very
- 52:20lean, right? You don't want to have to
- 52:22hire, you don't want to be, you don't
- 52:23want to become like meta, right? Like I
- 52:25mean like in terms of you don't want to
- 52:26become a 10,000 person engineering or
- 52:29because I mean that's a cool problem to
- 52:32have, but also you you have so much
- 52:34overhead. There's like planning, you
- 52:37know, like you it's it's a personally I
- 52:39I wouldn't uh it it's not super fun, but
- 52:43um I think you want to stay very nimble,
- 52:46right? And you want to you want to be
- 52:47like agents are all about allowing you
- 52:50to do things that you couldn't do
- 52:51before. That's really to me like the
- 52:53value of agents, you know, it's not just
- 52:56storing tokens on every single little
- 52:58thing, but um to me like the thing I
- 53:00couldn't do before is like enforce this
- 53:03level of constraints in a codebase by
- 53:06myself, right? Like I'm just a single
- 53:09person, you know? uh it would have taken
- 53:11me years to build this framework uh and
- 53:15do all the refactoring and test
- 53:18everything myself and verify you know
- 53:20like run imagine if there it was just me
- 53:22right no in in pre- agent era just like
- 53:25running you know by it would take me so
- 53:28long right and my salary is pretty high
- 53:30right like so you know the the question
- 53:34I think an engineering leader might have
- 53:35is just then you know like what is
- 53:38there's a trade-off of do Do you hire
- 53:40someone to do this or do you spend the
- 53:43tokens to set up a codebase so that even
- 53:46the the most naive, right, the dumbest
- 53:49agents can do a good job? And when you
- 53:52actually get to this point, like even
- 53:54agents that are not, you know, fable
- 53:56size do an excellent job of writing
- 53:58code. And this pays a lot of dividends
- 54:01as well for me personally where I've
- 54:04empowered not just myself but again like
- 54:07PMs, designers, engineers who are not
- 54:10familiar with Grockbot to just
- 54:12contribute in a way that is sustainable.
- 54:16So I think yeah it's definitely like a
- 54:18trade-off for sure. You know like
- 54:20nothing is like free for sure. Uh and
- 54:22tokens are pretty expensive. Uh but oh
- 54:25actually uh I I I don't know how many of
- 54:27you have seen this but we actually
- 54:29announced Grock 4.6 today. So very
- 54:32exciting finally out. Um so yeah graph
- 54:354.6 would be like a great it was very
- 54:37very smart. Uh it's really good on the
- 54:40on the benchmarks. Uh and it's the same
- 54:42the tokens uh well uh I hopefully I'm
- 54:45not saying this incorrectly but uh I
- 54:47believe the cost per token is the same
- 54:50as 4.5. So you're actually getting more
- 54:53intelligence for the same cost. Uh I
- 54:57think this is an area that cursor tries
- 54:59to cursor and SpaceX AI try to really
- 55:02optimize for like that heredto frontier
- 55:05of you know cost versus intelligence. Uh
- 55:08you know we don't necessarily want to
- 55:10build the biggest model ever because
- 55:12that is extremely expensive to run. It's
- 55:14really about like how do you find that
- 55:16sweet spot right? you don't you don't
- 55:18need a giant model, but it's just super
- 55:19smart, right? And it's not very
- 55:21expensive for inference.
- 55:24Uh but um yeah, I think to kind of round
- 55:27it up, um I think it's like a it's it's
- 55:30there's a if you do your own analysis, I
- 55:33feel like it's pretty positive. It it'll
- 55:36be pretty positive that the ROI you get
- 55:38from investing in stuff like this uh
- 55:41just empowers not just yourself, but
- 55:44your whole team to be so much more
- 55:46productive, right? right? Like imagine
- 55:47if you have an army of engineers like me
- 55:49who are shipping so much improvements
- 55:52and and bug fixes uh you know every day,
- 55:56right? Like that is pretty exciting.
- 56:00>> Cool. Uh one last question before we
- 56:01wrap up. This one is for the people in
- 56:03product on the on the call.
- 56:05>> So let's say we do have an army of
- 56:07engineers who are shipping like Lauren.
- 56:09I'm just curious like how is the product
- 56:11team or other functions of your company
- 56:13keeping up given that like if you're
- 56:15shipping so quickly have are they using
- 56:18AI more to do their jobs like as much as
- 56:20you can speak to that obviously you
- 56:21don't have like you're not in that role
- 56:23but just curious about how that works.
- 56:25Um I think this is where Grothbot has
- 56:28been actually exceedingly powerful. Uh
- 56:31where so before grabbot like you know uh
- 56:34obviously cursor only had cursor like we
- 56:37only had agents window we had a CLI we
- 56:39had an IDE and these are really like
- 56:42power user tools right like de they're
- 56:44designed for developers so it's very
- 56:46very developerentric you can do
- 56:48knowledge work in them but it like the
- 56:50UI is not really optimized for that. So
- 56:54we actually didn't really have uh well I
- 56:57think like a lot of people like you know
- 56:58GTM product like they might have used
- 57:01cursor uh to do their work but it
- 57:04definitely wasn't like a delightful
- 57:05experience for them. Um I think now with
- 57:08Grogbot
- 57:10uh it's become Grogbot is basically like
- 57:13the kusher moment for people who are not
- 57:16in tech in my opinion like it's like
- 57:18it's like a very very accessible way to
- 57:21use agents in a very comfortable very
- 57:24familiar interface. It looks like
- 57:25iMessage
- 57:27um and it's very fun to you know you can
- 57:29give your agent a fun name. uh you can
- 57:32have you can kind of do orchestration
- 57:34with in a very like natural way where
- 57:36you can sort of you know each agent is
- 57:38like a person right now you got a team
- 57:39of agents like working on you have one
- 57:41one agent per account that you manage as
- 57:43an example or if you're a PM you have
- 57:46you know you can have an agent that
- 57:47summarizes all the work that Lauren did
- 57:49last night and then now you know what I
- 57:51did right so I think our PMs are
- 57:53leveraging that a lot and they're
- 57:56shipping code too uh so you know like
- 57:58oftent times they will just say oh
- 57:59here's a bug I fixed can you look at it
- 58:01and then I'll go review it and actually
- 58:04it's just perfect. I'm like okay stamp.
- 58:06Uh so uh that I think that shows that
- 58:08you know the the Dune architecture is
- 58:10holding up right the all the the really
- 58:13strict constraints allow people who are
- 58:16not experts in engineering to contribute
- 58:18at a high level. Uh so I'm I feel like
- 58:21I'm already seeing that pay off a lot
- 58:23where uh you know designers and PMs are
- 58:26just able to to to ship features
- 58:29directly. Um and that just makes the
- 58:32Grockbot team super fast, right? Where
- 58:35we can ship so quickly. Um and we have a
- 58:40lot planned. So I'm very excited uh to
- 58:43you know uh to to ship more ship more
- 58:46stuff.
- 58:47>> Yeah, that's awesome. Uh well, we are at
- 58:49time. So, uh I guess Lauren, if if folks
- 58:52want to support you, maybe go try out
- 58:53Grockbot, try out uh 46 and uh you know,
- 58:57get provide some feedback. But yeah,
- 58:59this was awesome. Really appreciate you
- 59:00taking the time. Uh thanks everyone for
- 59:02all the messages in the chat. Lots of
- 59:03good questions. I know we didn't get
- 59:04through everything, but as kind of said
- 59:06at the top, way more questions than than
- 59:07we could get through, but uh yeah,
- 59:09really really thanks thanks for for
- 59:11joining. Thanks everyone for joining and
- 59:13hopefully you enjoyed the session.
- 59:15>> Yep.
- 59:16>> All right. Yeah, I see your man. Thanks
- 59:18for having me. And uh if you have any
- 59:19more questions yet, just DM me on
- 59:20Twitter. I'll I'll open them up. I guess
- 59:23I'll let the let the
- 59:24>> You're going to get a lot of DMs.
- 59:26>> Yeah, I'll open the updates. So, yeah,
- 59:28DM me. Maybe I'll do like a Twitter
- 59:30space at some point. That's all for more
- 59:32questions. But
- 59:33>> really appreciate everyone for showing
- 59:34up uh you know, taking an hour out of
- 59:36your day.
- 59:37>> Yeah. All right. Thanks all. I'll see
- 59:38you the next one.
- 59:39>> Okay. Thanks everyone. And bye.
About this transcript
This page contains the full transcript of [한영자막] Cursor 핵심 개발자 Lauren Tan: AI 에이전트를 실전에서 제대로 신뢰하는 법 (xAI GrokBot 워크숍) by Tech Bridge, generated from the public captions YouTube serves with the video. The transcript has 10,751 words across 1,473 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.