The best AI agents are simpler than you think — Transcript
Full transcript
- 0:00Agentic commerce will be bigger than
- 0:02e-commerce. There are cases where the
- 0:04Sierra agent is actually getting paid a
- 0:06commission on a sale.
- 0:08>> Today, I'm talking to Zack Reno Wedeen,
- 0:10head of product at Sierra, the platform
- 0:13powering customer experience agents for
- 0:15most of the Fortune 20.
- 0:16>> Coding agents are really good at file
- 0:17systems, they're really good at Git,
- 0:19they're really good at grep. Let's
- 0:21materialize everything into those
- 0:23structures so that coding agents can
- 0:25just, [music] you know, cook.
- 0:26>> He breaks down how Sierra builds for
- 0:28voice and why the architecture looks
- 0:30nothing like a standard agent harness.
- 0:32>> One of the big unlocks for Sierra agents
- 0:34was how to parallelize thinking,
- 0:36listening, and talking. And if this
- 0:38model says it's silent, you trust it. If
- 0:40this model does not say it's silent, you
- 0:42trust this one.
- 0:42>> We get into why a Sierra conversation is
- 0:45unlike a typical LLM call.
- 0:46>> You know, 10 or 15 different models
- 0:48might be invoked for a given
- 0:49conversation turn. So, sometimes you're
- 0:51classifying and you're responding at the
- 0:54same time.
- 0:54>> And Zack explains how and why Sierra
- 0:57built an entirely separate
- 0:58infrastructure layer for payments.
- 0:59>> We have isolated infrastructure where
- 1:02payment info doesn't go to an external
- 1:04large language model because none of the
- 1:05LLM providers are PCI certified in that
- 1:07way.
- 1:08>> Welcome to Max Agency, the podcast
- 1:10[music] that goes deep into how the best
- 1:11agents are being built by builders like
- 1:14you.
- 1:15>> [music]
- 1:17>> Most people are probably familiar with
- 1:18Sierra as a customer support platform.
- 1:21But, from what I understand, recently
- 1:22you guys are going broader than that.
- 1:24Could you talk a little bit about the
- 1:25types of agents that you help people
- 1:27build?
- 1:28>> Yeah, so this has been the vision since
- 1:30the beginning. I think because many
- 1:31companies have an RFP process where
- 1:34they're very specific about, "Hey, we
- 1:36want to solve customer service." That is
- 1:39often where we start, but we think of
- 1:41Sierra as the full engagement platform
- 1:44across all of the moments that matter
- 1:46for your customers. So, if you're an
- 1:48airline, that might be browsing for a
- 1:50flight, might be booking the flight,
- 1:52might be choosing your seat, might be in
- 1:54my case, I have a small dog, so adding a
- 1:57pet in cabin, then the flight might get
- 1:59rescheduled or delayed or
- 2:02cancelled, etc. etc. You need to get
- 2:04your bags there. There are just so many
- 2:06different things across that process.
- 2:08Some of them are sort of sales, some of
- 2:10them are more service, some of them are
- 2:12more loyalty, but they all kind of
- 2:14ladder up to the relationship between a
- 2:16business and its customers. And Sierra
- 2:19agents are present at all of these
- 2:21different parts of the customer life
- 2:22cycle. So, as an example,
- 2:25there are cases where the Sierra agent
- 2:27is actually, because of our
- 2:29outcome-based pricing model, getting
- 2:31paid a commission on a sale, which I
- 2:34think is quite different from how most
- 2:35people would imagine service. And so, we
- 2:37get really excited about those more
- 2:40exotic opportunities because they also
- 2:41give us an opportunity to push the
- 2:43platform forward and, you know, continue
- 2:45adding to what the platform can do, and
- 2:48kind of turn that into a product,
- 2:50package it up nicely, and bring more to
- 2:52all of our existing and future
- 2:54customers.
- 2:54>> How similar is the platform between
- 2:57these different use cases?
- 2:58>> I would say that
- 3:01it's very extensible, and so you can
- 3:03kind of take it in different directions.
- 3:06We like to say, I think it's originally
- 3:08attributed to the
- 3:09one of the creators of the programming
- 3:11language Pearl, but we try to make the
- 3:13easy things easy and the hard things
- 3:15possible. So, out of the box, pretty
- 3:17similar, the starting point, but you can
- 3:20kind of take it in any direction that
- 3:21you want. So, it's not that, oh, there's
- 3:23like a separate product for, you know,
- 3:25one company versus another, but the
- 3:28agents that you can build on it, and
- 3:30we'd like to think of these agents as
- 3:32products onto themselves, can be
- 3:34arbitrarily customized.
- 3:36>> What does it look like to build on the
- 3:38Sierra platform?
- 3:40>> So, we have a It's basically a web app.
- 3:42There's three main sections. There's
- 3:44analyze, build, and then there's
- 3:46release.
- 3:47Within the analyze section, you have
- 3:49things like our explorer agent, which is
- 3:53kind of the long-running ChatGPT deep
- 3:55research for all of your customer
- 3:57conversations and data. You have
- 4:00reports, you have monitors, which are
- 4:02kind of always-on evaluators of
- 4:04conversation data as well.
- 4:06And then within build, you have
- 4:08ghostwriter, which is the agent similar
- 4:10to Codex or Claude code for building
- 4:12agents.
- 4:14You also have journeys, kind of the
- 4:15underlying source code layer, although
- 4:17it's not really code. It's more like
- 4:19natural language or standard operating
- 4:21procedures. As well as kind of different
- 4:24variables
- 4:26and everything like that. On the release
- 4:28side, you have all of the collaboration
- 4:31and change management and governance
- 4:33procedures. And so Sierra, I think at
- 4:35this point we're working with most of
- 4:37the Fortune 20, something like 40 or 50%
- 4:41of the Fortune 50 or Fortune 100. So
- 4:43very much
- 4:45with a lot of the largest companies in
- 4:46the world. And they have needs around
- 4:49governance and release processes and
- 4:51change management that have just pushed
- 4:54us to develop from the very beginning,
- 4:56you know, very buttoned-up procedures
- 4:58and collaboration and review and all of
- 5:00this stuff. That's basically what it's
- 5:02like, I think, on the surface, but
- 5:04probably similar to a lot of other, you
- 5:06know, places that you go to build
- 5:08things, whether that's
- 5:09Figma or Claude code or these different
- 5:12places, but just very much optimized
- 5:14around no-code agent building and um
- 5:18giving you all those capabilities.
- 5:20>> And are those different steps intended
- 5:23to be done in that order? Like analyze,
- 5:24build, release. Can you analyze
- 5:26basically human transcripts before you
- 5:28build the AI agent? Or does analyze
- 5:30really come after you build and release
- 5:31the first version of the agent and now
- 5:32you're iterating on it?
- 5:34>> It's both. So, typically, you'll come in
- 5:38with some sort of resource of how you
- 5:40want the agent to be structured and
- 5:43architected, how you want it to behave.
- 5:45Maybe that's transcripts, maybe that's
- 5:47standard operating procedure, maybe
- 5:49that's a conversation that you have with
- 5:51Ghostwriter.
- 5:52And that will typically be how you build
- 5:54the agent. So, I'd say most people will
- 5:56start with build, but then once your
- 5:58agent is live and production
- 6:01conversations are happening, your daily
- 6:04routine probably starts more with
- 6:06analysis. You're probably thinking, how
- 6:08can I optimize the metric that I care
- 6:10about, whether that's customer
- 6:11satisfaction or resolution rate or in
- 6:15the case of the customer I mentioned,
- 6:16like sales converted.
- 6:18And so, you get those insights and then
- 6:21you want to make improvements to the
- 6:23agent, whether it's, you know, fixing an
- 6:25issue or finding a new opportunity to
- 6:27hill climb on a metric or please
- 6:30customers in one one more way.
- 6:32And so, that typically involves, you
- 6:34know, working with Ghostwriter. Often,
- 6:37Ghostwriter will actually proactively
- 6:38suggest an improvement on the insights
- 6:41to kind of close that loop and build
- 6:42that flywheel.
- 6:44Um, but I would say the day to day is
- 6:45more analyze, build, release.
- 6:47>> Who is doing that analyzing and that
- 6:50iterative improvement? Is this is this
- 6:52engineers? Is this product folks?
- 6:54>> It's primarily people that have the most
- 6:58depth and insight about the ideal
- 7:00customer experience, which tends to be
- 7:03operations uh people, so customer
- 7:05experience managers,
- 7:07um folks in that department at our
- 7:09customer companies.
- 7:11It's also a number of engineering teams
- 7:15will build either other agents that
- 7:17interface with Sierra agent or they can
- 7:20extend the platform
- 7:22uh via basically tools and packages that
- 7:25you can then kind of see and introspect
- 7:26on Sierra. So, it's very much kind of
- 7:29the same way that you have the person
- 7:31that knows everything about your
- 7:33knowledge base, we want them to be able
- 7:34to come in and self-serve on day one and
- 7:36just, you know, make the perfect
- 7:38instantiation of knowledge. The person
- 7:40who knows everything about the standard
- 7:42operating procedures should be able to
- 7:44just do that in the product. And so
- 7:45we're constantly kind of trying to sand
- 7:47down all of the barriers between the
- 7:50people with the most context and their
- 7:52ability to contribute directly to the
- 7:53platform.
- 7:54>> You've said no code a few times. So,
- 7:57what does this agent building experience
- 7:59look like? Is it truly no code? And And
- 8:01I'm assuming it's maybe something like
- 8:04quad code where you talk to it and it
- 8:06generates something under the hood. Is
- 8:07it generating code? Is it generating a
- 8:09custom DSL?
- 8:11>> Yeah, good question. So, the layers of
- 8:12the stack, you have kind of what we call
- 8:14Agent OS, which has our constellation of
- 8:17models. So, translating the tasks that
- 8:20need to be done on the platform
- 8:22into prompts, uh into data injection
- 8:26across, you know, 10 or 15 different
- 8:28models that might be invoked for a given
- 8:30conversation turn. Some of those might
- 8:32be frontier models that need to do, you
- 8:35know, top-tier reasoning. Some of them
- 8:37might be in-house models that are very
- 8:39good at a specific task, and some of
- 8:41them might be just, you know, classifier
- 8:43models that run really well on a
- 8:46uh a model that's a little bit cheaper
- 8:47and more performant. And so, that's kind
- 8:49of the base layer. On top of that, you
- 8:52have the Agent SDK, which is the
- 8:56code-based layer of agent orchestration
- 8:59and context management.
- 9:01That's kind of where Sierra started, but
- 9:03over the last 18 months, most of the
- 9:06agent development, um pretty much all of
- 9:08the agent development has shifted to our
- 9:10no code layer that we call journeys.
- 9:13It compiles down to Agent SDK code
- 9:16deterministically
- 9:17and uh isomorphically, which is a fancy
- 9:19word for you can turn it one way and
- 9:21then turn it back and it's the same. And
- 9:23so, you can have uh code that you
- 9:26transition over to no code, you can have
- 9:27no code that you transition over to
- 9:29code, but the language of specifying it
- 9:32is very much
- 9:34declarative. Here's how I want the agent
- 9:36behavior to be Uh when customers ask
- 9:39about this. We want to unlock these
- 9:41conditions and kind of flow in this
- 9:42direction and we find that that's pretty
- 9:44intuitive cuz it maps the type of
- 9:46document that we would write for someone
- 9:48joining the team in a customer
- 9:50experience role or sales role. You would
- 9:52explain to them how to do the job and
- 9:54that's kind of what you're doing here as
- 9:56well. But there is some DSL for
- 9:58journeys. It's not pure raw kind of like
- 10:01text.
- 10:01>> Correct.
- 10:03And it's very hard. I'm not sure we
- 10:05could get into a discussion about it. If
- 10:07you're just doing text, you have to
- 10:09choose between
- 10:11this is non-deterministically compiled,
- 10:14which all of the experiments we've done
- 10:16in that direction,
- 10:18you end up I think with more harm than
- 10:19good.
- 10:20Or this is a prompt engineering task,
- 10:23which then puts you in the realm of
- 10:25engineering teams. And we, you know, are
- 10:27very proud to be more in the realm of
- 10:30operations teams where a lot of that
- 10:32domain specific knowledge resides. The
- 10:34other big piece of it is that
- 10:35Ghostwriter has totally changed the
- 10:37learning curve for building agents. So
- 10:39you come in and you just say, "Hey, you
- 10:41know, I want to orchestrate order
- 10:42returns or I want to do flight booking
- 10:45or I want to do car rental or
- 10:47referral from primary care provider to a
- 10:49specialist." And Ghostwriter just kind
- 10:52of already knows those concepts and is
- 10:54an expert in journeys. But Ghostwriter
- 10:56is using the journeys product. So it's
- 10:58not writing code, it's writing journeys
- 11:01directly so that you can go inspect that
- 11:03after the fact as well.
- 11:05>> I imagine there's some there's some
- 11:07format that these journeys have to have
- 11:09to adhere to and I imagine that's not in
- 11:11the models training data at all. Was it
- 11:13hard to teach it that format or was it
- 11:15pretty easy?
- 11:16>> That's a really good question because at
- 11:18every point there's this conflict
- 11:20between here are the perfect
- 11:22abstractions for me
- 11:23and here are the abstractions that the
- 11:25models are most familiar with. And
- 11:26similar to in math how you're often
- 11:29taking one problem and reframing it in
- 11:31another problem to do a proof or
- 11:32something like that. You have to decide
- 11:34if you want to reframe this problem into
- 11:36something the models understand or build
- 11:38a skill and inject context in the right
- 11:40way so that the models can understand
- 11:42your way of thinking.
- 11:44The truth is that we do both. So there
- 11:47are cases where we'll say you know,
- 11:49coding agents are really good at file
- 11:50systems. They're really good at get.
- 11:52They're really good at grep. Let's
- 11:54materialize everything into those
- 11:56structures so that coding agents can
- 11:59just you know, cook. Then there are
- 12:01other cases where it's like, "No, no,
- 12:02no. Our way of thinking about this is
- 12:05the correct way of thinking about this
- 12:07and there's not really a way to shoehorn
- 12:09it into what it models are already good
- 12:11at. So let's do the investment to make
- 12:13the models good at this. My personal
- 12:15perspective is that 80% of the time you
- 12:18want to do the first thing
- 12:20and just meet the models of where they
- 12:22are on their turf and you should reserve
- 12:24the second one for that really special
- 12:27case. I'm curious uh if that's what been
- 12:29your experience as well.
- 12:30>> I think recently it's probably gotten to
- 12:33be we see a lot of people using the file
- 12:35system as an abstraction and so I think
- 12:37recently there's been a lot of talk
- 12:38especially as the labs talk about how
- 12:40they're RL-ing the models to be really
- 12:41good for their hardness to try to fit
- 12:43everything into a file system or this
- 12:45particular like edit file tool or things
- 12:48like that. I also think that the models
- 12:50are really good at writing certain
- 12:52packages. Like if you're in the training
- 12:53data, I think I think a lot of LangGraph
- 12:55is in the training data. So I think at
- 12:57least Anthropic models recommend
- 12:58LangGraph for a lot of use cases and
- 12:59that's great.
- 13:00But for newer things like deep agents,
- 13:02which is a new package we have, it's not
- 13:03in the training data at all. We spend a
- 13:05little bit of time, not maybe not as
- 13:06much as we should, but we spend a little
- 13:07bit of time thinking about what makes
- 13:09these models good at writing certain
- 13:10things. We really have no clue how to
- 13:12like affect know how to affect what goes
- 13:13in the training data, but it's a really
- 13:15interesting thing and so I think there's
- 13:16definitely been cases where we see that
- 13:19people choose technology because the
- 13:21models are really good at writing it.
- 13:23And so one question I was also going to
- 13:25ask for the agent SDK, I imagine that's
- 13:27you know, your own custom kind of like
- 13:28framework built in house. I don't know
- 13:30if you experimented with having It
- 13:32sounds like you didn't you you're not
- 13:34having Ghostwriter write that directly,
- 13:36but like that's obviously much more
- 13:37closer to code and so I was curious if
- 13:39you experimented with Ghostwriter or any
- 13:41model like writing agent SDK versus just
- 13:43writing like raw code. So, yes, um
- 13:48one of the things too is if you almost
- 13:50do the abstraction that the models are
- 13:52really good at, it can be overconfident
- 13:54or it can be familiar and successful.
- 13:56And so you have to be very thoughtful
- 13:58about going either all the way there or
- 14:00just not going there at all.
- 14:02>> Like you're saying agent SDK is
- 14:03somewhere in the middle and that might
- 14:04actually confuse it.
- 14:05>> Exactly. Exactly. Um
- 14:07we have kind of reinvented the agent SDK
- 14:10two or three times as models improve. So
- 14:13it used to be you had to have more
- 14:15deterministic guardrails in order to get
- 14:17the behavior that you want. Now there's
- 14:19more room for reasoning at each
- 14:22individual step and you can kind of push
- 14:23out the frontier of that reliability
- 14:26versus reasoning trade-off. So that's
- 14:29been very interesting. The reason for
- 14:31Ghostwriter primarily or entirely
- 14:34editing no code is just that that's
- 14:36where the vast majority of activity is
- 14:39on the platform today. So that's what
- 14:41our customers know and so making
- 14:43Ghostwriter good at it is really where
- 14:45all of the payoff is.
- 14:47I think if we tried to do it for code,
- 14:49it would be a a similarly scoped task,
- 14:51but it would be hard to get it to be
- 14:53really good at both at the same time.
- 14:54There's always going to be some
- 14:55trade-off.
- 14:56>> Do you still let users edit the agent
- 14:58SDK code if they want or is that now
- 15:01like completely abstracted away from
- 15:02them in terms of just pitch journeys?
- 15:05>> So the core agent SDK is part of the
- 15:07Sierra platform,
- 15:08but building agents in code is totally
- 15:11something you can do. One example is a
- 15:13number of our customers have CI/CD a
- 15:16continuous integration pipelines that
- 15:18they want to make sure their agent is
- 15:20released on. And so they need to get
- 15:23repository, which is where their agent
- 15:25lives. Another example is sometimes you
- 15:27have a particularly complex tool that
- 15:29has interacts with a streaming API or
- 15:31something in a way that is just easier
- 15:33to model in code than in no code. And
- 15:36so, the way that these work is because
- 15:38no code compiles down to code, you can
- 15:41kind of import or under the hood it will
- 15:43import code files and compiled no code
- 15:45files kind of all as though they're the
- 15:46same thing because they are.
- 15:48So, I think this is a benefit of
- 15:50starting out as a code-based platform is
- 15:52that we still support it. We have a
- 15:54number of customers that have dozens or
- 15:56in a few cases 100-plus developers
- 15:59building on the platform. And
- 16:01sometimes for that, you know, they work
- 16:03in Git, they release in Git. And so,
- 16:05being part of their enterprise change
- 16:07management protocol just means
- 16:09supporting Git.
- 16:10>> You mentioned that the agent SDK has
- 16:12changed over over the past few years as
- 16:14everything in the space has. Um,
- 16:16what does it look like now and how has
- 16:18it changed? What does that evolution
- 16:20look like?
- 16:21>> So, it started out we call this now
- 16:23flow-based, very much like, you know, do
- 16:26this. And it wasn't just like do these
- 16:28things. I think I'm a big fan of your
- 16:30not another workflow builder blog post.
- 16:32So, it wasn't that rigid, but it would
- 16:34be hey, you know, make sure you collect
- 16:36their email before you say that you're
- 16:40going to send them a confirmation email,
- 16:42right? Very clear to us, but you would
- 16:45want to do those things in that order.
- 16:47Now, I think if you think about just the
- 16:49way that agents can reason through tool
- 16:51calling,
- 16:52instead of having to specify that in the
- 16:54actual structure of your agent, you
- 16:56might just give it the context that in
- 16:59order to call this tool, you know, as a
- 17:01prerequisite you should have their
- 17:02email.
- 17:03And it will know how to ask for it. So,
- 17:05it's really like you'll just say that in
- 17:06the prompt. Yes.
- 17:08Or, you know, eventually all things end
- 17:10up in prompts, but it would be in the
- 17:12journey.
- 17:13And so, as you're kind of writing it
- 17:15out, you would specify that that's one
- 17:16of the rules or policies of the journey.
- 17:19And then the agent can take care of the
- 17:22rest. I think it's a mix of the models
- 17:24getting better and our orchestration
- 17:26platform becoming even more
- 17:28sophisticated and robust. So, when I
- 17:29talk about that constellation of models,
- 17:32there's a lot of We talked a little bit
- 17:34before about Linnaeus and Darwin.
- 17:36There's post-training that goes into
- 17:37that. There's model selection and eval
- 17:40and prompt engineering as well. And so,
- 17:43I think it's it's kind of equal parts
- 17:45model improvements and platform
- 17:47improvements.
- 17:47>> I want to talk about this model in a
- 17:48second, but I want to stay on the
- 17:49harness for a little bit. How similar
- 17:51does it look like in its current form to
- 17:53a coding agent harness? Does it have
- 17:55access to skills and sub-agents in the
- 17:57same way that someone using Claude code
- 17:59would would have?
- 18:00>> So, the one
- 18:02constraint we have that Claude code
- 18:05doesn't have is latency. Majority of
- 18:07Sierra conversations are voice. And if
- 18:10you're not responding in 1 or 2 seconds,
- 18:12then people wonder where you went. And
- 18:15so,
- 18:16we are highly optimized for these low
- 18:17latency use cases. There's a ton of
- 18:19parallelism.
- 18:21That being said, at a high level, it is
- 18:23using a lot of the same models. It has
- 18:26access to tools. Um so, there's a lot of
- 18:28similarities. There are also, you know,
- 18:31you can invoke other agents from the
- 18:33Sierra agent. So, the core What's best
- 18:34for the core conversation loop isn't
- 18:37typically what's best for software
- 18:39development. But you might want to say,
- 18:41"Hey, let me actually give you a
- 18:43callback in 20 minutes after I figure
- 18:45this out." And then you would have a
- 18:47type of loop that runs, you know, more
- 18:48like Claude code.
- 18:49>> And for those for those longer loops
- 18:51that might happen in the background, are
- 18:53those also built on the Sierra platform
- 18:55and are just a separate type of agent
- 18:57that remove the latency constraint?
- 18:58>> You can do it either way. So, you could
- 19:00have a Sierra agent calling out to
- 19:03another Sierra agent or you could also
- 19:05have a Sierra agent calling out to an
- 19:07in-house platform. And because so many
- 19:09of our customers have their own
- 19:11technology teams and a robust array of
- 19:15different AI projects internally.
- 19:17They might be experts in a particular
- 19:20area. Like they might have document
- 19:22generation handled themselves and that
- 19:24might be a long-running agent and then
- 19:25the Sierra agent can call out to it and
- 19:28wait for a response. So it's kind of up
- 19:29to you to choose and we find that
- 19:32enterprises are varied enough that they
- 19:34appreciate kind of having choice.
- 19:35>> When you do these agent-to-agent
- 19:37communications, are you are you using
- 19:40A2A or one of the protocols specifically
- 19:42for that or MCP or just an in-house, I
- 19:45don't know, REST API call?
- 19:47>> The most common is an API call. When you
- 19:49know who you're talking to in advance,
- 19:51often times you can save a lot of tokens
- 19:54and make sure that you're 100% accurate
- 19:56that way. That being said, Sierra agents
- 19:59support the MCP and agent-to-agent
- 20:02protocols. You can kind of install that
- 20:04integration and then your agent can be
- 20:06an MCP client. You can also set up your
- 20:08agent to be an MCP server. So this is
- 20:10how we support ChatGPT apps
- 20:13uh which rely on MCP servers. Basically,
- 20:16the tools of the agent can be made
- 20:18available to ChatGPT and then you can
- 20:20at-reference a Sierra agent. The example
- 20:23uh that would be most familiar is
- 20:24Redfin.
- 20:26Um if you do if you go to redfin.com and
- 20:28do their AI search, uh under the hood,
- 20:31it is a Sierra agent that is returning
- 20:33the home listings and having the
- 20:35conversation with you and that agent is
- 20:37also, I believe, available in ChatGPT.
- 20:40>> Interesting. I I didn't realize that
- 20:42Sierra agents could be ChatGPT apps. Is
- 20:45that the right terminology for them?
- 20:46>> It is. Yeah, exactly.
- 20:48>> Do you have Do you have an opinion or
- 20:50hot take on in the future, do you think
- 20:52people will be interacting with the
- 20:55agents that represent brands on
- 20:57dedicated chatbot websites or in uh
- 21:00ChatGPT or central chat engine?
- 21:03>> I think that agentic commerce will be
- 21:06bigger than e-commerce.
- 21:08So, if I think about how I get things
- 21:10done today, it used to be that I went to
- 21:12websites and clicked around. Now, I ask
- 21:16Codex or Claude to do things for me.
- 21:19And I don't see why I won't do that to
- 21:21manage my subscriptions, to order
- 21:23supplies to my home, to make dinner
- 21:25reservations. It just feels like that's
- 21:28where we're headed.
- 21:29And so, if that's happening, I think
- 21:31brands will want to be ready on the
- 21:33other side of that. So, we're very much
- 21:35planning for that world. We were
- 21:37investing in payments before it made
- 21:40sense, I think, and it's a long process,
- 21:43but a few months ago, we announced, you
- 21:46know, we're uh fully PCI DSS level one
- 21:49certified.
- 21:50>> no clue what that means. What does that
- 21:51mean?
- 21:52>> payment card industry.
- 21:54>> Okay.
- 21:54>> Um oh, man, you stump me on DSS. Uh
- 21:57[laughter]
- 21:58I have uh some of the other acronyms um
- 22:00in my head, but um
- 22:01>> We'll We'll put it in in post.
- 22:03>> Okay. Okay, thanks. Oh, man, it must be
- 22:05like digital
- 22:07I don't know. I know that the We were uh
- 22:09certified by a QSA, which is a qualified
- 22:11security assessor. And uh what that
- 22:14means is we're able to do the only voice
- 22:17payments platform, certainly at launch,
- 22:19I think still this is the case, where
- 22:21you don't need to transfer to another
- 22:23platform. So, it's a co- cohesive
- 22:25experience throughout checkout. And all
- 22:28the work that went into that, we have
- 22:29isolated infrastructure where the
- 22:32payment info doesn't go to a large
- 22:34language model, doesn't go to an
- 22:36external large language model, because
- 22:37we have none of the LLM providers are uh
- 22:39PCI certified in that way.
- 22:41And so, putting that all together is
- 22:43like spinning up a separate cluster, you
- 22:45know, getting certified, um making sure
- 22:47all of our operational rituals, you
- 22:49know, conform to what the security
- 22:51assessor is looking for.
- 22:53And we put in that work because we
- 22:55believe in this future where agentic
- 22:56commerce is actually bigger than
- 22:59e-commerce. And I think e-commerce is in
- 23:01the hundreds of billions of dollars at
- 23:03this point, couple percentage points of
- 23:04GDP or something like that just in the
- 23:06US. And so if you think about that
- 23:09space, um it's pretty big.
- 23:11>> And by agentic commerce, do you mean
- 23:13like chat GPT talking to uh Sierra agent
- 23:17that represents Redfin, or do you mean
- 23:19someone going to Redfin's agent no
- 23:21matter where it is and talking with it
- 23:23there?
- 23:24>> Both, both. I do think that the majority
- 23:27of this will be from personal agents.
- 23:30Just looking at user behavior, we spend
- 23:33so much time in Claude and chat GPT and
- 23:36Codex
- 23:38that you have to think that's where a
- 23:39lot of that behavior will accrue.
- 23:42>> And do you think those agents will
- 23:43interact with another agent? Why not
- 23:46just the raw APIs themselves?
- 23:48>> I do. I think that as you think about
- 23:51being ready for that world,
- 23:53um the same way that you might want to
- 23:57use Shopify or you might want to use
- 24:01certain software on your website to do
- 24:03product recommendations, to do checkout,
- 24:07uh you might want to use Stripe.
- 24:08Similarly, you'll want to use a platform
- 24:11that can make sure that you're
- 24:12presenting your products in the right
- 24:15way, that you're making checkout as easy
- 24:16as possible, and that you're showing up
- 24:19at your best, whether it's for a
- 24:21customer that's browsing or for an agent
- 24:23that's browsing. The one thing that I
- 24:24think is pretty different is the
- 24:26attention
- 24:28isn't necessarily valuable in the same
- 24:30way.
- 24:31Like our eyeballs are more valuable than
- 24:34an agent just, you know, spewing out
- 24:36tokens assuming that no one's ever going
- 24:37to look at it. At least I think that's
- 24:39true for now. At some point it lands up
- 24:41in some future training run and maybe
- 24:43has value, but I think that's de minimis
- 24:45relative to getting us to look at
- 24:47things. And so I do feel like maybe
- 24:49that's a bit different, but the
- 24:50presenting yourself in the right way,
- 24:52making it easy to check out, making it
- 24:54easy to understand what products are
- 24:55available, to express the preferences of
- 24:58whoever is responsible for that agent
- 25:01going off and doing something
- 25:02commercial. That all still feels
- 25:04relevant to me. I've seen some dev tools
- 25:06provider, I think Sentry, doing a
- 25:08similar thing where they have a bunch of
- 25:10APIs, obviously, for the underlying
- 25:11platform, but they also have an endpoint
- 25:14to just ask questions of the agent
- 25:15directly. And I think you could make a
- 25:17counterargument that like, great, brands
- 25:18should absolutely care about how the
- 25:20platform's being used and how it's being
- 25:22presented, but you could do that with
- 25:23skills or or some other mechanism to
- 25:25expose that to the agent. And I I
- 25:26honestly don't know which one's right,
- 25:28but it it has been interesting to see
- 25:30the whole space is so new, but
- 25:31increasingly so companies exposing
- 25:32agents as endpoints to interact with
- 25:34rather than the endpoints themselves.
- 25:36>> I agree. I think that all of this stuff
- 25:39you could try to do it yourself. It
- 25:40might be that certain companies, that's
- 25:42the best option.
- 25:44What we've seen is that because there's
- 25:47often tens or hundreds of millions of
- 25:49dollars on the line, in some cases
- 25:51billions of dollars on the line, you
- 25:53really want to make sure you're getting
- 25:54the best solution. And so if you're
- 25:56going to be 90% as good at it as you
- 25:59could be partnering with a company like
- 26:01Sierra,
- 26:02it still makes sense to partner and, you
- 26:05know, get that extra few billion
- 26:07dollars.
- 26:08>> One last question on this fun side
- 26:10tangent, payments. How early are we?
- 26:12>> I think we're really early.
- 26:13I personally still don't order paper
- 26:17towels with Codex. I don't know if you
- 26:20do.
- 26:20>> No, and that's why I asked. I'm glad
- 26:21that you said that cuz I know I'm not
- 26:23close to doing that. And so I was
- 26:24wondering how far behind I was.
- 26:26>> I mean, yeah, like
- 26:28I also didn't do it with Alexa. You
- 26:30know, I think for some of that really
- 26:31easy stuff, you could probably have done
- 26:33it already.
- 26:34The one that I think will definitely
- 26:36become a thing is there are a lot of
- 26:38apps that claim they can, you know, go
- 26:40through all of your subscriptions and
- 26:41cancel the ones that you're not using.
- 26:44That feels like as a consumer, that's a
- 26:45useful service. I'm definitely closer to
- 26:48doing that with Codex than I am
- 26:51with an app. You know, it would it would
- 26:53be so much work to tell it all the
- 26:55things and try to tell it which ones to
- 26:57cancel. It's a very manual process.
- 27:00If I gave Codex or Cloud Co-worker or
- 27:02something just access to my browser and
- 27:04said,
- 27:05"Hey, you know, go to all of the
- 27:06streaming apps and like the ones that
- 27:08I'm not logging into,
- 27:11just, you know, cancel those and let me
- 27:13know if you need my password." And
- 27:15obviously you have to figure out how to
- 27:16make that secure and everything. But I
- 27:18feel like that I would have demand for
- 27:20that product.
- 27:21>> Going back to the harness for a little
- 27:23bit, you said something earlier about
- 27:25things running in parallel. Is that like
- 27:26guardrails that you're running or
- 27:28retrieval steps or what's running in
- 27:30parallel in in this process?
- 27:32>> So many things. One example, knowledge.
- 27:35We will
- 27:37often look up answers before we know if
- 27:40we want them.
- 27:41So you'll, you know, before you decide
- 27:43whether this question needs an answer,
- 27:45you'll at least have the answer ready or
- 27:47in parallel with deciding. So sometimes
- 27:49you're classifying and you're responding
- 27:52at the same time. Basically speculative
- 27:53execution.
- 27:55Another example would be transcription,
- 27:57what we call ensembling.
- 27:59Um I think we might have published a
- 28:01blog post on this today, which is great.
- 28:03Go read it. We've learned so many things
- 28:06from being a
- 28:08uh modular architecture on voice. This
- 28:11was an early decision we made that I
- 28:13think has totally played out to our
- 28:15advantage where we have the ability for
- 28:18any language, for any customer, for any
- 28:21use case to multi-home providers
- 28:25uh across transcription, across
- 28:27synthesis, and across native uh
- 28:29voice-to-voice models. And so on the
- 28:31transcription side, for example, it just
- 28:33turns out uh when you have a thick UK
- 28:36accent from northern UK or at least
- 28:39parts of northern UK. I don't know
- 28:40exactly the region.
- 28:42There is one model that has the highest
- 28:44quality transcription,
- 28:46but it hallucinates during silence more
- 28:48than other models.
- 28:50So, we run two models in parallel.
- 28:52And if this model says it's silent, you
- 28:54trust it. If this model does not say
- 28:56it's silent, you trust this one. And so,
- 28:58that's just an example where we're
- 29:00running those in parallel. Uh we have
- 29:02logic for when you take the right one,
- 29:04and it's very specific. And if you had,
- 29:07you know, all your chips in with one
- 29:09provider or one system, uh or you
- 29:12weren't doing things in parallel, you
- 29:14would inevitably hit the limits of what
- 29:16that provider can do. So, the same way
- 29:18we use Claude and Gemini and the
- 29:20GPT-class models, we're also able to use
- 29:23all of the leading players on the
- 29:25transcription and synthesis and
- 29:27speech-to-speech side as well.
- 29:28>> You mentioned you have some in-house
- 29:29models as well. What do those models do,
- 29:32and why did you guys decide to build
- 29:34those in-house?
- 29:35>> So, uh knowledge is a great example. I
- 29:38think whenever we are pushing the limits
- 29:40of what's possible, we always consider
- 29:43whether we should build this in-house
- 29:45whenever it's limiting our ability to
- 29:47deliver more for our customers.
- 29:49So, an example where we're probably not
- 29:52the company is these, you know, many
- 29:54millions of dollar training runs that
- 29:56produce the GPT-5.5 class models. And,
- 30:00you know, Mythos, I'm sure is a many,
- 30:02many millions or tens or hundreds of
- 30:04millions of dollars training run um to
- 30:06get that produced. And that's stuff that
- 30:08OpenAI and Anthropic are just the best
- 30:10in the world at.
- 30:11I think what we're the best in the world
- 30:13at is going really deep with customers,
- 30:15understanding all of the process
- 30:17knowledge uh specific to their industry,
- 30:20specific to their company, specific to
- 30:22their customer base.
- 30:24And then having the products that can
- 30:26allow them to serve those customers as
- 30:28best as possible. And so, an example
- 30:30like knowledge where we were hitting the
- 30:32limits of the retrieval and reranking
- 30:35that we could do with out-of-the-box
- 30:36models, we asked the question of, you
- 30:39know, should we create our own models
- 30:40here
- 30:41um and eval them. And we have a research
- 30:43team that's pretty sizable and tightly
- 30:46integrated with our product teams. And
- 30:48so, we can flex that muscle when we need
- 30:50to, but we try not to be doing it just
- 30:52for the sake of doing it.
- 30:54>> You mentioned something like every run
- 30:55of the agent would have 10 to 15
- 30:57different model calls. If you had to
- 30:59guestimate like how many of those are
- 31:01frontier model calls versus like
- 31:02in-house fine-tune versus not frontier
- 31:05model but but third party?
- 31:07>> So, for a typical, um, turn of a
- 31:09conversation, I would guess that and
- 31:13this is just ballpark, but, you know,
- 31:15being, uh, precise rather than being
- 31:17accurate. Uh, I think maybe a couple
- 31:20frontier model, a handful of classifiers
- 31:24that probably don't require that, a
- 31:26handful of speculative execution in the
- 31:28case of voice in particular to make sure
- 31:30that it's low latency. Um, sometimes
- 31:33there will be an interim response that's
- 31:34generated to, you know, the same way
- 31:36you'd say, "Hold on a minute, I'm just
- 31:37pulling up your account." Like that kind
- 31:39of thing. Roughly like a third a third a
- 31:41third or a quarter quarter quarter, but
- 31:43I would say the frontier models just
- 31:45because they might be slower or more
- 31:49expensive would probably be, you know,
- 31:52more doing the bulk of the reasoning but
- 31:54in one or two inferences for a given
- 31:56conversation turn.
- 31:57>> Do you ever end up training models
- 31:59specific to a customer?
- 32:01>> It's not something that would be out of
- 32:02the question, but I can't think of a
- 32:04specific example. The reason I pause is
- 32:07because we do have our agent data
- 32:09platform
- 32:11and there are machine learning models
- 32:13that power strategies that are specific
- 32:16to customers. But in terms of like a,
- 32:18uh,
- 32:19you know, language model or generative
- 32:21model, we don't have cases of that.
- 32:23>> What's the agent data platform and what
- 32:25does it power?
- 32:26>> Basically, one thing we realized pretty
- 32:28early on is large language models are
- 32:30really good at in the moment empathy.
- 32:32Oftentimes better than we are of
- 32:33understanding, okay, you You I
- 32:35understand you're having a hard time.
- 32:36I'm really sorry about that. And it's
- 32:38the same way when you walk into
- 32:40a restaurant that has amazing service or
- 32:43a hotel that has amazing service, they
- 32:45recognize the moment you walk in, okay,
- 32:48this person just got off a really long
- 32:49flight or this person's 10 minutes late
- 32:52to their reservation and they were stuck
- 32:54in traffic and I'm just going to let
- 32:55them know that that is not a problem.
- 32:57Their table is ready.
- 32:59And large language models have that,
- 33:01especially on a platform like Sierra.
- 33:03But they don't necessarily know what you
- 33:05care about at a level deeper than that.
- 33:08And oftentimes the previous generation
- 33:10of AI or recommender systems have a
- 33:13better understanding of some of those
- 33:14things.
- 33:16And so what Agent Data Platform does is
- 33:18it can either integrate with your
- 33:20customer data platform, all your
- 33:21internal systems, or you know, you can
- 33:24have it just on Sierra or you can do
- 33:25sort of a zero copy integration.
- 33:28And it can take that structured data
- 33:31that knows what to recommend along with
- 33:33the here and now and now in the use
- 33:34those to
- 33:38generate uh better conversations, better
- 33:42uh orchestrations around how you want
- 33:44customers to feel and what you want to
- 33:46do for them.
- 33:47Um so that all sounds maybe a little bit
- 33:49abstract. Uh one example would be during
- 33:52sales. Oftentimes there's structured
- 33:54data that knows the right offer to
- 33:55present, but doing it just with that
- 33:58structured data with the previous
- 34:00generation of AI and before Sierra feels
- 34:02very stilted or it feels like you know,
- 34:05I don't know why you're doing this. And
- 34:07so large language models can really
- 34:09understand how to present an offer,
- 34:12uh how to attribute it, weigh two
- 34:14different offers based on conversation
- 34:16context and pick the right one for the
- 34:17moment and that kind of thing. So we see
- 34:19it a lot with sales, with loyalty and
- 34:21retention, those types of conversations.
- 34:23>> One of the last agents we haven't talked
- 34:25about too much is Explorer.
- 34:27>> Yes.
- 34:27>> What is Explorer? What does it look like
- 34:29under the hood?
- 34:30>> I've described it as chat GPT deep
- 34:32research uh for all of your customer
- 34:34context and conversations and all of the
- 34:36data on Sierra. And so what it allows
- 34:38you to do
- 34:39is basically instead of having to go
- 34:41spelunking for the specific insight in
- 34:44reports or in monitors, you can just ask
- 34:47the question. You can say, "Hey,
- 34:49I noticed my resolution rate dipped, you
- 34:51know, why was that?" Or "How can I
- 34:53generate more sales?" Or
- 34:55"I wish that more people were converting
- 34:57from trial to, you know, full-time paid
- 35:00plan. How come that's not happening?"
- 35:03And then more than that, you can set up
- 35:04automations so that, you know, on a
- 35:06daily basis, for example, Explorer can
- 35:09ask the same questions proactively.
- 35:11And then partner with Ghostwriter. We
- 35:14currently think of these as kind of two
- 35:15separate agents, the analysis agent and
- 35:18the authoring agent,
- 35:20to say, "Oh, here are some fixes that
- 35:22are suggested to improve." And you can
- 35:23chat with Ghostwriter and kind of pick
- 35:25it up from there. Under the hood where
- 35:27this is converging, I think is a shared
- 35:29harness that is an expert at using Agent
- 35:31Studio,
- 35:33Sierra's platform. Um and so that's kind
- 35:35of what we've been setting up in terms
- 35:37of what we talked about at the
- 35:38beginning, like,
- 35:39you know, figuring out the file system
- 35:41architecture that maps to the product.
- 35:44Um and so as we've exposed more and more
- 35:47tools,
- 35:48um you know, like building knowledge
- 35:49bases to these agents, they get more and
- 35:52more powerful and we see a lot of
- 35:54emergent behavior between both.
- 35:55>> Does this harness end up looking more
- 35:58similar to like a coding agent harness
- 36:00than the the harness that's part of
- 36:02Agent OS or Agent SDK?
- 36:04>> Yes. Um so this is less of a quick-turn
- 36:07conversational agent and more of a
- 36:09longer-term deep analysis agent. And so
- 36:13it ends up looking a lot more like a
- 36:15cloud coder or Codex.
- 36:16>> One of the things I've been thinking
- 36:17about, I'm curious if you have a take
- 36:18here.
- 36:19In a year, two years, three years, will
- 36:22there be the split just in terms of
- 36:24harness? Ones that's optimized for kind
- 36:26of like, yeah, lower latency, external
- 36:28facing, customer experience type things.
- 36:31Voice is maybe heavily involved. And
- 36:33another that's really focused on these
- 36:34deep research, maybe coding, like you're
- 36:36running in a sandbox, things like that.
- 36:38Or will just end up converging into one
- 36:39harness that, you know, depending on how
- 36:41you prompt it or, you know, has these
- 36:43async sub-agents in the background that
- 36:45can maybe run for longer periods of
- 36:46time.
- 36:47>> I think there will always be latency,
- 36:49performance, cost tradeoffs and
- 36:52different architectures that emerge
- 36:54because of that.
- 36:55I've actually been surprised by how many
- 36:58different types of model companies there
- 37:00still are.
- 37:01And when I talk to people who are
- 37:03particularly AGI filled about it and I
- 37:05say, "Hey, what's like a cool model
- 37:08opportunity that's not flying like so
- 37:10close to the sun that the labs will do
- 37:12it?" They'll say, "Oh, well, you'll just
- 37:14ask, you know, GPT-12 to make that
- 37:16model, so it's actually not a big
- 37:17opportunity." But I think in reality, at
- 37:20least up until now, I'll probably look
- 37:23dumb when AGI comes out.
- 37:25>> [laughter]
- 37:26>> You do see a lot of success in areas
- 37:29like voice models from transcription to
- 37:32synthesis. You don't always see
- 37:34leadership from the model labs. You see
- 37:36the model labs actually trying to focus
- 37:38more on specific problems. For
- 37:40Anthropic, I think it's coding. For
- 37:42OpenAI, it's been consumer, now maybe
- 37:44shifting a little bit more to
- 37:45enterprise. Um for Google, definitely
- 37:48consumer as well. And it really does
- 37:50feel like there are still tradeoffs to
- 37:52me. So, I expect there will still be
- 37:54multiple architectures up until that
- 37:57event horizon of AGI's here, so all bets
- 37:59are off.
- 38:00>> As you guys build your core agent
- 38:02harnesses, and I'm assuming one to build
- 38:04them in a model agnostic way, what do
- 38:06you need to change to go from an OpenAI
- 38:08model to an Anthropic model?
- 38:10>> Usually, you have the evals that are
- 38:12designed to work across both. So, if you
- 38:13have really good evals and a really good
- 38:15harness, a really good architecture,
- 38:18then you should be able to kind of hill
- 38:19climb toward eval performance without
- 38:22too much effort. Often what will happen
- 38:24is you'll learn the first time you're
- 38:26switching one task from one to another
- 38:28or making it possible to run on multiple
- 38:30systems that your eval wasn't quite as
- 38:33good as you thought. And so then you
- 38:34make your eval better and you continue
- 38:36to improve. But the short answer is that
- 38:38it's pretty simple for a given
- 38:42intelligence level of a model
- 38:44to run a task on one or the other. And
- 38:48so again it's that basically latency
- 38:50quality cost trade-off, but not more
- 38:52than that. And because we have customers
- 38:54that have very specific requirements
- 38:56around what clouds they can run in, what
- 38:59models they can use, and our
- 39:02company approach is to meet them, you
- 39:04know, on their terms. You don't serve
- 39:07most of the Fortune 20 without that
- 39:08approach. It's not really a choice.
- 39:11And so because that's our approach and
- 39:13we've built a lot of products around
- 39:14that, we also have made sure that we can
- 39:16kind of move between models of
- 39:18comparable intelligence without too much
- 39:20heartburn.
- 39:21>> And what do you end up changing when you
- 39:22hill climb? Is it just the prompts? Do
- 39:24you also change out some of the tools
- 39:26themselves?
- 39:27>> It depends on the case. And I might not
- 39:29be the expert on the exact history of
- 39:32each. I would I think that if you change
- 39:34the tools,
- 39:36it's pretty hard not to have downstream
- 39:38effects of that. And there might be
- 39:39certain tasks that can only run on
- 39:41certain models
- 39:43and other tasks that can run on other
- 39:45models. And so there's always kind of a
- 39:47set of eligible models for specific
- 39:49tasks. I don't know exactly how tools
- 39:52change, but I know the eval getting more
- 39:54robust and the prompt, you know,
- 39:56conforming to the quirks of each model
- 39:58is definitely part of the development.
- 40:01>> You guys recently wrote a blog around
- 40:03context engineering and I think you said
- 40:05it was the key to great agent building
- 40:07or something like that. How do you guys
- 40:09think about context engineering and and
- 40:11and what, you know, tips or tricks would
- 40:12you have for others?
- 40:14>> I think it's showing agents everything
- 40:16they need to do the right thing, but
- 40:18nothing more.
- 40:20And as models get smarter, you can be
- 40:24a little bit less precise with
- 40:26everything they need, and certainly less
- 40:27precise with nothing more. So, early on
- 40:30it was
- 40:31the agent SDK was really about only
- 40:33giving the model exactly what it needed
- 40:35and kind of spoon-feeding the context.
- 40:37Now, to extend the uh meal analogy, it's
- 40:40probably more like, you know, putting
- 40:42out the right dish. And maybe in the
- 40:43future, it might be something that is
- 40:46even less structured. One concept that I
- 40:48think is in that blog post is kind of
- 40:50progressive disclosure. You'll probably
- 40:52know more about this than I do, but
- 40:54when you bring something into the
- 40:56prompt,
- 40:58you don't want to do it before it's
- 40:59relevant, and then you also risk
- 41:02incoherence if you then yank it out of
- 41:04the prompt. So, when you do things like
- 41:06prompt compaction, you just want to be
- 41:08really thoughtful about not making it
- 41:10lossy, because if you keep something in
- 41:12the history that is
- 41:14incoherent with the rest of the system
- 41:16prompt, it's not going to end well. And
- 41:19so, I think to the degree, you know,
- 41:21when we're fixing uh
- 41:24issues, or when when we've seen
- 41:25hallucinations, it's often because one
- 41:28part of the prompt was this, and the
- 41:29other part was this. And actually, one
- 41:31of my main learnings from Sierra is
- 41:34anytime you think the model's being
- 41:36dumb, it's probably you.
- 41:39>> I like that. I think that I I think a
- 41:40lot of people have learned similar
- 41:42lessons from doubting the model.
- 41:43>> the model's too dumb, the model's
- 41:45actually too smart.
- 41:46>> How much you guys care about prompt
- 41:48caching and maintaining that cache? I've
- 41:50heard I've heard kind of like two
- 41:52mindsets on it. One is like, yeah, do
- 41:54everything you can to maintain the
- 41:55cache, like don't invalidate it until
- 41:57you like absolutely need to. And then
- 41:59I've heard another theory that's
- 42:00basically like, yeah, prompt caching's
- 42:01great, but like what matters most is
- 42:03like performance, and sometimes you need
- 42:05to just like break the cache in order to
- 42:07insert the right context, or give it a
- 42:09system reminder, or something like that.
- 42:11How how strictly do you guys try to
- 42:13adhere to prompt caching?
- 42:15>> Is the purpose for those who are prompt
- 42:17caching loyalists for uh speed or cost
- 42:22or quality?
- 42:23>> I think the first two mostly speed and
- 42:25cost.
- 42:25>> Speed and cost.
- 42:26>> Yeah. I haven't I haven't heard anyone
- 42:28argue that it's for quality, but maybe
- 42:30maybe it works better.
- 42:32>> It's a nice to have.
- 42:33Um we definitely don't want to
- 42:36invalidate a cache for no good reason,
- 42:38but quality comes first.
- 42:40So, we aren't uh zealots about it at
- 42:43all. I would also say that when the
- 42:46outcomes that your agents are delivering
- 42:49are very valuable, you have the luxury
- 42:51of not being extremely focused on cost
- 42:54in particular.
- 42:56Uh and so, that probably is part of the
- 42:59reason for that is that, you know,
- 43:01a conversation with a customer could
- 43:03sell a
- 43:05$100 product or a $1,000 lifetime value
- 43:09plan. And so, those are valuable enough
- 43:12that quality almost always comes first.
- 43:14>> We've talked a bunch about the agent
- 43:16itself. There's two topics that we were
- 43:17discussing earlier, and I'm curious
- 43:19We've talked a little bit about them,
- 43:20but I'm curious if you have any more
- 43:21thoughts. First being RL. When is RL
- 43:24good? When is it bad? How much have you
- 43:25guys explored it?
- 43:26>> We've explored it a lot, in part because
- 43:29it has two great promises, you know,
- 43:31increasing the ceiling of the quality of
- 43:33models, and then also making it so that
- 43:35you can do a similar task on more
- 43:38models.
- 43:39I'm curious for your take, but in
- 43:40practice I've seen a little bit more of
- 43:42the second one when it comes to like
- 43:44enterprise RL. It's taking an open model
- 43:47or open weights model and saying, "How
- 43:48can we get similar performance to a
- 43:51frontier model?" The two things that
- 43:53make it hard are, number one, uh the way
- 43:55that that gets delivered is
- 43:57non-deterministic and might include um
- 44:00you can't include any data that you
- 44:02don't want the model to regurgitate. So,
- 44:04we basically would never fine-tune a
- 44:05model on something when it could lead to
- 44:08regurgitation risk. That's just a
- 44:10non-starter. And then also uh in
- 44:12general, just the way that uh you would
- 44:16train the model, you have to think about
- 44:17preparing all of that data.
- 44:19The other big one is that the frontier
- 44:21models are improving so fast that you
- 44:24want to remain as agile as possible.
- 44:28And so in many cases, doing something
- 44:30like RL makes a ton of sense for
- 44:32something like knowledge where we feel
- 44:34like we are pushing the state of the
- 44:35art. But if we're not pushing the state
- 44:38of the art, we really want to be
- 44:39thinking about what's going to be true 3
- 44:41months from now, 6 months from now, and
- 44:43oftentimes RL is a rounding error
- 44:45against that.
- 44:46>> Yeah, I feel like to your point earlier,
- 44:47we've started to hear it a little bit
- 44:49more recently. I think because of cost.
- 44:51So I think like most people are
- 44:53interested in it when they're using
- 44:54these frontier models and the
- 44:55performance is good. But now, whether
- 44:58it's coding or other things, their cost
- 45:00is just going through the roof. I think
- 45:02the and we're starting to investigate
- 45:04this more, but I think the places where
- 45:05we're hearing it most are basically in
- 45:07those where performance is good, cost
- 45:09too high. How can I bring it down? Let's
- 45:10see if I can I can train a model to get
- 45:12similar cost at a fraction of the cost.
- 45:14Or similar performance, fraction of the
- 45:15cost.
- 45:16>> Yeah.
- 45:17Interestingly enough, a lot of our
- 45:19progress here has been driven by
- 45:20capacity, not cost. Where
- 45:24you know, we have a lot of customers
- 45:25that are in the retail space. And when
- 45:27we go into Black Friday, Cyber Monday,
- 45:30for example, you need a lot of capacity
- 45:32to deal with the spikes that they face.
- 45:35Uh we've also done load tests that are
- 45:37on the order of you know, if you were to
- 45:39have that rate of conversation over a
- 45:42year, it would be billions of
- 45:44conversations. And so that level of
- 45:46concurrency and those spikes just mean
- 45:49that we need to be resilient to downtime
- 45:53with a particular provider and ready
- 45:56for, you know, using whoever has the
- 45:58capacity to serve us. And so
- 46:00it's funny because it's useful in so
- 46:02many ways, but a lot of the reason why
- 46:04we have such good support for multiple
- 46:07providers is specifically preparing for
- 46:10Black Friday Cyber Monday and running
- 46:11load tests for really large customers.
- 46:13>> One other
- 46:15harness agent engineering topic,
- 46:17multi-agent systems. Where do you think
- 46:19they're useful, where they're not
- 46:20useful?
- 46:21>> I think they are often not as useful as
- 46:24people think. My thoughts on this would
- 46:26be people should be really thoughtful
- 46:28about why they want a multi-agent
- 46:30system. If you want a multi-agent system
- 46:33so that one team can work on one agent
- 46:35and one team can work on another agent,
- 46:37then you're shipping your org chart. If
- 46:38you want a multi-agent system because
- 46:40it's just makes you more comfortable to
- 46:42think about this problem over here and
- 46:44this other problem over here,
- 46:46then you're also not optimizing around
- 46:49impact.
- 46:50If, for example, you had an agent that
- 46:52does triage and another agent that does
- 46:54a task, by building it as a multi-agent
- 46:57system,
- 46:58you're often
- 47:00depriving of the agent doing the task of
- 47:02the information from the triage and
- 47:05depriving the agent doing the triage of
- 47:06all of the procedural information from
- 47:09the task. And that's typically
- 47:11destructive of value. And so, we are
- 47:14often just want to make sure that we're
- 47:16doing multi-agent systems for the right
- 47:18reason. If you're kicking off even a
- 47:20sub-agent, you want to make sure that it
- 47:22has everything it needs to do that task
- 47:25and that there's no reason why it
- 47:26shouldn't just be part of the main
- 47:28agent. Um and so, I think I've seen a
- 47:30lot of cases where people are reaching
- 47:34for multi-agent systems the same way you
- 47:36might reach for microservices
- 47:39uh before you're necessarily ready for
- 47:41that level of optimization and also for
- 47:44reasons that might not be just about
- 47:46building the best possible agent. And
- 47:48so, Sierra agents tend to be kind of one
- 47:52agent representing the brand. You
- 47:54certainly can have multiple agents and
- 47:56build a multi-agent system, but be if
- 47:59you're managing context correctly, if
- 48:00you're doing really really good context
- 48:02engineering, then typically it's just
- 48:04not a problem because you're not
- 48:05exposing the wrong context to the wrong
- 48:08agent.
- 48:08>> Is there a right time to build a
- 48:10multi-agent system?
- 48:11>> I think if you have truly separable
- 48:13jobs, right? Where there's not any
- 48:15purpose of the first context being part
- 48:17of the second context.
- 48:19I will say that
- 48:21in my personal opinion with you know,
- 48:24May 18th, 2026,
- 48:27there not a lot of great times for it.
- 48:29There might be times where it actually
- 48:31the organizational difficulties are
- 48:33worth the quality drop, but if you're
- 48:36doing it specifically for quality, I
- 48:38think it's it's
- 48:39pretty rare that you can't just solve it
- 48:41with better context engineering and I'm
- 48:43kind of a monolith loyalist on that.
- 48:45>> I feel like voice is one of the things
- 48:48that is getting more and more popular,
- 48:49but there still aren't a ton of people
- 48:52doing a lot of, but you guys are. Can
- 48:54you give me a voice 101 or 201? What
- 48:56what should I and other agent builders
- 48:57know about voice compared to just
- 48:59building, you know, simple chat agents?
- 49:02>> Voice has been maybe the most fun
- 49:04project that I've worked on in my whole
- 49:05career.
- 49:06So, and and I for context, I joined
- 49:09Sierra as an agent PM working on
- 49:12building agents specifically with
- 49:13customers in a forward deployed role.
- 49:16And one of the first customers I worked
- 49:17on
- 49:18is SiriusXM, the in-car streaming radio
- 49:21service. And so, I'm big SiriusXM fan
- 49:25before that and as a result. And they
- 49:29have a ton of volume over voice, even
- 49:32more than they have over chat. And so,
- 49:34many of their touchpoints with customers
- 49:36are over the phone. And so, early on it
- 49:37was very obvious that voice was going to
- 49:39be impactful for the business.
- 49:41And we got to think from first
- 49:42principles, basically from the ground
- 49:44up, what makes a voice experience great?
- 49:47How is that similar to chat? How is that
- 49:49similar How is that different from chat?
- 49:51And so, latency is probably the most
- 49:53obvious one.
- 49:55You need to be really thoughtful about
- 49:58parallelism. You need to be really
- 49:59thoughtful about what we call progress
- 50:01indicators, which is where you
- 50:03say, you know, hang on a second while I
- 50:05look up your account.
- 50:07That's number one. Number two is
- 50:08naturalism. This is a combination of a
- 50:10number of different things. So,
- 50:12oftentimes when something sounds a
- 50:14little bit robotic,
- 50:16I'll I'll read what the agent said, and
- 50:18I'm like, "Well, I sound robotic, too."
- 50:20So, it's a combination of what the agent
- 50:22is reading and then also the quality of
- 50:23the voice itself.
- 50:25Um there's multilingualism. It's very
- 50:27easy to speak different languages over
- 50:29chat using large language models. It's a
- 50:31lot harder to be fluent in I think it's
- 50:36almost about 60 languages on Sierra
- 50:38platform um and
- 50:41you know, than it is on chat. And each
- 50:42of those languages, you know, sometimes
- 50:44the very best transcription provider
- 50:47might have a 20% word error rate. I
- 50:49think that's true for a language like
- 50:51Hungarian, for example. And so, it's
- 50:52like, how can we ensemble multiple
- 50:54transcription providers in order to get
- 50:56that down and kind of be better than a
- 50:58single model is on its own. The other
- 51:01big factor is I think we all believe
- 51:03that a few years from now
- 51:06most voice agents will be running voice
- 51:08native models. So, you know, real time I
- 51:11think it they might be up to 2.5 at this
- 51:13point. They've had like three big
- 51:15real-time launches at OpenAI this year
- 51:17already. There was the really cool demo
- 51:19from Thinking Machines Labs as well. So,
- 51:22there's been a lot of increased momentum
- 51:24here.
- 51:25And as of a few months ago, we now have
- 51:28production agents live with the
- 51:30voice-to-voice models. Um and so,
- 51:32they're you know, it's you know, fully
- 51:34end-to-end doing that. You still need
- 51:35the transcript in order to like make API
- 51:37calls and that sort of thing, but the
- 51:39agent is responding you know,
- 51:41with audio as the input. The other big
- 51:44piece of it, I think the The Machines
- 51:45demo was a really good example.
- 51:48Up until now, we basically had like 50
- 51:51lines of Python. I think Silero is the
- 51:54most popular voice activity detection
- 51:56library
- 51:57deciding when to speak.
- 51:59And then a trillion parameters deciding
- 52:02what to say. And that balance feels very
- 52:04off to me. If you think about the
- 52:06conversation we're having right now,
- 52:09I'm actually using a lot of my brain
- 52:10power to decide when to speak
- 52:12uh in addition to decide what deciding
- 52:14what to say. And it's probably more like
- 52:1550/50.
- 52:17And so the one of the big unlocks for
- 52:19Sierra agents was deciding to think
- 52:23about not only how to parallelize a
- 52:24task, but how to parallelize thinking,
- 52:27listening, and talking.
- 52:29So that when I'm listening, I'm already
- 52:31thinking about what I might say next.
- 52:33When I'm talking, I'm listening for
- 52:35interruptions. And so that was a big
- 52:37unlock uh in terms of the product
- 52:38design. The other one I would say is
- 52:40just modularity, like I said earlier.
- 52:42Um no one is the best at everything in
- 52:44this space. And when there are, you
- 52:46know, 100-plus languages uh worldwide
- 52:49that, you know, really deliver
- 52:50meaningful results, when many of our
- 52:51customers are global brands, global
- 52:54companies, you need that flexibility to
- 52:57use one provider here and another
- 52:58provider there and to ensemble them
- 53:01together in a specific place as well.
- 53:03>> How much of that modularity and that
- 53:06parallelism and thinking about different
- 53:08things goes away when it's like a native
- 53:11voice-to-voice model?
- 53:13>> In one specific conversation,
- 53:16it goes away.
- 53:18But if you think about the businesses we
- 53:19serve, the voice-to-voice models today
- 53:22are just reaching a level of reliability
- 53:24where you would trust them for English.
- 53:27And so if you still want to support all
- 53:28the different languages, you need that
- 53:30modularity for the foreseeable future.
- 53:34Um the other thing is they're still
- 53:36almost an order of magnitude more
- 53:37expensive. They aren't quite as good at
- 53:40reasoning yet.
- 53:42And so the cases where they are live in
- 53:44production, they're not quite as
- 53:45reliable with tool calling and
- 53:46instruction following. The cases where
- 53:48they're live in production, it's cases
- 53:50where we know in advance that the
- 53:52journey is a little bit simpler.
- 53:54And where the naturalism matters even
- 53:57more than usual.
- 53:59And the procedure is not as complex as
- 54:01some other cases. And so it's still I
- 54:04would say a fraction of our market that
- 54:06we can use voice-to-voice models for.
- 54:09>> My perception also, and I I have never
- 54:11built a voice agent. So I know truly
- 54:13nothing here. But my perception here is
- 54:15for the voice-to-voice models, you you
- 54:17probably you have less control over what
- 54:18goes on inside of the loop basically of
- 54:21of tool calling and reasoning. Is that
- 54:23correct or are there pretty good
- 54:24controls for for what happens inside?
- 54:27>> You may not have built a voice model,
- 54:28but you're an expert in developer
- 54:30ergonomics. And I would say early on
- 54:33the APIs missed the mark on the
- 54:37ergonomics. And so they got the uh
- 54:40integration points wrong. And they were
- 54:42it was exactly what you said. It was
- 54:43hey, if you want our model
- 54:45you need our voice activity detection
- 54:47and you need the whole thing. There was
- 54:49still an underlying model that was
- 54:50available. So I'm, you know, dating
- 54:53myself in AI, but the GPT-4 audio model
- 54:56was extremely exciting. It did things
- 54:59that no model before it could do. I
- 55:00think maybe like people that are real AI
- 55:03OGs would say this about like GPT-2 or
- 55:06something. And so you could see that
- 55:08this was coming. And I think we all
- 55:10would have said 5 years from now this is
- 55:11where we're going to be.
- 55:13But the way that we wired that up in our
- 55:16system was basically
- 55:18using the entire Sierra pipeline.
- 55:21And then holding on to the input audio.
- 55:24And piping that in with all of the
- 55:27prompt context into the audio model to
- 55:30do the last mile. So we were basically
- 55:32still doing everything ourselves and
- 55:33using it for the last mile. I think
- 55:35you're right that over time there's more
- 55:37and that you can do with the audio model
- 55:39the same way there's more that you can
- 55:41do with the text models. The fallacy
- 55:44would be that okay, so then you don't
- 55:45need the harness or you don't need all
- 55:47of the orchestration and simulations and
- 55:49everything because
- 55:51you can make that choice. You can either
- 55:52do the same thing a little bit more
- 55:55easily or you can set your sights on new
- 55:58and more impressive things. Which I
- 56:00think not to get too philosophical, but
- 56:02that's kind of the direction of the
- 56:03industry in general. It's like are we
- 56:06all obsolete or are we going to find new
- 56:08things to do that raise our horizons
- 56:11even farther?
- 56:11>> If you had to guestimate a time, we're
- 56:14big into guestimating on the podcast
- 56:15apparently.
- 56:16>> Great.
- 56:16>> Um, when do you when do you think more
- 56:18than 50% of your either traffic or
- 56:21customers will be served by a
- 56:23voice-to-voice model as opposed to this
- 56:25this speech-to-text text-to-speech
- 56:27pipeline?
- 56:28>> I will be surprised if it happens in the
- 56:30next 18 months. I've been surprised
- 56:32before. I was surprised by Opus 4.5
- 56:36uh, late last year. Um, certainly
- 56:38surprised by ChatGPT. Like vividly
- 56:40remember the first time
- 56:42staying up till 3:00 a.m. just, you
- 56:44know,
- 56:45trying to jailbreak the prompt.
- 56:47>> [laughter]
- 56:47>> And so, you know, I I know you're a
- 56:50sports fan, too. So, if we're doing
- 56:51over/unders, it would be like 24 months
- 56:55and 1 day or something like that. You
- 56:56know, like or over/under 24 months would
- 56:59be probably my my personal guess.
- 57:02Demation.
- 57:03>> How, if at all, do you guys think about
- 57:06memory? Specifically long-term memory,
- 57:08specific It sounds like you've got users
- 57:10potentially interacting with multiple
- 57:12different agents that your your that a
- 57:14single brand can can be building. How do
- 57:16you think about the memory that's shared
- 57:18across them?
- 57:19>> Memory is very important to the
- 57:20platform.
- 57:21So, I mentioned the agent data platform
- 57:23earlier, which kind of brings together
- 57:26uh,
- 57:27machine learning data or, you know, big
- 57:29data as you might say about about
- 57:32customers and then marries that with in
- 57:35the moment context. That can only happen
- 57:38if you have a sense of identity and can
- 57:40also bring in
- 57:42memory from the past. So, in every
- 57:44Sierra conversation, there's the
- 57:47possibility of
- 57:49identifying the customer, saving
- 57:51memories either implicitly automatically
- 57:54or explicitly, and then extracting those
- 57:56memories at a future date for use in the
- 57:59agent. So, it's very much first-class
- 58:01primitive on the platform.
- 58:02I think you'll see that happen more and
- 58:04more over time as well. Just as these
- 58:07journeys get more complex, as we see
- 58:09more and more wins from the personal
- 58:11touch. We already have a number of cases
- 58:13where resolution rate has gone up
- 58:16meaningfully from memory, whether it's
- 58:18just greeting you by name, remembering
- 58:20what you called about last time, knowing
- 58:22that yesterday you were on the phone for
- 58:24an hour and it was really frustrating.
- 58:26And so,
- 58:27early on we had that memory through
- 58:30customer systems only, but we found just
- 58:33from customers asking over and over,
- 58:34"Hey, can you just have this first-class
- 58:36on the platform?" that it's helpful to
- 58:38have both. Seamless integrations with a
- 58:40CRM
- 58:41as well as on platform memory that
- 58:43really understands AI better than most
- 58:46CRM software does.
- 58:47>> How do you guys think about memory? I
- 58:49feel like you've got agents, multiple
- 58:51agents interacting with customers
- 58:53throughout various stages of their
- 58:56buying experience life cycle. So, I
- 58:58imagine memory must be important. How
- 59:00How do you guys think about it?
- 59:01>> So, memory is extremely important to the
- 59:02platform and since the agent data
- 59:05platform introduction, which we launched
- 59:07back in early November,
- 59:09it's been a first-class primitive on
- 59:12Sierra. So, if you call, for example, my
- 59:15wife lived in Hawaii for a year and so I
- 59:17was flying Hawaiian Airlines back and
- 59:19forth quite a bit and on a couple
- 59:21occasions, for anyone who's brought a
- 59:23dog to Hawaii, there's a lot of
- 59:25paperwork involved. I'm excited for the
- 59:26Sierra agent that can help with that.
- 59:29But, I would often add a pet in cabin,
- 59:31not that often, a couple times. And if I
- 59:33call back, you know, it's nice for them
- 59:35to remember why I'm calling, to know
- 59:37about me, to know I prefer aisle seats.
- 59:40I'm a big user of the in-flight
- 59:42internet. Hawaiian has Starlink back and
- 59:45forth from Hawaii.
- 59:46And so,
- 59:47these things just what we've seen in
- 59:49practice is that if you know who someone
- 59:51is, you greet them by name, you remember
- 59:53what's important to them, and you show
- 59:56empathy in the moment, it increases all
- 59:59of the metrics that are most important
- 1:00:00to businesses, from resolution rate to
- 1:00:02conversion rate, etc. And so, we've made
- 1:00:05memory first class on the Sierra
- 1:00:07platform, where during a conversation,
- 1:00:10implicitly or explicitly, you can
- 1:00:11basically store memories.
- 1:00:13And then, the agent, if the same person
- 1:00:16calls back, can extract those memories.
- 1:00:19The one thing to be aware of is you have
- 1:00:21to be really thoughtful about
- 1:00:23authentication, because often times, if
- 1:00:26someone calls over the phone,
- 1:00:28you don't necessarily know 100% from
- 1:00:31their phone number that it is this
- 1:00:32person. You know, some office networks
- 1:00:35all have the same phone number, maybe
- 1:00:36it's a family line, etc. And so, every
- 1:00:39business has to think about what the
- 1:00:41policy is for allowing the extraction of
- 1:00:43memories, and which memories are
- 1:00:45sensitive versus not so sensitive.
- 1:00:47Saying, "Hey Harrison, thanks for
- 1:00:49calling again." You know, that's
- 1:00:50probably fine. But, if it's like, "Hey
- 1:00:52Harrison, and are you calling about your
- 1:00:54social security number or the you know,
- 1:00:55that's like a definitely a different
- 1:00:56standard. Um and so, we try to be very
- 1:00:59thoughtful about that with our customers
- 1:01:00as well.
- 1:01:01>> When you say you can implicitly or
- 1:01:03explicitly save memories, what what
- 1:01:05exactly does that mean?
- 1:01:06>> So, there's kind of three layers of it.
- 1:01:08Number one is on a given conversation
- 1:01:11turn, you could say, "I want to save
- 1:01:13this to memory."
- 1:01:14Number two would be at the beginning of
- 1:01:16a conversation, you could say, "These
- 1:01:18are the things that are important to
- 1:01:20remember." You know, "Remember their
- 1:01:21birthday." That's always nice. Uh I
- 1:01:23remember Cold Stone Creamery growing up,
- 1:01:25they would give you a free scoop on your
- 1:01:27birthday. You know, it's a great
- 1:01:28opportunity for a brand loyalty.
- 1:01:30>> you could say, so would the brand say,
- 1:01:32would the customer say this in the
- 1:01:34system prompt, or would this be the end
- 1:01:36customer talking to the agent saying,
- 1:01:38"Hey, remember for future things that my
- 1:01:41birthday is on XYZ."
- 1:01:42>> So, the first one you said would be an
- 1:01:44example of journey building. An example
- 1:01:45of what I just said, and you would do it
- 1:01:47in journey building. You'd say, "I care
- 1:01:48about birthdays as an agent developer."
- 1:01:51Or an agent builder at any one of our
- 1:01:52customer companies. The second thing you
- 1:01:54said would be the third category of
- 1:01:56memory, which would be just sort of
- 1:01:58remember important things, and it would
- 1:02:00be an important thing if the customer
- 1:02:01said, "Hey, I want you to remember this
- 1:02:03when I call back in the future."
- 1:02:05Um so, whether you're deciding something
- 1:02:06in the moment this is important as an
- 1:02:08agent builder, I care about these
- 1:02:09things, or you know, let the agent
- 1:02:12decide. Uh those are kind of three ways
- 1:02:14to structure uh memory storage.
- 1:02:17>> And when you think about the structure
- 1:02:18of memory itself, do you guys think
- 1:02:20about it as a knowledge graph, a vector
- 1:02:22store, a file system, TBD?
- 1:02:24>> It's not super important. Um I guess I
- 1:02:28would say you want to optimize around
- 1:02:30retrieval.
- 1:02:32But, the reason why I said it's not
- 1:02:33super important is that typically your
- 1:02:35knowledge base is three orders of
- 1:02:36magnitude larger than the memories for
- 1:02:38an individual customer.
- 1:02:40And so,
- 1:02:42the retrieval and ranking problem is
- 1:02:43pretty simple, and I don't think it
- 1:02:45matters what structure you use, at least
- 1:02:47in our system today.
- 1:02:48>> I feel like memory is this really hot
- 1:02:51topic, and everyone loves to talk about
- 1:02:52it. And And there have been memory
- 1:02:54startups now for like 2 years, but but I
- 1:02:56don't see any of them
- 1:02:59being massive breakout successful. Why
- 1:03:02is that? Is Is memory not that important
- 1:03:04in the grand scheme of things? Is it so
- 1:03:06bespoke? Is it just too early on? Is it
- 1:03:08Is it really hard? Like why why isn't
- 1:03:10there a more established memory company
- 1:03:13or memory pattern?
- 1:03:14>> Do you have it turned on with Claude or
- 1:03:16ChatGPT?
- 1:03:16>> Not on purpose, although I think it is
- 1:03:18accidentally.
- 1:03:19>> And do you find it useful with your
- 1:03:21accidental turning it on?
- 1:03:22>> I don't really there.
- 1:03:23>> Okay. I ask because I would say that
- 1:03:25those are useful to me. I think what I
- 1:03:27said earlier about how when you're
- 1:03:29trusting us with memory, you're trusting
- 1:03:31us with authentication. That's part of
- 1:03:32it is that actually in order to pull off
- 1:03:35memory,
- 1:03:36you need to be trusted with something
- 1:03:39that has higher risk, you know, as well.
- 1:03:42And so, the reason I mentioned ChatGPT
- 1:03:45and Claude is those are products that
- 1:03:48you are already trusting. And so, I
- 1:03:50think they have more freedom than a B2B
- 1:03:53player would have where it's like, "Hey,
- 1:03:55if I want to buy memory from you,
- 1:03:57I also need to buy authentication or
- 1:04:00verification at least or identification
- 1:04:02at least from you. And I don't know
- 1:04:05exactly what the startups are in the
- 1:04:06space, but I would imagine like
- 1:04:08you're biting off more than you think
- 1:04:10when you sell memory."
- 1:04:12>> Talking about observability and evals
- 1:04:14for a bit. You guys have an interesting
- 1:04:16problem, I presume, where you have evals
- 1:04:18for your internal agents and for the
- 1:04:20maybe like general purpose agent SDK,
- 1:04:23but then I'm assuming your customers
- 1:04:24want to do evals themselves as well. Are
- 1:04:25those the same? Do you use the same
- 1:04:27tools for both? Or if they're different,
- 1:04:29why and how are they different?
- 1:04:31>> Typically not exactly the same. So,
- 1:04:34internally, um the Agent OS, you know,
- 1:04:37if you think about it as a series of
- 1:04:39tasks and some of those tasks might be
- 1:04:41very complex and some of those tasks
- 1:04:43might be more simple.
- 1:04:44The eval problem is more similar to the
- 1:04:48eval problem that any applied AI company
- 1:04:50has.
- 1:04:51When you think about our customers,
- 1:04:53I think the eval problem is much more
- 1:04:57complicated and involves things like
- 1:05:00what happens when there's background
- 1:05:01noise in voice and uh what if I have an
- 1:05:04adversarial user and I want to save
- 1:05:06these 20 personas and run all of my
- 1:05:08simulations against all 20 of the
- 1:05:10personas and make sure that it works.
- 1:05:13And so, you end up with just a more
- 1:05:15complicated topography because a
- 1:05:17conversation by nature is very
- 1:05:19complicated and can go in so many
- 1:05:21different directions. And so, we built a
- 1:05:23product specifically for our customers
- 1:05:25to eval agents called simulations. And
- 1:05:28it supports all of these different
- 1:05:29things. I think it is probably you can
- 1:05:32tell when someone's building an agent if
- 1:05:35they have good simulations, it's such a
- 1:05:37great unlock because you can make
- 1:05:39changes in a way that is constantly
- 1:05:42improving the agent and being sure that
- 1:05:44you're not regressing, especially as you
- 1:05:45get into big teams with complex agents
- 1:05:48that are doing so many things. I mean, I
- 1:05:49know you see this at LangChain as well.
- 1:05:51Like, having really good evals is such a
- 1:05:54great unlock. And so, we pride
- 1:05:56ourselves, in addition to, you know,
- 1:05:59government governance and collaboration
- 1:06:02and review and making sure that, you
- 1:06:04know, you have workspaces, so you can
- 1:06:06let Ghostwriter run free but still
- 1:06:08review it before you make any changes.
- 1:06:11We also have that simulation layer, so
- 1:06:14that every change you make is tested
- 1:06:16against all the assumptions of the
- 1:06:17platform across voice and chat and many
- 1:06:19languages and many personas in this
- 1:06:21high-dimensional space that you're going
- 1:06:23to experience in production.
- 1:06:24>> Going out from evals for just a second
- 1:06:26because you said something around
- 1:06:28continually improving the agent. I want
- 1:06:30to talk about continual learning. That
- 1:06:31also ties into memory, I guess, a little
- 1:06:33bit. Like, how do you how do you think
- 1:06:35about continual learning in general?
- 1:06:36Does the Sierra platform support it in a
- 1:06:39fully I'm assuming not like completely
- 1:06:41automated way, but like how how far
- 1:06:44along are you guys and and and what do
- 1:06:46you think the future in continual
- 1:06:47learning holds?
- 1:06:48>> Where we are today is
- 1:06:51you can automatically detect an issue
- 1:06:53with a monitor. Ghostwriter can
- 1:06:55automatically suggest a fix to an issue.
- 1:06:58And you can review that issue and push
- 1:07:01it to your agent. And so, you're still
- 1:07:04in the loop or people are still in the
- 1:07:05loop in all of the cases.
- 1:07:07But, it's
- 1:07:09as automated as it can be with still
- 1:07:11giving you authority over that.
- 1:07:13I think in the near future, you will
- 1:07:15start to see the first cases of Sierra
- 1:07:18agents improving themselves where they
- 1:07:20have a confidence level to the fix. For
- 1:07:22example, if there's an error in a
- 1:07:25knowledge article and it can tell that
- 1:07:26there's a contradiction and it can go
- 1:07:28check the website and, you know, for
- 1:07:30whatever reason, it's very clear what
- 1:07:32the true answer is, it could it could
- 1:07:34give you an FYI instead of needing
- 1:07:36approval. Same way, I do some work, I
- 1:07:39ask for approval, I do some other work,
- 1:07:40I give FYI. And so, all of the
- 1:07:42primitives are there, it's just around
- 1:07:45the confidence that people have in the
- 1:07:46level of control that they want to have.
- 1:07:49And so, we also don't want to get ahead
- 1:07:50of our skis there. Most of our
- 1:07:52customers, they want to review every
- 1:07:53change that goes into the agent. This is
- 1:07:55a really important part of their
- 1:07:56business. We don't want to pull the
- 1:07:58future forward too quickly. Um, we want
- 1:08:00to move at the pace our customers are
- 1:08:02excited about.
- 1:08:03>> One of the things you mentioned, going
- 1:08:04back to the eval's is monitors. What are
- 1:08:06monitors? And then you guys also wrote a
- 1:08:08blog called monitoring the monitors or
- 1:08:10something like that. I'd be curious to
- 1:08:11hear about that.
- 1:08:12>> Yeah, we have a saying in the company
- 1:08:14that the solution to all problems with
- 1:08:16AI is more AI. And so, often times, you
- 1:08:18have something that's 90% accurate and
- 1:08:20you figure out how to verify it 90% of
- 1:08:23the time. Figure out how to verify that
- 1:08:2590% of the time. And and so on and so on
- 1:08:28and you have something that's, you know,
- 1:08:29three or four nines of reliability. And
- 1:08:32I think with non-deterministic systems,
- 1:08:33that's just quite a bit about how it
- 1:08:35works. And so, similarly,
- 1:08:37uh, with a conversation platform, you
- 1:08:40can set up monitors that run on every
- 1:08:43conversation and look out for the things
- 1:08:45that you want to flag either for review
- 1:08:48or to create issues from, uh, etc. And
- 1:08:51it just basically gives you peace of
- 1:08:53mind, narrows the set of, "Hey, I don't
- 1:08:55have to wake up every morning and try to
- 1:08:57read 10,000 conversations. I can read
- 1:08:59five. And I can say, 'Okay, these five
- 1:09:01look good. I feel comfortable going on
- 1:09:03with my day."
- 1:09:04And so, that frees up a lot of our
- 1:09:06customers to think about how do I
- 1:09:09actually improve customer satisfaction
- 1:09:11or resolution rate or some of these more
- 1:09:13strategic levers
- 1:09:15as opposed to feeling like they need to
- 1:09:17review everything. So, I think that's
- 1:09:19why it's one of our more popular
- 1:09:20features.
- 1:09:21>> You guys released TaoBench, which is an
- 1:09:24eval for a few different agentic use
- 1:09:27cases. And I think you released a few
- 1:09:29other benches as well. Why do you guys
- 1:09:32invest in these and why should people
- 1:09:33check them out?
- 1:09:34>> So, I mentioned we have a research team
- 1:09:35and
- 1:09:37it's very exciting when you're building
- 1:09:40something to also think about how other
- 1:09:41people could use it. I mean, the
- 1:09:44distance that the AI space has come and
- 1:09:47how we've benefited just from all of the
- 1:09:50contributions to open source, you know,
- 1:09:52our knowledge engine as as I mentioned
- 1:09:54runs on open models that we fine-tuned.
- 1:09:58It has felt like one of the areas where
- 1:09:59we can contribute because we actually
- 1:10:02know a lot about what it takes to build
- 1:10:05a good voice agent. I don't think anyone
- 1:10:06knows more than we do. We know a lot
- 1:10:08about knowledge retrieval, we know a lot
- 1:10:10about tool calling and following
- 1:10:12process.
- 1:10:13Um and we know a lot about
- 1:10:15transcription. Um and so, we've released
- 1:10:17I think those are the four areas. There
- 1:10:19might be another one where we release
- 1:10:21benchmarks in the sort of Tao cinematic
- 1:10:23universe. There's TaoVoice,
- 1:10:25TaoKnowledge, TaoBench, and MuBench,
- 1:10:28which is the multilingual transcription
- 1:10:29benchmark.
- 1:10:30And so,
- 1:10:32it really just started because the first
- 1:10:34TaoBench was a lot more popular than we
- 1:10:35expected. We're like, "Oh, people trust
- 1:10:38us to kind of say what good looks like
- 1:10:40in this space." And so, we've continued
- 1:10:42to do more and more, and our research
- 1:10:44team has grown, and there's appetite. I
- 1:10:46think it also has this ancillary benefit
- 1:10:47of causing us to think about these
- 1:10:49problems
- 1:10:50in a very principled way. And you know,
- 1:10:52from kind of the first principles of
- 1:10:54what good looks like.
- 1:10:56And then we can evaluate our agents that
- 1:10:58way as well. So, I think it has that
- 1:10:59benefit, but it is very path dependent
- 1:11:01on Tow bench being a hit and you know,
- 1:11:04Tow is squared being the sequel being a
- 1:11:07hit as well and then us just deciding,
- 1:11:08okay, let's do more of this. People seem
- 1:11:10to like it.
- 1:11:11>> How much does the core agent team use
- 1:11:15these to guide their harness choices?
- 1:11:17>> Most of the benchmarks we use to
- 1:11:19evaluate providers
- 1:11:22more than to evaluate agents. And so,
- 1:11:25for example, we had there's a really
- 1:11:27exciting new transcription model that
- 1:11:29came by the office and presented it to
- 1:11:31us.
- 1:11:32And so, we were able to say, this looks
- 1:11:34really exciting, but we'd like you to
- 1:11:36run it against Mu bench and then it will
- 1:11:38be really exciting. And so, it really
- 1:11:40helps in the modular approach that we've
- 1:11:43taken. Like, the reason we discovered
- 1:11:45that this model works really well when
- 1:11:48there's silence in northern United
- 1:11:51Kingdom, but this other model works
- 1:11:53really well when there's speech is
- 1:11:54because of things like Mu bench in
- 1:11:56particular for that one. Internally,
- 1:11:59simulations is the main way that we eval
- 1:12:02the actual agents that are going out to
- 1:12:03production. So, it's just too customer
- 1:12:06specific for us to rely on something as
- 1:12:08general as a benchmark.
- 1:12:10>> How do you create these benchmarks? Are
- 1:12:11they synthetically generated? Do you do
- 1:12:13a lot of data labeling internally? Do
- 1:12:15you outsource it?
- 1:12:16>> I think it's a mix of all three.
- 1:12:19I don't know all of the details for all
- 1:12:21of the benchmarks, but I know that we
- 1:12:24do a lot of stuff internally just in
- 1:12:27terms of especially when you're kind of
- 1:12:30in the zero to one phase, just figuring
- 1:12:32out what the right shape of the data is.
- 1:12:34Even when you work with external
- 1:12:35companies, they often want to see some
- 1:12:37number of examples from you. And then I
- 1:12:39think also
- 1:12:41being able to synthesize data when scale
- 1:12:43matters a lot especially if you can do
- 1:12:46it in a reliable way, is very helpful,
- 1:12:48too.
- 1:12:49It's harder for something like
- 1:12:51transcription where audio synthesis
- 1:12:53might be, you know, already in the
- 1:12:55training set of the transcription and
- 1:12:56that kind of thing.
- 1:12:58Um, but for things like text, I think
- 1:12:59it's easier.
- 1:13:00>> One of the things that I think is pretty
- 1:13:01underrated in building agents is UX. So,
- 1:13:04we've already talked about voice as a
- 1:13:05modality. We've talked about actually
- 1:13:07showing up as a chat GPT app. How else
- 1:13:10do you guys think about modalities or
- 1:13:12UX's? Do you have you experimented with
- 1:13:14generative UI in any form?
- 1:13:17>> We have quite a bit. I think it's pretty
- 1:13:19vertical dependent as well.
- 1:13:22To give you an example, when you're
- 1:13:23checking in for a flight,
- 1:13:25if you have a hypothetically 12-letter
- 1:13:28last name with a hyphen in the middle of
- 1:13:30it, um, and a first name that's hard to
- 1:13:32spell as well,
- 1:13:33hypothetically, then it might be helpful
- 1:13:36to type that in while you're on the
- 1:13:37phone. And so, we see in industries like
- 1:13:40airlines appetite for multimodal
- 1:13:42experiences, especially when there's a
- 1:13:45lot of reservation retrieval or input.
- 1:13:48For something like retail, we see
- 1:13:50exactly what you described, where really
- 1:13:52polished UI around product discovery,
- 1:13:55um, and around recommendation moves the
- 1:13:57needle and makes a difference.
- 1:13:59I think where Sierra is particularly
- 1:14:01differentiated is going really deep with
- 1:14:03customers, especially in specific areas,
- 1:14:07and learning, you know, what does it
- 1:14:09mean to build an amazing
- 1:14:11retail discovery experience, and then
- 1:14:13just from first principles, what's the
- 1:14:15agent that could help drive that? versus
- 1:14:18what does it mean to build a great
- 1:14:20airline check-in experience or flight
- 1:14:22disruption experience, um, to the degree
- 1:14:25that can be great, it can be not
- 1:14:26terrible, I guess. Then, you know,
- 1:14:29what's the right form factor for that?
- 1:14:31We've seen, and I think one of the
- 1:14:32reasons vertical companies have been
- 1:14:35pretty successful lately is that
- 1:14:38understanding the contours of each
- 1:14:40industry and each company really makes a
- 1:14:42difference.
- 1:14:43>> One of the things that I think you guys
- 1:14:44are actually best known for is your
- 1:14:45revenue model, and charging for
- 1:14:47outcome-based pricing.
- 1:14:49How do you actually do that? How do you
- 1:14:51estimate the the value that an
- 1:14:53interaction has? And And is it specific
- 1:14:54to each customer?
- 1:14:56>> This, I think, is maybe the number one
- 1:15:00operational reason or business reason
- 1:15:02why ICR has been successful. It aligns
- 1:15:05the incentives between
- 1:15:07our company and our customers. And I
- 1:15:10think the phrase I like to use, which is
- 1:15:12a little bit cheeky, probably, is if you
- 1:15:15don't understand the value of
- 1:15:16outcome-based pricing, your outcomes are
- 1:15:19probably not that valuable. Because when
- 1:15:21you're delivering, you know, $100
- 1:15:23outcomes, and you get to keep a portion
- 1:15:26of it,
- 1:15:27everyone wants to row in the same
- 1:15:29direction, and it cuts through all of
- 1:15:31the prioritization and decision-making
- 1:15:34that often will cloud and resource
- 1:15:36allocation that often will will cloud
- 1:15:38enterprise partnerships. So, it's
- 1:15:39extremely valuable, and I think it's a
- 1:15:41big reason why we've been successful. I
- 1:15:42think it will just become the norm for
- 1:15:45companies that are doing differentiated
- 1:15:48high-value activities. If your product
- 1:15:50really like feels a little bit more like
- 1:15:52a commodity,
- 1:15:54you'll start to see more usage-based and
- 1:15:57seat-based pricing, cuz it's just
- 1:15:58simpler. An area, for example,
- 1:16:01knowledge-based lookups are a little bit
- 1:16:03more that way, just question answering.
- 1:16:06And so, in the case of question
- 1:16:08answering, that's not an area where you
- 1:16:10would have a high premium for an outcome
- 1:16:13of any particular sort. But if it's
- 1:16:16making a sale on a membership, or you
- 1:16:18know, selling someone a car,
- 1:16:20that's a really big outcome. Um, and so,
- 1:16:22companies will be more than happy to pay
- 1:16:24for that.
- 1:16:25Uh, I think where we're seeing things
- 1:16:27going is intra-conversation outcomes, to
- 1:16:31also thinking about more, you know, as I
- 1:16:33mentioned, kind of the moments that
- 1:16:34matter across the customer life cycle,
- 1:16:37and driving outcomes on top of our agent
- 1:16:39data platform that kind of span that
- 1:16:42whole life cycle. I think that's
- 1:16:44particularly interesting.
- 1:16:45>> You guys support multiple different
- 1:16:47outcomes. So, you've got customer
- 1:16:48support and you've got sales. How
- 1:16:50different is the pricing between those
- 1:16:52and how many different of these like
- 1:16:53categories or templates do you guys end
- 1:16:56up having?
- 1:16:56>> It really depends on the value. So, you
- 1:16:58asked if it was customer specific. The
- 1:17:01answer ends up being that it sort of has
- 1:17:02to be. In certain cases, you are
- 1:17:06troubleshooting very complex setup to a
- 1:17:10device or something and you have to try
- 1:17:1215 different things to get it to work
- 1:17:14and the average conversation might take
- 1:17:1520 turns and the amount of, you know,
- 1:17:18context engineering to make that work
- 1:17:20might be very high.
- 1:17:21In other cases, you might have something
- 1:17:23where, you know, you're just resetting
- 1:17:26the signal on your TV and it's very
- 1:17:28quick and easy or you're checking your
- 1:17:29balance with the bank and that's very
- 1:17:32easy. And so, you know, one outcome is
- 1:17:35very valuable and drives a lot of
- 1:17:37loyalty and one outcome is somewhat
- 1:17:39commoditized. You might have some cases
- 1:17:41where there's, you know, an outcome
- 1:17:44that's tens of dollars
- 1:17:47and in terms of the, you know, money
- 1:17:49that the agent would earn
- 1:17:51and then you might have some cases where
- 1:17:52it's, you know, much, much lower than
- 1:17:54that.
- 1:17:54>> And does that ever differ
- 1:17:57within a customer? So, like in your
- 1:17:59example, I could imagine you you could
- 1:18:01have an agent doing a really simple task
- 1:18:03of, oh, tell them to unplug the computer
- 1:18:05and plug it back in or something like
- 1:18:06that. And there's another one where
- 1:18:08like, oh my god, who knows what's going
- 1:18:09wrong and and it like is a miracle that
- 1:18:11it solves it at all. If it's the same
- 1:18:13customer, will it be charged the same
- 1:18:15amount or do you differentiate even
- 1:18:16within those different types of
- 1:18:18requests?
- 1:18:19>> There are cases where we differentiate.
- 1:18:21We're not dogmatic about it. What we
- 1:18:24found is that often times the benefits
- 1:18:27of having our incentives aligned
- 1:18:30are so high that it's not worth
- 1:18:33negotiating every detail of what counts
- 1:18:36for what and it kind of even out over
- 1:18:40time and you do right by your customers
- 1:18:42over time and you build trust and
- 1:18:44contracts aren't infinite and you want
- 1:18:46to have a really high renewal rate and
- 1:18:48have them trust you with more use cases
- 1:18:50and these kinds of things. So we make
- 1:18:52sure that incentives are deeply aligned.
- 1:18:55And then on top of that I think you can
- 1:18:58get really pedantic about the
- 1:18:59engineering of specific outcomes and
- 1:19:01maybe over time the market will move in
- 1:19:03that direction. But I think you're
- 1:19:05missing the forest for the trees in that
- 1:19:08case because of just how powerful the
- 1:19:10concept is. And so most of our customers
- 1:19:12are eager to find something simple that
- 1:19:14we all understand that feels fair. As
- 1:19:16opposed to trying to engineer like the
- 1:19:18perfect value for the for each outcome.
- 1:19:21>> Why don't you think there's more outcome
- 1:19:23based pricing right now? Is it because
- 1:19:24there's not enough agents doing valuable
- 1:19:26things or because it's so operationally
- 1:19:28intensive for now cuz it's early on that
- 1:19:31you guys have just a built up muscle of
- 1:19:32doing it and that's what allows you guys
- 1:19:34to do it so effectively.
- 1:19:35>> I think it's probably a bit of both. I
- 1:19:37think that there are a lot of products
- 1:19:40that
- 1:19:41probably as models have improved find
- 1:19:45themselves in a position of being more
- 1:19:47similar to
- 1:19:48what you could just buy tokens and
- 1:19:51create and then also there's just we're
- 1:19:53very early here.
- 1:19:55If I had to say though I would guess
- 1:19:58that the second one is more important
- 1:20:00and there will be a lot more of this the
- 1:20:03same way someone doesn't care how many
- 1:20:07hours I work as long as I produce you
- 1:20:10know new products that are good.
- 1:20:12And I think that that will become true
- 1:20:14of agents as well. There will be a mix
- 1:20:17of building agents in house on platforms
- 1:20:20like LangGraph and then there will be
- 1:20:22also
- 1:20:23you know buying
- 1:20:25products like Sierra to build agents on.
- 1:20:27>> Maybe switching to the last topic which
- 1:20:28is just the type of people that thrive
- 1:20:31at Sierra. I think you guys are also
- 1:20:33pretty famously known for your forward
- 1:20:35deployed engineering or agent builder
- 1:20:36approach. Could you talk a little bit
- 1:20:38about that both in terms of what those
- 1:20:40people do as well as the right persona
- 1:20:42to grow into that role?
- 1:20:44>> I joined Sierra about 2 and 1/2 years
- 1:20:46ago and it was my first B2B job ever.
- 1:20:48I'd only worked in consumer products and
- 1:20:51I love building consumer products. I
- 1:20:52love being like, oh, I could imagine,
- 1:20:54you know, my friends using this or my
- 1:20:56parents using this, but I'd never really
- 1:20:58loved growth
- 1:21:00uh and the idea of figuring out how to
- 1:21:02drive a couple percentage points of
- 1:21:06attention or a couple percentage points
- 1:21:08of usage. And what I learned when I
- 1:21:10joined Sierra is I love enterprise
- 1:21:12sales.
- 1:21:14Uh
- 1:21:15>> [laughter]
- 1:21:16>> I got a tattoo.
- 1:21:17Um so, basically, the the process of
- 1:21:20caring about each customer individually,
- 1:21:23saying one customer is upset, I'm going
- 1:21:25to call them right now and find out why
- 1:21:27and see how I can help. Just felt very
- 1:21:30empowering as a builder in a way where
- 1:21:33building for a billion users on Google
- 1:21:36Search, for example, you know, it was
- 1:21:38exciting in other ways, but it didn't
- 1:21:40feel like you could listen to each user
- 1:21:41and help them. And in many cases, we
- 1:21:44have customers of Sierra that have, you
- 1:21:45know, gotten promoted in their
- 1:21:47organizations. They're building careers
- 1:21:49because of the agents that they built on
- 1:21:51Sierra and so it's just feels very deep
- 1:21:54in terms of those relationships.
- 1:21:56What I love as well though is that the
- 1:21:57end user of a Sierra agent is still a
- 1:21:59consumer in the vast majority of cases.
- 1:22:01And I think it's pretty rare to have a
- 1:22:04product that needs to be consumer grade
- 1:22:07where the product that you're building,
- 1:22:09it's a it's a platform, but then the end
- 1:22:11user is really a consumer and you have
- 1:22:13to have them in your mind the whole
- 1:22:14time, but where you have kind of the
- 1:22:16enterprise sales process of building
- 1:22:19trust, of solving problems, of
- 1:22:21discovering value, and then delivering
- 1:22:23that value for people.
- 1:22:25Um and so, I think the people that
- 1:22:27really appreciate those two things, the
- 1:22:30customer obsession and the
- 1:22:32craftsmanship,
- 1:22:33uh, do very well.
- 1:22:35I think we've also discovered just with
- 1:22:37the rise of coding agents, certain
- 1:22:40things are more important than they used
- 1:22:41to be. Deep customer intuition, GPT-5.5
- 1:22:45doesn't really have that.
- 1:22:46Um, agency, the ability to say, "Why
- 1:22:49can't I do this?" Um, one of our uh,
- 1:22:52engineers that has really high degree of
- 1:22:54agency, her status message is just like,
- 1:22:56"Why not today?"
- 1:22:57Um, and so, yeah, having that mindset, I
- 1:23:00think is really important. And then the
- 1:23:02other thing just as someone with a
- 1:23:03product background is I think we kind of
- 1:23:06have
- 1:23:07a faster car than you need more pit
- 1:23:10stops, kind of thing. So, like a a
- 1:23:12Formula 1 car needs to get its tires
- 1:23:14changed more often than my Hyundai Kona.
- 1:23:17Uh, and the reason for that is, you
- 1:23:19know, it's driving faster, it's burning
- 1:23:21more rubber, uh, etc. And I think we
- 1:23:23have a similar thing building products
- 1:23:25as well now where coding agents have
- 1:23:27allowed us to write code a lot faster
- 1:23:30and even to review it faster now. But
- 1:23:32certain things like product judgment and
- 1:23:34customer intuition are therefore
- 1:23:36actually needed more often, not less
- 1:23:38often. And so, uh, people that can bring
- 1:23:41that to the table themselves are in this
- 1:23:44amazing loop of moving fast, but people
- 1:23:47where it's one person's job to bring
- 1:23:49that and another person's job to do
- 1:23:51engineering, they need even tighter
- 1:23:53collaboration and, you know, more daily
- 1:23:55stand-ups and that kind of thing to be
- 1:23:56successful.
- 1:23:57>> I really like that car analogy. I hadn't
- 1:23:59heard that before and totally resonates
- 1:24:00with what what what I'm seeing where
- 1:24:02product is becoming the bottleneck
- 1:24:04because it's so easy to code and you can
- 1:24:06make so much of things, but that doesn't
- 1:24:07mean you should. Who ends up fitting
- 1:24:10this agent builder profile the best? Is
- 1:24:12this product people then? Is this
- 1:24:13engineers with good product intuition?
- 1:24:16Like what does it look like practically?
- 1:24:17We're still figuring it out.
- 1:24:19>> I will say that people that have done
- 1:24:22both roles are often successful in the
- 1:24:24company. Our head of engineering, Arya,
- 1:24:26has been a product manager in the past.
- 1:24:28We have a number of engineers that have
- 1:24:29been product managers. I think those
- 1:24:32skills, knowing how to talk to
- 1:24:34customers, not just like what to say
- 1:24:36when you're in front of a customer, but
- 1:24:38how to find your way into the right
- 1:24:40conversations, having a high degree of
- 1:24:42agency, being really strong with
- 1:24:44communication, so that you're getting,
- 1:24:46you know, product isn't the bottleneck
- 1:24:48anymore. Uh those are really important
- 1:24:50skills.
- 1:24:51I still think kind of knowing the right
- 1:24:53questions to ask and the right things to
- 1:24:56tell coding agents is really important.
- 1:24:57So, the systems thinking and the
- 1:24:59architecture design are really
- 1:25:01important. And so, if you
- 1:25:03have not been an engineer before, uh it
- 1:25:05can be difficult. And so, I I think that
- 1:25:07the multi-disciplinary approach is more
- 1:25:10important than ever. My own personal
- 1:25:12rubric, which is like very much in beta,
- 1:25:15is kind of this customer intuition,
- 1:25:18agency, product judgment,
- 1:25:21technical depth,
- 1:25:23communication, intensity. Because when
- 1:25:26the car, you know, you need to be really
- 1:25:28locked in when you're driving a Formula
- 1:25:301 car. Um and then one which is a little
- 1:25:33harder to pin down, but it's just kind
- 1:25:34of leadership, where when there's more
- 1:25:36activity going on, the ability to to
- 1:25:39draw it into the correct direction is
- 1:25:41really important as well. So, this is
- 1:25:43kind of the working framework in my
- 1:25:45head, um but I'm sure there are lots of
- 1:25:47other things, too.
- 1:25:48>> How do you interview for agency? And I
- 1:25:50asked this because I think the the guest
- 1:25:52we had on in the previous episode said
- 1:25:54the exact same word agency for one of
- 1:25:56the traits that they look at. And I
- 1:25:58asked him the same question. So, now I'm
- 1:25:59going to ask you the same question. How
- 1:26:00do you How do you interview for agency?
- 1:26:02>> So, the most concrete way that we've
- 1:26:04changed our interviewing process is we
- 1:26:06have this AI native interview.
- 1:26:08>> And you wrote a great blog on it the
- 1:26:10other week.
- 1:26:10>> Yes. And so,
- 1:26:11you, by the way, is Vijay and Arya and
- 1:26:13our uh engineering leaders. But I've
- 1:26:16seen it done and participated in the
- 1:26:18interview panels. And basically it
- 1:26:21involves building a product end-to-end
- 1:26:24over the course of a few hours and then
- 1:26:26reviewing it with the team.
- 1:26:28I think in that environment you can see
- 1:26:31what people think is off-limits or
- 1:26:33what's their job and what's not their
- 1:26:34job and how far they extend sort of what
- 1:26:37they're allowed to do.
- 1:26:38And if they're able to
- 1:26:41find opportunities that you would have
- 1:26:43thought, oh, maybe that they would think
- 1:26:44that's out of scope, bring them into
- 1:26:46scope and build build great products on
- 1:26:47top of it. You kind of see agency. You
- 1:26:50see that they have a sense that a lot is
- 1:26:52in their control instead of feeling like
- 1:26:54certain things are not in their control.
- 1:26:56And if you think about coding agents,
- 1:26:58they bring so much more into I think the
- 1:27:02like the locus of control, right? And so
- 1:27:05you can do more things and if you
- 1:27:06appreciate that, I think it comes
- 1:27:09through in that AI-native interview.
- 1:27:11>> Thanks for listening to Max Agency.
- 1:27:13If you liked this episode, leave a
- 1:27:15review and subscribe. Send feedback or
- 1:27:17questions to [email protected].
- 1:27:18[music]
- 1:27:21We want to hear from you.
About this transcript
This page contains the full transcript of The best AI agents are simpler than you think by LangChain, generated from the public captions YouTube serves with the video. The transcript has 16,511 words across 2,545 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.