YouTube2Text

The best AI agents are simpler than you think — Transcript

by LangChain · 16,511 words · 2,545 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Agentic commerce will be bigger than
  2. 0:02e-commerce. There are cases where the
  3. 0:04Sierra agent is actually getting paid a
  4. 0:06commission on a sale.
  5. 0:08>> Today, I'm talking to Zack Reno Wedeen,
  6. 0:10head of product at Sierra, the platform
  7. 0:13powering customer experience agents for
  8. 0:15most of the Fortune 20.
  9. 0:16>> Coding agents are really good at file
  10. 0:17systems, they're really good at Git,
  11. 0:19they're really good at grep. Let's
  12. 0:21materialize everything into those
  13. 0:23structures so that coding agents can
  14. 0:25just, [music] you know, cook.
  15. 0:26>> He breaks down how Sierra builds for
  16. 0:28voice and why the architecture looks
  17. 0:30nothing like a standard agent harness.
  18. 0:32>> One of the big unlocks for Sierra agents
  19. 0:34was how to parallelize thinking,
  20. 0:36listening, and talking. And if this
  21. 0:38model says it's silent, you trust it. If
  22. 0:40this model does not say it's silent, you
  23. 0:42trust this one.
  24. 0:42>> We get into why a Sierra conversation is
  25. 0:45unlike a typical LLM call.
  26. 0:46>> You know, 10 or 15 different models
  27. 0:48might be invoked for a given
  28. 0:49conversation turn. So, sometimes you're
  29. 0:51classifying and you're responding at the
  30. 0:54same time.
  31. 0:54>> And Zack explains how and why Sierra
  32. 0:57built an entirely separate
  33. 0:58infrastructure layer for payments.
  34. 0:59>> We have isolated infrastructure where
  35. 1:02payment info doesn't go to an external
  36. 1:04large language model because none of the
  37. 1:05LLM providers are PCI certified in that
  38. 1:07way.
  39. 1:08>> Welcome to Max Agency, the podcast
  40. 1:10[music] that goes deep into how the best
  41. 1:11agents are being built by builders like
  42. 1:14you.
  43. 1:15>> [music]
  44. 1:17>> Most people are probably familiar with
  45. 1:18Sierra as a customer support platform.
  46. 1:21But, from what I understand, recently
  47. 1:22you guys are going broader than that.
  48. 1:24Could you talk a little bit about the
  49. 1:25types of agents that you help people
  50. 1:27build?
  51. 1:28>> Yeah, so this has been the vision since
  52. 1:30the beginning. I think because many
  53. 1:31companies have an RFP process where
  54. 1:34they're very specific about, "Hey, we
  55. 1:36want to solve customer service." That is
  56. 1:39often where we start, but we think of
  57. 1:41Sierra as the full engagement platform
  58. 1:44across all of the moments that matter
  59. 1:46for your customers. So, if you're an
  60. 1:48airline, that might be browsing for a
  61. 1:50flight, might be booking the flight,
  62. 1:52might be choosing your seat, might be in
  63. 1:54my case, I have a small dog, so adding a
  64. 1:57pet in cabin, then the flight might get
  65. 1:59rescheduled or delayed or
  66. 2:02cancelled, etc. etc. You need to get
  67. 2:04your bags there. There are just so many
  68. 2:06different things across that process.
  69. 2:08Some of them are sort of sales, some of
  70. 2:10them are more service, some of them are
  71. 2:12more loyalty, but they all kind of
  72. 2:14ladder up to the relationship between a
  73. 2:16business and its customers. And Sierra
  74. 2:19agents are present at all of these
  75. 2:21different parts of the customer life
  76. 2:22cycle. So, as an example,
  77. 2:25there are cases where the Sierra agent
  78. 2:27is actually, because of our
  79. 2:29outcome-based pricing model, getting
  80. 2:31paid a commission on a sale, which I
  81. 2:34think is quite different from how most
  82. 2:35people would imagine service. And so, we
  83. 2:37get really excited about those more
  84. 2:40exotic opportunities because they also
  85. 2:41give us an opportunity to push the
  86. 2:43platform forward and, you know, continue
  87. 2:45adding to what the platform can do, and
  88. 2:48kind of turn that into a product,
  89. 2:50package it up nicely, and bring more to
  90. 2:52all of our existing and future
  91. 2:54customers.
  92. 2:54>> How similar is the platform between
  93. 2:57these different use cases?
  94. 2:58>> I would say that
  95. 3:01it's very extensible, and so you can
  96. 3:03kind of take it in different directions.
  97. 3:06We like to say, I think it's originally
  98. 3:08attributed to the
  99. 3:09one of the creators of the programming
  100. 3:11language Pearl, but we try to make the
  101. 3:13easy things easy and the hard things
  102. 3:15possible. So, out of the box, pretty
  103. 3:17similar, the starting point, but you can
  104. 3:20kind of take it in any direction that
  105. 3:21you want. So, it's not that, oh, there's
  106. 3:23like a separate product for, you know,
  107. 3:25one company versus another, but the
  108. 3:28agents that you can build on it, and
  109. 3:30we'd like to think of these agents as
  110. 3:32products onto themselves, can be
  111. 3:34arbitrarily customized.
  112. 3:36>> What does it look like to build on the
  113. 3:38Sierra platform?
  114. 3:40>> So, we have a It's basically a web app.
  115. 3:42There's three main sections. There's
  116. 3:44analyze, build, and then there's
  117. 3:46release.
  118. 3:47Within the analyze section, you have
  119. 3:49things like our explorer agent, which is
  120. 3:53kind of the long-running ChatGPT deep
  121. 3:55research for all of your customer
  122. 3:57conversations and data. You have
  123. 4:00reports, you have monitors, which are
  124. 4:02kind of always-on evaluators of
  125. 4:04conversation data as well.
  126. 4:06And then within build, you have
  127. 4:08ghostwriter, which is the agent similar
  128. 4:10to Codex or Claude code for building
  129. 4:12agents.
  130. 4:14You also have journeys, kind of the
  131. 4:15underlying source code layer, although
  132. 4:17it's not really code. It's more like
  133. 4:19natural language or standard operating
  134. 4:21procedures. As well as kind of different
  135. 4:24variables
  136. 4:26and everything like that. On the release
  137. 4:28side, you have all of the collaboration
  138. 4:31and change management and governance
  139. 4:33procedures. And so Sierra, I think at
  140. 4:35this point we're working with most of
  141. 4:37the Fortune 20, something like 40 or 50%
  142. 4:41of the Fortune 50 or Fortune 100. So
  143. 4:43very much
  144. 4:45with a lot of the largest companies in
  145. 4:46the world. And they have needs around
  146. 4:49governance and release processes and
  147. 4:51change management that have just pushed
  148. 4:54us to develop from the very beginning,
  149. 4:56you know, very buttoned-up procedures
  150. 4:58and collaboration and review and all of
  151. 5:00this stuff. That's basically what it's
  152. 5:02like, I think, on the surface, but
  153. 5:04probably similar to a lot of other, you
  154. 5:06know, places that you go to build
  155. 5:08things, whether that's
  156. 5:09Figma or Claude code or these different
  157. 5:12places, but just very much optimized
  158. 5:14around no-code agent building and um
  159. 5:18giving you all those capabilities.
  160. 5:20>> And are those different steps intended
  161. 5:23to be done in that order? Like analyze,
  162. 5:24build, release. Can you analyze
  163. 5:26basically human transcripts before you
  164. 5:28build the AI agent? Or does analyze
  165. 5:30really come after you build and release
  166. 5:31the first version of the agent and now
  167. 5:32you're iterating on it?
  168. 5:34>> It's both. So, typically, you'll come in
  169. 5:38with some sort of resource of how you
  170. 5:40want the agent to be structured and
  171. 5:43architected, how you want it to behave.
  172. 5:45Maybe that's transcripts, maybe that's
  173. 5:47standard operating procedure, maybe
  174. 5:49that's a conversation that you have with
  175. 5:51Ghostwriter.
  176. 5:52And that will typically be how you build
  177. 5:54the agent. So, I'd say most people will
  178. 5:56start with build, but then once your
  179. 5:58agent is live and production
  180. 6:01conversations are happening, your daily
  181. 6:04routine probably starts more with
  182. 6:06analysis. You're probably thinking, how
  183. 6:08can I optimize the metric that I care
  184. 6:10about, whether that's customer
  185. 6:11satisfaction or resolution rate or in
  186. 6:15the case of the customer I mentioned,
  187. 6:16like sales converted.
  188. 6:18And so, you get those insights and then
  189. 6:21you want to make improvements to the
  190. 6:23agent, whether it's, you know, fixing an
  191. 6:25issue or finding a new opportunity to
  192. 6:27hill climb on a metric or please
  193. 6:30customers in one one more way.
  194. 6:32And so, that typically involves, you
  195. 6:34know, working with Ghostwriter. Often,
  196. 6:37Ghostwriter will actually proactively
  197. 6:38suggest an improvement on the insights
  198. 6:41to kind of close that loop and build
  199. 6:42that flywheel.
  200. 6:44Um, but I would say the day to day is
  201. 6:45more analyze, build, release.
  202. 6:47>> Who is doing that analyzing and that
  203. 6:50iterative improvement? Is this is this
  204. 6:52engineers? Is this product folks?
  205. 6:54>> It's primarily people that have the most
  206. 6:58depth and insight about the ideal
  207. 7:00customer experience, which tends to be
  208. 7:03operations uh people, so customer
  209. 7:05experience managers,
  210. 7:07um folks in that department at our
  211. 7:09customer companies.
  212. 7:11It's also a number of engineering teams
  213. 7:15will build either other agents that
  214. 7:17interface with Sierra agent or they can
  215. 7:20extend the platform
  216. 7:22uh via basically tools and packages that
  217. 7:25you can then kind of see and introspect
  218. 7:26on Sierra. So, it's very much kind of
  219. 7:29the same way that you have the person
  220. 7:31that knows everything about your
  221. 7:33knowledge base, we want them to be able
  222. 7:34to come in and self-serve on day one and
  223. 7:36just, you know, make the perfect
  224. 7:38instantiation of knowledge. The person
  225. 7:40who knows everything about the standard
  226. 7:42operating procedures should be able to
  227. 7:44just do that in the product. And so
  228. 7:45we're constantly kind of trying to sand
  229. 7:47down all of the barriers between the
  230. 7:50people with the most context and their
  231. 7:52ability to contribute directly to the
  232. 7:53platform.
  233. 7:54>> You've said no code a few times. So,
  234. 7:57what does this agent building experience
  235. 7:59look like? Is it truly no code? And And
  236. 8:01I'm assuming it's maybe something like
  237. 8:04quad code where you talk to it and it
  238. 8:06generates something under the hood. Is
  239. 8:07it generating code? Is it generating a
  240. 8:09custom DSL?
  241. 8:11>> Yeah, good question. So, the layers of
  242. 8:12the stack, you have kind of what we call
  243. 8:14Agent OS, which has our constellation of
  244. 8:17models. So, translating the tasks that
  245. 8:20need to be done on the platform
  246. 8:22into prompts, uh into data injection
  247. 8:26across, you know, 10 or 15 different
  248. 8:28models that might be invoked for a given
  249. 8:30conversation turn. Some of those might
  250. 8:32be frontier models that need to do, you
  251. 8:35know, top-tier reasoning. Some of them
  252. 8:37might be in-house models that are very
  253. 8:39good at a specific task, and some of
  254. 8:41them might be just, you know, classifier
  255. 8:43models that run really well on a
  256. 8:46uh a model that's a little bit cheaper
  257. 8:47and more performant. And so, that's kind
  258. 8:49of the base layer. On top of that, you
  259. 8:52have the Agent SDK, which is the
  260. 8:56code-based layer of agent orchestration
  261. 8:59and context management.
  262. 9:01That's kind of where Sierra started, but
  263. 9:03over the last 18 months, most of the
  264. 9:06agent development, um pretty much all of
  265. 9:08the agent development has shifted to our
  266. 9:10no code layer that we call journeys.
  267. 9:13It compiles down to Agent SDK code
  268. 9:16deterministically
  269. 9:17and uh isomorphically, which is a fancy
  270. 9:19word for you can turn it one way and
  271. 9:21then turn it back and it's the same. And
  272. 9:23so, you can have uh code that you
  273. 9:26transition over to no code, you can have
  274. 9:27no code that you transition over to
  275. 9:29code, but the language of specifying it
  276. 9:32is very much
  277. 9:34declarative. Here's how I want the agent
  278. 9:36behavior to be Uh when customers ask
  279. 9:39about this. We want to unlock these
  280. 9:41conditions and kind of flow in this
  281. 9:42direction and we find that that's pretty
  282. 9:44intuitive cuz it maps the type of
  283. 9:46document that we would write for someone
  284. 9:48joining the team in a customer
  285. 9:50experience role or sales role. You would
  286. 9:52explain to them how to do the job and
  287. 9:54that's kind of what you're doing here as
  288. 9:56well. But there is some DSL for
  289. 9:58journeys. It's not pure raw kind of like
  290. 10:01text.
  291. 10:01>> Correct.
  292. 10:03And it's very hard. I'm not sure we
  293. 10:05could get into a discussion about it. If
  294. 10:07you're just doing text, you have to
  295. 10:09choose between
  296. 10:11this is non-deterministically compiled,
  297. 10:14which all of the experiments we've done
  298. 10:16in that direction,
  299. 10:18you end up I think with more harm than
  300. 10:19good.
  301. 10:20Or this is a prompt engineering task,
  302. 10:23which then puts you in the realm of
  303. 10:25engineering teams. And we, you know, are
  304. 10:27very proud to be more in the realm of
  305. 10:30operations teams where a lot of that
  306. 10:32domain specific knowledge resides. The
  307. 10:34other big piece of it is that
  308. 10:35Ghostwriter has totally changed the
  309. 10:37learning curve for building agents. So
  310. 10:39you come in and you just say, "Hey, you
  311. 10:41know, I want to orchestrate order
  312. 10:42returns or I want to do flight booking
  313. 10:45or I want to do car rental or
  314. 10:47referral from primary care provider to a
  315. 10:49specialist." And Ghostwriter just kind
  316. 10:52of already knows those concepts and is
  317. 10:54an expert in journeys. But Ghostwriter
  318. 10:56is using the journeys product. So it's
  319. 10:58not writing code, it's writing journeys
  320. 11:01directly so that you can go inspect that
  321. 11:03after the fact as well.
  322. 11:05>> I imagine there's some there's some
  323. 11:07format that these journeys have to have
  324. 11:09to adhere to and I imagine that's not in
  325. 11:11the models training data at all. Was it
  326. 11:13hard to teach it that format or was it
  327. 11:15pretty easy?
  328. 11:16>> That's a really good question because at
  329. 11:18every point there's this conflict
  330. 11:20between here are the perfect
  331. 11:22abstractions for me
  332. 11:23and here are the abstractions that the
  333. 11:25models are most familiar with. And
  334. 11:26similar to in math how you're often
  335. 11:29taking one problem and reframing it in
  336. 11:31another problem to do a proof or
  337. 11:32something like that. You have to decide
  338. 11:34if you want to reframe this problem into
  339. 11:36something the models understand or build
  340. 11:38a skill and inject context in the right
  341. 11:40way so that the models can understand
  342. 11:42your way of thinking.
  343. 11:44The truth is that we do both. So there
  344. 11:47are cases where we'll say you know,
  345. 11:49coding agents are really good at file
  346. 11:50systems. They're really good at get.
  347. 11:52They're really good at grep. Let's
  348. 11:54materialize everything into those
  349. 11:56structures so that coding agents can
  350. 11:59just you know, cook. Then there are
  351. 12:01other cases where it's like, "No, no,
  352. 12:02no. Our way of thinking about this is
  353. 12:05the correct way of thinking about this
  354. 12:07and there's not really a way to shoehorn
  355. 12:09it into what it models are already good
  356. 12:11at. So let's do the investment to make
  357. 12:13the models good at this. My personal
  358. 12:15perspective is that 80% of the time you
  359. 12:18want to do the first thing
  360. 12:20and just meet the models of where they
  361. 12:22are on their turf and you should reserve
  362. 12:24the second one for that really special
  363. 12:27case. I'm curious uh if that's what been
  364. 12:29your experience as well.
  365. 12:30>> I think recently it's probably gotten to
  366. 12:33be we see a lot of people using the file
  367. 12:35system as an abstraction and so I think
  368. 12:37recently there's been a lot of talk
  369. 12:38especially as the labs talk about how
  370. 12:40they're RL-ing the models to be really
  371. 12:41good for their hardness to try to fit
  372. 12:43everything into a file system or this
  373. 12:45particular like edit file tool or things
  374. 12:48like that. I also think that the models
  375. 12:50are really good at writing certain
  376. 12:52packages. Like if you're in the training
  377. 12:53data, I think I think a lot of LangGraph
  378. 12:55is in the training data. So I think at
  379. 12:57least Anthropic models recommend
  380. 12:58LangGraph for a lot of use cases and
  381. 12:59that's great.
  382. 13:00But for newer things like deep agents,
  383. 13:02which is a new package we have, it's not
  384. 13:03in the training data at all. We spend a
  385. 13:05little bit of time, not maybe not as
  386. 13:06much as we should, but we spend a little
  387. 13:07bit of time thinking about what makes
  388. 13:09these models good at writing certain
  389. 13:10things. We really have no clue how to
  390. 13:12like affect know how to affect what goes
  391. 13:13in the training data, but it's a really
  392. 13:15interesting thing and so I think there's
  393. 13:16definitely been cases where we see that
  394. 13:19people choose technology because the
  395. 13:21models are really good at writing it.
  396. 13:23And so one question I was also going to
  397. 13:25ask for the agent SDK, I imagine that's
  398. 13:27you know, your own custom kind of like
  399. 13:28framework built in house. I don't know
  400. 13:30if you experimented with having It
  401. 13:32sounds like you didn't you you're not
  402. 13:34having Ghostwriter write that directly,
  403. 13:36but like that's obviously much more
  404. 13:37closer to code and so I was curious if
  405. 13:39you experimented with Ghostwriter or any
  406. 13:41model like writing agent SDK versus just
  407. 13:43writing like raw code. So, yes, um
  408. 13:48one of the things too is if you almost
  409. 13:50do the abstraction that the models are
  410. 13:52really good at, it can be overconfident
  411. 13:54or it can be familiar and successful.
  412. 13:56And so you have to be very thoughtful
  413. 13:58about going either all the way there or
  414. 14:00just not going there at all.
  415. 14:02>> Like you're saying agent SDK is
  416. 14:03somewhere in the middle and that might
  417. 14:04actually confuse it.
  418. 14:05>> Exactly. Exactly. Um
  419. 14:07we have kind of reinvented the agent SDK
  420. 14:10two or three times as models improve. So
  421. 14:13it used to be you had to have more
  422. 14:15deterministic guardrails in order to get
  423. 14:17the behavior that you want. Now there's
  424. 14:19more room for reasoning at each
  425. 14:22individual step and you can kind of push
  426. 14:23out the frontier of that reliability
  427. 14:26versus reasoning trade-off. So that's
  428. 14:29been very interesting. The reason for
  429. 14:31Ghostwriter primarily or entirely
  430. 14:34editing no code is just that that's
  431. 14:36where the vast majority of activity is
  432. 14:39on the platform today. So that's what
  433. 14:41our customers know and so making
  434. 14:43Ghostwriter good at it is really where
  435. 14:45all of the payoff is.
  436. 14:47I think if we tried to do it for code,
  437. 14:49it would be a a similarly scoped task,
  438. 14:51but it would be hard to get it to be
  439. 14:53really good at both at the same time.
  440. 14:54There's always going to be some
  441. 14:55trade-off.
  442. 14:56>> Do you still let users edit the agent
  443. 14:58SDK code if they want or is that now
  444. 15:01like completely abstracted away from
  445. 15:02them in terms of just pitch journeys?
  446. 15:05>> So the core agent SDK is part of the
  447. 15:07Sierra platform,
  448. 15:08but building agents in code is totally
  449. 15:11something you can do. One example is a
  450. 15:13number of our customers have CI/CD a
  451. 15:16continuous integration pipelines that
  452. 15:18they want to make sure their agent is
  453. 15:20released on. And so they need to get
  454. 15:23repository, which is where their agent
  455. 15:25lives. Another example is sometimes you
  456. 15:27have a particularly complex tool that
  457. 15:29has interacts with a streaming API or
  458. 15:31something in a way that is just easier
  459. 15:33to model in code than in no code. And
  460. 15:36so, the way that these work is because
  461. 15:38no code compiles down to code, you can
  462. 15:41kind of import or under the hood it will
  463. 15:43import code files and compiled no code
  464. 15:45files kind of all as though they're the
  465. 15:46same thing because they are.
  466. 15:48So, I think this is a benefit of
  467. 15:50starting out as a code-based platform is
  468. 15:52that we still support it. We have a
  469. 15:54number of customers that have dozens or
  470. 15:56in a few cases 100-plus developers
  471. 15:59building on the platform. And
  472. 16:01sometimes for that, you know, they work
  473. 16:03in Git, they release in Git. And so,
  474. 16:05being part of their enterprise change
  475. 16:07management protocol just means
  476. 16:09supporting Git.
  477. 16:10>> You mentioned that the agent SDK has
  478. 16:12changed over over the past few years as
  479. 16:14everything in the space has. Um,
  480. 16:16what does it look like now and how has
  481. 16:18it changed? What does that evolution
  482. 16:20look like?
  483. 16:21>> So, it started out we call this now
  484. 16:23flow-based, very much like, you know, do
  485. 16:26this. And it wasn't just like do these
  486. 16:28things. I think I'm a big fan of your
  487. 16:30not another workflow builder blog post.
  488. 16:32So, it wasn't that rigid, but it would
  489. 16:34be hey, you know, make sure you collect
  490. 16:36their email before you say that you're
  491. 16:40going to send them a confirmation email,
  492. 16:42right? Very clear to us, but you would
  493. 16:45want to do those things in that order.
  494. 16:47Now, I think if you think about just the
  495. 16:49way that agents can reason through tool
  496. 16:51calling,
  497. 16:52instead of having to specify that in the
  498. 16:54actual structure of your agent, you
  499. 16:56might just give it the context that in
  500. 16:59order to call this tool, you know, as a
  501. 17:01prerequisite you should have their
  502. 17:02email.
  503. 17:03And it will know how to ask for it. So,
  504. 17:05it's really like you'll just say that in
  505. 17:06the prompt. Yes.
  506. 17:08Or, you know, eventually all things end
  507. 17:10up in prompts, but it would be in the
  508. 17:12journey.
  509. 17:13And so, as you're kind of writing it
  510. 17:15out, you would specify that that's one
  511. 17:16of the rules or policies of the journey.
  512. 17:19And then the agent can take care of the
  513. 17:22rest. I think it's a mix of the models
  514. 17:24getting better and our orchestration
  515. 17:26platform becoming even more
  516. 17:28sophisticated and robust. So, when I
  517. 17:29talk about that constellation of models,
  518. 17:32there's a lot of We talked a little bit
  519. 17:34before about Linnaeus and Darwin.
  520. 17:36There's post-training that goes into
  521. 17:37that. There's model selection and eval
  522. 17:40and prompt engineering as well. And so,
  523. 17:43I think it's it's kind of equal parts
  524. 17:45model improvements and platform
  525. 17:47improvements.
  526. 17:47>> I want to talk about this model in a
  527. 17:48second, but I want to stay on the
  528. 17:49harness for a little bit. How similar
  529. 17:51does it look like in its current form to
  530. 17:53a coding agent harness? Does it have
  531. 17:55access to skills and sub-agents in the
  532. 17:57same way that someone using Claude code
  533. 17:59would would have?
  534. 18:00>> So, the one
  535. 18:02constraint we have that Claude code
  536. 18:05doesn't have is latency. Majority of
  537. 18:07Sierra conversations are voice. And if
  538. 18:10you're not responding in 1 or 2 seconds,
  539. 18:12then people wonder where you went. And
  540. 18:15so,
  541. 18:16we are highly optimized for these low
  542. 18:17latency use cases. There's a ton of
  543. 18:19parallelism.
  544. 18:21That being said, at a high level, it is
  545. 18:23using a lot of the same models. It has
  546. 18:26access to tools. Um so, there's a lot of
  547. 18:28similarities. There are also, you know,
  548. 18:31you can invoke other agents from the
  549. 18:33Sierra agent. So, the core What's best
  550. 18:34for the core conversation loop isn't
  551. 18:37typically what's best for software
  552. 18:39development. But you might want to say,
  553. 18:41"Hey, let me actually give you a
  554. 18:43callback in 20 minutes after I figure
  555. 18:45this out." And then you would have a
  556. 18:47type of loop that runs, you know, more
  557. 18:48like Claude code.
  558. 18:49>> And for those for those longer loops
  559. 18:51that might happen in the background, are
  560. 18:53those also built on the Sierra platform
  561. 18:55and are just a separate type of agent
  562. 18:57that remove the latency constraint?
  563. 18:58>> You can do it either way. So, you could
  564. 19:00have a Sierra agent calling out to
  565. 19:03another Sierra agent or you could also
  566. 19:05have a Sierra agent calling out to an
  567. 19:07in-house platform. And because so many
  568. 19:09of our customers have their own
  569. 19:11technology teams and a robust array of
  570. 19:15different AI projects internally.
  571. 19:17They might be experts in a particular
  572. 19:20area. Like they might have document
  573. 19:22generation handled themselves and that
  574. 19:24might be a long-running agent and then
  575. 19:25the Sierra agent can call out to it and
  576. 19:28wait for a response. So it's kind of up
  577. 19:29to you to choose and we find that
  578. 19:32enterprises are varied enough that they
  579. 19:34appreciate kind of having choice.
  580. 19:35>> When you do these agent-to-agent
  581. 19:37communications, are you are you using
  582. 19:40A2A or one of the protocols specifically
  583. 19:42for that or MCP or just an in-house, I
  584. 19:45don't know, REST API call?
  585. 19:47>> The most common is an API call. When you
  586. 19:49know who you're talking to in advance,
  587. 19:51often times you can save a lot of tokens
  588. 19:54and make sure that you're 100% accurate
  589. 19:56that way. That being said, Sierra agents
  590. 19:59support the MCP and agent-to-agent
  591. 20:02protocols. You can kind of install that
  592. 20:04integration and then your agent can be
  593. 20:06an MCP client. You can also set up your
  594. 20:08agent to be an MCP server. So this is
  595. 20:10how we support ChatGPT apps
  596. 20:13uh which rely on MCP servers. Basically,
  597. 20:16the tools of the agent can be made
  598. 20:18available to ChatGPT and then you can
  599. 20:20at-reference a Sierra agent. The example
  600. 20:23uh that would be most familiar is
  601. 20:24Redfin.
  602. 20:26Um if you do if you go to redfin.com and
  603. 20:28do their AI search, uh under the hood,
  604. 20:31it is a Sierra agent that is returning
  605. 20:33the home listings and having the
  606. 20:35conversation with you and that agent is
  607. 20:37also, I believe, available in ChatGPT.
  608. 20:40>> Interesting. I I didn't realize that
  609. 20:42Sierra agents could be ChatGPT apps. Is
  610. 20:45that the right terminology for them?
  611. 20:46>> It is. Yeah, exactly.
  612. 20:48>> Do you have Do you have an opinion or
  613. 20:50hot take on in the future, do you think
  614. 20:52people will be interacting with the
  615. 20:55agents that represent brands on
  616. 20:57dedicated chatbot websites or in uh
  617. 21:00ChatGPT or central chat engine?
  618. 21:03>> I think that agentic commerce will be
  619. 21:06bigger than e-commerce.
  620. 21:08So, if I think about how I get things
  621. 21:10done today, it used to be that I went to
  622. 21:12websites and clicked around. Now, I ask
  623. 21:16Codex or Claude to do things for me.
  624. 21:19And I don't see why I won't do that to
  625. 21:21manage my subscriptions, to order
  626. 21:23supplies to my home, to make dinner
  627. 21:25reservations. It just feels like that's
  628. 21:28where we're headed.
  629. 21:29And so, if that's happening, I think
  630. 21:31brands will want to be ready on the
  631. 21:33other side of that. So, we're very much
  632. 21:35planning for that world. We were
  633. 21:37investing in payments before it made
  634. 21:40sense, I think, and it's a long process,
  635. 21:43but a few months ago, we announced, you
  636. 21:46know, we're uh fully PCI DSS level one
  637. 21:49certified.
  638. 21:50>> no clue what that means. What does that
  639. 21:51mean?
  640. 21:52>> payment card industry.
  641. 21:54>> Okay.
  642. 21:54>> Um oh, man, you stump me on DSS. Uh
  643. 21:57[laughter]
  644. 21:58I have uh some of the other acronyms um
  645. 22:00in my head, but um
  646. 22:01>> We'll We'll put it in in post.
  647. 22:03>> Okay. Okay, thanks. Oh, man, it must be
  648. 22:05like digital
  649. 22:07I don't know. I know that the We were uh
  650. 22:09certified by a QSA, which is a qualified
  651. 22:11security assessor. And uh what that
  652. 22:14means is we're able to do the only voice
  653. 22:17payments platform, certainly at launch,
  654. 22:19I think still this is the case, where
  655. 22:21you don't need to transfer to another
  656. 22:23platform. So, it's a co- cohesive
  657. 22:25experience throughout checkout. And all
  658. 22:28the work that went into that, we have
  659. 22:29isolated infrastructure where the
  660. 22:32payment info doesn't go to a large
  661. 22:34language model, doesn't go to an
  662. 22:36external large language model, because
  663. 22:37we have none of the LLM providers are uh
  664. 22:39PCI certified in that way.
  665. 22:41And so, putting that all together is
  666. 22:43like spinning up a separate cluster, you
  667. 22:45know, getting certified, um making sure
  668. 22:47all of our operational rituals, you
  669. 22:49know, conform to what the security
  670. 22:51assessor is looking for.
  671. 22:53And we put in that work because we
  672. 22:55believe in this future where agentic
  673. 22:56commerce is actually bigger than
  674. 22:59e-commerce. And I think e-commerce is in
  675. 23:01the hundreds of billions of dollars at
  676. 23:03this point, couple percentage points of
  677. 23:04GDP or something like that just in the
  678. 23:06US. And so if you think about that
  679. 23:09space, um it's pretty big.
  680. 23:11>> And by agentic commerce, do you mean
  681. 23:13like chat GPT talking to uh Sierra agent
  682. 23:17that represents Redfin, or do you mean
  683. 23:19someone going to Redfin's agent no
  684. 23:21matter where it is and talking with it
  685. 23:23there?
  686. 23:24>> Both, both. I do think that the majority
  687. 23:27of this will be from personal agents.
  688. 23:30Just looking at user behavior, we spend
  689. 23:33so much time in Claude and chat GPT and
  690. 23:36Codex
  691. 23:38that you have to think that's where a
  692. 23:39lot of that behavior will accrue.
  693. 23:42>> And do you think those agents will
  694. 23:43interact with another agent? Why not
  695. 23:46just the raw APIs themselves?
  696. 23:48>> I do. I think that as you think about
  697. 23:51being ready for that world,
  698. 23:53um the same way that you might want to
  699. 23:57use Shopify or you might want to use
  700. 24:01certain software on your website to do
  701. 24:03product recommendations, to do checkout,
  702. 24:07uh you might want to use Stripe.
  703. 24:08Similarly, you'll want to use a platform
  704. 24:11that can make sure that you're
  705. 24:12presenting your products in the right
  706. 24:15way, that you're making checkout as easy
  707. 24:16as possible, and that you're showing up
  708. 24:19at your best, whether it's for a
  709. 24:21customer that's browsing or for an agent
  710. 24:23that's browsing. The one thing that I
  711. 24:24think is pretty different is the
  712. 24:26attention
  713. 24:28isn't necessarily valuable in the same
  714. 24:30way.
  715. 24:31Like our eyeballs are more valuable than
  716. 24:34an agent just, you know, spewing out
  717. 24:36tokens assuming that no one's ever going
  718. 24:37to look at it. At least I think that's
  719. 24:39true for now. At some point it lands up
  720. 24:41in some future training run and maybe
  721. 24:43has value, but I think that's de minimis
  722. 24:45relative to getting us to look at
  723. 24:47things. And so I do feel like maybe
  724. 24:49that's a bit different, but the
  725. 24:50presenting yourself in the right way,
  726. 24:52making it easy to check out, making it
  727. 24:54easy to understand what products are
  728. 24:55available, to express the preferences of
  729. 24:58whoever is responsible for that agent
  730. 25:01going off and doing something
  731. 25:02commercial. That all still feels
  732. 25:04relevant to me. I've seen some dev tools
  733. 25:06provider, I think Sentry, doing a
  734. 25:08similar thing where they have a bunch of
  735. 25:10APIs, obviously, for the underlying
  736. 25:11platform, but they also have an endpoint
  737. 25:14to just ask questions of the agent
  738. 25:15directly. And I think you could make a
  739. 25:17counterargument that like, great, brands
  740. 25:18should absolutely care about how the
  741. 25:20platform's being used and how it's being
  742. 25:22presented, but you could do that with
  743. 25:23skills or or some other mechanism to
  744. 25:25expose that to the agent. And I I
  745. 25:26honestly don't know which one's right,
  746. 25:28but it it has been interesting to see
  747. 25:30the whole space is so new, but
  748. 25:31increasingly so companies exposing
  749. 25:32agents as endpoints to interact with
  750. 25:34rather than the endpoints themselves.
  751. 25:36>> I agree. I think that all of this stuff
  752. 25:39you could try to do it yourself. It
  753. 25:40might be that certain companies, that's
  754. 25:42the best option.
  755. 25:44What we've seen is that because there's
  756. 25:47often tens or hundreds of millions of
  757. 25:49dollars on the line, in some cases
  758. 25:51billions of dollars on the line, you
  759. 25:53really want to make sure you're getting
  760. 25:54the best solution. And so if you're
  761. 25:56going to be 90% as good at it as you
  762. 25:59could be partnering with a company like
  763. 26:01Sierra,
  764. 26:02it still makes sense to partner and, you
  765. 26:05know, get that extra few billion
  766. 26:07dollars.
  767. 26:08>> One last question on this fun side
  768. 26:10tangent, payments. How early are we?
  769. 26:12>> I think we're really early.
  770. 26:13I personally still don't order paper
  771. 26:17towels with Codex. I don't know if you
  772. 26:20do.
  773. 26:20>> No, and that's why I asked. I'm glad
  774. 26:21that you said that cuz I know I'm not
  775. 26:23close to doing that. And so I was
  776. 26:24wondering how far behind I was.
  777. 26:26>> I mean, yeah, like
  778. 26:28I also didn't do it with Alexa. You
  779. 26:30know, I think for some of that really
  780. 26:31easy stuff, you could probably have done
  781. 26:33it already.
  782. 26:34The one that I think will definitely
  783. 26:36become a thing is there are a lot of
  784. 26:38apps that claim they can, you know, go
  785. 26:40through all of your subscriptions and
  786. 26:41cancel the ones that you're not using.
  787. 26:44That feels like as a consumer, that's a
  788. 26:45useful service. I'm definitely closer to
  789. 26:48doing that with Codex than I am
  790. 26:51with an app. You know, it would it would
  791. 26:53be so much work to tell it all the
  792. 26:55things and try to tell it which ones to
  793. 26:57cancel. It's a very manual process.
  794. 27:00If I gave Codex or Cloud Co-worker or
  795. 27:02something just access to my browser and
  796. 27:04said,
  797. 27:05"Hey, you know, go to all of the
  798. 27:06streaming apps and like the ones that
  799. 27:08I'm not logging into,
  800. 27:11just, you know, cancel those and let me
  801. 27:13know if you need my password." And
  802. 27:15obviously you have to figure out how to
  803. 27:16make that secure and everything. But I
  804. 27:18feel like that I would have demand for
  805. 27:20that product.
  806. 27:21>> Going back to the harness for a little
  807. 27:23bit, you said something earlier about
  808. 27:25things running in parallel. Is that like
  809. 27:26guardrails that you're running or
  810. 27:28retrieval steps or what's running in
  811. 27:30parallel in in this process?
  812. 27:32>> So many things. One example, knowledge.
  813. 27:35We will
  814. 27:37often look up answers before we know if
  815. 27:40we want them.
  816. 27:41So you'll, you know, before you decide
  817. 27:43whether this question needs an answer,
  818. 27:45you'll at least have the answer ready or
  819. 27:47in parallel with deciding. So sometimes
  820. 27:49you're classifying and you're responding
  821. 27:52at the same time. Basically speculative
  822. 27:53execution.
  823. 27:55Another example would be transcription,
  824. 27:57what we call ensembling.
  825. 27:59Um I think we might have published a
  826. 28:01blog post on this today, which is great.
  827. 28:03Go read it. We've learned so many things
  828. 28:06from being a
  829. 28:08uh modular architecture on voice. This
  830. 28:11was an early decision we made that I
  831. 28:13think has totally played out to our
  832. 28:15advantage where we have the ability for
  833. 28:18any language, for any customer, for any
  834. 28:21use case to multi-home providers
  835. 28:25uh across transcription, across
  836. 28:27synthesis, and across native uh
  837. 28:29voice-to-voice models. And so on the
  838. 28:31transcription side, for example, it just
  839. 28:33turns out uh when you have a thick UK
  840. 28:36accent from northern UK or at least
  841. 28:39parts of northern UK. I don't know
  842. 28:40exactly the region.
  843. 28:42There is one model that has the highest
  844. 28:44quality transcription,
  845. 28:46but it hallucinates during silence more
  846. 28:48than other models.
  847. 28:50So, we run two models in parallel.
  848. 28:52And if this model says it's silent, you
  849. 28:54trust it. If this model does not say
  850. 28:56it's silent, you trust this one. And so,
  851. 28:58that's just an example where we're
  852. 29:00running those in parallel. Uh we have
  853. 29:02logic for when you take the right one,
  854. 29:04and it's very specific. And if you had,
  855. 29:07you know, all your chips in with one
  856. 29:09provider or one system, uh or you
  857. 29:12weren't doing things in parallel, you
  858. 29:14would inevitably hit the limits of what
  859. 29:16that provider can do. So, the same way
  860. 29:18we use Claude and Gemini and the
  861. 29:20GPT-class models, we're also able to use
  862. 29:23all of the leading players on the
  863. 29:25transcription and synthesis and
  864. 29:27speech-to-speech side as well.
  865. 29:28>> You mentioned you have some in-house
  866. 29:29models as well. What do those models do,
  867. 29:32and why did you guys decide to build
  868. 29:34those in-house?
  869. 29:35>> So, uh knowledge is a great example. I
  870. 29:38think whenever we are pushing the limits
  871. 29:40of what's possible, we always consider
  872. 29:43whether we should build this in-house
  873. 29:45whenever it's limiting our ability to
  874. 29:47deliver more for our customers.
  875. 29:49So, an example where we're probably not
  876. 29:52the company is these, you know, many
  877. 29:54millions of dollar training runs that
  878. 29:56produce the GPT-5.5 class models. And,
  879. 30:00you know, Mythos, I'm sure is a many,
  880. 30:02many millions or tens or hundreds of
  881. 30:04millions of dollars training run um to
  882. 30:06get that produced. And that's stuff that
  883. 30:08OpenAI and Anthropic are just the best
  884. 30:10in the world at.
  885. 30:11I think what we're the best in the world
  886. 30:13at is going really deep with customers,
  887. 30:15understanding all of the process
  888. 30:17knowledge uh specific to their industry,
  889. 30:20specific to their company, specific to
  890. 30:22their customer base.
  891. 30:24And then having the products that can
  892. 30:26allow them to serve those customers as
  893. 30:28best as possible. And so, an example
  894. 30:30like knowledge where we were hitting the
  895. 30:32limits of the retrieval and reranking
  896. 30:35that we could do with out-of-the-box
  897. 30:36models, we asked the question of, you
  898. 30:39know, should we create our own models
  899. 30:40here
  900. 30:41um and eval them. And we have a research
  901. 30:43team that's pretty sizable and tightly
  902. 30:46integrated with our product teams. And
  903. 30:48so, we can flex that muscle when we need
  904. 30:50to, but we try not to be doing it just
  905. 30:52for the sake of doing it.
  906. 30:54>> You mentioned something like every run
  907. 30:55of the agent would have 10 to 15
  908. 30:57different model calls. If you had to
  909. 30:59guestimate like how many of those are
  910. 31:01frontier model calls versus like
  911. 31:02in-house fine-tune versus not frontier
  912. 31:05model but but third party?
  913. 31:07>> So, for a typical, um, turn of a
  914. 31:09conversation, I would guess that and
  915. 31:13this is just ballpark, but, you know,
  916. 31:15being, uh, precise rather than being
  917. 31:17accurate. Uh, I think maybe a couple
  918. 31:20frontier model, a handful of classifiers
  919. 31:24that probably don't require that, a
  920. 31:26handful of speculative execution in the
  921. 31:28case of voice in particular to make sure
  922. 31:30that it's low latency. Um, sometimes
  923. 31:33there will be an interim response that's
  924. 31:34generated to, you know, the same way
  925. 31:36you'd say, "Hold on a minute, I'm just
  926. 31:37pulling up your account." Like that kind
  927. 31:39of thing. Roughly like a third a third a
  928. 31:41third or a quarter quarter quarter, but
  929. 31:43I would say the frontier models just
  930. 31:45because they might be slower or more
  931. 31:49expensive would probably be, you know,
  932. 31:52more doing the bulk of the reasoning but
  933. 31:54in one or two inferences for a given
  934. 31:56conversation turn.
  935. 31:57>> Do you ever end up training models
  936. 31:59specific to a customer?
  937. 32:01>> It's not something that would be out of
  938. 32:02the question, but I can't think of a
  939. 32:04specific example. The reason I pause is
  940. 32:07because we do have our agent data
  941. 32:09platform
  942. 32:11and there are machine learning models
  943. 32:13that power strategies that are specific
  944. 32:16to customers. But in terms of like a,
  945. 32:18uh,
  946. 32:19you know, language model or generative
  947. 32:21model, we don't have cases of that.
  948. 32:23>> What's the agent data platform and what
  949. 32:25does it power?
  950. 32:26>> Basically, one thing we realized pretty
  951. 32:28early on is large language models are
  952. 32:30really good at in the moment empathy.
  953. 32:32Oftentimes better than we are of
  954. 32:33understanding, okay, you You I
  955. 32:35understand you're having a hard time.
  956. 32:36I'm really sorry about that. And it's
  957. 32:38the same way when you walk into
  958. 32:40a restaurant that has amazing service or
  959. 32:43a hotel that has amazing service, they
  960. 32:45recognize the moment you walk in, okay,
  961. 32:48this person just got off a really long
  962. 32:49flight or this person's 10 minutes late
  963. 32:52to their reservation and they were stuck
  964. 32:54in traffic and I'm just going to let
  965. 32:55them know that that is not a problem.
  966. 32:57Their table is ready.
  967. 32:59And large language models have that,
  968. 33:01especially on a platform like Sierra.
  969. 33:03But they don't necessarily know what you
  970. 33:05care about at a level deeper than that.
  971. 33:08And oftentimes the previous generation
  972. 33:10of AI or recommender systems have a
  973. 33:13better understanding of some of those
  974. 33:14things.
  975. 33:16And so what Agent Data Platform does is
  976. 33:18it can either integrate with your
  977. 33:20customer data platform, all your
  978. 33:21internal systems, or you know, you can
  979. 33:24have it just on Sierra or you can do
  980. 33:25sort of a zero copy integration.
  981. 33:28And it can take that structured data
  982. 33:31that knows what to recommend along with
  983. 33:33the here and now and now in the use
  984. 33:34those to
  985. 33:38generate uh better conversations, better
  986. 33:42uh orchestrations around how you want
  987. 33:44customers to feel and what you want to
  988. 33:46do for them.
  989. 33:47Um so that all sounds maybe a little bit
  990. 33:49abstract. Uh one example would be during
  991. 33:52sales. Oftentimes there's structured
  992. 33:54data that knows the right offer to
  993. 33:55present, but doing it just with that
  994. 33:58structured data with the previous
  995. 34:00generation of AI and before Sierra feels
  996. 34:02very stilted or it feels like you know,
  997. 34:05I don't know why you're doing this. And
  998. 34:07so large language models can really
  999. 34:09understand how to present an offer,
  1000. 34:12uh how to attribute it, weigh two
  1001. 34:14different offers based on conversation
  1002. 34:16context and pick the right one for the
  1003. 34:17moment and that kind of thing. So we see
  1004. 34:19it a lot with sales, with loyalty and
  1005. 34:21retention, those types of conversations.
  1006. 34:23>> One of the last agents we haven't talked
  1007. 34:25about too much is Explorer.
  1008. 34:27>> Yes.
  1009. 34:27>> What is Explorer? What does it look like
  1010. 34:29under the hood?
  1011. 34:30>> I've described it as chat GPT deep
  1012. 34:32research uh for all of your customer
  1013. 34:34context and conversations and all of the
  1014. 34:36data on Sierra. And so what it allows
  1015. 34:38you to do
  1016. 34:39is basically instead of having to go
  1017. 34:41spelunking for the specific insight in
  1018. 34:44reports or in monitors, you can just ask
  1019. 34:47the question. You can say, "Hey,
  1020. 34:49I noticed my resolution rate dipped, you
  1021. 34:51know, why was that?" Or "How can I
  1022. 34:53generate more sales?" Or
  1023. 34:55"I wish that more people were converting
  1024. 34:57from trial to, you know, full-time paid
  1025. 35:00plan. How come that's not happening?"
  1026. 35:03And then more than that, you can set up
  1027. 35:04automations so that, you know, on a
  1028. 35:06daily basis, for example, Explorer can
  1029. 35:09ask the same questions proactively.
  1030. 35:11And then partner with Ghostwriter. We
  1031. 35:14currently think of these as kind of two
  1032. 35:15separate agents, the analysis agent and
  1033. 35:18the authoring agent,
  1034. 35:20to say, "Oh, here are some fixes that
  1035. 35:22are suggested to improve." And you can
  1036. 35:23chat with Ghostwriter and kind of pick
  1037. 35:25it up from there. Under the hood where
  1038. 35:27this is converging, I think is a shared
  1039. 35:29harness that is an expert at using Agent
  1040. 35:31Studio,
  1041. 35:33Sierra's platform. Um and so that's kind
  1042. 35:35of what we've been setting up in terms
  1043. 35:37of what we talked about at the
  1044. 35:38beginning, like,
  1045. 35:39you know, figuring out the file system
  1046. 35:41architecture that maps to the product.
  1047. 35:44Um and so as we've exposed more and more
  1048. 35:47tools,
  1049. 35:48um you know, like building knowledge
  1050. 35:49bases to these agents, they get more and
  1051. 35:52more powerful and we see a lot of
  1052. 35:54emergent behavior between both.
  1053. 35:55>> Does this harness end up looking more
  1054. 35:58similar to like a coding agent harness
  1055. 36:00than the the harness that's part of
  1056. 36:02Agent OS or Agent SDK?
  1057. 36:04>> Yes. Um so this is less of a quick-turn
  1058. 36:07conversational agent and more of a
  1059. 36:09longer-term deep analysis agent. And so
  1060. 36:13it ends up looking a lot more like a
  1061. 36:15cloud coder or Codex.
  1062. 36:16>> One of the things I've been thinking
  1063. 36:17about, I'm curious if you have a take
  1064. 36:18here.
  1065. 36:19In a year, two years, three years, will
  1066. 36:22there be the split just in terms of
  1067. 36:24harness? Ones that's optimized for kind
  1068. 36:26of like, yeah, lower latency, external
  1069. 36:28facing, customer experience type things.
  1070. 36:31Voice is maybe heavily involved. And
  1071. 36:33another that's really focused on these
  1072. 36:34deep research, maybe coding, like you're
  1073. 36:36running in a sandbox, things like that.
  1074. 36:38Or will just end up converging into one
  1075. 36:39harness that, you know, depending on how
  1076. 36:41you prompt it or, you know, has these
  1077. 36:43async sub-agents in the background that
  1078. 36:45can maybe run for longer periods of
  1079. 36:46time.
  1080. 36:47>> I think there will always be latency,
  1081. 36:49performance, cost tradeoffs and
  1082. 36:52different architectures that emerge
  1083. 36:54because of that.
  1084. 36:55I've actually been surprised by how many
  1085. 36:58different types of model companies there
  1086. 37:00still are.
  1087. 37:01And when I talk to people who are
  1088. 37:03particularly AGI filled about it and I
  1089. 37:05say, "Hey, what's like a cool model
  1090. 37:08opportunity that's not flying like so
  1091. 37:10close to the sun that the labs will do
  1092. 37:12it?" They'll say, "Oh, well, you'll just
  1093. 37:14ask, you know, GPT-12 to make that
  1094. 37:16model, so it's actually not a big
  1095. 37:17opportunity." But I think in reality, at
  1096. 37:20least up until now, I'll probably look
  1097. 37:23dumb when AGI comes out.
  1098. 37:25>> [laughter]
  1099. 37:26>> You do see a lot of success in areas
  1100. 37:29like voice models from transcription to
  1101. 37:32synthesis. You don't always see
  1102. 37:34leadership from the model labs. You see
  1103. 37:36the model labs actually trying to focus
  1104. 37:38more on specific problems. For
  1105. 37:40Anthropic, I think it's coding. For
  1106. 37:42OpenAI, it's been consumer, now maybe
  1107. 37:44shifting a little bit more to
  1108. 37:45enterprise. Um for Google, definitely
  1109. 37:48consumer as well. And it really does
  1110. 37:50feel like there are still tradeoffs to
  1111. 37:52me. So, I expect there will still be
  1112. 37:54multiple architectures up until that
  1113. 37:57event horizon of AGI's here, so all bets
  1114. 37:59are off.
  1115. 38:00>> As you guys build your core agent
  1116. 38:02harnesses, and I'm assuming one to build
  1117. 38:04them in a model agnostic way, what do
  1118. 38:06you need to change to go from an OpenAI
  1119. 38:08model to an Anthropic model?
  1120. 38:10>> Usually, you have the evals that are
  1121. 38:12designed to work across both. So, if you
  1122. 38:13have really good evals and a really good
  1123. 38:15harness, a really good architecture,
  1124. 38:18then you should be able to kind of hill
  1125. 38:19climb toward eval performance without
  1126. 38:22too much effort. Often what will happen
  1127. 38:24is you'll learn the first time you're
  1128. 38:26switching one task from one to another
  1129. 38:28or making it possible to run on multiple
  1130. 38:30systems that your eval wasn't quite as
  1131. 38:33good as you thought. And so then you
  1132. 38:34make your eval better and you continue
  1133. 38:36to improve. But the short answer is that
  1134. 38:38it's pretty simple for a given
  1135. 38:42intelligence level of a model
  1136. 38:44to run a task on one or the other. And
  1137. 38:48so again it's that basically latency
  1138. 38:50quality cost trade-off, but not more
  1139. 38:52than that. And because we have customers
  1140. 38:54that have very specific requirements
  1141. 38:56around what clouds they can run in, what
  1142. 38:59models they can use, and our
  1143. 39:02company approach is to meet them, you
  1144. 39:04know, on their terms. You don't serve
  1145. 39:07most of the Fortune 20 without that
  1146. 39:08approach. It's not really a choice.
  1147. 39:11And so because that's our approach and
  1148. 39:13we've built a lot of products around
  1149. 39:14that, we also have made sure that we can
  1150. 39:16kind of move between models of
  1151. 39:18comparable intelligence without too much
  1152. 39:20heartburn.
  1153. 39:21>> And what do you end up changing when you
  1154. 39:22hill climb? Is it just the prompts? Do
  1155. 39:24you also change out some of the tools
  1156. 39:26themselves?
  1157. 39:27>> It depends on the case. And I might not
  1158. 39:29be the expert on the exact history of
  1159. 39:32each. I would I think that if you change
  1160. 39:34the tools,
  1161. 39:36it's pretty hard not to have downstream
  1162. 39:38effects of that. And there might be
  1163. 39:39certain tasks that can only run on
  1164. 39:41certain models
  1165. 39:43and other tasks that can run on other
  1166. 39:45models. And so there's always kind of a
  1167. 39:47set of eligible models for specific
  1168. 39:49tasks. I don't know exactly how tools
  1169. 39:52change, but I know the eval getting more
  1170. 39:54robust and the prompt, you know,
  1171. 39:56conforming to the quirks of each model
  1172. 39:58is definitely part of the development.
  1173. 40:01>> You guys recently wrote a blog around
  1174. 40:03context engineering and I think you said
  1175. 40:05it was the key to great agent building
  1176. 40:07or something like that. How do you guys
  1177. 40:09think about context engineering and and
  1178. 40:11and what, you know, tips or tricks would
  1179. 40:12you have for others?
  1180. 40:14>> I think it's showing agents everything
  1181. 40:16they need to do the right thing, but
  1182. 40:18nothing more.
  1183. 40:20And as models get smarter, you can be
  1184. 40:24a little bit less precise with
  1185. 40:26everything they need, and certainly less
  1186. 40:27precise with nothing more. So, early on
  1187. 40:30it was
  1188. 40:31the agent SDK was really about only
  1189. 40:33giving the model exactly what it needed
  1190. 40:35and kind of spoon-feeding the context.
  1191. 40:37Now, to extend the uh meal analogy, it's
  1192. 40:40probably more like, you know, putting
  1193. 40:42out the right dish. And maybe in the
  1194. 40:43future, it might be something that is
  1195. 40:46even less structured. One concept that I
  1196. 40:48think is in that blog post is kind of
  1197. 40:50progressive disclosure. You'll probably
  1198. 40:52know more about this than I do, but
  1199. 40:54when you bring something into the
  1200. 40:56prompt,
  1201. 40:58you don't want to do it before it's
  1202. 40:59relevant, and then you also risk
  1203. 41:02incoherence if you then yank it out of
  1204. 41:04the prompt. So, when you do things like
  1205. 41:06prompt compaction, you just want to be
  1206. 41:08really thoughtful about not making it
  1207. 41:10lossy, because if you keep something in
  1208. 41:12the history that is
  1209. 41:14incoherent with the rest of the system
  1210. 41:16prompt, it's not going to end well. And
  1211. 41:19so, I think to the degree, you know,
  1212. 41:21when we're fixing uh
  1213. 41:24issues, or when when we've seen
  1214. 41:25hallucinations, it's often because one
  1215. 41:28part of the prompt was this, and the
  1216. 41:29other part was this. And actually, one
  1217. 41:31of my main learnings from Sierra is
  1218. 41:34anytime you think the model's being
  1219. 41:36dumb, it's probably you.
  1220. 41:39>> I like that. I think that I I think a
  1221. 41:40lot of people have learned similar
  1222. 41:42lessons from doubting the model.
  1223. 41:43>> the model's too dumb, the model's
  1224. 41:45actually too smart.
  1225. 41:46>> How much you guys care about prompt
  1226. 41:48caching and maintaining that cache? I've
  1227. 41:50heard I've heard kind of like two
  1228. 41:52mindsets on it. One is like, yeah, do
  1229. 41:54everything you can to maintain the
  1230. 41:55cache, like don't invalidate it until
  1231. 41:57you like absolutely need to. And then
  1232. 41:59I've heard another theory that's
  1233. 42:00basically like, yeah, prompt caching's
  1234. 42:01great, but like what matters most is
  1235. 42:03like performance, and sometimes you need
  1236. 42:05to just like break the cache in order to
  1237. 42:07insert the right context, or give it a
  1238. 42:09system reminder, or something like that.
  1239. 42:11How how strictly do you guys try to
  1240. 42:13adhere to prompt caching?
  1241. 42:15>> Is the purpose for those who are prompt
  1242. 42:17caching loyalists for uh speed or cost
  1243. 42:22or quality?
  1244. 42:23>> I think the first two mostly speed and
  1245. 42:25cost.
  1246. 42:25>> Speed and cost.
  1247. 42:26>> Yeah. I haven't I haven't heard anyone
  1248. 42:28argue that it's for quality, but maybe
  1249. 42:30maybe it works better.
  1250. 42:32>> It's a nice to have.
  1251. 42:33Um we definitely don't want to
  1252. 42:36invalidate a cache for no good reason,
  1253. 42:38but quality comes first.
  1254. 42:40So, we aren't uh zealots about it at
  1255. 42:43all. I would also say that when the
  1256. 42:46outcomes that your agents are delivering
  1257. 42:49are very valuable, you have the luxury
  1258. 42:51of not being extremely focused on cost
  1259. 42:54in particular.
  1260. 42:56Uh and so, that probably is part of the
  1261. 42:59reason for that is that, you know,
  1262. 43:01a conversation with a customer could
  1263. 43:03sell a
  1264. 43:05$100 product or a $1,000 lifetime value
  1265. 43:09plan. And so, those are valuable enough
  1266. 43:12that quality almost always comes first.
  1267. 43:14>> We've talked a bunch about the agent
  1268. 43:16itself. There's two topics that we were
  1269. 43:17discussing earlier, and I'm curious
  1270. 43:19We've talked a little bit about them,
  1271. 43:20but I'm curious if you have any more
  1272. 43:21thoughts. First being RL. When is RL
  1273. 43:24good? When is it bad? How much have you
  1274. 43:25guys explored it?
  1275. 43:26>> We've explored it a lot, in part because
  1276. 43:29it has two great promises, you know,
  1277. 43:31increasing the ceiling of the quality of
  1278. 43:33models, and then also making it so that
  1279. 43:35you can do a similar task on more
  1280. 43:38models.
  1281. 43:39I'm curious for your take, but in
  1282. 43:40practice I've seen a little bit more of
  1283. 43:42the second one when it comes to like
  1284. 43:44enterprise RL. It's taking an open model
  1285. 43:47or open weights model and saying, "How
  1286. 43:48can we get similar performance to a
  1287. 43:51frontier model?" The two things that
  1288. 43:53make it hard are, number one, uh the way
  1289. 43:55that that gets delivered is
  1290. 43:57non-deterministic and might include um
  1291. 44:00you can't include any data that you
  1292. 44:02don't want the model to regurgitate. So,
  1293. 44:04we basically would never fine-tune a
  1294. 44:05model on something when it could lead to
  1295. 44:08regurgitation risk. That's just a
  1296. 44:10non-starter. And then also uh in
  1297. 44:12general, just the way that uh you would
  1298. 44:16train the model, you have to think about
  1299. 44:17preparing all of that data.
  1300. 44:19The other big one is that the frontier
  1301. 44:21models are improving so fast that you
  1302. 44:24want to remain as agile as possible.
  1303. 44:28And so in many cases, doing something
  1304. 44:30like RL makes a ton of sense for
  1305. 44:32something like knowledge where we feel
  1306. 44:34like we are pushing the state of the
  1307. 44:35art. But if we're not pushing the state
  1308. 44:38of the art, we really want to be
  1309. 44:39thinking about what's going to be true 3
  1310. 44:41months from now, 6 months from now, and
  1311. 44:43oftentimes RL is a rounding error
  1312. 44:45against that.
  1313. 44:46>> Yeah, I feel like to your point earlier,
  1314. 44:47we've started to hear it a little bit
  1315. 44:49more recently. I think because of cost.
  1316. 44:51So I think like most people are
  1317. 44:53interested in it when they're using
  1318. 44:54these frontier models and the
  1319. 44:55performance is good. But now, whether
  1320. 44:58it's coding or other things, their cost
  1321. 45:00is just going through the roof. I think
  1322. 45:02the and we're starting to investigate
  1323. 45:04this more, but I think the places where
  1324. 45:05we're hearing it most are basically in
  1325. 45:07those where performance is good, cost
  1326. 45:09too high. How can I bring it down? Let's
  1327. 45:10see if I can I can train a model to get
  1328. 45:12similar cost at a fraction of the cost.
  1329. 45:14Or similar performance, fraction of the
  1330. 45:15cost.
  1331. 45:16>> Yeah.
  1332. 45:17Interestingly enough, a lot of our
  1333. 45:19progress here has been driven by
  1334. 45:20capacity, not cost. Where
  1335. 45:24you know, we have a lot of customers
  1336. 45:25that are in the retail space. And when
  1337. 45:27we go into Black Friday, Cyber Monday,
  1338. 45:30for example, you need a lot of capacity
  1339. 45:32to deal with the spikes that they face.
  1340. 45:35Uh we've also done load tests that are
  1341. 45:37on the order of you know, if you were to
  1342. 45:39have that rate of conversation over a
  1343. 45:42year, it would be billions of
  1344. 45:44conversations. And so that level of
  1345. 45:46concurrency and those spikes just mean
  1346. 45:49that we need to be resilient to downtime
  1347. 45:53with a particular provider and ready
  1348. 45:56for, you know, using whoever has the
  1349. 45:58capacity to serve us. And so
  1350. 46:00it's funny because it's useful in so
  1351. 46:02many ways, but a lot of the reason why
  1352. 46:04we have such good support for multiple
  1353. 46:07providers is specifically preparing for
  1354. 46:10Black Friday Cyber Monday and running
  1355. 46:11load tests for really large customers.
  1356. 46:13>> One other
  1357. 46:15harness agent engineering topic,
  1358. 46:17multi-agent systems. Where do you think
  1359. 46:19they're useful, where they're not
  1360. 46:20useful?
  1361. 46:21>> I think they are often not as useful as
  1362. 46:24people think. My thoughts on this would
  1363. 46:26be people should be really thoughtful
  1364. 46:28about why they want a multi-agent
  1365. 46:30system. If you want a multi-agent system
  1366. 46:33so that one team can work on one agent
  1367. 46:35and one team can work on another agent,
  1368. 46:37then you're shipping your org chart. If
  1369. 46:38you want a multi-agent system because
  1370. 46:40it's just makes you more comfortable to
  1371. 46:42think about this problem over here and
  1372. 46:44this other problem over here,
  1373. 46:46then you're also not optimizing around
  1374. 46:49impact.
  1375. 46:50If, for example, you had an agent that
  1376. 46:52does triage and another agent that does
  1377. 46:54a task, by building it as a multi-agent
  1378. 46:57system,
  1379. 46:58you're often
  1380. 47:00depriving of the agent doing the task of
  1381. 47:02the information from the triage and
  1382. 47:05depriving the agent doing the triage of
  1383. 47:06all of the procedural information from
  1384. 47:09the task. And that's typically
  1385. 47:11destructive of value. And so, we are
  1386. 47:14often just want to make sure that we're
  1387. 47:16doing multi-agent systems for the right
  1388. 47:18reason. If you're kicking off even a
  1389. 47:20sub-agent, you want to make sure that it
  1390. 47:22has everything it needs to do that task
  1391. 47:25and that there's no reason why it
  1392. 47:26shouldn't just be part of the main
  1393. 47:28agent. Um and so, I think I've seen a
  1394. 47:30lot of cases where people are reaching
  1395. 47:34for multi-agent systems the same way you
  1396. 47:36might reach for microservices
  1397. 47:39uh before you're necessarily ready for
  1398. 47:41that level of optimization and also for
  1399. 47:44reasons that might not be just about
  1400. 47:46building the best possible agent. And
  1401. 47:48so, Sierra agents tend to be kind of one
  1402. 47:52agent representing the brand. You
  1403. 47:54certainly can have multiple agents and
  1404. 47:56build a multi-agent system, but be if
  1405. 47:59you're managing context correctly, if
  1406. 48:00you're doing really really good context
  1407. 48:02engineering, then typically it's just
  1408. 48:04not a problem because you're not
  1409. 48:05exposing the wrong context to the wrong
  1410. 48:08agent.
  1411. 48:08>> Is there a right time to build a
  1412. 48:10multi-agent system?
  1413. 48:11>> I think if you have truly separable
  1414. 48:13jobs, right? Where there's not any
  1415. 48:15purpose of the first context being part
  1416. 48:17of the second context.
  1417. 48:19I will say that
  1418. 48:21in my personal opinion with you know,
  1419. 48:24May 18th, 2026,
  1420. 48:27there not a lot of great times for it.
  1421. 48:29There might be times where it actually
  1422. 48:31the organizational difficulties are
  1423. 48:33worth the quality drop, but if you're
  1424. 48:36doing it specifically for quality, I
  1425. 48:38think it's it's
  1426. 48:39pretty rare that you can't just solve it
  1427. 48:41with better context engineering and I'm
  1428. 48:43kind of a monolith loyalist on that.
  1429. 48:45>> I feel like voice is one of the things
  1430. 48:48that is getting more and more popular,
  1431. 48:49but there still aren't a ton of people
  1432. 48:52doing a lot of, but you guys are. Can
  1433. 48:54you give me a voice 101 or 201? What
  1434. 48:56what should I and other agent builders
  1435. 48:57know about voice compared to just
  1436. 48:59building, you know, simple chat agents?
  1437. 49:02>> Voice has been maybe the most fun
  1438. 49:04project that I've worked on in my whole
  1439. 49:05career.
  1440. 49:06So, and and I for context, I joined
  1441. 49:09Sierra as an agent PM working on
  1442. 49:12building agents specifically with
  1443. 49:13customers in a forward deployed role.
  1444. 49:16And one of the first customers I worked
  1445. 49:17on
  1446. 49:18is SiriusXM, the in-car streaming radio
  1447. 49:21service. And so, I'm big SiriusXM fan
  1448. 49:25before that and as a result. And they
  1449. 49:29have a ton of volume over voice, even
  1450. 49:32more than they have over chat. And so,
  1451. 49:34many of their touchpoints with customers
  1452. 49:36are over the phone. And so, early on it
  1453. 49:37was very obvious that voice was going to
  1454. 49:39be impactful for the business.
  1455. 49:41And we got to think from first
  1456. 49:42principles, basically from the ground
  1457. 49:44up, what makes a voice experience great?
  1458. 49:47How is that similar to chat? How is that
  1459. 49:49similar How is that different from chat?
  1460. 49:51And so, latency is probably the most
  1461. 49:53obvious one.
  1462. 49:55You need to be really thoughtful about
  1463. 49:58parallelism. You need to be really
  1464. 49:59thoughtful about what we call progress
  1465. 50:01indicators, which is where you
  1466. 50:03say, you know, hang on a second while I
  1467. 50:05look up your account.
  1468. 50:07That's number one. Number two is
  1469. 50:08naturalism. This is a combination of a
  1470. 50:10number of different things. So,
  1471. 50:12oftentimes when something sounds a
  1472. 50:14little bit robotic,
  1473. 50:16I'll I'll read what the agent said, and
  1474. 50:18I'm like, "Well, I sound robotic, too."
  1475. 50:20So, it's a combination of what the agent
  1476. 50:22is reading and then also the quality of
  1477. 50:23the voice itself.
  1478. 50:25Um there's multilingualism. It's very
  1479. 50:27easy to speak different languages over
  1480. 50:29chat using large language models. It's a
  1481. 50:31lot harder to be fluent in I think it's
  1482. 50:36almost about 60 languages on Sierra
  1483. 50:38platform um and
  1484. 50:41you know, than it is on chat. And each
  1485. 50:42of those languages, you know, sometimes
  1486. 50:44the very best transcription provider
  1487. 50:47might have a 20% word error rate. I
  1488. 50:49think that's true for a language like
  1489. 50:51Hungarian, for example. And so, it's
  1490. 50:52like, how can we ensemble multiple
  1491. 50:54transcription providers in order to get
  1492. 50:56that down and kind of be better than a
  1493. 50:58single model is on its own. The other
  1494. 51:01big factor is I think we all believe
  1495. 51:03that a few years from now
  1496. 51:06most voice agents will be running voice
  1497. 51:08native models. So, you know, real time I
  1498. 51:11think it they might be up to 2.5 at this
  1499. 51:13point. They've had like three big
  1500. 51:15real-time launches at OpenAI this year
  1501. 51:17already. There was the really cool demo
  1502. 51:19from Thinking Machines Labs as well. So,
  1503. 51:22there's been a lot of increased momentum
  1504. 51:24here.
  1505. 51:25And as of a few months ago, we now have
  1506. 51:28production agents live with the
  1507. 51:30voice-to-voice models. Um and so,
  1508. 51:32they're you know, it's you know, fully
  1509. 51:34end-to-end doing that. You still need
  1510. 51:35the transcript in order to like make API
  1511. 51:37calls and that sort of thing, but the
  1512. 51:39agent is responding you know,
  1513. 51:41with audio as the input. The other big
  1514. 51:44piece of it, I think the The Machines
  1515. 51:45demo was a really good example.
  1516. 51:48Up until now, we basically had like 50
  1517. 51:51lines of Python. I think Silero is the
  1518. 51:54most popular voice activity detection
  1519. 51:56library
  1520. 51:57deciding when to speak.
  1521. 51:59And then a trillion parameters deciding
  1522. 52:02what to say. And that balance feels very
  1523. 52:04off to me. If you think about the
  1524. 52:06conversation we're having right now,
  1525. 52:09I'm actually using a lot of my brain
  1526. 52:10power to decide when to speak
  1527. 52:12uh in addition to decide what deciding
  1528. 52:14what to say. And it's probably more like
  1529. 52:1550/50.
  1530. 52:17And so the one of the big unlocks for
  1531. 52:19Sierra agents was deciding to think
  1532. 52:23about not only how to parallelize a
  1533. 52:24task, but how to parallelize thinking,
  1534. 52:27listening, and talking.
  1535. 52:29So that when I'm listening, I'm already
  1536. 52:31thinking about what I might say next.
  1537. 52:33When I'm talking, I'm listening for
  1538. 52:35interruptions. And so that was a big
  1539. 52:37unlock uh in terms of the product
  1540. 52:38design. The other one I would say is
  1541. 52:40just modularity, like I said earlier.
  1542. 52:42Um no one is the best at everything in
  1543. 52:44this space. And when there are, you
  1544. 52:46know, 100-plus languages uh worldwide
  1545. 52:49that, you know, really deliver
  1546. 52:50meaningful results, when many of our
  1547. 52:51customers are global brands, global
  1548. 52:54companies, you need that flexibility to
  1549. 52:57use one provider here and another
  1550. 52:58provider there and to ensemble them
  1551. 53:01together in a specific place as well.
  1552. 53:03>> How much of that modularity and that
  1553. 53:06parallelism and thinking about different
  1554. 53:08things goes away when it's like a native
  1555. 53:11voice-to-voice model?
  1556. 53:13>> In one specific conversation,
  1557. 53:16it goes away.
  1558. 53:18But if you think about the businesses we
  1559. 53:19serve, the voice-to-voice models today
  1560. 53:22are just reaching a level of reliability
  1561. 53:24where you would trust them for English.
  1562. 53:27And so if you still want to support all
  1563. 53:28the different languages, you need that
  1564. 53:30modularity for the foreseeable future.
  1565. 53:34Um the other thing is they're still
  1566. 53:36almost an order of magnitude more
  1567. 53:37expensive. They aren't quite as good at
  1568. 53:40reasoning yet.
  1569. 53:42And so the cases where they are live in
  1570. 53:44production, they're not quite as
  1571. 53:45reliable with tool calling and
  1572. 53:46instruction following. The cases where
  1573. 53:48they're live in production, it's cases
  1574. 53:50where we know in advance that the
  1575. 53:52journey is a little bit simpler.
  1576. 53:54And where the naturalism matters even
  1577. 53:57more than usual.
  1578. 53:59And the procedure is not as complex as
  1579. 54:01some other cases. And so it's still I
  1580. 54:04would say a fraction of our market that
  1581. 54:06we can use voice-to-voice models for.
  1582. 54:09>> My perception also, and I I have never
  1583. 54:11built a voice agent. So I know truly
  1584. 54:13nothing here. But my perception here is
  1585. 54:15for the voice-to-voice models, you you
  1586. 54:17probably you have less control over what
  1587. 54:18goes on inside of the loop basically of
  1588. 54:21of tool calling and reasoning. Is that
  1589. 54:23correct or are there pretty good
  1590. 54:24controls for for what happens inside?
  1591. 54:27>> You may not have built a voice model,
  1592. 54:28but you're an expert in developer
  1593. 54:30ergonomics. And I would say early on
  1594. 54:33the APIs missed the mark on the
  1595. 54:37ergonomics. And so they got the uh
  1596. 54:40integration points wrong. And they were
  1597. 54:42it was exactly what you said. It was
  1598. 54:43hey, if you want our model
  1599. 54:45you need our voice activity detection
  1600. 54:47and you need the whole thing. There was
  1601. 54:49still an underlying model that was
  1602. 54:50available. So I'm, you know, dating
  1603. 54:53myself in AI, but the GPT-4 audio model
  1604. 54:56was extremely exciting. It did things
  1605. 54:59that no model before it could do. I
  1606. 55:00think maybe like people that are real AI
  1607. 55:03OGs would say this about like GPT-2 or
  1608. 55:06something. And so you could see that
  1609. 55:08this was coming. And I think we all
  1610. 55:10would have said 5 years from now this is
  1611. 55:11where we're going to be.
  1612. 55:13But the way that we wired that up in our
  1613. 55:16system was basically
  1614. 55:18using the entire Sierra pipeline.
  1615. 55:21And then holding on to the input audio.
  1616. 55:24And piping that in with all of the
  1617. 55:27prompt context into the audio model to
  1618. 55:30do the last mile. So we were basically
  1619. 55:32still doing everything ourselves and
  1620. 55:33using it for the last mile. I think
  1621. 55:35you're right that over time there's more
  1622. 55:37and that you can do with the audio model
  1623. 55:39the same way there's more that you can
  1624. 55:41do with the text models. The fallacy
  1625. 55:44would be that okay, so then you don't
  1626. 55:45need the harness or you don't need all
  1627. 55:47of the orchestration and simulations and
  1628. 55:49everything because
  1629. 55:51you can make that choice. You can either
  1630. 55:52do the same thing a little bit more
  1631. 55:55easily or you can set your sights on new
  1632. 55:58and more impressive things. Which I
  1633. 56:00think not to get too philosophical, but
  1634. 56:02that's kind of the direction of the
  1635. 56:03industry in general. It's like are we
  1636. 56:06all obsolete or are we going to find new
  1637. 56:08things to do that raise our horizons
  1638. 56:11even farther?
  1639. 56:11>> If you had to guestimate a time, we're
  1640. 56:14big into guestimating on the podcast
  1641. 56:15apparently.
  1642. 56:16>> Great.
  1643. 56:16>> Um, when do you when do you think more
  1644. 56:18than 50% of your either traffic or
  1645. 56:21customers will be served by a
  1646. 56:23voice-to-voice model as opposed to this
  1647. 56:25this speech-to-text text-to-speech
  1648. 56:27pipeline?
  1649. 56:28>> I will be surprised if it happens in the
  1650. 56:30next 18 months. I've been surprised
  1651. 56:32before. I was surprised by Opus 4.5
  1652. 56:36uh, late last year. Um, certainly
  1653. 56:38surprised by ChatGPT. Like vividly
  1654. 56:40remember the first time
  1655. 56:42staying up till 3:00 a.m. just, you
  1656. 56:44know,
  1657. 56:45trying to jailbreak the prompt.
  1658. 56:47>> [laughter]
  1659. 56:47>> And so, you know, I I know you're a
  1660. 56:50sports fan, too. So, if we're doing
  1661. 56:51over/unders, it would be like 24 months
  1662. 56:55and 1 day or something like that. You
  1663. 56:56know, like or over/under 24 months would
  1664. 56:59be probably my my personal guess.
  1665. 57:02Demation.
  1666. 57:03>> How, if at all, do you guys think about
  1667. 57:06memory? Specifically long-term memory,
  1668. 57:08specific It sounds like you've got users
  1669. 57:10potentially interacting with multiple
  1670. 57:12different agents that your your that a
  1671. 57:14single brand can can be building. How do
  1672. 57:16you think about the memory that's shared
  1673. 57:18across them?
  1674. 57:19>> Memory is very important to the
  1675. 57:20platform.
  1676. 57:21So, I mentioned the agent data platform
  1677. 57:23earlier, which kind of brings together
  1678. 57:26uh,
  1679. 57:27machine learning data or, you know, big
  1680. 57:29data as you might say about about
  1681. 57:32customers and then marries that with in
  1682. 57:35the moment context. That can only happen
  1683. 57:38if you have a sense of identity and can
  1684. 57:40also bring in
  1685. 57:42memory from the past. So, in every
  1686. 57:44Sierra conversation, there's the
  1687. 57:47possibility of
  1688. 57:49identifying the customer, saving
  1689. 57:51memories either implicitly automatically
  1690. 57:54or explicitly, and then extracting those
  1691. 57:56memories at a future date for use in the
  1692. 57:59agent. So, it's very much first-class
  1693. 58:01primitive on the platform.
  1694. 58:02I think you'll see that happen more and
  1695. 58:04more over time as well. Just as these
  1696. 58:07journeys get more complex, as we see
  1697. 58:09more and more wins from the personal
  1698. 58:11touch. We already have a number of cases
  1699. 58:13where resolution rate has gone up
  1700. 58:16meaningfully from memory, whether it's
  1701. 58:18just greeting you by name, remembering
  1702. 58:20what you called about last time, knowing
  1703. 58:22that yesterday you were on the phone for
  1704. 58:24an hour and it was really frustrating.
  1705. 58:26And so,
  1706. 58:27early on we had that memory through
  1707. 58:30customer systems only, but we found just
  1708. 58:33from customers asking over and over,
  1709. 58:34"Hey, can you just have this first-class
  1710. 58:36on the platform?" that it's helpful to
  1711. 58:38have both. Seamless integrations with a
  1712. 58:40CRM
  1713. 58:41as well as on platform memory that
  1714. 58:43really understands AI better than most
  1715. 58:46CRM software does.
  1716. 58:47>> How do you guys think about memory? I
  1717. 58:49feel like you've got agents, multiple
  1718. 58:51agents interacting with customers
  1719. 58:53throughout various stages of their
  1720. 58:56buying experience life cycle. So, I
  1721. 58:58imagine memory must be important. How
  1722. 59:00How do you guys think about it?
  1723. 59:01>> So, memory is extremely important to the
  1724. 59:02platform and since the agent data
  1725. 59:05platform introduction, which we launched
  1726. 59:07back in early November,
  1727. 59:09it's been a first-class primitive on
  1728. 59:12Sierra. So, if you call, for example, my
  1729. 59:15wife lived in Hawaii for a year and so I
  1730. 59:17was flying Hawaiian Airlines back and
  1731. 59:19forth quite a bit and on a couple
  1732. 59:21occasions, for anyone who's brought a
  1733. 59:23dog to Hawaii, there's a lot of
  1734. 59:25paperwork involved. I'm excited for the
  1735. 59:26Sierra agent that can help with that.
  1736. 59:29But, I would often add a pet in cabin,
  1737. 59:31not that often, a couple times. And if I
  1738. 59:33call back, you know, it's nice for them
  1739. 59:35to remember why I'm calling, to know
  1740. 59:37about me, to know I prefer aisle seats.
  1741. 59:40I'm a big user of the in-flight
  1742. 59:42internet. Hawaiian has Starlink back and
  1743. 59:45forth from Hawaii.
  1744. 59:46And so,
  1745. 59:47these things just what we've seen in
  1746. 59:49practice is that if you know who someone
  1747. 59:51is, you greet them by name, you remember
  1748. 59:53what's important to them, and you show
  1749. 59:56empathy in the moment, it increases all
  1750. 59:59of the metrics that are most important
  1751. 1:00:00to businesses, from resolution rate to
  1752. 1:00:02conversion rate, etc. And so, we've made
  1753. 1:00:05memory first class on the Sierra
  1754. 1:00:07platform, where during a conversation,
  1755. 1:00:10implicitly or explicitly, you can
  1756. 1:00:11basically store memories.
  1757. 1:00:13And then, the agent, if the same person
  1758. 1:00:16calls back, can extract those memories.
  1759. 1:00:19The one thing to be aware of is you have
  1760. 1:00:21to be really thoughtful about
  1761. 1:00:23authentication, because often times, if
  1762. 1:00:26someone calls over the phone,
  1763. 1:00:28you don't necessarily know 100% from
  1764. 1:00:31their phone number that it is this
  1765. 1:00:32person. You know, some office networks
  1766. 1:00:35all have the same phone number, maybe
  1767. 1:00:36it's a family line, etc. And so, every
  1768. 1:00:39business has to think about what the
  1769. 1:00:41policy is for allowing the extraction of
  1770. 1:00:43memories, and which memories are
  1771. 1:00:45sensitive versus not so sensitive.
  1772. 1:00:47Saying, "Hey Harrison, thanks for
  1773. 1:00:49calling again." You know, that's
  1774. 1:00:50probably fine. But, if it's like, "Hey
  1775. 1:00:52Harrison, and are you calling about your
  1776. 1:00:54social security number or the you know,
  1777. 1:00:55that's like a definitely a different
  1778. 1:00:56standard. Um and so, we try to be very
  1779. 1:00:59thoughtful about that with our customers
  1780. 1:01:00as well.
  1781. 1:01:01>> When you say you can implicitly or
  1782. 1:01:03explicitly save memories, what what
  1783. 1:01:05exactly does that mean?
  1784. 1:01:06>> So, there's kind of three layers of it.
  1785. 1:01:08Number one is on a given conversation
  1786. 1:01:11turn, you could say, "I want to save
  1787. 1:01:13this to memory."
  1788. 1:01:14Number two would be at the beginning of
  1789. 1:01:16a conversation, you could say, "These
  1790. 1:01:18are the things that are important to
  1791. 1:01:20remember." You know, "Remember their
  1792. 1:01:21birthday." That's always nice. Uh I
  1793. 1:01:23remember Cold Stone Creamery growing up,
  1794. 1:01:25they would give you a free scoop on your
  1795. 1:01:27birthday. You know, it's a great
  1796. 1:01:28opportunity for a brand loyalty.
  1797. 1:01:30>> you could say, so would the brand say,
  1798. 1:01:32would the customer say this in the
  1799. 1:01:34system prompt, or would this be the end
  1800. 1:01:36customer talking to the agent saying,
  1801. 1:01:38"Hey, remember for future things that my
  1802. 1:01:41birthday is on XYZ."
  1803. 1:01:42>> So, the first one you said would be an
  1804. 1:01:44example of journey building. An example
  1805. 1:01:45of what I just said, and you would do it
  1806. 1:01:47in journey building. You'd say, "I care
  1807. 1:01:48about birthdays as an agent developer."
  1808. 1:01:51Or an agent builder at any one of our
  1809. 1:01:52customer companies. The second thing you
  1810. 1:01:54said would be the third category of
  1811. 1:01:56memory, which would be just sort of
  1812. 1:01:58remember important things, and it would
  1813. 1:02:00be an important thing if the customer
  1814. 1:02:01said, "Hey, I want you to remember this
  1815. 1:02:03when I call back in the future."
  1816. 1:02:05Um so, whether you're deciding something
  1817. 1:02:06in the moment this is important as an
  1818. 1:02:08agent builder, I care about these
  1819. 1:02:09things, or you know, let the agent
  1820. 1:02:12decide. Uh those are kind of three ways
  1821. 1:02:14to structure uh memory storage.
  1822. 1:02:17>> And when you think about the structure
  1823. 1:02:18of memory itself, do you guys think
  1824. 1:02:20about it as a knowledge graph, a vector
  1825. 1:02:22store, a file system, TBD?
  1826. 1:02:24>> It's not super important. Um I guess I
  1827. 1:02:28would say you want to optimize around
  1828. 1:02:30retrieval.
  1829. 1:02:32But, the reason why I said it's not
  1830. 1:02:33super important is that typically your
  1831. 1:02:35knowledge base is three orders of
  1832. 1:02:36magnitude larger than the memories for
  1833. 1:02:38an individual customer.
  1834. 1:02:40And so,
  1835. 1:02:42the retrieval and ranking problem is
  1836. 1:02:43pretty simple, and I don't think it
  1837. 1:02:45matters what structure you use, at least
  1838. 1:02:47in our system today.
  1839. 1:02:48>> I feel like memory is this really hot
  1840. 1:02:51topic, and everyone loves to talk about
  1841. 1:02:52it. And And there have been memory
  1842. 1:02:54startups now for like 2 years, but but I
  1843. 1:02:56don't see any of them
  1844. 1:02:59being massive breakout successful. Why
  1845. 1:03:02is that? Is Is memory not that important
  1846. 1:03:04in the grand scheme of things? Is it so
  1847. 1:03:06bespoke? Is it just too early on? Is it
  1848. 1:03:08Is it really hard? Like why why isn't
  1849. 1:03:10there a more established memory company
  1850. 1:03:13or memory pattern?
  1851. 1:03:14>> Do you have it turned on with Claude or
  1852. 1:03:16ChatGPT?
  1853. 1:03:16>> Not on purpose, although I think it is
  1854. 1:03:18accidentally.
  1855. 1:03:19>> And do you find it useful with your
  1856. 1:03:21accidental turning it on?
  1857. 1:03:22>> I don't really there.
  1858. 1:03:23>> Okay. I ask because I would say that
  1859. 1:03:25those are useful to me. I think what I
  1860. 1:03:27said earlier about how when you're
  1861. 1:03:29trusting us with memory, you're trusting
  1862. 1:03:31us with authentication. That's part of
  1863. 1:03:32it is that actually in order to pull off
  1864. 1:03:35memory,
  1865. 1:03:36you need to be trusted with something
  1866. 1:03:39that has higher risk, you know, as well.
  1867. 1:03:42And so, the reason I mentioned ChatGPT
  1868. 1:03:45and Claude is those are products that
  1869. 1:03:48you are already trusting. And so, I
  1870. 1:03:50think they have more freedom than a B2B
  1871. 1:03:53player would have where it's like, "Hey,
  1872. 1:03:55if I want to buy memory from you,
  1873. 1:03:57I also need to buy authentication or
  1874. 1:04:00verification at least or identification
  1875. 1:04:02at least from you. And I don't know
  1876. 1:04:05exactly what the startups are in the
  1877. 1:04:06space, but I would imagine like
  1878. 1:04:08you're biting off more than you think
  1879. 1:04:10when you sell memory."
  1880. 1:04:12>> Talking about observability and evals
  1881. 1:04:14for a bit. You guys have an interesting
  1882. 1:04:16problem, I presume, where you have evals
  1883. 1:04:18for your internal agents and for the
  1884. 1:04:20maybe like general purpose agent SDK,
  1885. 1:04:23but then I'm assuming your customers
  1886. 1:04:24want to do evals themselves as well. Are
  1887. 1:04:25those the same? Do you use the same
  1888. 1:04:27tools for both? Or if they're different,
  1889. 1:04:29why and how are they different?
  1890. 1:04:31>> Typically not exactly the same. So,
  1891. 1:04:34internally, um the Agent OS, you know,
  1892. 1:04:37if you think about it as a series of
  1893. 1:04:39tasks and some of those tasks might be
  1894. 1:04:41very complex and some of those tasks
  1895. 1:04:43might be more simple.
  1896. 1:04:44The eval problem is more similar to the
  1897. 1:04:48eval problem that any applied AI company
  1898. 1:04:50has.
  1899. 1:04:51When you think about our customers,
  1900. 1:04:53I think the eval problem is much more
  1901. 1:04:57complicated and involves things like
  1902. 1:05:00what happens when there's background
  1903. 1:05:01noise in voice and uh what if I have an
  1904. 1:05:04adversarial user and I want to save
  1905. 1:05:06these 20 personas and run all of my
  1906. 1:05:08simulations against all 20 of the
  1907. 1:05:10personas and make sure that it works.
  1908. 1:05:13And so, you end up with just a more
  1909. 1:05:15complicated topography because a
  1910. 1:05:17conversation by nature is very
  1911. 1:05:19complicated and can go in so many
  1912. 1:05:21different directions. And so, we built a
  1913. 1:05:23product specifically for our customers
  1914. 1:05:25to eval agents called simulations. And
  1915. 1:05:28it supports all of these different
  1916. 1:05:29things. I think it is probably you can
  1917. 1:05:32tell when someone's building an agent if
  1918. 1:05:35they have good simulations, it's such a
  1919. 1:05:37great unlock because you can make
  1920. 1:05:39changes in a way that is constantly
  1921. 1:05:42improving the agent and being sure that
  1922. 1:05:44you're not regressing, especially as you
  1923. 1:05:45get into big teams with complex agents
  1924. 1:05:48that are doing so many things. I mean, I
  1925. 1:05:49know you see this at LangChain as well.
  1926. 1:05:51Like, having really good evals is such a
  1927. 1:05:54great unlock. And so, we pride
  1928. 1:05:56ourselves, in addition to, you know,
  1929. 1:05:59government governance and collaboration
  1930. 1:06:02and review and making sure that, you
  1931. 1:06:04know, you have workspaces, so you can
  1932. 1:06:06let Ghostwriter run free but still
  1933. 1:06:08review it before you make any changes.
  1934. 1:06:11We also have that simulation layer, so
  1935. 1:06:14that every change you make is tested
  1936. 1:06:16against all the assumptions of the
  1937. 1:06:17platform across voice and chat and many
  1938. 1:06:19languages and many personas in this
  1939. 1:06:21high-dimensional space that you're going
  1940. 1:06:23to experience in production.
  1941. 1:06:24>> Going out from evals for just a second
  1942. 1:06:26because you said something around
  1943. 1:06:28continually improving the agent. I want
  1944. 1:06:30to talk about continual learning. That
  1945. 1:06:31also ties into memory, I guess, a little
  1946. 1:06:33bit. Like, how do you how do you think
  1947. 1:06:35about continual learning in general?
  1948. 1:06:36Does the Sierra platform support it in a
  1949. 1:06:39fully I'm assuming not like completely
  1950. 1:06:41automated way, but like how how far
  1951. 1:06:44along are you guys and and and what do
  1952. 1:06:46you think the future in continual
  1953. 1:06:47learning holds?
  1954. 1:06:48>> Where we are today is
  1955. 1:06:51you can automatically detect an issue
  1956. 1:06:53with a monitor. Ghostwriter can
  1957. 1:06:55automatically suggest a fix to an issue.
  1958. 1:06:58And you can review that issue and push
  1959. 1:07:01it to your agent. And so, you're still
  1960. 1:07:04in the loop or people are still in the
  1961. 1:07:05loop in all of the cases.
  1962. 1:07:07But, it's
  1963. 1:07:09as automated as it can be with still
  1964. 1:07:11giving you authority over that.
  1965. 1:07:13I think in the near future, you will
  1966. 1:07:15start to see the first cases of Sierra
  1967. 1:07:18agents improving themselves where they
  1968. 1:07:20have a confidence level to the fix. For
  1969. 1:07:22example, if there's an error in a
  1970. 1:07:25knowledge article and it can tell that
  1971. 1:07:26there's a contradiction and it can go
  1972. 1:07:28check the website and, you know, for
  1973. 1:07:30whatever reason, it's very clear what
  1974. 1:07:32the true answer is, it could it could
  1975. 1:07:34give you an FYI instead of needing
  1976. 1:07:36approval. Same way, I do some work, I
  1977. 1:07:39ask for approval, I do some other work,
  1978. 1:07:40I give FYI. And so, all of the
  1979. 1:07:42primitives are there, it's just around
  1980. 1:07:45the confidence that people have in the
  1981. 1:07:46level of control that they want to have.
  1982. 1:07:49And so, we also don't want to get ahead
  1983. 1:07:50of our skis there. Most of our
  1984. 1:07:52customers, they want to review every
  1985. 1:07:53change that goes into the agent. This is
  1986. 1:07:55a really important part of their
  1987. 1:07:56business. We don't want to pull the
  1988. 1:07:58future forward too quickly. Um, we want
  1989. 1:08:00to move at the pace our customers are
  1990. 1:08:02excited about.
  1991. 1:08:03>> One of the things you mentioned, going
  1992. 1:08:04back to the eval's is monitors. What are
  1993. 1:08:06monitors? And then you guys also wrote a
  1994. 1:08:08blog called monitoring the monitors or
  1995. 1:08:10something like that. I'd be curious to
  1996. 1:08:11hear about that.
  1997. 1:08:12>> Yeah, we have a saying in the company
  1998. 1:08:14that the solution to all problems with
  1999. 1:08:16AI is more AI. And so, often times, you
  2000. 1:08:18have something that's 90% accurate and
  2001. 1:08:20you figure out how to verify it 90% of
  2002. 1:08:23the time. Figure out how to verify that
  2003. 1:08:2590% of the time. And and so on and so on
  2004. 1:08:28and you have something that's, you know,
  2005. 1:08:29three or four nines of reliability. And
  2006. 1:08:32I think with non-deterministic systems,
  2007. 1:08:33that's just quite a bit about how it
  2008. 1:08:35works. And so, similarly,
  2009. 1:08:37uh, with a conversation platform, you
  2010. 1:08:40can set up monitors that run on every
  2011. 1:08:43conversation and look out for the things
  2012. 1:08:45that you want to flag either for review
  2013. 1:08:48or to create issues from, uh, etc. And
  2014. 1:08:51it just basically gives you peace of
  2015. 1:08:53mind, narrows the set of, "Hey, I don't
  2016. 1:08:55have to wake up every morning and try to
  2017. 1:08:57read 10,000 conversations. I can read
  2018. 1:08:59five. And I can say, 'Okay, these five
  2019. 1:09:01look good. I feel comfortable going on
  2020. 1:09:03with my day."
  2021. 1:09:04And so, that frees up a lot of our
  2022. 1:09:06customers to think about how do I
  2023. 1:09:09actually improve customer satisfaction
  2024. 1:09:11or resolution rate or some of these more
  2025. 1:09:13strategic levers
  2026. 1:09:15as opposed to feeling like they need to
  2027. 1:09:17review everything. So, I think that's
  2028. 1:09:19why it's one of our more popular
  2029. 1:09:20features.
  2030. 1:09:21>> You guys released TaoBench, which is an
  2031. 1:09:24eval for a few different agentic use
  2032. 1:09:27cases. And I think you released a few
  2033. 1:09:29other benches as well. Why do you guys
  2034. 1:09:32invest in these and why should people
  2035. 1:09:33check them out?
  2036. 1:09:34>> So, I mentioned we have a research team
  2037. 1:09:35and
  2038. 1:09:37it's very exciting when you're building
  2039. 1:09:40something to also think about how other
  2040. 1:09:41people could use it. I mean, the
  2041. 1:09:44distance that the AI space has come and
  2042. 1:09:47how we've benefited just from all of the
  2043. 1:09:50contributions to open source, you know,
  2044. 1:09:52our knowledge engine as as I mentioned
  2045. 1:09:54runs on open models that we fine-tuned.
  2046. 1:09:58It has felt like one of the areas where
  2047. 1:09:59we can contribute because we actually
  2048. 1:10:02know a lot about what it takes to build
  2049. 1:10:05a good voice agent. I don't think anyone
  2050. 1:10:06knows more than we do. We know a lot
  2051. 1:10:08about knowledge retrieval, we know a lot
  2052. 1:10:10about tool calling and following
  2053. 1:10:12process.
  2054. 1:10:13Um and we know a lot about
  2055. 1:10:15transcription. Um and so, we've released
  2056. 1:10:17I think those are the four areas. There
  2057. 1:10:19might be another one where we release
  2058. 1:10:21benchmarks in the sort of Tao cinematic
  2059. 1:10:23universe. There's TaoVoice,
  2060. 1:10:25TaoKnowledge, TaoBench, and MuBench,
  2061. 1:10:28which is the multilingual transcription
  2062. 1:10:29benchmark.
  2063. 1:10:30And so,
  2064. 1:10:32it really just started because the first
  2065. 1:10:34TaoBench was a lot more popular than we
  2066. 1:10:35expected. We're like, "Oh, people trust
  2067. 1:10:38us to kind of say what good looks like
  2068. 1:10:40in this space." And so, we've continued
  2069. 1:10:42to do more and more, and our research
  2070. 1:10:44team has grown, and there's appetite. I
  2071. 1:10:46think it also has this ancillary benefit
  2072. 1:10:47of causing us to think about these
  2073. 1:10:49problems
  2074. 1:10:50in a very principled way. And you know,
  2075. 1:10:52from kind of the first principles of
  2076. 1:10:54what good looks like.
  2077. 1:10:56And then we can evaluate our agents that
  2078. 1:10:58way as well. So, I think it has that
  2079. 1:10:59benefit, but it is very path dependent
  2080. 1:11:01on Tow bench being a hit and you know,
  2081. 1:11:04Tow is squared being the sequel being a
  2082. 1:11:07hit as well and then us just deciding,
  2083. 1:11:08okay, let's do more of this. People seem
  2084. 1:11:10to like it.
  2085. 1:11:11>> How much does the core agent team use
  2086. 1:11:15these to guide their harness choices?
  2087. 1:11:17>> Most of the benchmarks we use to
  2088. 1:11:19evaluate providers
  2089. 1:11:22more than to evaluate agents. And so,
  2090. 1:11:25for example, we had there's a really
  2091. 1:11:27exciting new transcription model that
  2092. 1:11:29came by the office and presented it to
  2093. 1:11:31us.
  2094. 1:11:32And so, we were able to say, this looks
  2095. 1:11:34really exciting, but we'd like you to
  2096. 1:11:36run it against Mu bench and then it will
  2097. 1:11:38be really exciting. And so, it really
  2098. 1:11:40helps in the modular approach that we've
  2099. 1:11:43taken. Like, the reason we discovered
  2100. 1:11:45that this model works really well when
  2101. 1:11:48there's silence in northern United
  2102. 1:11:51Kingdom, but this other model works
  2103. 1:11:53really well when there's speech is
  2104. 1:11:54because of things like Mu bench in
  2105. 1:11:56particular for that one. Internally,
  2106. 1:11:59simulations is the main way that we eval
  2107. 1:12:02the actual agents that are going out to
  2108. 1:12:03production. So, it's just too customer
  2109. 1:12:06specific for us to rely on something as
  2110. 1:12:08general as a benchmark.
  2111. 1:12:10>> How do you create these benchmarks? Are
  2112. 1:12:11they synthetically generated? Do you do
  2113. 1:12:13a lot of data labeling internally? Do
  2114. 1:12:15you outsource it?
  2115. 1:12:16>> I think it's a mix of all three.
  2116. 1:12:19I don't know all of the details for all
  2117. 1:12:21of the benchmarks, but I know that we
  2118. 1:12:24do a lot of stuff internally just in
  2119. 1:12:27terms of especially when you're kind of
  2120. 1:12:30in the zero to one phase, just figuring
  2121. 1:12:32out what the right shape of the data is.
  2122. 1:12:34Even when you work with external
  2123. 1:12:35companies, they often want to see some
  2124. 1:12:37number of examples from you. And then I
  2125. 1:12:39think also
  2126. 1:12:41being able to synthesize data when scale
  2127. 1:12:43matters a lot especially if you can do
  2128. 1:12:46it in a reliable way, is very helpful,
  2129. 1:12:48too.
  2130. 1:12:49It's harder for something like
  2131. 1:12:51transcription where audio synthesis
  2132. 1:12:53might be, you know, already in the
  2133. 1:12:55training set of the transcription and
  2134. 1:12:56that kind of thing.
  2135. 1:12:58Um, but for things like text, I think
  2136. 1:12:59it's easier.
  2137. 1:13:00>> One of the things that I think is pretty
  2138. 1:13:01underrated in building agents is UX. So,
  2139. 1:13:04we've already talked about voice as a
  2140. 1:13:05modality. We've talked about actually
  2141. 1:13:07showing up as a chat GPT app. How else
  2142. 1:13:10do you guys think about modalities or
  2143. 1:13:12UX's? Do you have you experimented with
  2144. 1:13:14generative UI in any form?
  2145. 1:13:17>> We have quite a bit. I think it's pretty
  2146. 1:13:19vertical dependent as well.
  2147. 1:13:22To give you an example, when you're
  2148. 1:13:23checking in for a flight,
  2149. 1:13:25if you have a hypothetically 12-letter
  2150. 1:13:28last name with a hyphen in the middle of
  2151. 1:13:30it, um, and a first name that's hard to
  2152. 1:13:32spell as well,
  2153. 1:13:33hypothetically, then it might be helpful
  2154. 1:13:36to type that in while you're on the
  2155. 1:13:37phone. And so, we see in industries like
  2156. 1:13:40airlines appetite for multimodal
  2157. 1:13:42experiences, especially when there's a
  2158. 1:13:45lot of reservation retrieval or input.
  2159. 1:13:48For something like retail, we see
  2160. 1:13:50exactly what you described, where really
  2161. 1:13:52polished UI around product discovery,
  2162. 1:13:55um, and around recommendation moves the
  2163. 1:13:57needle and makes a difference.
  2164. 1:13:59I think where Sierra is particularly
  2165. 1:14:01differentiated is going really deep with
  2166. 1:14:03customers, especially in specific areas,
  2167. 1:14:07and learning, you know, what does it
  2168. 1:14:09mean to build an amazing
  2169. 1:14:11retail discovery experience, and then
  2170. 1:14:13just from first principles, what's the
  2171. 1:14:15agent that could help drive that? versus
  2172. 1:14:18what does it mean to build a great
  2173. 1:14:20airline check-in experience or flight
  2174. 1:14:22disruption experience, um, to the degree
  2175. 1:14:25that can be great, it can be not
  2176. 1:14:26terrible, I guess. Then, you know,
  2177. 1:14:29what's the right form factor for that?
  2178. 1:14:31We've seen, and I think one of the
  2179. 1:14:32reasons vertical companies have been
  2180. 1:14:35pretty successful lately is that
  2181. 1:14:38understanding the contours of each
  2182. 1:14:40industry and each company really makes a
  2183. 1:14:42difference.
  2184. 1:14:43>> One of the things that I think you guys
  2185. 1:14:44are actually best known for is your
  2186. 1:14:45revenue model, and charging for
  2187. 1:14:47outcome-based pricing.
  2188. 1:14:49How do you actually do that? How do you
  2189. 1:14:51estimate the the value that an
  2190. 1:14:53interaction has? And And is it specific
  2191. 1:14:54to each customer?
  2192. 1:14:56>> This, I think, is maybe the number one
  2193. 1:15:00operational reason or business reason
  2194. 1:15:02why ICR has been successful. It aligns
  2195. 1:15:05the incentives between
  2196. 1:15:07our company and our customers. And I
  2197. 1:15:10think the phrase I like to use, which is
  2198. 1:15:12a little bit cheeky, probably, is if you
  2199. 1:15:15don't understand the value of
  2200. 1:15:16outcome-based pricing, your outcomes are
  2201. 1:15:19probably not that valuable. Because when
  2202. 1:15:21you're delivering, you know, $100
  2203. 1:15:23outcomes, and you get to keep a portion
  2204. 1:15:26of it,
  2205. 1:15:27everyone wants to row in the same
  2206. 1:15:29direction, and it cuts through all of
  2207. 1:15:31the prioritization and decision-making
  2208. 1:15:34that often will cloud and resource
  2209. 1:15:36allocation that often will will cloud
  2210. 1:15:38enterprise partnerships. So, it's
  2211. 1:15:39extremely valuable, and I think it's a
  2212. 1:15:41big reason why we've been successful. I
  2213. 1:15:42think it will just become the norm for
  2214. 1:15:45companies that are doing differentiated
  2215. 1:15:48high-value activities. If your product
  2216. 1:15:50really like feels a little bit more like
  2217. 1:15:52a commodity,
  2218. 1:15:54you'll start to see more usage-based and
  2219. 1:15:57seat-based pricing, cuz it's just
  2220. 1:15:58simpler. An area, for example,
  2221. 1:16:01knowledge-based lookups are a little bit
  2222. 1:16:03more that way, just question answering.
  2223. 1:16:06And so, in the case of question
  2224. 1:16:08answering, that's not an area where you
  2225. 1:16:10would have a high premium for an outcome
  2226. 1:16:13of any particular sort. But if it's
  2227. 1:16:16making a sale on a membership, or you
  2228. 1:16:18know, selling someone a car,
  2229. 1:16:20that's a really big outcome. Um, and so,
  2230. 1:16:22companies will be more than happy to pay
  2231. 1:16:24for that.
  2232. 1:16:25Uh, I think where we're seeing things
  2233. 1:16:27going is intra-conversation outcomes, to
  2234. 1:16:31also thinking about more, you know, as I
  2235. 1:16:33mentioned, kind of the moments that
  2236. 1:16:34matter across the customer life cycle,
  2237. 1:16:37and driving outcomes on top of our agent
  2238. 1:16:39data platform that kind of span that
  2239. 1:16:42whole life cycle. I think that's
  2240. 1:16:44particularly interesting.
  2241. 1:16:45>> You guys support multiple different
  2242. 1:16:47outcomes. So, you've got customer
  2243. 1:16:48support and you've got sales. How
  2244. 1:16:50different is the pricing between those
  2245. 1:16:52and how many different of these like
  2246. 1:16:53categories or templates do you guys end
  2247. 1:16:56up having?
  2248. 1:16:56>> It really depends on the value. So, you
  2249. 1:16:58asked if it was customer specific. The
  2250. 1:17:01answer ends up being that it sort of has
  2251. 1:17:02to be. In certain cases, you are
  2252. 1:17:06troubleshooting very complex setup to a
  2253. 1:17:10device or something and you have to try
  2254. 1:17:1215 different things to get it to work
  2255. 1:17:14and the average conversation might take
  2256. 1:17:1520 turns and the amount of, you know,
  2257. 1:17:18context engineering to make that work
  2258. 1:17:20might be very high.
  2259. 1:17:21In other cases, you might have something
  2260. 1:17:23where, you know, you're just resetting
  2261. 1:17:26the signal on your TV and it's very
  2262. 1:17:28quick and easy or you're checking your
  2263. 1:17:29balance with the bank and that's very
  2264. 1:17:32easy. And so, you know, one outcome is
  2265. 1:17:35very valuable and drives a lot of
  2266. 1:17:37loyalty and one outcome is somewhat
  2267. 1:17:39commoditized. You might have some cases
  2268. 1:17:41where there's, you know, an outcome
  2269. 1:17:44that's tens of dollars
  2270. 1:17:47and in terms of the, you know, money
  2271. 1:17:49that the agent would earn
  2272. 1:17:51and then you might have some cases where
  2273. 1:17:52it's, you know, much, much lower than
  2274. 1:17:54that.
  2275. 1:17:54>> And does that ever differ
  2276. 1:17:57within a customer? So, like in your
  2277. 1:17:59example, I could imagine you you could
  2278. 1:18:01have an agent doing a really simple task
  2279. 1:18:03of, oh, tell them to unplug the computer
  2280. 1:18:05and plug it back in or something like
  2281. 1:18:06that. And there's another one where
  2282. 1:18:08like, oh my god, who knows what's going
  2283. 1:18:09wrong and and it like is a miracle that
  2284. 1:18:11it solves it at all. If it's the same
  2285. 1:18:13customer, will it be charged the same
  2286. 1:18:15amount or do you differentiate even
  2287. 1:18:16within those different types of
  2288. 1:18:18requests?
  2289. 1:18:19>> There are cases where we differentiate.
  2290. 1:18:21We're not dogmatic about it. What we
  2291. 1:18:24found is that often times the benefits
  2292. 1:18:27of having our incentives aligned
  2293. 1:18:30are so high that it's not worth
  2294. 1:18:33negotiating every detail of what counts
  2295. 1:18:36for what and it kind of even out over
  2296. 1:18:40time and you do right by your customers
  2297. 1:18:42over time and you build trust and
  2298. 1:18:44contracts aren't infinite and you want
  2299. 1:18:46to have a really high renewal rate and
  2300. 1:18:48have them trust you with more use cases
  2301. 1:18:50and these kinds of things. So we make
  2302. 1:18:52sure that incentives are deeply aligned.
  2303. 1:18:55And then on top of that I think you can
  2304. 1:18:58get really pedantic about the
  2305. 1:18:59engineering of specific outcomes and
  2306. 1:19:01maybe over time the market will move in
  2307. 1:19:03that direction. But I think you're
  2308. 1:19:05missing the forest for the trees in that
  2309. 1:19:08case because of just how powerful the
  2310. 1:19:10concept is. And so most of our customers
  2311. 1:19:12are eager to find something simple that
  2312. 1:19:14we all understand that feels fair. As
  2313. 1:19:16opposed to trying to engineer like the
  2314. 1:19:18perfect value for the for each outcome.
  2315. 1:19:21>> Why don't you think there's more outcome
  2316. 1:19:23based pricing right now? Is it because
  2317. 1:19:24there's not enough agents doing valuable
  2318. 1:19:26things or because it's so operationally
  2319. 1:19:28intensive for now cuz it's early on that
  2320. 1:19:31you guys have just a built up muscle of
  2321. 1:19:32doing it and that's what allows you guys
  2322. 1:19:34to do it so effectively.
  2323. 1:19:35>> I think it's probably a bit of both. I
  2324. 1:19:37think that there are a lot of products
  2325. 1:19:40that
  2326. 1:19:41probably as models have improved find
  2327. 1:19:45themselves in a position of being more
  2328. 1:19:47similar to
  2329. 1:19:48what you could just buy tokens and
  2330. 1:19:51create and then also there's just we're
  2331. 1:19:53very early here.
  2332. 1:19:55If I had to say though I would guess
  2333. 1:19:58that the second one is more important
  2334. 1:20:00and there will be a lot more of this the
  2335. 1:20:03same way someone doesn't care how many
  2336. 1:20:07hours I work as long as I produce you
  2337. 1:20:10know new products that are good.
  2338. 1:20:12And I think that that will become true
  2339. 1:20:14of agents as well. There will be a mix
  2340. 1:20:17of building agents in house on platforms
  2341. 1:20:20like LangGraph and then there will be
  2342. 1:20:22also
  2343. 1:20:23you know buying
  2344. 1:20:25products like Sierra to build agents on.
  2345. 1:20:27>> Maybe switching to the last topic which
  2346. 1:20:28is just the type of people that thrive
  2347. 1:20:31at Sierra. I think you guys are also
  2348. 1:20:33pretty famously known for your forward
  2349. 1:20:35deployed engineering or agent builder
  2350. 1:20:36approach. Could you talk a little bit
  2351. 1:20:38about that both in terms of what those
  2352. 1:20:40people do as well as the right persona
  2353. 1:20:42to grow into that role?
  2354. 1:20:44>> I joined Sierra about 2 and 1/2 years
  2355. 1:20:46ago and it was my first B2B job ever.
  2356. 1:20:48I'd only worked in consumer products and
  2357. 1:20:51I love building consumer products. I
  2358. 1:20:52love being like, oh, I could imagine,
  2359. 1:20:54you know, my friends using this or my
  2360. 1:20:56parents using this, but I'd never really
  2361. 1:20:58loved growth
  2362. 1:21:00uh and the idea of figuring out how to
  2363. 1:21:02drive a couple percentage points of
  2364. 1:21:06attention or a couple percentage points
  2365. 1:21:08of usage. And what I learned when I
  2366. 1:21:10joined Sierra is I love enterprise
  2367. 1:21:12sales.
  2368. 1:21:14Uh
  2369. 1:21:15>> [laughter]
  2370. 1:21:16>> I got a tattoo.
  2371. 1:21:17Um so, basically, the the process of
  2372. 1:21:20caring about each customer individually,
  2373. 1:21:23saying one customer is upset, I'm going
  2374. 1:21:25to call them right now and find out why
  2375. 1:21:27and see how I can help. Just felt very
  2376. 1:21:30empowering as a builder in a way where
  2377. 1:21:33building for a billion users on Google
  2378. 1:21:36Search, for example, you know, it was
  2379. 1:21:38exciting in other ways, but it didn't
  2380. 1:21:40feel like you could listen to each user
  2381. 1:21:41and help them. And in many cases, we
  2382. 1:21:44have customers of Sierra that have, you
  2383. 1:21:45know, gotten promoted in their
  2384. 1:21:47organizations. They're building careers
  2385. 1:21:49because of the agents that they built on
  2386. 1:21:51Sierra and so it's just feels very deep
  2387. 1:21:54in terms of those relationships.
  2388. 1:21:56What I love as well though is that the
  2389. 1:21:57end user of a Sierra agent is still a
  2390. 1:21:59consumer in the vast majority of cases.
  2391. 1:22:01And I think it's pretty rare to have a
  2392. 1:22:04product that needs to be consumer grade
  2393. 1:22:07where the product that you're building,
  2394. 1:22:09it's a it's a platform, but then the end
  2395. 1:22:11user is really a consumer and you have
  2396. 1:22:13to have them in your mind the whole
  2397. 1:22:14time, but where you have kind of the
  2398. 1:22:16enterprise sales process of building
  2399. 1:22:19trust, of solving problems, of
  2400. 1:22:21discovering value, and then delivering
  2401. 1:22:23that value for people.
  2402. 1:22:25Um and so, I think the people that
  2403. 1:22:27really appreciate those two things, the
  2404. 1:22:30customer obsession and the
  2405. 1:22:32craftsmanship,
  2406. 1:22:33uh, do very well.
  2407. 1:22:35I think we've also discovered just with
  2408. 1:22:37the rise of coding agents, certain
  2409. 1:22:40things are more important than they used
  2410. 1:22:41to be. Deep customer intuition, GPT-5.5
  2411. 1:22:45doesn't really have that.
  2412. 1:22:46Um, agency, the ability to say, "Why
  2413. 1:22:49can't I do this?" Um, one of our uh,
  2414. 1:22:52engineers that has really high degree of
  2415. 1:22:54agency, her status message is just like,
  2416. 1:22:56"Why not today?"
  2417. 1:22:57Um, and so, yeah, having that mindset, I
  2418. 1:23:00think is really important. And then the
  2419. 1:23:02other thing just as someone with a
  2420. 1:23:03product background is I think we kind of
  2421. 1:23:06have
  2422. 1:23:07a faster car than you need more pit
  2423. 1:23:10stops, kind of thing. So, like a a
  2424. 1:23:12Formula 1 car needs to get its tires
  2425. 1:23:14changed more often than my Hyundai Kona.
  2426. 1:23:17Uh, and the reason for that is, you
  2427. 1:23:19know, it's driving faster, it's burning
  2428. 1:23:21more rubber, uh, etc. And I think we
  2429. 1:23:23have a similar thing building products
  2430. 1:23:25as well now where coding agents have
  2431. 1:23:27allowed us to write code a lot faster
  2432. 1:23:30and even to review it faster now. But
  2433. 1:23:32certain things like product judgment and
  2434. 1:23:34customer intuition are therefore
  2435. 1:23:36actually needed more often, not less
  2436. 1:23:38often. And so, uh, people that can bring
  2437. 1:23:41that to the table themselves are in this
  2438. 1:23:44amazing loop of moving fast, but people
  2439. 1:23:47where it's one person's job to bring
  2440. 1:23:49that and another person's job to do
  2441. 1:23:51engineering, they need even tighter
  2442. 1:23:53collaboration and, you know, more daily
  2443. 1:23:55stand-ups and that kind of thing to be
  2444. 1:23:56successful.
  2445. 1:23:57>> I really like that car analogy. I hadn't
  2446. 1:23:59heard that before and totally resonates
  2447. 1:24:00with what what what I'm seeing where
  2448. 1:24:02product is becoming the bottleneck
  2449. 1:24:04because it's so easy to code and you can
  2450. 1:24:06make so much of things, but that doesn't
  2451. 1:24:07mean you should. Who ends up fitting
  2452. 1:24:10this agent builder profile the best? Is
  2453. 1:24:12this product people then? Is this
  2454. 1:24:13engineers with good product intuition?
  2455. 1:24:16Like what does it look like practically?
  2456. 1:24:17We're still figuring it out.
  2457. 1:24:19>> I will say that people that have done
  2458. 1:24:22both roles are often successful in the
  2459. 1:24:24company. Our head of engineering, Arya,
  2460. 1:24:26has been a product manager in the past.
  2461. 1:24:28We have a number of engineers that have
  2462. 1:24:29been product managers. I think those
  2463. 1:24:32skills, knowing how to talk to
  2464. 1:24:34customers, not just like what to say
  2465. 1:24:36when you're in front of a customer, but
  2466. 1:24:38how to find your way into the right
  2467. 1:24:40conversations, having a high degree of
  2468. 1:24:42agency, being really strong with
  2469. 1:24:44communication, so that you're getting,
  2470. 1:24:46you know, product isn't the bottleneck
  2471. 1:24:48anymore. Uh those are really important
  2472. 1:24:50skills.
  2473. 1:24:51I still think kind of knowing the right
  2474. 1:24:53questions to ask and the right things to
  2475. 1:24:56tell coding agents is really important.
  2476. 1:24:57So, the systems thinking and the
  2477. 1:24:59architecture design are really
  2478. 1:25:01important. And so, if you
  2479. 1:25:03have not been an engineer before, uh it
  2480. 1:25:05can be difficult. And so, I I think that
  2481. 1:25:07the multi-disciplinary approach is more
  2482. 1:25:10important than ever. My own personal
  2483. 1:25:12rubric, which is like very much in beta,
  2484. 1:25:15is kind of this customer intuition,
  2485. 1:25:18agency, product judgment,
  2486. 1:25:21technical depth,
  2487. 1:25:23communication, intensity. Because when
  2488. 1:25:26the car, you know, you need to be really
  2489. 1:25:28locked in when you're driving a Formula
  2490. 1:25:301 car. Um and then one which is a little
  2491. 1:25:33harder to pin down, but it's just kind
  2492. 1:25:34of leadership, where when there's more
  2493. 1:25:36activity going on, the ability to to
  2494. 1:25:39draw it into the correct direction is
  2495. 1:25:41really important as well. So, this is
  2496. 1:25:43kind of the working framework in my
  2497. 1:25:45head, um but I'm sure there are lots of
  2498. 1:25:47other things, too.
  2499. 1:25:48>> How do you interview for agency? And I
  2500. 1:25:50asked this because I think the the guest
  2501. 1:25:52we had on in the previous episode said
  2502. 1:25:54the exact same word agency for one of
  2503. 1:25:56the traits that they look at. And I
  2504. 1:25:58asked him the same question. So, now I'm
  2505. 1:25:59going to ask you the same question. How
  2506. 1:26:00do you How do you interview for agency?
  2507. 1:26:02>> So, the most concrete way that we've
  2508. 1:26:04changed our interviewing process is we
  2509. 1:26:06have this AI native interview.
  2510. 1:26:08>> And you wrote a great blog on it the
  2511. 1:26:10other week.
  2512. 1:26:10>> Yes. And so,
  2513. 1:26:11you, by the way, is Vijay and Arya and
  2514. 1:26:13our uh engineering leaders. But I've
  2515. 1:26:16seen it done and participated in the
  2516. 1:26:18interview panels. And basically it
  2517. 1:26:21involves building a product end-to-end
  2518. 1:26:24over the course of a few hours and then
  2519. 1:26:26reviewing it with the team.
  2520. 1:26:28I think in that environment you can see
  2521. 1:26:31what people think is off-limits or
  2522. 1:26:33what's their job and what's not their
  2523. 1:26:34job and how far they extend sort of what
  2524. 1:26:37they're allowed to do.
  2525. 1:26:38And if they're able to
  2526. 1:26:41find opportunities that you would have
  2527. 1:26:43thought, oh, maybe that they would think
  2528. 1:26:44that's out of scope, bring them into
  2529. 1:26:46scope and build build great products on
  2530. 1:26:47top of it. You kind of see agency. You
  2531. 1:26:50see that they have a sense that a lot is
  2532. 1:26:52in their control instead of feeling like
  2533. 1:26:54certain things are not in their control.
  2534. 1:26:56And if you think about coding agents,
  2535. 1:26:58they bring so much more into I think the
  2536. 1:27:02like the locus of control, right? And so
  2537. 1:27:05you can do more things and if you
  2538. 1:27:06appreciate that, I think it comes
  2539. 1:27:09through in that AI-native interview.
  2540. 1:27:11>> Thanks for listening to Max Agency.
  2541. 1:27:13If you liked this episode, leave a
  2542. 1:27:15review and subscribe. Send feedback or
  2543. 1:27:17questions to [email protected].
  2544. 1:27:18[music]
  2545. 1:27:21We want to hear from you.

About this transcript

This page contains the full transcript of The best AI agents are simpler than you think by LangChain, generated from the public captions YouTube serves with the video. The transcript has 16,511 words across 2,545 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.