Ramp: Lessons from Building a New AI Product - The Pragmatic Summit — Transcript
Full transcript
- 0:05Today we're going to talk about AI at
- 0:07Ramp and uh
- 0:10I'm going to give an intro quick
- 0:11introduction into what Ramp is.
- 0:13Um
- 0:14really briefly, we're going to walk
- 0:15through the simplest possible expense
- 0:17use case that you guys can all resonate
- 0:19cuz I see everybody's drinking coffee.
- 0:22And
- 0:23then we're going to
- 0:24talk quickly about a lesson that we
- 0:27learned
- 0:28this year while we were building
- 0:29Brazilian agents. Um and sort of the
- 0:32pivot in the paradigm that's happening
- 0:34especially after February 6th.
- 0:36And
- 0:38then we're going to double click onto
- 0:39how we built one of our most popular
- 0:41agents, the policy agent.
- 0:43Um and then finally, we'll dig in into
- 0:46the infrastructure built that this is
- 0:49requiring requiring to do on our side.
- 0:52And in my mind most importantly, the
- 0:54culture shift that needs to happen on
- 0:57everyone's teams in order to be able to
- 0:59operate in a way that delivers products
- 1:02into the hands of your customers in the
- 1:04fastest and most impactful way.
- 1:06Uh so without further ado,
- 1:09quick intro about Ramp. We are number
- 1:11one finance platform for modern
- 1:13businesses with 50,000 plus customers
- 1:17and we're in the business of saving you
- 1:19time and money.
- 1:21Uh we have uh
- 1:23I've seen some of the some of those
- 1:25names on the on the name tags here. So
- 1:26thank you for being Ramp customers. Uh
- 1:30Really exciting. Uh really quickly, so
- 1:33cup of coffee
- 1:34takes
- 1:36usually about 15 minutes of your time
- 1:40cuz you got to do these three simple
- 1:42things which unfortunately take minutes.
- 1:45This compounds through the company.
- 1:47And what Ramp does in the simplest
- 1:50possible way, we just condense time and
- 1:53return money back.
- 1:55Uh so a simple story over a transaction
- 1:58from tapping the card to writing a memo
- 2:00to classifying the transaction according
- 2:02to your GL to sourcing the receipt,
- 2:05attaching the receipt, um,
- 2:07normalizing the merchant to your, um,
- 2:10inventory of merchants is all done
- 2:12agentically at Ramp. And this was our
- 2:14first foray, uh, probably by now,
- 2:17uh, you guys still feel here? Yeah.
- 2:19Probably by now about 3 years ago, we
- 2:20started doing this one-shot things with
- 2:22AI. Uh, normalize merchant, write a
- 2:25memo, and it's been working really,
- 2:27really well as the models get better.
- 2:29Uh, what else is going on at the
- 2:31company? Well, literally every persona,
- 2:35uh, at the company is wasting time on a
- 2:40lot of manual work. Uh, so from AP
- 2:43clerks to your finance team, from your
- 2:45purchasing teams, uh, keep going to more
- 2:48finance work, your data teams, uh, at
- 2:51Ramp we used to have a channel called
- 2:53help data where somebody will ask for a
- 2:55CSV and a poor person will go and write
- 2:57a SQL query.
- 2:58Uh, it we replaced it about, uh, a year
- 3:01and a half ago.
- 3:02Uh, so a lot of time being spent and the
- 3:05complexity has a ramp shape. It only
- 3:07increases as you go through different
- 3:09jobs to be done. Um, so if you guys
- 3:12watch Super Bowl, uh, you might be
- 3:13familiar with Brian, um, our agent. Uh,
- 3:16so we've been writing a lot of agents
- 3:18literally for every job to be done to
- 3:20cover the entirety in the end state, the
- 3:23entirety of what admins, employees, and
- 3:26finance teams are doing that is not
- 3:29directly related to making the money. We
- 3:31want you all to be making money and
- 3:33focus on your customers, not on how to
- 3:35close the books.
- 3:37Uh,
- 3:38but what's been happening for the past
- 3:40few weeks is, uh, that we're living
- 3:43through the most exciting paradigm shift
- 3:45in software, um, and it requires
- 3:48complete rethink. And with rethink,
- 3:51simplification of your stack. Uh So,
- 3:54what we learned is you don't need to
- 3:56build a thousand agents. We
- 3:57intentionally last year allowed each
- 3:59individual team to go and experiment.
- 4:01And we ended up maybe with four
- 4:03different ways of doing the same thing
- 4:06both for synchronous agents as well as
- 4:07for background agents. Um but instead
- 4:11you want to drive your framework towards
- 4:15a single agent with a thousand skills.
- 4:19Uh So, let's talk about what the
- 4:21software traditionally used to focus on.
- 4:23So, every process, um especially in the
- 4:26modern modern AI stack, boils down to
- 4:28having an event. Um so, a prompt you can
- 4:31receive an invoice and you want to pay
- 4:32it. Um some prompt instructions of what
- 4:36you want to do with it and some
- 4:37guardrails like a policy, uh like an
- 4:39expense policy or your payables policy.
- 4:41Um context, what is the data that the
- 4:44agent should consider. And then finally,
- 4:47tools. These are APIs and actions that
- 4:50you can do. And traditionally software
- 4:51would focus on only four and five.
- 4:54Uh In the new paradigm
- 4:56software is doing everything. So, you
- 5:00want to focus on building an autonomous
- 5:03system of action that can react, reason,
- 5:06and act without a human or with very
- 5:08little human supervision.
- 5:11Um so, what does it mean in terms of
- 5:13what we're building?
- 5:15So, first, uh
- 5:17we decided we go into consolidate the
- 5:20interactions.
- 5:22Um verbal interactions uh
- 5:25with the agents to a single
- 5:27conversational UX. Uh We literally at
- 5:29the end of last year we had about five
- 5:31different conversational UXs. We now
- 5:33have consolidated it into what we call
- 5:35an OmniChat. Omni meaning for
- 5:37omnipresent. It is now being deployed to
- 5:39every surface of the product. And it
- 5:42works well with the traditional UX
- 5:43because you still need tables and
- 5:44buttons. And uh you don't always want to
- 5:47be talking uh to your software.
- 5:50But this is a good example what Omni
- 5:52chat looks like. Please onboard a new
- 5:54employee
- 5:55Omni chat can resolve an employee to an
- 5:59employee ID and look up through an HRIS
- 6:01tool
- 6:02their corporate structure and it found a
- 6:05workflow agentic workflow that we
- 6:07created previously called the new hire
- 6:09playbook. And the agent is asking, would
- 6:11you like me to to onboard the person
- 6:13using this playbook?
- 6:14How is this possible? We built a
- 6:16in-house lightweight agent framework
- 6:18that provides orchestration with tools
- 6:22that engineers are very quickly building
- 6:24and most recently we have one product
- 6:26manager Vibe coded about 20 tools so
- 6:28engineers are no longer needed to build
- 6:30these tools. Um
- 6:32and sometimes your workflows are
- 6:34involved such as employee onboarding
- 6:37consists of four steps. So you can just
- 6:39go and ramp and describe what do you
- 6:41want to happen when a new employee
- 6:42joins, give them a card, make sure they
- 6:46get receipts for every transaction,
- 6:47congratulate them on on Slack and check
- 6:49in with them in two weeks.
- 6:51We now are able to compile this into a
- 6:54runnable deterministic workflow
- 6:56and then give it to the agent to
- 6:58execute. Playbooks make use of tools
- 7:02and how this all comes together
- 7:05this is an example which Viral is going
- 7:07to double click next is
- 7:10upon swiping the card
- 7:12there's a real-time policy review that's
- 7:14happening directly in the software
- 7:17and policy agent enforces
- 7:20your company requirements with regard to
- 7:23spend.
- 7:24Therefore, it's very safe to give Ramp
- 7:26cards to literally every employee in
- 7:27your company. And there's a handoff
- 7:29happening with an accounting code and
- 7:31agent that classifies this transaction,
- 7:34applies the rules of your back office
- 7:37team of your finance team. As an
- 7:39employee I have no idea how certain
- 7:41transaction should match to our GL, and
- 7:43that's what typical traditional products
- 7:45would do. They will expose it to you. Um
- 7:47so the agent is much better doing it
- 7:49because it has the full context of your
- 7:51chart of accounts, it understands your
- 7:52ERP,
- 7:54and then it can either auto-approve or
- 7:56in the worst-case scenario, it will
- 7:57involve uh the human in the loop to
- 7:59review my materiality or notify that
- 8:02there is an out-of-policy spend.
- 8:04Um with that, uh please welcome Viral,
- 8:07who will
- 8:08dive deeper into the policy agent.
- 8:12Thanks, Nick.
- 8:17Oops.
- 8:19Awesome. So, a lot of finance teams are
- 8:22looking at receipts like this basically
- 8:24every day, and maybe they might have
- 8:26hundreds or thousands of these. If you
- 8:28told me to look at this and decide if I
- 8:30should approve or reject this
- 8:31transaction, I'm probably going to make
- 8:32a mistake.
- 8:34So, policy agent basically reasons on
- 8:36this image and all the transaction data
- 8:38that we have and told me that there were
- 8:40eight guests in the receipt. I could
- 8:42barely see that when I was looking at
- 8:43it. Uh it was below the $80 a person cap
- 8:46that we have internally.
- 8:48Uh they were going for team welcome
- 8:50dinner. Uh and so because the amount was
- 8:53verified as well and the merchant, uh
- 8:55policy agent told me to approve this
- 8:56transaction.
- 8:59Similarly, for this open AI transaction,
- 9:01Anan was testing out um some some
- 9:03ChatGPT features, and so policy agent
- 9:06told me this was a valid uh business
- 9:08expense and told me to approve it. And
- 9:10then this $3 bakery charge was told uh
- 9:13was was uh rejected because uh it wasn't
- 9:17uh part of an overtime purchase and it
- 9:19didn't happen on the weekend.
- 9:22So, really we looked at this as an
- 9:24opportunity to rethink how Ramp was set
- 9:27up. Um controllers and finance teams are
- 9:30looking at transactions like these and
- 9:32and making these decisions every day.
- 9:34And a Fortune 500 company that is one of
- 9:36our customers was coming to us and
- 9:38saying, "Hey, can you uh uh make sure
- 9:40that you approve these types of expenses
- 9:42and reject these types of expenses?" And
- 9:43they basically had a list of all the
- 9:45rules that uh Ramp uh should should
- 9:47follow. And we kind of saw this as an
- 9:49opportunity not to kind of add more
- 9:52incremental deterministic rules that
- 9:55kind of define our product. And I worked
- 9:56on some of the first versions of these,
- 9:58um but actually kind of take out uh a
- 10:01page from Andrej Karpathy saying that
- 10:03English is the new programming language
- 10:05and kind of turn the expense policy into
- 10:07the rules themselves. So,
- 10:09um
- 10:10you can you can see Ramp's expense
- 10:11policy on the left and and this is a
- 10:13screenshot from our production
- 10:14environment, but we are seeing really
- 10:16great uh use out of our policy agent
- 10:19product. And it kind of needed to start
- 10:23it kind of needed to start really um
- 10:25organically. So, we kind of operated
- 10:27like a early-stage startup. We're
- 10:28already very incremental and and and
- 10:30fast at Ramp, but uh we found some
- 10:32design partners like that Fortune 500
- 10:34company. We iterated really quickly, and
- 10:37we had weekly weekly meetings with all
- 10:39of them to kind of understand exactly
- 10:40what uh feedback we wanted to hear and
- 10:43what what we could improve.
- 10:46I think one of the main important um
- 10:49I guess things that we realized across
- 10:51uh Ramp is that we really needed to lean
- 10:54into the fact that AI products cannot be
- 10:56one-shotted. You need to start with
- 10:58something simple. And so, as long as
- 11:00everyone on your team, PMs, designers,
- 11:03engineers are aligned that you're not
- 11:04going to have perfection on day one, I
- 11:07think that was actually one of the main
- 11:08like cultural learnings. Um and so, we
- 11:11dogfooded a lot of this work internally
- 11:13uh and started with an even more
- 11:14constrained problem of trying to decide
- 11:17whether our coffee with a colleague
- 11:18transaction should be approved or
- 11:20rejected. These are single uh
- 11:22uh
- 11:23dollar amount transactions that are low
- 11:24risk um
- 11:26according to our finance team. And so,
- 11:28we started uh with these transactions.
- 11:30And uh one of the early learnings,
- 11:32especially as we kind of release this
- 11:35into production, was that a lot of the
- 11:37reason that policy agent would be wrong
- 11:39would be less on the models themselves
- 11:41and more about the context that we were
- 11:42giving
- 11:43to to LLM's themselves. So, we we could
- 11:46have sat down and thought about all the
- 11:48context in the beginning before we even
- 11:50kicked off any engineering work, but we
- 11:52realized actually the best thing would
- 11:53be to learn from some of our live
- 11:55internal data. And so, for example, we
- 11:58learned that the role in the title of an
- 12:00employee is super important when looking
- 12:02at expense policy docs or in level
- 12:04C-suite, for example, might have higher
- 12:06limits. Maybe they can fly on first
- 12:07class for for certain flights. And so,
- 12:09we started extracting more information
- 12:11from receipts, started pulling in
- 12:13information from HRS fields that are
- 12:15already on ramp. And so,
- 12:17Will is going to kind of talk you
- 12:19through exactly the iterations that we
- 12:21went through to implement policy agent
- 12:23and and some of the learnings along the
- 12:25way.
- 12:32Is this down?
- 12:33It's down. Yeah. Okay.
- 12:37All right, cool. Um
- 12:39awesome. So, when we first started
- 12:41building the policy agent internally,
- 12:43we dream we went big. We're like, "Hey,
- 12:45let's automate all of finance. Let's
- 12:47automate all reviews." But when it came
- 12:49down to it, we actually have to start
- 12:50small. Is that cup of coffee, you know,
- 12:53in your expense policy? And the reason
- 12:55that we did that was because even though
- 12:56the problem sounds simple
- 12:58to automate, you know, is this a simple
- 13:01question, is this in policy or not?
- 13:04It was going to grow to be complex. Kind
- 13:06of like Vimal said, we could have gone
- 13:07down and we could have figured out what
- 13:08context do we have, how can we add it,
- 13:10how can we put it all together in a way
- 13:11that LLM can understand, and you know,
- 13:13put it all together from the get-go. But
- 13:16we knew that even if we aimed and got
- 13:19everything right the first time, it was
- 13:20probably going to be wrong once you
- 13:21applied and generalized it and then to
- 13:23another business.
- 13:24Um
- 13:25so,
- 13:27the simpler the system, I think the
- 13:29easier it is to iterate on top of it.
- 13:31And once you iterate, you know what's
- 13:32going to work, you know what's not, and
- 13:33you can kind of layer complexity on top
- 13:34of that. And I think that's pretty
- 13:35important to um keep in mind when you're
- 13:37building a um
- 13:39LM or an agent starter. So, for us,
- 13:42we started really simple, very very um
- 13:45kind of the classic, you know, we have
- 13:46an expense come in, retrieve the context
- 13:48around it, we pass it through a series
- 13:50of LM calls that are very well defined
- 13:52of like, "Hey, is this in policy? Why is
- 13:54it in policy? How can we show the user
- 13:55that's in policy?" And then give an
- 13:57output that uh makes sense in this way
- 13:59to the user.
- 14:00Eventually, we learned that each expense
- 14:02is kind of different. We can classify an
- 14:04expense based on is it travel? Is it a
- 14:05meal? Is it entertainment? Do
- 14:07conditional prompting, and then retrieve
- 14:09context based on that, and then pass it
- 14:11through a series of LM calls, and give
- 14:12it some tools so that it can also
- 14:14autonomously decide, "Hey, um I need
- 14:16flight information actually, or I need
- 14:17this employee's level." Um and kind of
- 14:19layer that on top.
- 14:21And a few iterations later, we came to a
- 14:23full-on agentic workflow. Um we ended up
- 14:26with um complex tools to read across all
- 14:30of our platform, and these tools are
- 14:31shared across our all of our agents.
- 14:33It's not just for policy agent. We have
- 14:35a company internal toolbox that all of
- 14:37our agents are easily can, you know,
- 14:39reach into and use. And we gave it the
- 14:41um we gave it the um capability to write
- 14:44as well. So, it's now writing decisions,
- 14:46it's writing uh reasoning, it's writing
- 14:48auto-proving expenses on users' behalf.
- 14:51Um and it goes in a loop. So, um you
- 14:53know, now it's more of a black box, and
- 14:55that's kind of the trade-off you get.
- 14:57Um
- 14:58as you go from simple to complex
- 15:00systems, um your capability goes up,
- 15:03your uh autonomy goes up, your agents
- 15:04are able to do more, your AI can do
- 15:06more, your AI seems smarter. But, in
- 15:08exchange, you're going to be able to
- 15:10you're losing traceability and
- 15:11explainability. Uh we look at it now, we
- 15:13can kind of look at the reasoning tokens
- 15:15that the LM gives us, but in the end, we
- 15:16have no control over it. It's going to
- 15:18do what it thinks it's right, it's going
- 15:19to make the tool calls, it's going to
- 15:20tell you it's right or wrong. So, a
- 15:22smaller black box becomes a bigger black
- 15:24box as the system becomes more complex.
- 15:29So, one thing that is really important
- 15:31when doing something like this is from
- 15:32the beginning, you need really good
- 15:33auditability.
- 15:34Um assume even if you know how it it
- 15:37works, assume that your inputs and
- 15:39outputs are all you know, and make sure
- 15:40that it's correct. Um
- 15:42so
- 15:44if it was a black box system and you
- 15:45only saw the input output, can you
- 15:47verify that it did the right thing? And
- 15:48even if that black box changes, you
- 15:50should be able to reason about whether
- 15:51the output is correct.
- 15:53Um
- 15:54as with many products that we built at
- 15:56Ramp and across, you know, other
- 15:58companies, we thought that the users
- 15:59would be correct. Uh you know, if the
- 16:01user says approve, the agent should
- 16:02approve. If the user says reject, the
- 16:04agent should reject. But turns out
- 16:07the users are actually incorrect.
- 16:08They're wrong. They are sometimes, you
- 16:10know, they don't know the expense
- 16:10policy, you know, they trust their
- 16:12employees, they're lazy, it's a Sunday,
- 16:14who knows. Um so, turns out we can't
- 16:17always do what the users are doing cuz
- 16:19sometimes that's where our finance teams
- 16:21come back to you and are like, "Hey,
- 16:22this is wrong. This shouldn't be on the
- 16:24uh company card."
- 16:25So, we have to define our own definition
- 16:28of correctness.
- 16:29Um and to do that, um we had a weekly
- 16:31labeling session with across functions
- 16:33that are working on this product. Um and
- 16:35that had two um kind of really good
- 16:37outcomes. One was that we had a ground
- 16:40truth data set that we could always test
- 16:42against and we knew that this was
- 16:43correct. And two was that everyone was
- 16:45on the same page. If our agent got
- 16:47something wrong, everyone knew that it
- 16:48got it wrong. Or you know, our agent is
- 16:50missing context, everyone knew that it's
- 16:52missing that context. So, there was less
- 16:54communication, everyone's on the same
- 16:55page, and um they could focus on what's
- 16:57really priority and kind of have
- 16:59alignment on that.
- 17:02Initially, um
- 17:04getting all those people together in a
- 17:05room every week, giving them homework to
- 17:07label 100 data points, it's expensive.
- 17:09You know, that everyone everyone has
- 17:10things to do and it sometimes they don't
- 17:12come back with their homework done. It's
- 17:14just a kind of like almost becomes
- 17:16tedious even though it's so important.
- 17:17So, we wanted to make it as simple as
- 17:19possible, and the way we did that was
- 17:20that we looked for third-party vendors
- 17:23that could provide us the tools to label
- 17:24data and collect the data.
- 17:26But, turns out some tools are too
- 17:28specific to a use case, some tools are
- 17:29too general, and we could have spent
- 17:31weeks trying out different tools, but we
- 17:33decided let's just build our own. Um so,
- 17:35we used Clockwork using Streamlit. We
- 17:38basically one-shotted all of this, and
- 17:40the greatest part of it all is that it's
- 17:41low maintenance, um low risk. It's in a
- 17:44particle base that it breaks,
- 17:46we can fix it right away. Deploy's
- 17:47happening like instant seconds. And
- 17:49non-engineers can go and personalize it.
- 17:50They can they can vibe code it. They can
- 17:52clock code it. And this was at Opus 4.
- 17:53So, now at Opus 4.6, I expect it's even
- 17:56better, and uh with something like that,
- 17:58it's definitely easier and cheaper
- 17:59sometimes to do something one-off like
- 18:00this.
- 18:04And
- 18:06with that with the ground truth data
- 18:07set, we were able to make quick
- 18:08iterations. We're able to find out,
- 18:10"Hey, we need employee levels. Add that.
- 18:11How does that work?" Running it against
- 18:13this data set, does it actually catch
- 18:14it? And now say accept or approve.
- 18:17Um and we're able to make really quick
- 18:18iterations, and that was kind of the key
- 18:20um
- 18:21that was actually kind of a key point in
- 18:22developing this. Uh we had really early
- 18:24confidence that this could actually
- 18:26work, and we were able to actually buy
- 18:27get a lot of buy-in, um get a lot of
- 18:30customers on board it and that kind of
- 18:31try it out as a design partner.
- 18:33Um
- 18:34and
- 18:36as part of like doing that iteration
- 18:38with the data set, you had evals, and I
- 18:40feel like evals are very you know,
- 18:41obviously everyone I think in the zoom
- 18:43now knows about evals and what they
- 18:44mean, but um it's pretty important to
- 18:46have them early on. I wouldn't say that,
- 18:48you know, don't let perfectionism, you
- 18:49know, get in the way. You don't need a
- 18:51full data set of a thousand data points
- 18:52that you're testing against every
- 18:53iteration. We started with five. You
- 18:55know, and we knew that those five we
- 18:57were not going to fail. We kept adding
- 18:58and adding and adding. And
- 19:01you know, make sure it's easy to run.
- 19:03Anyone could go and just run that
- 19:04command. And then make sure that the
- 19:05results are really easy to understand.
- 19:07Um they're able to look at it, get
- 19:09instant, you know, output like and
- 19:10understand like, "Hey, this is what the
- 19:11model's doing. This is like good, this
- 19:13is bad, and like if you want to do it as
- 19:15part of your CI, then everyone now can
- 19:17just hopefully merge in code because
- 19:20whenever um
- 19:22whenever you think you're doing
- 19:23something right for the LLMs or agent,
- 19:25giving more context, giving it tools,
- 19:27more likely than not it's probably going
- 19:28to have some kind of bad, you know,
- 19:29consequence that you didn't see
- 19:30happening. Context was wrong, um whether
- 19:33it be the tool instructions were wrong,
- 19:35or maybe the docstring was like a little
- 19:36confusing and conflicting. Um so it
- 19:37might have consequences. You just want
- 19:39to You just want to make sure you're
- 19:40catching against those. Um and then I'll
- 19:42touch on it briefly, but online evals
- 19:44are also great. So these are offline.
- 19:46You have a data set, it's historical,
- 19:47you're testing it, but if you can,
- 19:49online evals can be a little more
- 19:50confusing and uh harder to kind of
- 19:52measure, but if you can measure anything
- 19:54that as your users are interacting with
- 19:55the system, definitely as a leading
- 19:57metric also set them up. And for us,
- 19:59part of that was, hey, how many our
- 20:00rates of like decisions. We had an
- 20:02unsure decision, which is which just
- 20:04meant that the agent didn't have enough
- 20:05information. So we can measure that
- 20:07online. So it's much simpler eval, but
- 20:09that also gave us a pretty good health
- 20:10check um as our system was running.
- 20:15Cool. And another great part about evals
- 20:17is that with evals, you can make
- 20:19confident model changes uh whenever a
- 20:21new model comes out, Opus 46, GPT-53,
- 20:24you want to make sure that you can
- 20:26leverage those new models because
- 20:27sometimes that could be the difference
- 20:28between, you know, your system getting
- 20:30one part of the problem right to wrong,
- 20:32but it could also be the opposite. It
- 20:33could have It could actually be
- 20:35not good without any prompt changes or
- 20:36changing how your system works. So um
- 20:38having evals really set it up and being
- 20:40able to benchmark really helps um make
- 20:42confident model changes.
- 20:46Cool. Um so now that policy agent, we've
- 20:49been developing this for a while, it's
- 20:50available for everyone on the Man
- 20:52platform. Some of the things that we
- 20:53learned along the way is that
- 20:55um cloud code as engineers is very
- 20:57exciting. We have full control, we get
- 20:58to modify our cloud MD, we get to make
- 21:00sure, you know, tell it to not leave
- 21:01comments, it won't leave comments,
- 21:02hopefully. Um
- 21:04turns out it's It's just us, um finance
- 21:06people also really like to have you
- 21:07know, modify their cloud MD, which is
- 21:09their expense policy. So, if something
- 21:11went wrong with the decision, then we
- 21:13just like tell them, "Hey, go update
- 21:14your policy doc." Which to them it's a
- 21:16little scary concept to begin. Like this
- 21:17is a document. Like you know, you don't
- 21:19mess with that. Um, you have to go
- 21:20through a lot of hoops if you're going
- 21:21to mess with that. Um, but then it turns
- 21:23out if you get them really excited about
- 21:24the feedback loop, "Hey, change that.
- 21:26You'll see it right away." Turns out
- 21:27they'll be like really excited to do
- 21:28this.
- 21:29Um, and then trust builds over time. So,
- 21:32some of the earlier customers that we
- 21:33had were some of the Fortune 500. We
- 21:36actually started with a really big, you
- 21:37know, um,
- 21:38enterprise customers that we had cuz we
- 21:39thought that they would have the most
- 21:40value. They have the most expenses
- 21:42coming in. They have the most time to
- 21:43spend on reviewing
- 21:45coffee expenses.
- 21:47Um, so,
- 21:48you know, roll it out to them. Let them
- 21:50have the trust. Don't We didn't do any
- 21:51autonomous action. We're just like,
- 21:52"Hey, we're going to give you a
- 21:53suggestion." That's That's how That's
- 21:55kind of how we phrased it. Suggestions.
- 21:57And then eventually they came to us and
- 21:58were like, "Okay, you know what? I want
- 22:00to go from suggestions to auto
- 22:02approvals. Like anything under $200, you
- 22:04guys are mostly right. I don't care
- 22:05about this. Let me just go auto approve
- 22:07it." So, we gave them the autonomy
- 22:08slider. We gave them a way to just like
- 22:10turn it on and then they actually could
- 22:12do it themselves.
- 22:13And then,
- 22:14last but not least, um, similar to LLMs,
- 22:17users thrive in, you know, in-product
- 22:18feedback loops. Um, so, you know, when
- 22:20you're building an AI product and you
- 22:22have a full way of like LLMs can test if
- 22:25it's code was right and still to go
- 22:27iterate, users are the same way. Um,
- 22:29give them in-product ways to improve the
- 22:31expense policy doc, improve the agent
- 22:33and how it operates. And, um, they're
- 22:35more than excited to kind of take it
- 22:36over themselves and um, kind of improve
- 22:38it and personalize it for them. So,
- 22:41um,
- 22:42from here I'll pass it on to Ian who's
- 22:43going to have to kind of talk about the
- 22:44infrastructure and the culture that we
- 22:45have at Ramp that kind of, you know, led
- 22:47us to building the policy agent.
- 22:54Hey, everybody.
- 22:56So, you've heard a little a little bit
- 22:58about like how we're kind of getting
- 22:59leverage to all of the different finance
- 23:01teams as we operate on top of their
- 23:03financial infrastructure and really try
- 23:05to get leverage for our customers. Um,
- 23:07but I think a big thing that we also
- 23:08spend a lot of time thinking about is
- 23:11how can we get leverage for Ramp itself,
- 23:14the engineers, our XFN works, all the
- 23:16people that we work with um, every
- 23:18single day. And this slide is this
- 23:20section is pretty intentionally named AI
- 23:22infrastructure and culture cuz we think
- 23:24that this is both like a really
- 23:26challenging infrastructure problem, but
- 23:27it's also a really challenging culture
- 23:29problem and changing how you work as
- 23:31well is a big part of the story.
- 23:35And so to kind of start on the
- 23:36infrastructure side, the core of how
- 23:38most of applied AI happens at Ramp is
- 23:41our applied AI surface a service. And at
- 23:44like a 10,000 ft view, this looks
- 23:46something kind of like an LLM proxy
- 23:49or something like light LLM, but there's
- 23:50really three kind of main extensions
- 23:52that we've invested in to make this a
- 23:54lot more powerful for a lot of our use
- 23:55cases.
- 23:56The first is like structured output and
- 23:58consistent API and SDKs across different
- 24:01model providers. This can be pretty
- 24:02tricky to do especially with how quickly
- 24:04the APIs are changing, but it's a
- 24:05problem that we don't want downstream
- 24:07product teams to have to think about. So
- 24:09if you have an idea of I want to switch
- 24:10from
- 24:11GPT 5.3 to Opus or I want to try Gemini
- 24:163 Pro, you should be able to do that
- 24:17with a config change and really quickly
- 24:19be able to iterate on semantic
- 24:21similarity and trying to do a bunch of
- 24:23different, um, you know, code sandboxing
- 24:25and structured output calls that way.
- 24:27The other thing that we've spent a ton
- 24:29of time thinking about is kind of batch
- 24:30processing and workflow handling. This
- 24:32is really useful for evals or if you're
- 24:33doing like bulk for us bulk document or
- 24:36data analysis,
- 24:37um, and that's something that we also
- 24:39don't want teams to have to spend a
- 24:40bunch of time on of how do you want to
- 24:41batch this and handle it with rate
- 24:42limits and do we want to do this on an
- 24:44offline or online job with something
- 24:46like Anthropic? We just want to handle
- 24:48that for downstream consumers so they
- 24:49can just focus on providing value for
- 24:51downstream customers.
- 24:53And then the last which is a pretty big
- 24:54deal is the ability to trace different
- 24:56costs across teams and against products
- 24:58as well. And this allows us to kind of
- 25:00identify the Pareto, you know, curve of
- 25:02like what is the best kind of model
- 25:04performance for cost, how are these
- 25:05evolving over time, what teams are
- 25:07actually not, you know, building
- 25:09something that's going to be sustainable
- 25:10long-term for different product
- 25:11services. And this can be really, really
- 25:13important to just remove all this work
- 25:15from internal teams having to think
- 25:17about this.
- 25:18And the last thing that's kind of, I
- 25:20think, funny to think about and we often
- 25:21joke about that, you know, our customers
- 25:24are actually using the front more of a
- 25:25frontier model than they may even know
- 25:27even is out yet, is it allows us to stay
- 25:29at the frontier when a new model comes
- 25:30out, it's a one-line config change that
- 25:32impacts every single SDK downstream. And
- 25:35so, rather than teams having to learn
- 25:36the SDK or go into 12 or dozens of
- 25:39different call sites, they can just
- 25:41change it in one place for their
- 25:42specific team, and they now get the
- 25:44benefit of being on the latest and
- 25:45greatest models um that we've kind of
- 25:47vetted and built into the rest of the
- 25:48system.
- 25:52Our product, as you've kind of heard
- 25:54earlier, earlier, works on a lot of like
- 25:56very sensitive data and very sensitive
- 25:59workflows. And I think often times, uh
- 26:01you know, something that I hear from
- 26:02engineers in the space is this kind of
- 26:04concept of hallucination and safety, and
- 26:07how are you actually going to be able to
- 26:08produce a lot of these things to have
- 26:09benefits to downstream finance teams?
- 26:12And we're pretty big believers that it
- 26:13all comes down to the catalog of tools
- 26:15that teams are building and integrating
- 26:16with on a daily basis.
- 26:18And so, what you're seeing here is our
- 26:20internal tool catalog. So, an example
- 26:23would be like get a policy snippet or
- 26:25per diem rate or a recent transactions.
- 26:28And these are built alongside of product
- 26:30teams to really understand a lot of the
- 26:31nuances in the data and the use case.
- 26:34And what's really cool about this is not
- 26:35only can you see where there's gaps in
- 26:36our offering that, oh, we actually don't
- 26:38have a tool for this specific use case.
- 26:40These can be used both in internal repos
- 26:42and our core product. And so, if you
- 26:44have an idea of I want to do a cool
- 26:45reimbursement agent idea, here are the
- 26:48different ways to integrate the tools,
- 26:49the different APIs and systems that they
- 26:50integrate with, and now you can
- 26:51prototype that on a totally new product
- 26:53and Vibe coded surface area without
- 26:56having to worry about like learning all
- 26:57of these things from scratch or building
- 26:59the tools on your own.
- 27:01We're up to like many hundreds of these
- 27:02tools today and we, as Nick mentioned
- 27:04earlier, thinks that this could be like
- 27:06multiple thousands over time.
- 27:11On the topic of context, another big
- 27:14thing we think about is context for our
- 27:15customers of how do we actually
- 27:17integrate the financial stack and allow
- 27:18them to be a little more productive.
- 27:20What we noticed like a very similar
- 27:21problem internally
- 27:23on our engineering team. And I think
- 27:25something that's like not as always
- 27:26obvious is that, you know, even if
- 27:28you're using something like Cloud Code
- 27:30or Codex, there's all this fragmentation
- 27:32of actually what you do on a daily basis
- 27:34to get work done in your company that
- 27:35that's not integrated, too.
- 27:37There's logs in DataDog, there's a
- 27:39production, you know, database that has
- 27:41a bunch of things going on, there's
- 27:43different alerting systems, there's
- 27:44incident.io, there's a Slack message you
- 27:46have to pull in, there's a Notion doc,
- 27:48and then there's a lot of like knowledge
- 27:50that those actual specific product teams
- 27:51have of how they actually need to get
- 27:53work done as well.
- 27:55And so, at the end of last year, we
- 27:58decided to start out and try to solve
- 28:00this problem of how can we actually
- 28:01integrate all this context and build our
- 28:03own internal background coding agent,
- 28:05which we've called Ramp Inspect. You may
- 28:07have seen this on LinkedIn or X. We
- 28:09actually have open-sourced the blueprint
- 28:10of how we built this, and at the end I
- 28:12can definitely show you guys a link of
- 28:13where to find that. And the the progress
- 28:16has been pretty phenomenal of actually
- 28:18integrating this into a background agent
- 28:20that can run autonomously as people are
- 28:22in meetings, if as bug fixes come up,
- 28:24and things like that. And currently this
- 28:27month, Ramp Inspect is responsible for
- 28:29over 50% of PRs that we merge to
- 28:31production. I have some interesting,
- 28:33we're like really big nerds with stats
- 28:35and numbers and things like that, so we
- 28:37have this dashboard to kind of create
- 28:38this like interesting, one like subtle
- 28:41healthy competition, but also inspire
- 28:44people that they can actually use this
- 28:45as well. And as you can engineering
- 28:48has a huge lead of the amount of
- 28:50sessions, but you also have product, you
- 28:51also have design, there's risk, legal,
- 28:54corporate finance, and even marketing
- 28:55and CX teams using Ramp Inspect. And
- 28:57they're doing things like simple copy
- 28:59changes, they're doing logic fixes,
- 29:02they're trying to respond to incidents
- 29:03or bugs. And what's been really cool to
- 29:06see as this has evolved over time,
- 29:08whoop,
- 29:11is how we've actually designed a couple
- 29:13of these things with some core
- 29:15principles to be really powerful. So
- 29:17what you're seeing here is a Ramp
- 29:18Inspect session. I think this is an
- 29:19example of like a query that we were
- 29:21trying to fix.
- 29:22This spins up in the background really
- 29:24fast
- 29:25modal code sandbox. This allows us to
- 29:27like resume, spin up, and spin down
- 29:28these containers in an isolated
- 29:30environment which has the same
- 29:31environment that you would have if
- 29:32you're developing a ramp.
- 29:34There's a series of tasks to keep it on
- 29:35track and it creates a GitHub branch and
- 29:37integrates with all of the context
- 29:39documents, our data dog, our read
- 29:41replica so we can actually write
- 29:43queries, and different context documents
- 29:45that product teams have
- 29:47have put together.
- 29:48And what's really I think subtle about
- 29:49how we've designed this is we've
- 29:51designed it to be multiplayer first. And
- 29:53that means that as you integrate or you
- 29:55try to pair with like a designer or
- 29:57somebody on the PM team, they you can
- 29:59actually help them like level up their
- 30:01own prompting skills. They can give us
- 30:03feedback of hey, click on this link,
- 30:04this actually failed in a way that I
- 30:06wasn't expecting. And so that can be a
- 30:08really great source of like
- 30:09cross-functional collaboration. That was
- 30:11a very subtle design choice that we made
- 30:13that ended up being a really big impact
- 30:15for the company.
- 30:17And then these can be kicked off either
- 30:18via a Kanban UI, we have an API, and
- 30:21then also a Slack thread. And we can
- 30:23take the full context of the Slack
- 30:24thread when it is actually kicked off so
- 30:26you don't have to re-prompt it with a
- 30:27bunch of conversation that happened
- 30:28earlier.
- 30:32What you see here is we also have a full
- 30:34VS Code environment. We run VNC inside
- 30:36of a modal sandbox as well so this
- 30:38allows us to have Chrome DevTools and
- 30:40MCP so we can actually do full stack
- 30:42work, which is pretty cool. And it has
- 30:44access to the 150 plus thousand tests
- 30:46that we have. So, it also knows if
- 30:47things are broken, can respond to the CI
- 30:50inside of GitHub, and actually patch
- 30:52fixes before it actually pings you that
- 30:54the PR is done.
- 30:56Um the link for this is
- 30:58builders.ramp.com.
- 31:00I think it's like one of the first uh
- 31:01blog post that we have uh or the most
- 31:03recent blog post that we have, and we
- 31:05open source like the whole blueprint of
- 31:06how to build this and put this together
- 31:07as well. I think there's also a GitHub
- 31:09repo called open inspect, which is an
- 31:11open source implementation of this as
- 31:13well.
- 31:17So, it's been pretty interesting to see
- 31:19the impact that Ramp Inspect has had.
- 31:20We're over 50% of PRs that we merge on a
- 31:23weekly basis goes through the system.
- 31:25And so, with all this time not spent on
- 31:28thinking about these really low-level
- 31:29firefighting tasks or really low-level
- 31:32small fixes or tweaks that can be kind
- 31:34of democratized across the company,
- 31:36we're really rethinking like how our
- 31:37engineering teams operate and think
- 31:39about their job and how they can
- 31:41actually be really impactful in this new
- 31:44kind of AI-native future.
- 31:47And so, as a thought experiment, um we
- 31:49let's pretend we have two different
- 31:50teams. I'm sure everyone in this room
- 31:52has worked with like their handful of
- 31:54extraordinary teams, maybe teams that
- 31:55are finding their footing.
- 31:57And you'll notice that there's like a
- 31:58couple of different qualities that may
- 32:00sound that may resonate.
- 32:02So, we have team A on the left here. And
- 32:04let's say that they really care about
- 32:05impact, they handle ambiguous problems,
- 32:07they understand the product, business,
- 32:08and data, they adopt new tools, they can
- 32:11find creative solutions, and they obsess
- 32:13over like the user experience.
- 32:15And then team B may also resonate with
- 32:17some people. You know, they debate
- 32:19libraries, they add process when they
- 32:21things start to feel chaotic. They
- 32:22constantly complain about head count.
- 32:25They bike shed the details instead of
- 32:26actually focusing on the user
- 32:28experience. Like, hey, should we use you
- 32:30know, functional programming paradigm
- 32:31here or what version of, you know,
- 32:33different TypeScript libraries do we
- 32:34want to use?
- 32:36And then they build before understanding
- 32:37the
- 32:38right? They just say, "Hey, we're going
- 32:39to just five code this, bro. Don't
- 32:40worry." Or they focus on, you know,
- 32:42performative code quality or nitpicks
- 32:44that may not actually They may be very
- 32:46much like a subjective kind of matter of
- 32:48fact, as well. I've worked on both of
- 32:50these teams, and I think the argument
- 32:52that I'm going to make today is that
- 32:54there's going to be a divergence, I
- 32:55think, depending on what side of the
- 32:56aisle you land there.
- 32:58This is a study from Harvard that was
- 32:59out, uh, I think the end of last year.
- 33:02And it was very much geared towards
- 33:03juniors and and seniors in terms of
- 33:05what's actually happening with hiring
- 33:06trends in uh in engineering since AI
- 33:09tools have accelerated. And I think what
- 33:11this glosses over is that I don't think
- 33:12it's just a years of experience problem.
- 33:14I actually think it's very much, um, all
- 33:17of the different qualities that I said
- 33:18in team A versus team B that really make
- 33:21it apparent that like coding was never
- 33:23really the hardest part of a lot of jobs
- 33:25for a lot for a long time. There's all
- 33:27these other engineering principles that
- 33:29become really important than just raw
- 33:31coding speed.
- 33:32So, when you think about like a staff or
- 33:34a staff plus engineer, you're really
- 33:36compensating those people more for a lot
- 33:39of the judgment that they bring to the
- 33:40table, the context, the ability to see
- 33:42around corners, all the learning that
- 33:43they have, the actual like scar tissue.
- 33:46And so, if, you know, you ask Opus 46 to
- 33:48do something, they'll have the knowledge
- 33:50to actually know if that is not going to
- 33:52work or that's actually a bad idea. And
- 33:54I think one thing that a lot of the
- 33:56narratives that we see in the media gets
- 33:57get wrong about coding agents is they
- 33:59don't really identify the fact that you
- 34:01could still build the wrong thing just a
- 34:03lot faster, and you can build like
- 34:04bigger messes. And I think that having a
- 34:07lot of these skills of a team A and
- 34:09really focusing on like what is the
- 34:10context and reason behind this will only
- 34:12become more important, um, in AI.
- 34:16And so, what does that actually look
- 34:18like? We hit on some of these things.
- 34:20Figuring out what to build and
- 34:22understanding users well enough.
- 34:24Selling an idea to skeptical
- 34:25stakeholders. This is still something
- 34:27when we decided to build a a background
- 34:29coding agent, this was not something
- 34:30that was obvious that we should be
- 34:31spending time on this.
- 34:33Having good design design decisions with
- 34:35incomplete information and maintaining
- 34:38momentum through the long middle of this
- 34:40project, which can be really gnarly. And
- 34:42I think this last
- 34:43bit, you know, everyone in this room I'm
- 34:44sure is painfully aware of you know, the
- 34:47conversation around SaaS and and the
- 34:49stock market and things like that. And I
- 34:51think this is like a big element that
- 34:52they gloss over, which is that yes, it's
- 34:54easy to vibe code something, but
- 34:56actually going through that middle
- 34:57process is like why you need really good
- 34:59engineers to actually get something
- 35:00deployed that has product market fit,
- 35:03that people are really excited about. Um
- 35:05and I think not enough people recognize
- 35:07that.
- 35:10And so where does that leave us?
- 35:11Personally, I think there's a lot of
- 35:13kind of doomerism and and scariness
- 35:14around a lot of the AI narratives, but I
- 35:17think it's also a really exciting time
- 35:19to be building.
- 35:20Unlike maybe factory work or farming,
- 35:23software's never done. We have this uh
- 35:26really
- 35:27kind of like meme internally where we
- 35:29say, you know, jobs not finished. You've
- 35:30probably seen in the marketing as well.
- 35:32I think software's perpetually not
- 35:34finished. And so with all this extra
- 35:36capacity, with people focusing less on
- 35:38this kind of low-level work and more on
- 35:40high-leverage engineering tasks, I think
- 35:42four things are going to really happen.
- 35:45I think companies are just going to
- 35:46chase opportunities they couldn't afford
- 35:47to pursue. I don't know if we would be
- 35:49chasing these like agentic workflows and
- 35:52really thinking about bigger scale
- 35:54problems in the financial stack if this
- 35:56technology didn't exist.
- 35:58People are going to enter adjacent
- 35:59markets. They're going to try to stitch
- 36:00together more value for customers. It's
- 36:02not going to be like because everyone's
- 36:042x more productive, you need two less or
- 36:06half the people.
- 36:08You're going to rebuild systems that are
- 36:09too expensive to touch. I think building
- 36:11an internal background coding agent uh
- 36:14for a company that does financial
- 36:15operations um software felt like
- 36:17probably a pretty crazy idea, but now
- 36:19that makes a ton of sense.
- 36:21And raise the bar for what good enough
- 36:22means. I think, you know, being able to
- 36:24kind of build more mind-blowing
- 36:26experiences for users, provide a lot
- 36:28more value is going to be the narrative
- 36:30of the next decade. And I'm super
- 36:32excited to be able to build some of
- 36:34these things and see what everyone in
- 36:35this room is going to build, too. So,
- 36:37thank you.
About this transcript
This page contains the full transcript of Ramp: Lessons from Building a New AI Product - The Pragmatic Summit by The Pragmatic Engineer, generated from the public captions YouTube serves with the video. The transcript has 7,385 words across 1,128 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.