Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber — Transcript
Full transcript
- 0:01[music]
- 0:17>> My name is Jay and I'm here with Sonya.
- 0:19We are part of the computer vision team
- 0:22at Aruba. We're going to talk to you
- 0:24about a real world production use.
- 0:27Oh, my son done. Okay.
- 0:30Try again.
- 0:32Okay. Don't worry. I'll I'll manage. You
- 0:34hear me now?
- 0:36Okay, so we're going to talk to you
- 0:37today about a real world production use
- 0:40case
- 0:41and specifically we're going to dive
- 0:43into how we design the e-bows and the
- 0:45e-bow loops. So
- 0:54All right, cool. So just before we get
- 0:56into the agent design,
- 0:58we're going to talk about a little bit
- 1:00about the use case. So our delivery
- 1:02marketplace Uber Eats, we do about 90
- 1:04billion
- 1:06run rate per year at the moment.
- 1:08We were adding millions of items to the
- 1:11marketplace each and every year.
- 1:14Sorry, every every month. We're growing
- 1:16at 20% year-on-year and and we operate
- 1:19in 10,000 cities globally. So not many
- 1:21people actually know this but our
- 1:23delivery marketplace is just as big as
- 1:26the mobility side on Uber today.
- 1:30Visual content actually plays a really
- 1:32important role for the user experience.
- 1:35So a photo is quite often the first
- 1:38signal that a customer gets that gives
- 1:41them that initial impression about a
- 1:43merchant.
- 1:44So a good photo can make the difference
- 1:47between someone scrolling through the
- 1:48feed and actually clicking on an item
- 1:50and adding to the cart. And more and
- 1:52more we're seeing different modalities
- 1:54on Uber Eats uh especially video
- 1:57content.
- 1:59But this is a problem.
- 2:01So, our smaller independent merchants
- 2:03simply just don't have the level of
- 2:05quality for their photos that reflect
- 2:08what the eater is actually going to get.
- 2:11And when we speak to our merchants,
- 2:13there are three themes that kind of
- 2:14emerge.
- 2:15Lack of time,
- 2:17lack of know-how, and costs cuz these
- 2:20professional um photo shoots actually
- 2:22cost a lot of money.
- 2:24And this can be especially [snorts]
- 2:25problematic if the merchant is updating
- 2:28their menu over time.
- 2:32So, this problem is actually pretty
- 2:34challenging to solve for at scale,
- 2:37right? Because our consumers, they want
- 2:38authentic, real-looking photos,
- 2:41um but a meaningful fraction of uh
- 2:44consumers actually distrust anything
- 2:46that is AI-generated. So, if you open up
- 2:49the Uber Eats app, the last thing that
- 2:50you want is to be scrolling through uh
- 2:53you know, food photography that looks
- 2:54like AI's lock.
- 2:56So, we're threading the needle here. We
- 2:58need to be able to stay faithful to the
- 3:00original image, preserve the brand of
- 3:03the merchant, and avoid everything
- 3:05looking the same. If we have the same
- 3:07prompt for every photo that we're
- 3:08editing, the diversity of the
- 3:10marketplace is going to collapse.
- 3:15We also, because we operate globally, we
- 3:18also have this long-tail distribution of
- 3:20different quality that we see it across
- 3:22the marketplace.
- 3:24Um so, we've got some examples here. You
- 3:26might see food photography that, you
- 3:28know, has poor sharpness, poor
- 3:30composition, not centered, uh or or poor
- 3:33colors as well. We also have a wide
- 3:36range of spectrum of user-generated
- 3:37content on the platform as well.
- 3:42So, what are our goals when we're
- 3:43designing these agents? When you think
- 3:45through these goals, you might actually
- 3:46be thinking through, you know, your own
- 3:48agents that you're building yourself.
- 3:50But for us, it's about one, preserving
- 3:52authenticity and trust.
- 3:54Two, improving the quality when we need
- 3:56to. So, we want to be able to improve
- 3:58quality selectively.
- 4:00We want to optimize globally for for the
- 4:02entire marketplace. We don't
- 4:04cannibalize certain merchants. We want
- 4:07to ship safely, and this is going to be
- 4:09an important theme throughout the talk.
- 4:11We want to learn continuously,
- 4:13and we want to operate at scale in a
- 4:15cost-efficient manner.
- 4:18So, agents are actually really
- 4:20well-suited to solve this problem.
- 4:23So, if you imagine a spectrum, on the
- 4:24one side, you've got something that's
- 4:27more deterministic. It's more
- 4:28rules-based. Um and uh you you have more
- 4:32control over it, but it's fairly it's a
- 4:34brittle system. It's not actually going
- 4:36to be able to scale for the entire
- 4:38marketplace.
- 4:40Imagine the other side, you provide an
- 4:41agent with obviously a lot of
- 4:43creativity, it has a lot of agency. Um
- 4:46and that's actually what we want to lean
- 4:48into,
- 4:49but we can't leave that unconstrained,
- 4:51right? Cuz we have certain safety and
- 4:53certain guardrails in place that we need
- 4:55to adhere to.
- 4:56So, we want to find a balancing act.
- 4:59Uh and that's kind of set the principle
- 5:01for the way that we think and design
- 5:03around agents and evals.
- 5:05So, now we're going to actually like
- 5:06dive a little bit deeper into a
- 5:09simplified but representative example of
- 5:12what we have in production.
- 5:14And we're going to go through each stage
- 5:16and how we eval it, and then talk
- 5:17through some continuous learning loops
- 5:19as well.
- 5:21So, first up, we have what we call an
- 5:23image understanding and routing agents.
- 5:26So, this is where multimodality is is
- 5:28pretty important. We actually ask the
- 5:30LLM to describe what it sees in the
- 5:32photo.
- 5:33Um and then we we create a structured
- 5:35output from that, and we send it to a
- 5:37router.
- 5:39The router will then determine, do we
- 5:40enhance it, or do we skip it? We skip
- 5:42it, we will keep the original.
- 5:45If we enhance it, we send it to our next
- 5:47agent,
- 5:48which is an image editing agent. And
- 5:50this can actually run in a loop. So, it
- 5:53gets feedback from a QA agent. Um it can
- 5:56edit uh
- 5:58in the in this loop and self-correct and
- 6:00fix things
- 6:02um as it goes.
- 6:04If it goes through a number of loops and
- 6:06it still fails, we we don't publish it.
- 6:09Then we actually send it to a final
- 6:10post-processing and QA step.
- 6:14If that's all good, we'll publish it to
- 6:16the menu.
- 6:17And the last thing that's really
- 6:19critical is we log everything.
- 6:22Just a quick note about logging.
- 6:25Don't know if you can actually read the
- 6:27JSON here, but you might notice that all
- 6:29of the agents in this end-to-end
- 6:31orchestration is within one It's It's
- 6:34basically a flat structure in this JSON.
- 6:37Um and so, this is actually incredibly
- 6:40useful for the entire team
- 6:42because anyone, be it non-technical uh
- 6:44technical um folks on engineering
- 6:46product, can actually dive in um and
- 6:49look at specific cases to diagnose and
- 6:51also roll up things to look in
- 6:53aggregates.
- 6:54Um and it's important to note here that,
- 6:56you know, we think this is important to
- 6:58start with. You want to start with your
- 7:01logging cuz if you don't start with it,
- 7:03you have nothing to optimize for, let
- 7:05alone set up a self-learning loop. And
- 7:07at Uber, we um we use our eyes.
- 7:11Cool. We're going to dive um a bit
- 7:13deeper into the router.
- 7:16So, the router's actually pretty
- 7:17straightforward. If you remember, we you
- 7:19know, we have this multimodality input.
- 7:21We look at certain text description
- 7:23metadata, the image itself. We ask it to
- 7:25to to to under um describe what it's
- 7:28seeing. We create structured output from
- 7:30that. With that structured output, we
- 7:32can then grade against a rubric. So, we
- 7:35have these pass and fail criteria. The
- 7:37last step is we want to decide whether
- 7:39or not we should enhance or skip.
- 7:43How do we actually eval this?
- 7:45This is you could think of this as a
- 7:46more sort of traditional classifier. So,
- 7:49here we we have a confusion matrix. You
- 7:51know, many of you are probably pretty
- 7:53familiar with this.
- 7:54Um but we can look at things like the
- 7:56true positive cases, the false negative
- 7:58negative cases, and so on and so forth.
- 7:59Essentially, what we're doing is we're
- 8:01measuring the precision recall.
- 8:03In practice, your routers might actually
- 8:05be much more sophisticated. So, for
- 8:07example, we might want to route an image
- 8:10to a lower latency smaller model to be
- 8:12able to save on cost and improve the
- 8:14user experience at the trade-off of
- 8:16quality.
- 8:17And if that's the case, instead of
- 8:19having a 2 by 2 matrix for your
- 8:21confusion matrix, you might actually
- 8:22have an n by n matrix.
- 8:25Where each grid is actually telling you
- 8:28whether or not you're correctly routing
- 8:30to that specific branch.
- 8:32So, I'm going to now hand over to Somya
- 8:35who's going to dive a little bit deeper
- 8:36into how we handle drift and human
- 8:38alignment.
- 8:48>> So, now that we spoke about how we eval
- 8:50the routing, I want to talk about how do
- 8:52you get the first version of the model
- 8:53out.
- 8:54For our use case, we consider human
- 8:56labels as the golden source of truth.
- 8:58And this is what we want to align our
- 9:00models to.
- 9:01The way we do about this is we go
- 9:03collect a dataset which is
- 9:04representative. So, you know, different
- 9:06cuts, geographies, dish type, image
- 9:08quality type. Send it to our human
- 9:10labelers and give them a very objective
- 9:12guideline to label on.
- 9:14This is to remove any subjective biases
- 9:16or any noise coming in from human
- 9:17labelers.
- 9:18Once we've got that system set up is
- 9:20when we start tuning our model. We take
- 9:22our agent, we go ahead get output from
- 9:24the agent, compare it to your golden
- 9:26dataset, evaluate if it's good enough to
- 9:28ship, if it meets your guardrail
- 9:29metrics, you go ahead and ship it. If
- 9:31not, then you go tune and you keep doing
- 9:33this until you meet your guardrail
- 9:34metrics.
- 9:36For routing, our guardrail metric is
- 9:37recall. We don't want any bad image to
- 9:40slip through our system.
- 9:44Here are some examples of the failures
- 9:45we've seen.
- 9:47Uh on your left you see a very good
- 9:48image of cheeseburger. Uh on the right
- 9:51you notice that the routing agent
- 9:52actually failed this. It said the
- 9:53technical is low ball and it will go
- 9:55send this image for enhancement. Now
- 9:57there's two challenges when you send
- 9:59this image for enhancement. Firstly, you
- 10:01pay the compute cost for a zero quality
- 10:03lift from this image. And secondly, uh
- 10:06there is a risk of degrading this image
- 10:08given it's already such a high quality
- 10:09image.
- 10:13And on the other end of the spectrum,
- 10:14you have a recall miss. So on your left
- 10:16you have an image with six chicken wings
- 10:19and on your right if you notice the dish
- 10:20name, it says eight pieces chicken
- 10:22wings.
- 10:23And your routing agent approved this
- 10:25image. That means So now there's a risk
- 10:27here if you send up send this image for
- 10:29enhancement and you only see six chicken
- 10:31wings, there's a chance your model's
- 10:33going to hallucinate these two extra
- 10:34wings to to match the description.
- 10:36And that's also an that's a the
- 10:39cut we take at our faithfulness metric
- 10:41that Jay earlier showed us.
- 10:43So the meta point I'm trying to get here
- 10:45is you've trained your offline model,
- 10:46but there will be long cases where your
- 10:48model is going to continue to fail and
- 10:50the static model will not work in the
- 10:51real system. You need a way such that
- 10:53your prompts, agents, system itself is
- 10:55evolving over time.
- 10:57And that's what we've done
- 10:59uh for our system as well. And I'm
- 11:01talking more from the routing
- 11:02perspective, but every component in our
- 11:04system is able to tune itself uh for any
- 11:06drift online.
- 11:08So what we do is we sample production
- 11:10data at regular cadence,
- 11:12uh send this to the human labelers with
- 11:13the same guidelines that we have seen
- 11:15before. Once you've got that data, we
- 11:17compare our agents' output with the
- 11:19output we got from the labelers and see
- 11:21if there's a mismatch. If there's a
- 11:23mismatch, we have an umbrella diagnosis
- 11:25agent which takes in the feedback,
- 11:27localizes where this issue is happening,
- 11:29and and and triggers our auto-tuning
- 11:31pipeline.
- 11:33Once we tune this agent, we go and
- 11:34benchmark it against our golden data set
- 11:36that we saw earlier, and if we pass our
- 11:38golden data set on the metrics that we
- 11:40had designed, we go ahead and ship this
- 11:42model.
- 11:43Uh if not, then you kind of keep
- 11:44iterating. And this happens on a regular
- 11:46basis on production data set.
- 11:48Um the beauty of this is this is
- 11:50completely config driven and doesn't
- 11:52require human in the loop. Your
- 11:53diagnoser agent can write your config
- 11:56and trigger the auto-tuning pipeline
- 11:57here. And this is what will keep your
- 11:59model sharp over time. You will have one
- 12:01static model with the offline, but this
- 12:03is what is going to keep your system
- 12:04alive.
- 12:08Um so Jay is going to spend more time on
- 12:10the diagnosis side of it. What I want to
- 12:12do is zoom into the auto-tuning bit. And
- 12:15again, we're looking at routing, but
- 12:16this is how we tune every agent in our
- 12:18system.
- 12:19Uh so we start with a target agent, and
- 12:22we've already got these uh unseen eval
- 12:24samples from our humans.
- 12:25We go find out the mismatch and matches
- 12:27and call a prompt optimizer agent. Now,
- 12:30this itself is two sub-agents. There's
- 12:32the reflect agent and the up synthesize
- 12:34agent. What reflect does is it it just
- 12:37looks at the mismatches, tries to find
- 12:39remove any noise, find any systemic
- 12:41issues that might be in your data set,
- 12:43and
- 12:44reflect on it and send that feedback to
- 12:46the synthesize agent. Now, the
- 12:48synthesize agent takes this feedback. It
- 12:50has your agent config. It goes and
- 12:52updates your agent with the new config
- 12:54based on the feedback it's getting. And
- 12:56goes and benchmarks again. If this
- 12:57benchmark is passed, you actually
- 12:59register this new agent in the new agent
- 13:01store. And next time your production
- 13:03runs, you pick up the new version of the
- 13:05agent.
- 13:07And this is a closed-loop system as I
- 13:08mentioned, no human in the loop. We
- 13:10definitely have observability on the
- 13:12guardrails, quick rollback built in in
- 13:14case of any issues with the system
- 13:16itself.
- 13:20Moving on to the next step of our
- 13:22orchestration flow. So we spoke about
- 13:23routing, moving on to the enhancement
- 13:25bit of it. It's a three-step process.
- 13:28What we do is the first step, we
- 13:29generate a prompt specific to this
- 13:31image. We take in the description, we
- 13:33take in the directives we were getting
- 13:34from our routing agent, and we go ahead
- 13:36and generate a prompt for this image.
- 13:38What needs improvement in this image
- 13:39specifically? And we go ahead and
- 13:41enhance this image. Then you've got the
- 13:43QA gate, which is a multi-dimensional
- 13:45gate, looks at multiple things like
- 13:46plating, faithfulness, colors. And if it
- 13:49passes is when you actually go ahead and
- 13:51publish this. If it doesn't pass, you
- 13:53take the feedback back from the QA gate,
- 13:55push it back to your generate prompt
- 13:56along with the initial inputs you sent
- 13:57it, and go ahead and enhance it again.
- 14:00So, there's two end results here. You
- 14:01either keep enhancing for K iterations
- 14:03and you pass your QA gate and you
- 14:04publish, or you take a coverage hit and
- 14:06you never enhance this image.
- 14:11Here's an example. On your left, you see
- 14:13a bowl of sweet potato fries. We send it
- 14:15up for the first iteration and our QA
- 14:16agent rejects it because the portion
- 14:18size is incorrect, the plating is very
- 14:21unrealistic. We take that feedback in,
- 14:23go for the second iteration, and we're
- 14:25actually able to pass it the second
- 14:26iteration. So, the metric we are
- 14:27measuring here is pass at K. Pass at K
- 14:30is essentially the pass rate at Kth
- 14:32iteration. And ideally with the more the
- 14:34iterations, your pass rate will increase
- 14:37because you're getting more feedback in.
- 14:40Now, I'll pass it on back to Jay to
- 14:42cover the rest of this.
- 14:47>> Thanks Thanks, Somya.
- 14:49Um
- 14:50So, yeah, just before we end here on the
- 14:53um on on the generation of the vowels,
- 14:56we use what's called pairwise
- 14:58comparison,
- 14:59right, for our pass at K. So, it's
- 15:01looking at the input image and the the
- 15:03output image, and it's assessing whether
- 15:05or not it's better.
- 15:07But how do we actually find what's
- 15:09better? So,
- 15:11um we're not going to dive into too much
- 15:13of the details here cuz this is kind of
- 15:14like proprietary stuff, and so we'll
- 15:16just mention it at a higher level that
- 15:18this is where you sort of For least for
- 15:20us at Rue Ba, we have to make sure that
- 15:22we're aligning with product, design,
- 15:24policy, legal. And this is where we're
- 15:27baking in what we define as a better
- 15:29image on the platform into our Evals.
- 15:33Um so, examples here, is it faithful? Is
- 15:35it complete? Is it natural? Is it
- 15:37realistic? And there's a bunch of other
- 15:39things as well. The output of this is
- 15:41then uh a yes, no, or unsure.
- 15:45So, here are some examples of failure
- 15:47modes.
- 15:48So, input and output on the right. The
- 15:50inputs on the left, outputs on the
- 15:52right-hand side. This might be a little
- 15:54bit uh difficult to to see at the first
- 15:56pass. We actually added shrimp here and
- 15:58we shouldn't be. So, we failed
- 16:00faithfulness.
- 16:03This is where we go the other way.
- 16:05So, the input um
- 16:07has some source at the bottom of the
- 16:08sushi. We actually remove it.
- 16:11So, we failed completeness.
- 16:15Here's actually a pretty interesting
- 16:16example where the agent actually
- 16:19attempted a more creative edit in the
- 16:21first iteration.
- 16:22Um and then the QA said, "Nope, it's not
- 16:24good enough."
- 16:26Uh and then it actually oversteers the
- 16:28other other way.
- 16:29And it becomes overly conservative.
- 16:32Sort of falls back to this generic
- 16:34ceramic plate uh ceramic bowl, sorry.
- 16:37So, this is an example of a reward
- 16:38hacking actually. And and this is a
- 16:40nugatory change, but something that we
- 16:42don't think is a meaningful or
- 16:44influential change despite the actual
- 16:46raw pixels of the input and output being
- 16:48pretty different.
- 16:50Here's another example where in the
- 16:52output the plate is covering the sauce.
- 16:55This is an example where for some some
- 16:58of the frontier models that we're using
- 16:59for the actual image editing, some of
- 17:01their um
- 17:02some of their problems will actually
- 17:03sort of leak up into our applied use
- 17:05case.
- 17:06Um and so, so object coherence and
- 17:09physics plausibility of the Evals that
- 17:11sometimes will coordinate with the
- 17:12frontier teams and and let them know
- 17:14about these problems and work together
- 17:15with them.
- 17:18Here's uh an example of why
- 17:20multimodality is is pretty important. In
- 17:22the input and the output, we we can't
- 17:24actually see that there are eight pieces
- 17:26here of of the wontons.
- 17:28So, we're not confident, actually. We're
- 17:30not sure. And so, this is an example
- 17:32where we would actually reject it in
- 17:33production and and it wouldn't it
- 17:35wouldn't go through.
- 17:39So, the last step after all of that is a
- 17:42post-processing and what we refer to as
- 17:45the publish-ready QA. This is the final
- 17:47gate before we decide we want to publish
- 17:50something to production.
- 17:53Here, we do some policy checks.
- 17:56We also do some more quality checks.
- 17:58Um and you might be wondering, like,
- 18:00we've already done some QA. Like, why
- 18:02are we going to do another step of QA?
- 18:04The reason is because we think of this
- 18:06like a Swiss cheese model.
- 18:09So, we want to try and optimize for
- 18:11reducing the chance of a failure getting
- 18:14into production. And so, there is some
- 18:16redundancy here or there.
- 18:18And that's okay.
- 18:20Um and so, this QA gate is is a little
- 18:23bit more holistic. It captures more
- 18:24things. But, it also will will try and
- 18:27flag things that we should have caught
- 18:28upstream, as well.
- 18:32All right. So, we've talked about
- 18:35a couple of uh
- 18:36feedback loops here. So, to summarize,
- 18:39we talked about predominantly this first
- 18:41one here, which is the model loop. And
- 18:44this is accounting for drifts and
- 18:46aligning with human labeled data set
- 18:49that we have and we've established
- 18:51offline.
- 18:52But, we actually have more feedback
- 18:53loops.
- 18:55So, we we have that Uber what we we have
- 18:57is a is a great sort of dog dog feeding
- 18:59culture.
- 19:00Um we will test apps before they go
- 19:02live. Um but, we also have when it goes
- 19:04live in production, how do we get that
- 19:07feedback back into our agent to be able
- 19:09to steer it appropriately?
- 19:11So, as we're adding more of these
- 19:12feedback loops, we want to be able to
- 19:15generalize the system.
- 19:17So, this is where we've actually created
- 19:19um a higher level of abstraction on top,
- 19:22which we call the diagnoser.
- 19:24So, the diagnoser can take in any input
- 19:26from these different feedback loops that
- 19:28we're we're capturing. It can reflect on
- 19:31what actual agent within the overall
- 19:33system needs to be optimized, and it can
- 19:36route that agent to be able to fix that
- 19:38configuration specifically. It could be
- 19:40one agent, it could be multiple agents.
- 19:45So, here's an example of internal dog
- 19:47fooding. You might see these in sort of
- 19:49different apps that you've got where you
- 19:50got the thumbs down and the thumbs up.
- 19:52We also take some free form feedback as
- 19:54well.
- 19:55Uh and this is actually great cuz we'll
- 19:57get feedback from merchants directly.
- 19:59We'll get feedback from, you know,
- 20:01design teams, other product teams uh at
- 20:03Uber. And we'll incorporate that
- 20:05feedback back into our diagnosis step
- 20:07and tune the system over time.
- 20:11Again, similar sort of workflow pattern
- 20:13here. We'll replay the examples that we
- 20:15know are those ones that have been
- 20:17flagged, be it good examples, be it bad
- 20:20examples, uh and then we'll benchmark
- 20:22the metrics before we push the latest
- 20:24config version.
- 20:27The last step is is actually getting
- 20:29this into production.
- 20:31And and this is where we're looking for
- 20:33a whole heap of different metrics we
- 20:35track for for the marketplace quality
- 20:37and health. Uh I've just called out one
- 20:39here, which is conversion. So, we're
- 20:40looking for improvements in people
- 20:43adding to cart, converting, completing
- 20:45their orders.
- 20:46Um I think this one's actually an
- 20:48interesting one to call out because now
- 20:50at I mean, at least at Uber, but
- 20:51especially in production um settings at
- 20:53scale, you have a wide um
- 20:56uh
- 20:57you have a lot of data that you can
- 20:59actually slice and dice.
- 21:00So, in this area as opposed to the
- 21:02others, what we can do is sort of slice
- 21:04by geos, by device type, by dish type,
- 21:07etc. And we can look at where things are
- 21:09improving in different segments and
- 21:11actually tune on certain segments as
- 21:13well.
- 21:17Cool, and that's it for our
- 21:18presentation. Appreciate it.
- 21:20>> [applause]
- 21:36[music]
About this transcript
This page contains the full transcript of Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber by AI Engineer, generated from the public captions YouTube serves with the video. The transcript has 3,775 words across 605 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.