YouTube2Text

Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber — Transcript

by AI Engineer · 3,775 words · 605 segments · language en · Watch on YouTube

Full transcript

  1. 0:01[music]
  2. 0:17>> My name is Jay and I'm here with Sonya.
  3. 0:19We are part of the computer vision team
  4. 0:22at Aruba. We're going to talk to you
  5. 0:24about a real world production use.
  6. 0:27Oh, my son done. Okay.
  7. 0:30Try again.
  8. 0:32Okay. Don't worry. I'll I'll manage. You
  9. 0:34hear me now?
  10. 0:36Okay, so we're going to talk to you
  11. 0:37today about a real world production use
  12. 0:40case
  13. 0:41and specifically we're going to dive
  14. 0:43into how we design the e-bows and the
  15. 0:45e-bow loops. So
  16. 0:54All right, cool. So just before we get
  17. 0:56into the agent design,
  18. 0:58we're going to talk about a little bit
  19. 1:00about the use case. So our delivery
  20. 1:02marketplace Uber Eats, we do about 90
  21. 1:04billion
  22. 1:06run rate per year at the moment.
  23. 1:08We were adding millions of items to the
  24. 1:11marketplace each and every year.
  25. 1:14Sorry, every every month. We're growing
  26. 1:16at 20% year-on-year and and we operate
  27. 1:19in 10,000 cities globally. So not many
  28. 1:21people actually know this but our
  29. 1:23delivery marketplace is just as big as
  30. 1:26the mobility side on Uber today.
  31. 1:30Visual content actually plays a really
  32. 1:32important role for the user experience.
  33. 1:35So a photo is quite often the first
  34. 1:38signal that a customer gets that gives
  35. 1:41them that initial impression about a
  36. 1:43merchant.
  37. 1:44So a good photo can make the difference
  38. 1:47between someone scrolling through the
  39. 1:48feed and actually clicking on an item
  40. 1:50and adding to the cart. And more and
  41. 1:52more we're seeing different modalities
  42. 1:54on Uber Eats uh especially video
  43. 1:57content.
  44. 1:59But this is a problem.
  45. 2:01So, our smaller independent merchants
  46. 2:03simply just don't have the level of
  47. 2:05quality for their photos that reflect
  48. 2:08what the eater is actually going to get.
  49. 2:11And when we speak to our merchants,
  50. 2:13there are three themes that kind of
  51. 2:14emerge.
  52. 2:15Lack of time,
  53. 2:17lack of know-how, and costs cuz these
  54. 2:20professional um photo shoots actually
  55. 2:22cost a lot of money.
  56. 2:24And this can be especially [snorts]
  57. 2:25problematic if the merchant is updating
  58. 2:28their menu over time.
  59. 2:32So, this problem is actually pretty
  60. 2:34challenging to solve for at scale,
  61. 2:37right? Because our consumers, they want
  62. 2:38authentic, real-looking photos,
  63. 2:41um but a meaningful fraction of uh
  64. 2:44consumers actually distrust anything
  65. 2:46that is AI-generated. So, if you open up
  66. 2:49the Uber Eats app, the last thing that
  67. 2:50you want is to be scrolling through uh
  68. 2:53you know, food photography that looks
  69. 2:54like AI's lock.
  70. 2:56So, we're threading the needle here. We
  71. 2:58need to be able to stay faithful to the
  72. 3:00original image, preserve the brand of
  73. 3:03the merchant, and avoid everything
  74. 3:05looking the same. If we have the same
  75. 3:07prompt for every photo that we're
  76. 3:08editing, the diversity of the
  77. 3:10marketplace is going to collapse.
  78. 3:15We also, because we operate globally, we
  79. 3:18also have this long-tail distribution of
  80. 3:20different quality that we see it across
  81. 3:22the marketplace.
  82. 3:24Um so, we've got some examples here. You
  83. 3:26might see food photography that, you
  84. 3:28know, has poor sharpness, poor
  85. 3:30composition, not centered, uh or or poor
  86. 3:33colors as well. We also have a wide
  87. 3:36range of spectrum of user-generated
  88. 3:37content on the platform as well.
  89. 3:42So, what are our goals when we're
  90. 3:43designing these agents? When you think
  91. 3:45through these goals, you might actually
  92. 3:46be thinking through, you know, your own
  93. 3:48agents that you're building yourself.
  94. 3:50But for us, it's about one, preserving
  95. 3:52authenticity and trust.
  96. 3:54Two, improving the quality when we need
  97. 3:56to. So, we want to be able to improve
  98. 3:58quality selectively.
  99. 4:00We want to optimize globally for for the
  100. 4:02entire marketplace. We don't
  101. 4:04cannibalize certain merchants. We want
  102. 4:07to ship safely, and this is going to be
  103. 4:09an important theme throughout the talk.
  104. 4:11We want to learn continuously,
  105. 4:13and we want to operate at scale in a
  106. 4:15cost-efficient manner.
  107. 4:18So, agents are actually really
  108. 4:20well-suited to solve this problem.
  109. 4:23So, if you imagine a spectrum, on the
  110. 4:24one side, you've got something that's
  111. 4:27more deterministic. It's more
  112. 4:28rules-based. Um and uh you you have more
  113. 4:32control over it, but it's fairly it's a
  114. 4:34brittle system. It's not actually going
  115. 4:36to be able to scale for the entire
  116. 4:38marketplace.
  117. 4:40Imagine the other side, you provide an
  118. 4:41agent with obviously a lot of
  119. 4:43creativity, it has a lot of agency. Um
  120. 4:46and that's actually what we want to lean
  121. 4:48into,
  122. 4:49but we can't leave that unconstrained,
  123. 4:51right? Cuz we have certain safety and
  124. 4:53certain guardrails in place that we need
  125. 4:55to adhere to.
  126. 4:56So, we want to find a balancing act.
  127. 4:59Uh and that's kind of set the principle
  128. 5:01for the way that we think and design
  129. 5:03around agents and evals.
  130. 5:05So, now we're going to actually like
  131. 5:06dive a little bit deeper into a
  132. 5:09simplified but representative example of
  133. 5:12what we have in production.
  134. 5:14And we're going to go through each stage
  135. 5:16and how we eval it, and then talk
  136. 5:17through some continuous learning loops
  137. 5:19as well.
  138. 5:21So, first up, we have what we call an
  139. 5:23image understanding and routing agents.
  140. 5:26So, this is where multimodality is is
  141. 5:28pretty important. We actually ask the
  142. 5:30LLM to describe what it sees in the
  143. 5:32photo.
  144. 5:33Um and then we we create a structured
  145. 5:35output from that, and we send it to a
  146. 5:37router.
  147. 5:39The router will then determine, do we
  148. 5:40enhance it, or do we skip it? We skip
  149. 5:42it, we will keep the original.
  150. 5:45If we enhance it, we send it to our next
  151. 5:47agent,
  152. 5:48which is an image editing agent. And
  153. 5:50this can actually run in a loop. So, it
  154. 5:53gets feedback from a QA agent. Um it can
  155. 5:56edit uh
  156. 5:58in the in this loop and self-correct and
  157. 6:00fix things
  158. 6:02um as it goes.
  159. 6:04If it goes through a number of loops and
  160. 6:06it still fails, we we don't publish it.
  161. 6:09Then we actually send it to a final
  162. 6:10post-processing and QA step.
  163. 6:14If that's all good, we'll publish it to
  164. 6:16the menu.
  165. 6:17And the last thing that's really
  166. 6:19critical is we log everything.
  167. 6:22Just a quick note about logging.
  168. 6:25Don't know if you can actually read the
  169. 6:27JSON here, but you might notice that all
  170. 6:29of the agents in this end-to-end
  171. 6:31orchestration is within one It's It's
  172. 6:34basically a flat structure in this JSON.
  173. 6:37Um and so, this is actually incredibly
  174. 6:40useful for the entire team
  175. 6:42because anyone, be it non-technical uh
  176. 6:44technical um folks on engineering
  177. 6:46product, can actually dive in um and
  178. 6:49look at specific cases to diagnose and
  179. 6:51also roll up things to look in
  180. 6:53aggregates.
  181. 6:54Um and it's important to note here that,
  182. 6:56you know, we think this is important to
  183. 6:58start with. You want to start with your
  184. 7:01logging cuz if you don't start with it,
  185. 7:03you have nothing to optimize for, let
  186. 7:05alone set up a self-learning loop. And
  187. 7:07at Uber, we um we use our eyes.
  188. 7:11Cool. We're going to dive um a bit
  189. 7:13deeper into the router.
  190. 7:16So, the router's actually pretty
  191. 7:17straightforward. If you remember, we you
  192. 7:19know, we have this multimodality input.
  193. 7:21We look at certain text description
  194. 7:23metadata, the image itself. We ask it to
  195. 7:25to to to under um describe what it's
  196. 7:28seeing. We create structured output from
  197. 7:30that. With that structured output, we
  198. 7:32can then grade against a rubric. So, we
  199. 7:35have these pass and fail criteria. The
  200. 7:37last step is we want to decide whether
  201. 7:39or not we should enhance or skip.
  202. 7:43How do we actually eval this?
  203. 7:45This is you could think of this as a
  204. 7:46more sort of traditional classifier. So,
  205. 7:49here we we have a confusion matrix. You
  206. 7:51know, many of you are probably pretty
  207. 7:53familiar with this.
  208. 7:54Um but we can look at things like the
  209. 7:56true positive cases, the false negative
  210. 7:58negative cases, and so on and so forth.
  211. 7:59Essentially, what we're doing is we're
  212. 8:01measuring the precision recall.
  213. 8:03In practice, your routers might actually
  214. 8:05be much more sophisticated. So, for
  215. 8:07example, we might want to route an image
  216. 8:10to a lower latency smaller model to be
  217. 8:12able to save on cost and improve the
  218. 8:14user experience at the trade-off of
  219. 8:16quality.
  220. 8:17And if that's the case, instead of
  221. 8:19having a 2 by 2 matrix for your
  222. 8:21confusion matrix, you might actually
  223. 8:22have an n by n matrix.
  224. 8:25Where each grid is actually telling you
  225. 8:28whether or not you're correctly routing
  226. 8:30to that specific branch.
  227. 8:32So, I'm going to now hand over to Somya
  228. 8:35who's going to dive a little bit deeper
  229. 8:36into how we handle drift and human
  230. 8:38alignment.
  231. 8:48>> So, now that we spoke about how we eval
  232. 8:50the routing, I want to talk about how do
  233. 8:52you get the first version of the model
  234. 8:53out.
  235. 8:54For our use case, we consider human
  236. 8:56labels as the golden source of truth.
  237. 8:58And this is what we want to align our
  238. 9:00models to.
  239. 9:01The way we do about this is we go
  240. 9:03collect a dataset which is
  241. 9:04representative. So, you know, different
  242. 9:06cuts, geographies, dish type, image
  243. 9:08quality type. Send it to our human
  244. 9:10labelers and give them a very objective
  245. 9:12guideline to label on.
  246. 9:14This is to remove any subjective biases
  247. 9:16or any noise coming in from human
  248. 9:17labelers.
  249. 9:18Once we've got that system set up is
  250. 9:20when we start tuning our model. We take
  251. 9:22our agent, we go ahead get output from
  252. 9:24the agent, compare it to your golden
  253. 9:26dataset, evaluate if it's good enough to
  254. 9:28ship, if it meets your guardrail
  255. 9:29metrics, you go ahead and ship it. If
  256. 9:31not, then you go tune and you keep doing
  257. 9:33this until you meet your guardrail
  258. 9:34metrics.
  259. 9:36For routing, our guardrail metric is
  260. 9:37recall. We don't want any bad image to
  261. 9:40slip through our system.
  262. 9:44Here are some examples of the failures
  263. 9:45we've seen.
  264. 9:47Uh on your left you see a very good
  265. 9:48image of cheeseburger. Uh on the right
  266. 9:51you notice that the routing agent
  267. 9:52actually failed this. It said the
  268. 9:53technical is low ball and it will go
  269. 9:55send this image for enhancement. Now
  270. 9:57there's two challenges when you send
  271. 9:59this image for enhancement. Firstly, you
  272. 10:01pay the compute cost for a zero quality
  273. 10:03lift from this image. And secondly, uh
  274. 10:06there is a risk of degrading this image
  275. 10:08given it's already such a high quality
  276. 10:09image.
  277. 10:13And on the other end of the spectrum,
  278. 10:14you have a recall miss. So on your left
  279. 10:16you have an image with six chicken wings
  280. 10:19and on your right if you notice the dish
  281. 10:20name, it says eight pieces chicken
  282. 10:22wings.
  283. 10:23And your routing agent approved this
  284. 10:25image. That means So now there's a risk
  285. 10:27here if you send up send this image for
  286. 10:29enhancement and you only see six chicken
  287. 10:31wings, there's a chance your model's
  288. 10:33going to hallucinate these two extra
  289. 10:34wings to to match the description.
  290. 10:36And that's also an that's a the
  291. 10:39cut we take at our faithfulness metric
  292. 10:41that Jay earlier showed us.
  293. 10:43So the meta point I'm trying to get here
  294. 10:45is you've trained your offline model,
  295. 10:46but there will be long cases where your
  296. 10:48model is going to continue to fail and
  297. 10:50the static model will not work in the
  298. 10:51real system. You need a way such that
  299. 10:53your prompts, agents, system itself is
  300. 10:55evolving over time.
  301. 10:57And that's what we've done
  302. 10:59uh for our system as well. And I'm
  303. 11:01talking more from the routing
  304. 11:02perspective, but every component in our
  305. 11:04system is able to tune itself uh for any
  306. 11:06drift online.
  307. 11:08So what we do is we sample production
  308. 11:10data at regular cadence,
  309. 11:12uh send this to the human labelers with
  310. 11:13the same guidelines that we have seen
  311. 11:15before. Once you've got that data, we
  312. 11:17compare our agents' output with the
  313. 11:19output we got from the labelers and see
  314. 11:21if there's a mismatch. If there's a
  315. 11:23mismatch, we have an umbrella diagnosis
  316. 11:25agent which takes in the feedback,
  317. 11:27localizes where this issue is happening,
  318. 11:29and and and triggers our auto-tuning
  319. 11:31pipeline.
  320. 11:33Once we tune this agent, we go and
  321. 11:34benchmark it against our golden data set
  322. 11:36that we saw earlier, and if we pass our
  323. 11:38golden data set on the metrics that we
  324. 11:40had designed, we go ahead and ship this
  325. 11:42model.
  326. 11:43Uh if not, then you kind of keep
  327. 11:44iterating. And this happens on a regular
  328. 11:46basis on production data set.
  329. 11:48Um the beauty of this is this is
  330. 11:50completely config driven and doesn't
  331. 11:52require human in the loop. Your
  332. 11:53diagnoser agent can write your config
  333. 11:56and trigger the auto-tuning pipeline
  334. 11:57here. And this is what will keep your
  335. 11:59model sharp over time. You will have one
  336. 12:01static model with the offline, but this
  337. 12:03is what is going to keep your system
  338. 12:04alive.
  339. 12:08Um so Jay is going to spend more time on
  340. 12:10the diagnosis side of it. What I want to
  341. 12:12do is zoom into the auto-tuning bit. And
  342. 12:15again, we're looking at routing, but
  343. 12:16this is how we tune every agent in our
  344. 12:18system.
  345. 12:19Uh so we start with a target agent, and
  346. 12:22we've already got these uh unseen eval
  347. 12:24samples from our humans.
  348. 12:25We go find out the mismatch and matches
  349. 12:27and call a prompt optimizer agent. Now,
  350. 12:30this itself is two sub-agents. There's
  351. 12:32the reflect agent and the up synthesize
  352. 12:34agent. What reflect does is it it just
  353. 12:37looks at the mismatches, tries to find
  354. 12:39remove any noise, find any systemic
  355. 12:41issues that might be in your data set,
  356. 12:43and
  357. 12:44reflect on it and send that feedback to
  358. 12:46the synthesize agent. Now, the
  359. 12:48synthesize agent takes this feedback. It
  360. 12:50has your agent config. It goes and
  361. 12:52updates your agent with the new config
  362. 12:54based on the feedback it's getting. And
  363. 12:56goes and benchmarks again. If this
  364. 12:57benchmark is passed, you actually
  365. 12:59register this new agent in the new agent
  366. 13:01store. And next time your production
  367. 13:03runs, you pick up the new version of the
  368. 13:05agent.
  369. 13:07And this is a closed-loop system as I
  370. 13:08mentioned, no human in the loop. We
  371. 13:10definitely have observability on the
  372. 13:12guardrails, quick rollback built in in
  373. 13:14case of any issues with the system
  374. 13:16itself.
  375. 13:20Moving on to the next step of our
  376. 13:22orchestration flow. So we spoke about
  377. 13:23routing, moving on to the enhancement
  378. 13:25bit of it. It's a three-step process.
  379. 13:28What we do is the first step, we
  380. 13:29generate a prompt specific to this
  381. 13:31image. We take in the description, we
  382. 13:33take in the directives we were getting
  383. 13:34from our routing agent, and we go ahead
  384. 13:36and generate a prompt for this image.
  385. 13:38What needs improvement in this image
  386. 13:39specifically? And we go ahead and
  387. 13:41enhance this image. Then you've got the
  388. 13:43QA gate, which is a multi-dimensional
  389. 13:45gate, looks at multiple things like
  390. 13:46plating, faithfulness, colors. And if it
  391. 13:49passes is when you actually go ahead and
  392. 13:51publish this. If it doesn't pass, you
  393. 13:53take the feedback back from the QA gate,
  394. 13:55push it back to your generate prompt
  395. 13:56along with the initial inputs you sent
  396. 13:57it, and go ahead and enhance it again.
  397. 14:00So, there's two end results here. You
  398. 14:01either keep enhancing for K iterations
  399. 14:03and you pass your QA gate and you
  400. 14:04publish, or you take a coverage hit and
  401. 14:06you never enhance this image.
  402. 14:11Here's an example. On your left, you see
  403. 14:13a bowl of sweet potato fries. We send it
  404. 14:15up for the first iteration and our QA
  405. 14:16agent rejects it because the portion
  406. 14:18size is incorrect, the plating is very
  407. 14:21unrealistic. We take that feedback in,
  408. 14:23go for the second iteration, and we're
  409. 14:25actually able to pass it the second
  410. 14:26iteration. So, the metric we are
  411. 14:27measuring here is pass at K. Pass at K
  412. 14:30is essentially the pass rate at Kth
  413. 14:32iteration. And ideally with the more the
  414. 14:34iterations, your pass rate will increase
  415. 14:37because you're getting more feedback in.
  416. 14:40Now, I'll pass it on back to Jay to
  417. 14:42cover the rest of this.
  418. 14:47>> Thanks Thanks, Somya.
  419. 14:49Um
  420. 14:50So, yeah, just before we end here on the
  421. 14:53um on on the generation of the vowels,
  422. 14:56we use what's called pairwise
  423. 14:58comparison,
  424. 14:59right, for our pass at K. So, it's
  425. 15:01looking at the input image and the the
  426. 15:03output image, and it's assessing whether
  427. 15:05or not it's better.
  428. 15:07But how do we actually find what's
  429. 15:09better? So,
  430. 15:11um we're not going to dive into too much
  431. 15:13of the details here cuz this is kind of
  432. 15:14like proprietary stuff, and so we'll
  433. 15:16just mention it at a higher level that
  434. 15:18this is where you sort of For least for
  435. 15:20us at Rue Ba, we have to make sure that
  436. 15:22we're aligning with product, design,
  437. 15:24policy, legal. And this is where we're
  438. 15:27baking in what we define as a better
  439. 15:29image on the platform into our Evals.
  440. 15:33Um so, examples here, is it faithful? Is
  441. 15:35it complete? Is it natural? Is it
  442. 15:37realistic? And there's a bunch of other
  443. 15:39things as well. The output of this is
  444. 15:41then uh a yes, no, or unsure.
  445. 15:45So, here are some examples of failure
  446. 15:47modes.
  447. 15:48So, input and output on the right. The
  448. 15:50inputs on the left, outputs on the
  449. 15:52right-hand side. This might be a little
  450. 15:54bit uh difficult to to see at the first
  451. 15:56pass. We actually added shrimp here and
  452. 15:58we shouldn't be. So, we failed
  453. 16:00faithfulness.
  454. 16:03This is where we go the other way.
  455. 16:05So, the input um
  456. 16:07has some source at the bottom of the
  457. 16:08sushi. We actually remove it.
  458. 16:11So, we failed completeness.
  459. 16:15Here's actually a pretty interesting
  460. 16:16example where the agent actually
  461. 16:19attempted a more creative edit in the
  462. 16:21first iteration.
  463. 16:22Um and then the QA said, "Nope, it's not
  464. 16:24good enough."
  465. 16:26Uh and then it actually oversteers the
  466. 16:28other other way.
  467. 16:29And it becomes overly conservative.
  468. 16:32Sort of falls back to this generic
  469. 16:34ceramic plate uh ceramic bowl, sorry.
  470. 16:37So, this is an example of a reward
  471. 16:38hacking actually. And and this is a
  472. 16:40nugatory change, but something that we
  473. 16:42don't think is a meaningful or
  474. 16:44influential change despite the actual
  475. 16:46raw pixels of the input and output being
  476. 16:48pretty different.
  477. 16:50Here's another example where in the
  478. 16:52output the plate is covering the sauce.
  479. 16:55This is an example where for some some
  480. 16:58of the frontier models that we're using
  481. 16:59for the actual image editing, some of
  482. 17:01their um
  483. 17:02some of their problems will actually
  484. 17:03sort of leak up into our applied use
  485. 17:05case.
  486. 17:06Um and so, so object coherence and
  487. 17:09physics plausibility of the Evals that
  488. 17:11sometimes will coordinate with the
  489. 17:12frontier teams and and let them know
  490. 17:14about these problems and work together
  491. 17:15with them.
  492. 17:18Here's uh an example of why
  493. 17:20multimodality is is pretty important. In
  494. 17:22the input and the output, we we can't
  495. 17:24actually see that there are eight pieces
  496. 17:26here of of the wontons.
  497. 17:28So, we're not confident, actually. We're
  498. 17:30not sure. And so, this is an example
  499. 17:32where we would actually reject it in
  500. 17:33production and and it wouldn't it
  501. 17:35wouldn't go through.
  502. 17:39So, the last step after all of that is a
  503. 17:42post-processing and what we refer to as
  504. 17:45the publish-ready QA. This is the final
  505. 17:47gate before we decide we want to publish
  506. 17:50something to production.
  507. 17:53Here, we do some policy checks.
  508. 17:56We also do some more quality checks.
  509. 17:58Um and you might be wondering, like,
  510. 18:00we've already done some QA. Like, why
  511. 18:02are we going to do another step of QA?
  512. 18:04The reason is because we think of this
  513. 18:06like a Swiss cheese model.
  514. 18:09So, we want to try and optimize for
  515. 18:11reducing the chance of a failure getting
  516. 18:14into production. And so, there is some
  517. 18:16redundancy here or there.
  518. 18:18And that's okay.
  519. 18:20Um and so, this QA gate is is a little
  520. 18:23bit more holistic. It captures more
  521. 18:24things. But, it also will will try and
  522. 18:27flag things that we should have caught
  523. 18:28upstream, as well.
  524. 18:32All right. So, we've talked about
  525. 18:35a couple of uh
  526. 18:36feedback loops here. So, to summarize,
  527. 18:39we talked about predominantly this first
  528. 18:41one here, which is the model loop. And
  529. 18:44this is accounting for drifts and
  530. 18:46aligning with human labeled data set
  531. 18:49that we have and we've established
  532. 18:51offline.
  533. 18:52But, we actually have more feedback
  534. 18:53loops.
  535. 18:55So, we we have that Uber what we we have
  536. 18:57is a is a great sort of dog dog feeding
  537. 18:59culture.
  538. 19:00Um we will test apps before they go
  539. 19:02live. Um but, we also have when it goes
  540. 19:04live in production, how do we get that
  541. 19:07feedback back into our agent to be able
  542. 19:09to steer it appropriately?
  543. 19:11So, as we're adding more of these
  544. 19:12feedback loops, we want to be able to
  545. 19:15generalize the system.
  546. 19:17So, this is where we've actually created
  547. 19:19um a higher level of abstraction on top,
  548. 19:22which we call the diagnoser.
  549. 19:24So, the diagnoser can take in any input
  550. 19:26from these different feedback loops that
  551. 19:28we're we're capturing. It can reflect on
  552. 19:31what actual agent within the overall
  553. 19:33system needs to be optimized, and it can
  554. 19:36route that agent to be able to fix that
  555. 19:38configuration specifically. It could be
  556. 19:40one agent, it could be multiple agents.
  557. 19:45So, here's an example of internal dog
  558. 19:47fooding. You might see these in sort of
  559. 19:49different apps that you've got where you
  560. 19:50got the thumbs down and the thumbs up.
  561. 19:52We also take some free form feedback as
  562. 19:54well.
  563. 19:55Uh and this is actually great cuz we'll
  564. 19:57get feedback from merchants directly.
  565. 19:59We'll get feedback from, you know,
  566. 20:01design teams, other product teams uh at
  567. 20:03Uber. And we'll incorporate that
  568. 20:05feedback back into our diagnosis step
  569. 20:07and tune the system over time.
  570. 20:11Again, similar sort of workflow pattern
  571. 20:13here. We'll replay the examples that we
  572. 20:15know are those ones that have been
  573. 20:17flagged, be it good examples, be it bad
  574. 20:20examples, uh and then we'll benchmark
  575. 20:22the metrics before we push the latest
  576. 20:24config version.
  577. 20:27The last step is is actually getting
  578. 20:29this into production.
  579. 20:31And and this is where we're looking for
  580. 20:33a whole heap of different metrics we
  581. 20:35track for for the marketplace quality
  582. 20:37and health. Uh I've just called out one
  583. 20:39here, which is conversion. So, we're
  584. 20:40looking for improvements in people
  585. 20:43adding to cart, converting, completing
  586. 20:45their orders.
  587. 20:46Um I think this one's actually an
  588. 20:48interesting one to call out because now
  589. 20:50at I mean, at least at Uber, but
  590. 20:51especially in production um settings at
  591. 20:53scale, you have a wide um
  592. 20:56uh
  593. 20:57you have a lot of data that you can
  594. 20:59actually slice and dice.
  595. 21:00So, in this area as opposed to the
  596. 21:02others, what we can do is sort of slice
  597. 21:04by geos, by device type, by dish type,
  598. 21:07etc. And we can look at where things are
  599. 21:09improving in different segments and
  600. 21:11actually tune on certain segments as
  601. 21:13well.
  602. 21:17Cool, and that's it for our
  603. 21:18presentation. Appreciate it.
  604. 21:20>> [applause]
  605. 21:36[music]

About this transcript

This page contains the full transcript of Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber by AI Engineer, generated from the public captions YouTube serves with the video. The transcript has 3,775 words across 605 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.