YouTube2Text

Inside OpenAI's Internal AI Data Agent — Bonnie Xu, OpenAI (Summit '26) — Transcript

by Collate · 3,638 words · 594 segments · language en · Watch on YouTube

Full transcript

  1. 0:00I'm really excited to
  2. 0:02soon welcome our next speaker. We have
  3. 0:04Bonnie Shi who will be walking us
  4. 0:06through OpenAI's internal data agent
  5. 0:08Kepler which Suresh mentioned. And that
  6. 0:11supports thousands of internal users
  7. 0:13within OpenAI.
  8. 0:15Uh and just to give a little background
  9. 0:17on Bonnie, she's a software engineer and
  10. 0:19the tech lead of the data productivity
  11. 0:21team at OpenAI where she built an
  12. 0:23AI-powered data analytics agent from the
  13. 0:25ground up to help other teams at OpenAI
  14. 0:27explore and understand data more
  15. 0:29intelligently.
  16. 0:31Before joining OpenAI, OpenAI, she spent
  17. 0:33four years at Stripe working on the data
  18. 0:36platform and previously held engineering
  19. 0:38roles at both Meta and Google.
  20. 0:41Her work focuses on building scalable
  21. 0:43systems that integrate AI and data to
  22. 0:46make analysis faster and more
  23. 0:48accessible. So, if you would, please
  24. 0:50join me in welcoming Bonnie Shi to the
  25. 0:53stage.
  26. 0:57Drumroll, please. Bonnie, there she is.
  27. 1:00All right.
  28. 1:02So good to see you, Bonnie. Thanks for
  29. 1:03joining us today.
  30. 1:05>> Hello. Thanks for the kind introduction,
  31. 1:07Steve. So good to be here with you all
  32. 1:09today.
  33. 1:09>> Absolutely. Thanks for making the time
  34. 1:11to share your perspective with us. I'm
  35. 1:12going to jump off. I'll turn things over
  36. 1:14to you to take it away. Thanks, Bonnie.
  37. 1:16>> Awesome. Um so, hello everyone. My name
  38. 1:19is Bonnie and today I'll be talking
  39. 1:21about how OpenAI deployed AI data agents
  40. 1:24to help our teams answer questions about
  41. 1:27their data.
  42. 1:30So, let me paint you a picture. Your
  43. 1:31business lead comes to you and asks you
  44. 1:33the question, "How many ChatGPT Pro
  45. 1:35users do we have in Italy?" You consult
  46. 1:38a data scientist, but they're like,
  47. 1:40"Actually, this is kind of hard. Let me
  48. 1:42get back to you."
  49. 1:44They don't know what table to look at,
  50. 1:45so they ask another engineer. But the
  51. 1:47other engineer doesn't know which table
  52. 1:49is the right one.
  53. 1:50So, after three code deep dives, two
  54. 1:52quick meetings, and five Slack threads,
  55. 1:55we finally have an answer.
  56. 1:57Simple questions like these shouldn't be
  57. 1:59this consuming and this time
  58. 2:02you know, time taking, but they are.
  59. 2:05And the reason why this is so hard is
  60. 2:07because there's so much data.
  61. 2:09So, I'm Bonnie and I'm here today to
  62. 2:11talk to you about how we solved this
  63. 2:12problem.
  64. 2:16Let me start off with an overview of the
  65. 2:18scale of our data platform to illustrate
  66. 2:20why we need an AI data agent.
  67. 2:22Then I'll go into implementation
  68. 2:24specifics, the learnings we had, and the
  69. 2:27next steps we're planning to take.
  70. 2:31At OpenAI, 80% of the company directly
  71. 2:34uses our data platform. That's 80% of
  72. 2:37the company using our team's 15 tools to
  73. 2:40process over 600 petabytes of data a day
  74. 2:43or about 70K data sets.
  75. 2:46The data is growing rapidly.
  76. 2:48This means we have so many more
  77. 2:50questions to answer.
  78. 2:52But there's more and more data to sift
  79. 2:53through to get the right result.
  80. 2:56When ChatGPT launched in 2022, we were
  81. 2:58asking ourselves, "How many users do we
  82. 3:01have?"
  83. 3:02As the product has evolved, we've
  84. 3:04introduced more regions, different
  85. 3:06plans, more features, and now the
  86. 3:07question is, "How many daily active
  87. 3:10instant checkout users do we have in New
  88. 3:12York?"
  89. 3:13That's a much harder question to answer
  90. 3:15now, but fundamentally, we're looking
  91. 3:17for the same type of answer.
  92. 3:21One of the reasons why this becomes a
  93. 3:23lot harder is because table discovery is
  94. 3:25a lot harder at scale.
  95. 3:27A common problem is having difficulty
  96. 3:29finding the right table to use because
  97. 3:31there's a lot of similarly sounding
  98. 3:32tables and it's unclear what data is in
  99. 3:35them.
  100. 3:36Another common issue is having trouble
  101. 3:38understanding the subtle nuances of each
  102. 3:40table.
  103. 3:41This is a hard problem because some
  104. 3:43tables have encrypted IDs, some tables
  105. 3:46have unencrypted IDs, but we still might
  106. 3:48want to join them.
  107. 3:50Some tables have columns that adjust for
  108. 3:51fraud rates, some tables don't.
  109. 3:54Some tables are pre-filtered by
  110. 3:56feedback, some are not.
  111. 3:58Missing one nuance can lead to an answer
  112. 4:00that is wrong by an order of magnitude.
  113. 4:03This can be catastrophic when making
  114. 4:04important business decisions.
  115. 4:07Not to mention, writing SQL is hard. Who
  116. 4:10can remember all the different ways we
  117. 4:11date format, or write performing
  118. 4:13queries, or aggravating fact that Trino
  119. 4:16arrays are one indexed?
  120. 4:20So, what's a better way of doing this?
  121. 4:25At OpenAI, we built an internal AI data
  122. 4:27analyst that takes the full context of
  123. 4:29data platform and answers these
  124. 4:31questions for you.
  125. 4:36At its core, the agent back-end service
  126. 4:38leverages the model to produce
  127. 4:40AI-powered results wherever you need
  128. 4:42them. This could be in Slack for
  129. 4:44answering your business questions, or in
  130. 4:46your IDE when you're building your data
  131. 4:48pipelines, or in the web agents when
  132. 4:50you're looking up different tables, or
  133. 4:52when you're writing workflows and need
  134. 4:54data to inform the next action.
  135. 4:56Let's go through an example.
  136. 5:00Suppose I'm a data scientist and I'm
  137. 5:02investigating a huge upward spike in
  138. 5:04user growth in ChatGPT late March 2025.
  139. 5:10The first step in reasoning is to check
  140. 5:12that the spike is real and as I
  141. 5:14described. So, as we can see here, the
  142. 5:17agent looks at the actual table that
  143. 5:18stores weekly active user data to get
  144. 5:21the exact numbers pre- and post-spike.
  145. 5:26To make sure it's getting the right
  146. 5:27table, the agent looks at internal
  147. 5:29knowledge to understand it. In this
  148. 5:31case, it was able to find extra context
  149. 5:33on the table from a Notion doc. It also
  150. 5:35cross-checked against a dashboard.
  151. 5:40As the agent is interactively exploring
  152. 5:42the data, it's checking the different
  153. 5:44types of data it can get.
  154. 5:46Here, you can tell it's looking at the
  155. 5:47different data dimensions like plan type
  156. 5:49and off status and running queries
  157. 5:51against those data different data
  158. 5:53dimensions.
  159. 5:54This helps the agent get more data and
  160. 5:56needed for analysis. Kind of slicing and
  161. 5:58dicing the data uh just like you might
  162. 6:00if you're doing this exploration.
  163. 6:04As part of the analysis, the agent is
  164. 6:06trying to come up with reasons why this
  165. 6:08is happening. In this particular case,
  166. 6:11it's exploring whether the spike was
  167. 6:12actually a logging bug due to
  168. 6:14duplication. These happen.
  169. 6:20After some analysis, the agent arrives
  170. 6:22at the conclusion that it was due to the
  171. 6:24image gen release.
  172. 6:26It validates this conclusion by using
  173. 6:28web search to pull up release stocks.
  174. 6:32And woohoo, this was in fact the right
  175. 6:34reason. The spike was due to the viral
  176. 6:37uh self anime image generation trend.
  177. 6:42So, how does this magic happen?
  178. 6:47This is a big overall picture diagram of
  179. 6:50how things happen behind the scenes.
  180. 6:52Let's go through each part.
  181. 6:54On the left are the entry points. On the
  182. 6:56top are the offline curated knowledge
  183. 6:58pieces. On the bottom are the sync API
  184. 7:01calls that the agent can make to our
  185. 7:02data platform tools. And in the middle
  186. 7:05is the most important piece. It's the
  187. 7:07agent talking to the model.
  188. 7:12Agents without context can give wildly
  189. 7:14wrong answers.
  190. 7:16As pointed out in the amazing earlier
  191. 7:17presentation.
  192. 7:20Take this example where somehow the
  193. 7:22agent thought there were 5 million chat
  194. 7:24GPT users compared to the actual 800
  195. 7:27million answer announced at Dev Day.
  196. 7:29Just a minor rounding detail. That's
  197. 7:32all.
  198. 7:35Here are the different layers of context
  199. 7:37that our internal data agent uses.
  200. 7:40Let's start at the bottom with table
  201. 7:42metadata context at layer one.
  202. 7:48As you might have guessed, fitting all
  203. 7:5070k tables with their schemas and query
  204. 7:52history is way too much data to fit in a
  205. 7:55model's context window. So, we need to
  206. 7:57do some preprocessing ahead of time.
  207. 7:59Table schema information is inserted so
  208. 8:02the model knows how to query each table,
  209. 8:04what columns are available with what
  210. 8:06types.
  211. 8:07All this rich table metadata information
  212. 8:09is surfaced from Open Metadata and is
  213. 8:11made available to the agent.
  214. 8:14Schemas alone aren't enough to
  215. 8:15understand the semantics and
  216. 8:16relationships between the data.
  217. 8:18So, we use table query history, table
  218. 8:20lineage, and code derived table
  219. 8:22information to provide this extra
  220. 8:24context.
  221. 8:25All of this is transformed into an
  222. 8:27embedding so that when the agent is
  223. 8:28writing, we do rag for every table.
  224. 8:31The agent can get information for a
  225. 8:33specific table or just semantic search
  226. 8:35over keywords.
  227. 8:36So, that's layer one.
  228. 8:38Layer two is human annotations, which
  229. 8:40can be incredibly especially when
  230. 8:42starting out and for key tables and can
  231. 8:44be collaboratively done via editing
  232. 8:46table descriptions, for example, in the
  233. 8:48Open Metadata UI. So, here you we you
  234. 8:51can see there's all these great table
  235. 8:52bits like the schema, you know, the
  236. 8:54queries, lineage, so on and so forth.
  237. 9:00We use Open Metadata's APIs for both
  238. 9:02online and offline retrieval.
  239. 9:04In our internal knowledge indexing, we
  240. 9:06pull useful bits like query history,
  241. 9:08descriptions, tags, usage, and in online
  242. 9:11cases, the agent can do table search or
  243. 9:13live pull groups needed for permissions.
  244. 9:16It's been really useful for us to have
  245. 9:18Open Metadata sit on top of our data
  246. 9:19warehouse for all this extra table
  247. 9:21context that is easy to find for both
  248. 9:24agents and humans.
  249. 9:27Another thing that people often point
  250. 9:30out about table metadata is that
  251. 9:32descriptions can get easily outdated and
  252. 9:34often become a burden to maintain.
  253. 9:37And this leads to the agent getting bad
  254. 9:38results.
  255. 9:39But it can be really tedious if you're
  256. 9:41manually updating them, especially when
  257. 9:43you have so many tables.
  258. 9:45So, we solve this by auto generating as
  259. 9:47much as possible.
  260. 9:49And we also include information beyond
  261. 9:51what might be available in just
  262. 9:52metadata, and that's what makes
  263. 9:53generations even better.
  264. 9:56And this part is layer three, the Codex
  265. 9:57enrichment.
  266. 9:59It's not enough to look at a table by
  267. 10:01itself as is. You need to really
  268. 10:03understand how a table is created, where
  269. 10:05it came from, and this is really the
  270. 10:07secret to the agent truly understanding
  271. 10:09the differences between tables, knowing
  272. 10:11that a table is filtered down because it
  273. 10:13only came from a subset of blogs.
  274. 10:15And we achieve this by running an
  275. 10:17offline job that generates Codex tasks
  276. 10:19for certain tables.
  277. 10:21These Codex tasks are launched in
  278. 10:22parallel and crawl the code base to
  279. 10:24understand things like a table's
  280. 10:25purpose, downstream usage patterns, the
  281. 10:28exact table grain, and primary keys.
  282. 10:30So, instead of only knowing that a table
  283. 10:32is for ChatGPT analytics, you know that
  284. 10:34it's a table that includes only
  285. 10:36first-party ChatGPT traffic, and it's
  286. 10:38enriched by safety sources. You know
  287. 10:40that these fields are null and when
  288. 10:42these signals are missing.
  289. 10:46For layer four, institutional knowledge,
  290. 10:48we have an internal knowledge service at
  291. 10:50OpenAI that ingests internal knowledge
  292. 10:52like Slack threads and docs.
  293. 10:54When fetching, we cache and do
  294. 10:56permissions checking so that results are
  295. 10:58fast and gated correctly.
  296. 11:00This way, we can efficiently get all the
  297. 11:02right company-specific context to
  298. 11:04accompany the data.
  299. 11:06Concretely, this means that when you see
  300. 11:08a dip in weekly active users, the agent
  301. 11:10can pull the Slack thread and instant
  302. 11:12channel that talks about the outage.
  303. 11:14And this is really helpful when doing
  304. 11:16the analysis for providing the extra
  305. 11:17context on the why rather than just the
  306. 11:20what.
  307. 11:21Level five is memory. Memory is useful
  308. 11:24for things like corrections and
  309. 11:26learnings.
  310. 11:27Here's an example.
  311. 11:28Without memory, the agent takes way
  312. 11:31longer to figure out how to answer a
  313. 11:33question correctly and has to work for
  314. 11:35much longer.
  315. 11:37Memory is the mechanism that helps the
  316. 11:39agent continuously improve and learn.
  317. 11:42Context will get you 80 to 90% of the
  318. 11:44way there, but sometimes you need these
  319. 11:46final little corrections to get the
  320. 11:47right answer.
  321. 11:49For example, you just need to know that
  322. 11:50a particular type of tea product is
  323. 11:52filtered by this random string. That's
  324. 11:54just the value and that's just how it
  325. 11:56is.
  326. 11:57We make memories easily available and
  327. 11:59accessible, so anyone can contribute to
  328. 12:01make the agent better.
  329. 12:03In a conversation, the agent itself or a
  330. 12:05human can save a memory. And there's two
  331. 12:07scopes, global or user. These scopes are
  332. 12:10important because users might want their
  333. 12:12own customizations or the data is
  334. 12:14potentially sensitive. And based on a
  335. 12:16conversation, the agent will also
  336. 12:18proactively suggest memories and users
  337. 12:20can confirm their insertions.
  338. 12:22And if all of this offline information
  339. 12:24isn't enough, we also have our last
  340. 12:27layer, which is runtime context. The
  341. 12:29agent can make live calls to our data
  342. 12:31warehouse or any other data platform
  343. 12:33service it needs. So we can talk to
  344. 12:35Spark or Airflow or open metadata. Uh
  345. 12:37it's just a short API call away.
  346. 12:40One example of when you might use this
  347. 12:42is if a table is new. So, you know, you
  348. 12:44generated this new testing table and you
  349. 12:46want to compare it with an existing
  350. 12:47table and the agent can just directly
  351. 12:50run those queries live.
  352. 12:52And I just really want to drive home
  353. 12:53this point that models are really smart,
  354. 12:56but they're not the full answer and
  355. 12:58context is really what makes the
  356. 13:00difference.
  357. 13:02The next important factor to consider is
  358. 13:04how we ensure we don't cause
  359. 13:06regressions. And let me talk about how
  360. 13:08we measure response quality.
  361. 13:10Our evals consists of sets of
  362. 13:13question-answer pairs. Question is
  363. 13:15usually some important metric we want to
  364. 13:17get right and then we have manually
  365. 13:19curated um expected SQLs that will
  366. 13:22generate the correct answer. And so we
  367. 13:24hit our agent query generation endpoint
  368. 13:26to it a natural language question to
  369. 13:28generate SQL and run the query and then
  370. 13:30we do the same with the expected SQL and
  371. 13:32we compare the results after.
  372. 13:34All of this is fed into the OpenAI
  373. 13:36emails grader.
  374. 13:37And a lot of times even though the
  375. 13:39generated SQL can be slightly different,
  376. 13:41it can produce the same results or you
  377. 13:42know, an extra column doesn't
  378. 13:43meaningfully change the answer, but all
  379. 13:46of this is taken into account with the
  380. 13:47final grader at that unit score and
  381. 13:49reason.
  382. 13:52Here are some key takeaways specific to
  383. 13:54our emails process.
  384. 13:56Exact SQL text equality is not a good
  385. 13:58representation of whether SQL email
  386. 14:00passed. So we normalize functions
  387. 14:03aliases by converting everything to its
  388. 14:05ESC representation.
  389. 14:07This helps us get around SQL syntax
  390. 14:08things like different date filtering.
  391. 14:11And when comparing result sets, we also
  392. 14:13give wiggle room for things that might
  393. 14:14not meaningfully change the answer. Like
  394. 14:16you know, sometimes a float or an int
  395. 14:18it's not meaningfully
  396. 14:20different.
  397. 14:22LLMs are also really good about
  398. 14:24reasoning about failures. So instead of
  399. 14:26being prescriptive about a certain
  400. 14:28thing, the model grader does a much
  401. 14:30better job of finding the actual
  402. 14:31differences that matter in the context
  403. 14:33of a question. And looking at the chain
  404. 14:35of thought for emails was really useful
  405. 14:36to help us more easily debug failures.
  406. 14:41Data security is something we also take
  407. 14:42very seriously at OpenAI.
  408. 14:45Users should only be accessing the data
  409. 14:47that they have a legitimate purpose
  410. 14:49business purpose to do so.
  411. 14:51When we ingest internal knowledge, we
  412. 14:52ingest sanitized queries that important
  413. 14:54IDs aren't accidentally leaked.
  414. 14:57Sometimes the users who rightfully
  415. 14:59should have access do need to see the
  416. 15:00sensitive results.
  417. 15:02And in this case, we link a web page
  418. 15:04where we check that the users have
  419. 15:06access to the underlying table. And the
  420. 15:08same permissions checking model is used
  421. 15:10when users share an agent chat.
  422. 15:12Sometimes when a user needs to access a
  423. 15:14certain table they don't have access to,
  424. 15:16they can see the right groups in the
  425. 15:17open in the open metadata UI.
  426. 15:21And just like a human, an agent can also
  427. 15:23make mistakes.
  428. 15:25This is why we stream the agent's chain
  429. 15:26of thought as it is answering a
  430. 15:28question, so at the end it can
  431. 15:31provide assumptions and the steps it
  432. 15:33could take, so you can send a check it's
  433. 15:35every move.
  434. 15:36If the agent ran any queries that
  435. 15:38resulted in the data you're seeing,
  436. 15:39those will also be linked and you can
  437. 15:40always directly click into the raw
  438. 15:42results.
  439. 15:45Now I'm going to talk about some lessons
  440. 15:46we learned along the way.
  441. 15:50It turns out that if you give the model
  442. 15:51too much information, it gets confused.
  443. 15:54For example, we have a lot of tools that
  444. 15:55we expose to users, but some are doing
  445. 15:57similar things and have overlapping
  446. 15:59functionalities, which is okay for a
  447. 16:00human who's just directly calling them
  448. 16:02via an agent, but when the agent's doing
  449. 16:04it, you know, it becomes a little bit
  450. 16:06more tricky. So restricting tool calls
  451. 16:08really helped.
  452. 16:10We also found that the model wasn't very
  453. 16:12good at consistently calling all the
  454. 16:13tools it had available, despite it being
  455. 16:15mentioned in the prompt. So combining
  456. 16:17multiple tool calls to force this also
  457. 16:19helped a lot, too.
  458. 16:22We also realized that specific
  459. 16:24instructions actually yielded when worse
  460. 16:26results. There are so many different
  461. 16:28types of questions you will ask, and
  462. 16:30while there is a similar general overall
  463. 16:32path, there's a lot of little branches
  464. 16:34and logic.
  465. 16:35So being overly prescriptive actually
  466. 16:37hurt us because the model would try to
  467. 16:39follow these exact set of instructions
  468. 16:41that didn't really make sense for the
  469. 16:42question.
  470. 16:43So changing our system prompting to be a
  471. 16:46little more general so that the agent
  472. 16:48gets a rough starting points, but then,
  473. 16:51you know, we leaving we leave the
  474. 16:52reasoning to GPT-5 really helped us
  475. 16:55because at the end of the day, it does
  476. 16:57have a lot of great context.
  477. 17:01So how can you get started if you're
  478. 17:02trying to do something similar at your
  479. 17:03own company?
  480. 17:05Firstly, the most helpful thing to us
  481. 17:07starting out was leveraging existing
  482. 17:08APIs.
  483. 17:10A lot of model platforms like OpenAI
  484. 17:12have this tool in just directly
  485. 17:13available. No need to reinvent the wheel
  486. 17:15of already exists.
  487. 17:17Uh using open metadata for feeding rich
  488. 17:19table information as part of our context
  489. 17:21layer was also helpful.
  490. 17:23Another thing we leveraged heavily when
  491. 17:24building this out was Codex. With the
  492. 17:26small team we have we had
  493. 17:29Codex was crucial in getting something
  494. 17:30out the door and iterating quickly with
  495. 17:32users and this sped up development so
  496. 17:34much and has continued to be really
  497. 17:36useful in accelerating our engineering
  498. 17:38workflows.
  499. 17:41Um so that was a lot of me talking.
  500. 17:42Thanks for not falling asleep.
  501. 17:44Um if there's anything you take away
  502. 17:46from this talk, I hope these three
  503. 17:47things stick with you. Remember that
  504. 17:49it's really important to have context
  505. 17:50beyond just table metadata. Code and
  506. 17:52rich context and company product context
  507. 17:54really goes a long way. Uh also
  508. 17:57incorporating memories really valuable
  509. 17:58so your agent can continuously improve
  510. 18:01and finally evals are the way to make
  511. 18:03sure that your model can stay
  512. 18:05consistently good.
  513. 18:07And that wraps up my presentation. Thank
  514. 18:09you so much for listening.
  515. 18:13>> That was awesome. Thanks, Bi. I
  516. 18:14appreciate you going into all the detail
  517. 18:15on that.
  518. 18:16Um we do have a couple minutes. I I've
  519. 18:19maybe a couple questions I could ask you
  520. 18:21if that's all right. So uh you talked
  521. 18:23about taking repeat queries from 22
  522. 18:26minutes to under 90 seconds and you
  523. 18:28talked about the memory and the
  524. 18:29self-learning that was crucial to that.
  525. 18:32What was the hardest part to get right?
  526. 18:33Was it
  527. 18:34the memory layer and getting the
  528. 18:35technical side set up or was like there
  529. 18:38any I don't know education you had to do
  530. 18:40with users for them to understand what
  531. 18:42was going on in the back end so that
  532. 18:43they felt comfortable getting the
  533. 18:45results. I'm just curious how much was
  534. 18:47enablement versus like the technical
  535. 18:49piece and what which was hardest for
  536. 18:50you?
  537. 18:51>> Yeah. Uh honestly there was a little bit
  538. 18:53of both. I think we definitely, you
  539. 18:55know, to have people even just adding
  540. 18:56memories and contributing there was like
  541. 18:58a little bit of user education we'd had
  542. 19:00to do just uh so that people were aware
  543. 19:02this even existed and that they could,
  544. 19:04you know, meaningfully make the model
  545. 19:06generations better. But I think uh after
  546. 19:08they understood that a lot of the heavy
  547. 19:10lifting was more in the actual uh the
  548. 19:13framework for retrieval and you know
  549. 19:14getting that right. And so I think the
  550. 19:17the combination of the two really helps
  551. 19:18with the generations.
  552. 19:20>> Awesome. Awesome.
  553. 19:22Uh last question for a data team that's
  554. 19:25nowhere near the maybe the size and
  555. 19:27scale
  556. 19:28of open AI, is there one design decision
  557. 19:31for Kepler that you'd tell that team to
  558. 19:34to copy first? Like what's the most
  559. 19:36important thing to think about if people
  560. 19:37are trying to build their own internal
  561. 19:39AI data agent?
  562. 19:41>> Yeah, that's a great question. Um I
  563. 19:44think you know I I went every company's
  564. 19:45data platform sure is is so different.
  565. 19:47And so maybe I want to maybe like one
  566. 19:49piece of advice that could be generally
  567. 19:51useful is really about um I think what
  568. 19:54was most useful for us in the beginning
  569. 19:56was quickly iterating. I think like uh
  570. 19:58you know when we started out we didn't
  571. 19:59we didn't actually know like what would
  572. 20:01have stuck. It was a lot of um
  573. 20:03iteration, trial and error. But I think
  574. 20:05the most important thing was okay, we
  575. 20:07knew that we had to get you know this
  576. 20:09There's so much data, there's a lot of
  577. 20:10curated contacts that has to be in
  578. 20:12there. Um and then from there about you
  579. 20:14know tweaking about how we got it or you
  580. 20:16know trying out different experiments is
  581. 20:18really helpful.
  582. 20:19>> Awesome. Thank you for sharing that.
  583. 20:22All right. Um
  584. 20:23thanks again for sharing this. Uh it's
  585. 20:25great to see that you know
  586. 20:28uh what a leading frontier lab does. Um
  587. 20:31saw the news about open AI filing to go
  588. 20:33public. So uh congratulations and good
  589. 20:36luck with all that. Uh so thanks again
  590. 20:38for joining us today and we'll move on
  591. 20:39to the next session and uh look forward
  592. 20:41to chatting with you later today.
  593. 20:44>> Thanks so much, Steve.
  594. 20:45Bye, folks.

About this transcript

This page contains the full transcript of Inside OpenAI's Internal AI Data Agent — Bonnie Xu, OpenAI (Summit '26) by Collate, generated from the public captions YouTube serves with the video. The transcript has 3,638 words across 594 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.