Inside OpenAI's Internal AI Data Agent — Bonnie Xu, OpenAI (Summit '26) — Transcript
Full transcript
- 0:00I'm really excited to
- 0:02soon welcome our next speaker. We have
- 0:04Bonnie Shi who will be walking us
- 0:06through OpenAI's internal data agent
- 0:08Kepler which Suresh mentioned. And that
- 0:11supports thousands of internal users
- 0:13within OpenAI.
- 0:15Uh and just to give a little background
- 0:17on Bonnie, she's a software engineer and
- 0:19the tech lead of the data productivity
- 0:21team at OpenAI where she built an
- 0:23AI-powered data analytics agent from the
- 0:25ground up to help other teams at OpenAI
- 0:27explore and understand data more
- 0:29intelligently.
- 0:31Before joining OpenAI, OpenAI, she spent
- 0:33four years at Stripe working on the data
- 0:36platform and previously held engineering
- 0:38roles at both Meta and Google.
- 0:41Her work focuses on building scalable
- 0:43systems that integrate AI and data to
- 0:46make analysis faster and more
- 0:48accessible. So, if you would, please
- 0:50join me in welcoming Bonnie Shi to the
- 0:53stage.
- 0:57Drumroll, please. Bonnie, there she is.
- 1:00All right.
- 1:02So good to see you, Bonnie. Thanks for
- 1:03joining us today.
- 1:05>> Hello. Thanks for the kind introduction,
- 1:07Steve. So good to be here with you all
- 1:09today.
- 1:09>> Absolutely. Thanks for making the time
- 1:11to share your perspective with us. I'm
- 1:12going to jump off. I'll turn things over
- 1:14to you to take it away. Thanks, Bonnie.
- 1:16>> Awesome. Um so, hello everyone. My name
- 1:19is Bonnie and today I'll be talking
- 1:21about how OpenAI deployed AI data agents
- 1:24to help our teams answer questions about
- 1:27their data.
- 1:30So, let me paint you a picture. Your
- 1:31business lead comes to you and asks you
- 1:33the question, "How many ChatGPT Pro
- 1:35users do we have in Italy?" You consult
- 1:38a data scientist, but they're like,
- 1:40"Actually, this is kind of hard. Let me
- 1:42get back to you."
- 1:44They don't know what table to look at,
- 1:45so they ask another engineer. But the
- 1:47other engineer doesn't know which table
- 1:49is the right one.
- 1:50So, after three code deep dives, two
- 1:52quick meetings, and five Slack threads,
- 1:55we finally have an answer.
- 1:57Simple questions like these shouldn't be
- 1:59this consuming and this time
- 2:02you know, time taking, but they are.
- 2:05And the reason why this is so hard is
- 2:07because there's so much data.
- 2:09So, I'm Bonnie and I'm here today to
- 2:11talk to you about how we solved this
- 2:12problem.
- 2:16Let me start off with an overview of the
- 2:18scale of our data platform to illustrate
- 2:20why we need an AI data agent.
- 2:22Then I'll go into implementation
- 2:24specifics, the learnings we had, and the
- 2:27next steps we're planning to take.
- 2:31At OpenAI, 80% of the company directly
- 2:34uses our data platform. That's 80% of
- 2:37the company using our team's 15 tools to
- 2:40process over 600 petabytes of data a day
- 2:43or about 70K data sets.
- 2:46The data is growing rapidly.
- 2:48This means we have so many more
- 2:50questions to answer.
- 2:52But there's more and more data to sift
- 2:53through to get the right result.
- 2:56When ChatGPT launched in 2022, we were
- 2:58asking ourselves, "How many users do we
- 3:01have?"
- 3:02As the product has evolved, we've
- 3:04introduced more regions, different
- 3:06plans, more features, and now the
- 3:07question is, "How many daily active
- 3:10instant checkout users do we have in New
- 3:12York?"
- 3:13That's a much harder question to answer
- 3:15now, but fundamentally, we're looking
- 3:17for the same type of answer.
- 3:21One of the reasons why this becomes a
- 3:23lot harder is because table discovery is
- 3:25a lot harder at scale.
- 3:27A common problem is having difficulty
- 3:29finding the right table to use because
- 3:31there's a lot of similarly sounding
- 3:32tables and it's unclear what data is in
- 3:35them.
- 3:36Another common issue is having trouble
- 3:38understanding the subtle nuances of each
- 3:40table.
- 3:41This is a hard problem because some
- 3:43tables have encrypted IDs, some tables
- 3:46have unencrypted IDs, but we still might
- 3:48want to join them.
- 3:50Some tables have columns that adjust for
- 3:51fraud rates, some tables don't.
- 3:54Some tables are pre-filtered by
- 3:56feedback, some are not.
- 3:58Missing one nuance can lead to an answer
- 4:00that is wrong by an order of magnitude.
- 4:03This can be catastrophic when making
- 4:04important business decisions.
- 4:07Not to mention, writing SQL is hard. Who
- 4:10can remember all the different ways we
- 4:11date format, or write performing
- 4:13queries, or aggravating fact that Trino
- 4:16arrays are one indexed?
- 4:20So, what's a better way of doing this?
- 4:25At OpenAI, we built an internal AI data
- 4:27analyst that takes the full context of
- 4:29data platform and answers these
- 4:31questions for you.
- 4:36At its core, the agent back-end service
- 4:38leverages the model to produce
- 4:40AI-powered results wherever you need
- 4:42them. This could be in Slack for
- 4:44answering your business questions, or in
- 4:46your IDE when you're building your data
- 4:48pipelines, or in the web agents when
- 4:50you're looking up different tables, or
- 4:52when you're writing workflows and need
- 4:54data to inform the next action.
- 4:56Let's go through an example.
- 5:00Suppose I'm a data scientist and I'm
- 5:02investigating a huge upward spike in
- 5:04user growth in ChatGPT late March 2025.
- 5:10The first step in reasoning is to check
- 5:12that the spike is real and as I
- 5:14described. So, as we can see here, the
- 5:17agent looks at the actual table that
- 5:18stores weekly active user data to get
- 5:21the exact numbers pre- and post-spike.
- 5:26To make sure it's getting the right
- 5:27table, the agent looks at internal
- 5:29knowledge to understand it. In this
- 5:31case, it was able to find extra context
- 5:33on the table from a Notion doc. It also
- 5:35cross-checked against a dashboard.
- 5:40As the agent is interactively exploring
- 5:42the data, it's checking the different
- 5:44types of data it can get.
- 5:46Here, you can tell it's looking at the
- 5:47different data dimensions like plan type
- 5:49and off status and running queries
- 5:51against those data different data
- 5:53dimensions.
- 5:54This helps the agent get more data and
- 5:56needed for analysis. Kind of slicing and
- 5:58dicing the data uh just like you might
- 6:00if you're doing this exploration.
- 6:04As part of the analysis, the agent is
- 6:06trying to come up with reasons why this
- 6:08is happening. In this particular case,
- 6:11it's exploring whether the spike was
- 6:12actually a logging bug due to
- 6:14duplication. These happen.
- 6:20After some analysis, the agent arrives
- 6:22at the conclusion that it was due to the
- 6:24image gen release.
- 6:26It validates this conclusion by using
- 6:28web search to pull up release stocks.
- 6:32And woohoo, this was in fact the right
- 6:34reason. The spike was due to the viral
- 6:37uh self anime image generation trend.
- 6:42So, how does this magic happen?
- 6:47This is a big overall picture diagram of
- 6:50how things happen behind the scenes.
- 6:52Let's go through each part.
- 6:54On the left are the entry points. On the
- 6:56top are the offline curated knowledge
- 6:58pieces. On the bottom are the sync API
- 7:01calls that the agent can make to our
- 7:02data platform tools. And in the middle
- 7:05is the most important piece. It's the
- 7:07agent talking to the model.
- 7:12Agents without context can give wildly
- 7:14wrong answers.
- 7:16As pointed out in the amazing earlier
- 7:17presentation.
- 7:20Take this example where somehow the
- 7:22agent thought there were 5 million chat
- 7:24GPT users compared to the actual 800
- 7:27million answer announced at Dev Day.
- 7:29Just a minor rounding detail. That's
- 7:32all.
- 7:35Here are the different layers of context
- 7:37that our internal data agent uses.
- 7:40Let's start at the bottom with table
- 7:42metadata context at layer one.
- 7:48As you might have guessed, fitting all
- 7:5070k tables with their schemas and query
- 7:52history is way too much data to fit in a
- 7:55model's context window. So, we need to
- 7:57do some preprocessing ahead of time.
- 7:59Table schema information is inserted so
- 8:02the model knows how to query each table,
- 8:04what columns are available with what
- 8:06types.
- 8:07All this rich table metadata information
- 8:09is surfaced from Open Metadata and is
- 8:11made available to the agent.
- 8:14Schemas alone aren't enough to
- 8:15understand the semantics and
- 8:16relationships between the data.
- 8:18So, we use table query history, table
- 8:20lineage, and code derived table
- 8:22information to provide this extra
- 8:24context.
- 8:25All of this is transformed into an
- 8:27embedding so that when the agent is
- 8:28writing, we do rag for every table.
- 8:31The agent can get information for a
- 8:33specific table or just semantic search
- 8:35over keywords.
- 8:36So, that's layer one.
- 8:38Layer two is human annotations, which
- 8:40can be incredibly especially when
- 8:42starting out and for key tables and can
- 8:44be collaboratively done via editing
- 8:46table descriptions, for example, in the
- 8:48Open Metadata UI. So, here you we you
- 8:51can see there's all these great table
- 8:52bits like the schema, you know, the
- 8:54queries, lineage, so on and so forth.
- 9:00We use Open Metadata's APIs for both
- 9:02online and offline retrieval.
- 9:04In our internal knowledge indexing, we
- 9:06pull useful bits like query history,
- 9:08descriptions, tags, usage, and in online
- 9:11cases, the agent can do table search or
- 9:13live pull groups needed for permissions.
- 9:16It's been really useful for us to have
- 9:18Open Metadata sit on top of our data
- 9:19warehouse for all this extra table
- 9:21context that is easy to find for both
- 9:24agents and humans.
- 9:27Another thing that people often point
- 9:30out about table metadata is that
- 9:32descriptions can get easily outdated and
- 9:34often become a burden to maintain.
- 9:37And this leads to the agent getting bad
- 9:38results.
- 9:39But it can be really tedious if you're
- 9:41manually updating them, especially when
- 9:43you have so many tables.
- 9:45So, we solve this by auto generating as
- 9:47much as possible.
- 9:49And we also include information beyond
- 9:51what might be available in just
- 9:52metadata, and that's what makes
- 9:53generations even better.
- 9:56And this part is layer three, the Codex
- 9:57enrichment.
- 9:59It's not enough to look at a table by
- 10:01itself as is. You need to really
- 10:03understand how a table is created, where
- 10:05it came from, and this is really the
- 10:07secret to the agent truly understanding
- 10:09the differences between tables, knowing
- 10:11that a table is filtered down because it
- 10:13only came from a subset of blogs.
- 10:15And we achieve this by running an
- 10:17offline job that generates Codex tasks
- 10:19for certain tables.
- 10:21These Codex tasks are launched in
- 10:22parallel and crawl the code base to
- 10:24understand things like a table's
- 10:25purpose, downstream usage patterns, the
- 10:28exact table grain, and primary keys.
- 10:30So, instead of only knowing that a table
- 10:32is for ChatGPT analytics, you know that
- 10:34it's a table that includes only
- 10:36first-party ChatGPT traffic, and it's
- 10:38enriched by safety sources. You know
- 10:40that these fields are null and when
- 10:42these signals are missing.
- 10:46For layer four, institutional knowledge,
- 10:48we have an internal knowledge service at
- 10:50OpenAI that ingests internal knowledge
- 10:52like Slack threads and docs.
- 10:54When fetching, we cache and do
- 10:56permissions checking so that results are
- 10:58fast and gated correctly.
- 11:00This way, we can efficiently get all the
- 11:02right company-specific context to
- 11:04accompany the data.
- 11:06Concretely, this means that when you see
- 11:08a dip in weekly active users, the agent
- 11:10can pull the Slack thread and instant
- 11:12channel that talks about the outage.
- 11:14And this is really helpful when doing
- 11:16the analysis for providing the extra
- 11:17context on the why rather than just the
- 11:20what.
- 11:21Level five is memory. Memory is useful
- 11:24for things like corrections and
- 11:26learnings.
- 11:27Here's an example.
- 11:28Without memory, the agent takes way
- 11:31longer to figure out how to answer a
- 11:33question correctly and has to work for
- 11:35much longer.
- 11:37Memory is the mechanism that helps the
- 11:39agent continuously improve and learn.
- 11:42Context will get you 80 to 90% of the
- 11:44way there, but sometimes you need these
- 11:46final little corrections to get the
- 11:47right answer.
- 11:49For example, you just need to know that
- 11:50a particular type of tea product is
- 11:52filtered by this random string. That's
- 11:54just the value and that's just how it
- 11:56is.
- 11:57We make memories easily available and
- 11:59accessible, so anyone can contribute to
- 12:01make the agent better.
- 12:03In a conversation, the agent itself or a
- 12:05human can save a memory. And there's two
- 12:07scopes, global or user. These scopes are
- 12:10important because users might want their
- 12:12own customizations or the data is
- 12:14potentially sensitive. And based on a
- 12:16conversation, the agent will also
- 12:18proactively suggest memories and users
- 12:20can confirm their insertions.
- 12:22And if all of this offline information
- 12:24isn't enough, we also have our last
- 12:27layer, which is runtime context. The
- 12:29agent can make live calls to our data
- 12:31warehouse or any other data platform
- 12:33service it needs. So we can talk to
- 12:35Spark or Airflow or open metadata. Uh
- 12:37it's just a short API call away.
- 12:40One example of when you might use this
- 12:42is if a table is new. So, you know, you
- 12:44generated this new testing table and you
- 12:46want to compare it with an existing
- 12:47table and the agent can just directly
- 12:50run those queries live.
- 12:52And I just really want to drive home
- 12:53this point that models are really smart,
- 12:56but they're not the full answer and
- 12:58context is really what makes the
- 13:00difference.
- 13:02The next important factor to consider is
- 13:04how we ensure we don't cause
- 13:06regressions. And let me talk about how
- 13:08we measure response quality.
- 13:10Our evals consists of sets of
- 13:13question-answer pairs. Question is
- 13:15usually some important metric we want to
- 13:17get right and then we have manually
- 13:19curated um expected SQLs that will
- 13:22generate the correct answer. And so we
- 13:24hit our agent query generation endpoint
- 13:26to it a natural language question to
- 13:28generate SQL and run the query and then
- 13:30we do the same with the expected SQL and
- 13:32we compare the results after.
- 13:34All of this is fed into the OpenAI
- 13:36emails grader.
- 13:37And a lot of times even though the
- 13:39generated SQL can be slightly different,
- 13:41it can produce the same results or you
- 13:42know, an extra column doesn't
- 13:43meaningfully change the answer, but all
- 13:46of this is taken into account with the
- 13:47final grader at that unit score and
- 13:49reason.
- 13:52Here are some key takeaways specific to
- 13:54our emails process.
- 13:56Exact SQL text equality is not a good
- 13:58representation of whether SQL email
- 14:00passed. So we normalize functions
- 14:03aliases by converting everything to its
- 14:05ESC representation.
- 14:07This helps us get around SQL syntax
- 14:08things like different date filtering.
- 14:11And when comparing result sets, we also
- 14:13give wiggle room for things that might
- 14:14not meaningfully change the answer. Like
- 14:16you know, sometimes a float or an int
- 14:18it's not meaningfully
- 14:20different.
- 14:22LLMs are also really good about
- 14:24reasoning about failures. So instead of
- 14:26being prescriptive about a certain
- 14:28thing, the model grader does a much
- 14:30better job of finding the actual
- 14:31differences that matter in the context
- 14:33of a question. And looking at the chain
- 14:35of thought for emails was really useful
- 14:36to help us more easily debug failures.
- 14:41Data security is something we also take
- 14:42very seriously at OpenAI.
- 14:45Users should only be accessing the data
- 14:47that they have a legitimate purpose
- 14:49business purpose to do so.
- 14:51When we ingest internal knowledge, we
- 14:52ingest sanitized queries that important
- 14:54IDs aren't accidentally leaked.
- 14:57Sometimes the users who rightfully
- 14:59should have access do need to see the
- 15:00sensitive results.
- 15:02And in this case, we link a web page
- 15:04where we check that the users have
- 15:06access to the underlying table. And the
- 15:08same permissions checking model is used
- 15:10when users share an agent chat.
- 15:12Sometimes when a user needs to access a
- 15:14certain table they don't have access to,
- 15:16they can see the right groups in the
- 15:17open in the open metadata UI.
- 15:21And just like a human, an agent can also
- 15:23make mistakes.
- 15:25This is why we stream the agent's chain
- 15:26of thought as it is answering a
- 15:28question, so at the end it can
- 15:31provide assumptions and the steps it
- 15:33could take, so you can send a check it's
- 15:35every move.
- 15:36If the agent ran any queries that
- 15:38resulted in the data you're seeing,
- 15:39those will also be linked and you can
- 15:40always directly click into the raw
- 15:42results.
- 15:45Now I'm going to talk about some lessons
- 15:46we learned along the way.
- 15:50It turns out that if you give the model
- 15:51too much information, it gets confused.
- 15:54For example, we have a lot of tools that
- 15:55we expose to users, but some are doing
- 15:57similar things and have overlapping
- 15:59functionalities, which is okay for a
- 16:00human who's just directly calling them
- 16:02via an agent, but when the agent's doing
- 16:04it, you know, it becomes a little bit
- 16:06more tricky. So restricting tool calls
- 16:08really helped.
- 16:10We also found that the model wasn't very
- 16:12good at consistently calling all the
- 16:13tools it had available, despite it being
- 16:15mentioned in the prompt. So combining
- 16:17multiple tool calls to force this also
- 16:19helped a lot, too.
- 16:22We also realized that specific
- 16:24instructions actually yielded when worse
- 16:26results. There are so many different
- 16:28types of questions you will ask, and
- 16:30while there is a similar general overall
- 16:32path, there's a lot of little branches
- 16:34and logic.
- 16:35So being overly prescriptive actually
- 16:37hurt us because the model would try to
- 16:39follow these exact set of instructions
- 16:41that didn't really make sense for the
- 16:42question.
- 16:43So changing our system prompting to be a
- 16:46little more general so that the agent
- 16:48gets a rough starting points, but then,
- 16:51you know, we leaving we leave the
- 16:52reasoning to GPT-5 really helped us
- 16:55because at the end of the day, it does
- 16:57have a lot of great context.
- 17:01So how can you get started if you're
- 17:02trying to do something similar at your
- 17:03own company?
- 17:05Firstly, the most helpful thing to us
- 17:07starting out was leveraging existing
- 17:08APIs.
- 17:10A lot of model platforms like OpenAI
- 17:12have this tool in just directly
- 17:13available. No need to reinvent the wheel
- 17:15of already exists.
- 17:17Uh using open metadata for feeding rich
- 17:19table information as part of our context
- 17:21layer was also helpful.
- 17:23Another thing we leveraged heavily when
- 17:24building this out was Codex. With the
- 17:26small team we have we had
- 17:29Codex was crucial in getting something
- 17:30out the door and iterating quickly with
- 17:32users and this sped up development so
- 17:34much and has continued to be really
- 17:36useful in accelerating our engineering
- 17:38workflows.
- 17:41Um so that was a lot of me talking.
- 17:42Thanks for not falling asleep.
- 17:44Um if there's anything you take away
- 17:46from this talk, I hope these three
- 17:47things stick with you. Remember that
- 17:49it's really important to have context
- 17:50beyond just table metadata. Code and
- 17:52rich context and company product context
- 17:54really goes a long way. Uh also
- 17:57incorporating memories really valuable
- 17:58so your agent can continuously improve
- 18:01and finally evals are the way to make
- 18:03sure that your model can stay
- 18:05consistently good.
- 18:07And that wraps up my presentation. Thank
- 18:09you so much for listening.
- 18:13>> That was awesome. Thanks, Bi. I
- 18:14appreciate you going into all the detail
- 18:15on that.
- 18:16Um we do have a couple minutes. I I've
- 18:19maybe a couple questions I could ask you
- 18:21if that's all right. So uh you talked
- 18:23about taking repeat queries from 22
- 18:26minutes to under 90 seconds and you
- 18:28talked about the memory and the
- 18:29self-learning that was crucial to that.
- 18:32What was the hardest part to get right?
- 18:33Was it
- 18:34the memory layer and getting the
- 18:35technical side set up or was like there
- 18:38any I don't know education you had to do
- 18:40with users for them to understand what
- 18:42was going on in the back end so that
- 18:43they felt comfortable getting the
- 18:45results. I'm just curious how much was
- 18:47enablement versus like the technical
- 18:49piece and what which was hardest for
- 18:50you?
- 18:51>> Yeah. Uh honestly there was a little bit
- 18:53of both. I think we definitely, you
- 18:55know, to have people even just adding
- 18:56memories and contributing there was like
- 18:58a little bit of user education we'd had
- 19:00to do just uh so that people were aware
- 19:02this even existed and that they could,
- 19:04you know, meaningfully make the model
- 19:06generations better. But I think uh after
- 19:08they understood that a lot of the heavy
- 19:10lifting was more in the actual uh the
- 19:13framework for retrieval and you know
- 19:14getting that right. And so I think the
- 19:17the combination of the two really helps
- 19:18with the generations.
- 19:20>> Awesome. Awesome.
- 19:22Uh last question for a data team that's
- 19:25nowhere near the maybe the size and
- 19:27scale
- 19:28of open AI, is there one design decision
- 19:31for Kepler that you'd tell that team to
- 19:34to copy first? Like what's the most
- 19:36important thing to think about if people
- 19:37are trying to build their own internal
- 19:39AI data agent?
- 19:41>> Yeah, that's a great question. Um I
- 19:44think you know I I went every company's
- 19:45data platform sure is is so different.
- 19:47And so maybe I want to maybe like one
- 19:49piece of advice that could be generally
- 19:51useful is really about um I think what
- 19:54was most useful for us in the beginning
- 19:56was quickly iterating. I think like uh
- 19:58you know when we started out we didn't
- 19:59we didn't actually know like what would
- 20:01have stuck. It was a lot of um
- 20:03iteration, trial and error. But I think
- 20:05the most important thing was okay, we
- 20:07knew that we had to get you know this
- 20:09There's so much data, there's a lot of
- 20:10curated contacts that has to be in
- 20:12there. Um and then from there about you
- 20:14know tweaking about how we got it or you
- 20:16know trying out different experiments is
- 20:18really helpful.
- 20:19>> Awesome. Thank you for sharing that.
- 20:22All right. Um
- 20:23thanks again for sharing this. Uh it's
- 20:25great to see that you know
- 20:28uh what a leading frontier lab does. Um
- 20:31saw the news about open AI filing to go
- 20:33public. So uh congratulations and good
- 20:36luck with all that. Uh so thanks again
- 20:38for joining us today and we'll move on
- 20:39to the next session and uh look forward
- 20:41to chatting with you later today.
- 20:44>> Thanks so much, Steve.
- 20:45Bye, folks.
About this transcript
This page contains the full transcript of Inside OpenAI's Internal AI Data Agent — Bonnie Xu, OpenAI (Summit '26) by Collate, generated from the public captions YouTube serves with the video. The transcript has 3,638 words across 594 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.