Graph RAG with Iceberg — Transcript
Full transcript
- 0:04Uh so like a quick introduction my name
- 0:07is Rajib San Gupta I am director uh
- 0:10systems engineering in AMD and I'm today
- 0:12with my colleague Amlan and Prem um so
- 0:17uh couple of years back uh you know we
- 0:20started this iceberg journey at that
- 0:22time it was very hard to push back all
- 0:23other different silos and we'll talk
- 0:26about the story but today's main
- 0:28presentation is more about the use of
- 0:31like there was a question that what you
- 0:33can do why you need an AI uh so maybe we
- 0:36will touch on that more right so we'll
- 0:38explain what we are doing right so um
- 0:42maybe a very quick introduction I will
- 0:44not bore you with AMD but one thing I
- 0:47realize that a lot of people don't know
- 0:49about the AMD's full length of we touch
- 0:52from space station space satellites to
- 0:55cars to your laptops to your data
- 0:58centers everywhere there you know just
- 1:00to give a glimpse of what uh just on the
- 1:03AI side not on everybody you know you
- 1:05have Xbox or you know your PlayStation
- 1:07is also on AMD platform but you know
- 1:10just look into the AI uh spread right um
- 1:14you have cloud which is all the GPUs
- 1:16instinct based and then HPC which is
- 1:18basically a GPU and a CPU combined is
- 1:21called the APU that's that's on the for
- 1:24the HPC and you know the the biggest one
- 1:27L capon which is uh still on the AMD
- 1:30platform, right? Actually, the number
- 1:32two is also on the AMD platform. uh to
- 1:34be honest and then in the enterprise
- 1:37side this portion you probably all know
- 1:39like that that starting from the edge to
- 1:41the compute like you know we have the
- 1:44epic based generation 5 now tins and the
- 1:48uh the uh on the other side is the
- 1:50instinct uh side of the house which I
- 1:52already covered and and of course PC and
- 1:54in the PC you know you know about the
- 1:57co-pilot PC or AIPC right so AIPC has
- 2:01like certain tops you need to have a
- 2:03co-pilot to certain tops which means is
- 2:05you need a CPU you need a GPU and you
- 2:08need an NPU then only you can reach that
- 2:11you know 55 tops or whatever you want to
- 2:13uh do for the copilot PC so that's
- 2:16that's all and then one thing which is
- 2:20like lot of people doesn't know is the
- 2:22adaptive computer and accelerate
- 2:24computing right platform what is that is
- 2:26it's a system the versel series of the
- 2:29uh chip is a system in it itself so it
- 2:32has a you know general purpose CPU which
- 2:37we call it the ARM cortex so yes we are
- 2:39also in ARMS so ARM cortex 72 then we
- 2:43have on top of it the FPGA or the logic
- 2:47uh fabric so it's it's an adaptive
- 2:49because think of it if you need
- 2:51something to accelerate you can actually
- 2:54use because it will accelerate so I will
- 2:56give you an example let's say you are
- 2:58doing a query query engine so query
- 3:00planning and all you did but now you
- 3:02have execute the query. When your
- 3:04execution plan is ready, you have to
- 3:06speed up the query engine. What you do
- 3:09today is you can create multiple
- 3:11executors, right? Instead, you can also
- 3:14use the adaptive computing because even
- 3:16though you know it the FPGF fabric works
- 3:19in a much much lower frequency, but it
- 3:21is highly optimized for that particular
- 3:23purpose. So it will accelerate you and
- 3:25then I I know that someone I met today
- 3:27from Saturn data or somebody he's doing
- 3:29a lot of acceleration of the pipelines
- 3:32using uh FPGS plus it has the AI
- 3:36accelerator which is is it has an NPO it
- 3:39has a DSP block. So in that small chip
- 3:41it is a it is everything. So just to
- 3:44give a context set if you are interested
- 3:46let me know. Okay. So now going to the
- 3:50topic right I think I will not repeat
- 3:53much of it lot of of you have already
- 3:55heard by the way uh these these slides
- 3:57are not I created these are created by
- 4:00claude when I prompted it so uh you know
- 4:02it's not mine so you will see some of
- 4:05the lags because it cannot see itself
- 4:07but at least if you give compar prompt
- 4:09it adjust the uh space so these are but
- 4:12it creates images so I was worried that
- 4:14when you present in a bigger slide it
- 4:16will kind of uh you You know the
- 4:18resolution will not be high but it looks
- 4:20okay here right. So what we are talking
- 4:22is which you all probably know by now
- 4:25that we have these silos of data and
- 4:28four years back when we started this
- 4:29iceberg journey I was being asked in the
- 4:32company hey is it a snowflake is it
- 4:34another tool you are bringing in why we
- 4:36need to change that concept of
- 4:38democratization of the iceberg was you
- 4:41know the the most important part which
- 4:42we need to tell lot of stories to kind
- 4:45of democratize it and today we are with
- 4:48the help of my team we are able to
- 4:50consider not all the different part
- 4:52wherever it's analytical requirement I
- 4:54know there are different kind of
- 4:56database requirement but if it's an
- 4:57analytical data this is the place where
- 5:00we you you kind of consolidate
- 5:02everything now the the beauty beauty is
- 5:05that once you consolate to iceberg and
- 5:08we've been using nessie rest catalog for
- 5:10a while uh it serves our purpose I know
- 5:12we talked about polaris and there are
- 5:14many other unity and and other cataloges
- 5:17but uh you know and I think it is in my
- 5:21opinion you choose anything which is
- 5:23iceberg rest as long as you it is open
- 5:25source and you get all the support you
- 5:27are good at good at it okay but you know
- 5:30the and on top of it any analytical
- 5:32engine which understand rest can work
- 5:35right so it could be you know spark dio
- 5:38like we use all these four actually you
- 5:41know um data bricks snowflake iceberg
- 5:43and spark engines as well
- 5:46going to the next I think this slide
- 5:48also I will take um sorry yeah this
- 5:52slide is kind of uh give a very high
- 5:55level that why we adopted iceberg and I
- 5:57don't have to explain this forum because
- 5:59this program is already understand that
- 6:01that's why they came in but one of the
- 6:03two major pieces is the kind of you know
- 6:07the the scalability iceberg provides is
- 6:10huge right and then the asset compliance
- 6:12and the you know fine grain access
- 6:14control yes actually in our data it's
- 6:16not that you control at the higher level
- 6:18in our data every row of data have
- 6:21access of a because we are into a
- 6:23semiconductor industry it is very
- 6:25important everything is group based so
- 6:27you you add that fine grain access
- 6:30control at the level of every row so
- 6:32that any engine which reads on top of it
- 6:35just honors it all the UDFs are defined
- 6:37so that it honor that particular it's
- 6:39like a uh you know you you member of or
- 6:42um kind of uh udfs right and it will
- 6:45honor that and it your your data will
- 6:47always be secure
- 6:49So
- 6:52uh you know now when you have a iceberg
- 6:55you need to optimize also right so the
- 6:57partitioning strategies are important
- 6:59you can have two level of you know or
- 7:02multi-level of partitioning strategy you
- 7:04need to do so that's one area your query
- 7:06optimization I'm I'm sure you all know
- 7:09all this engine we talked about
- 7:10including spark dramo snowflake they
- 7:13actually do the predicate push down
- 7:15which is basically makes you perform
- 7:18format and since these data are columnar
- 7:20park is a columnar data so you can just
- 7:23you know it's very fast if you select
- 7:25because you don't need all the data you
- 7:26just need probably one column or two
- 7:28column and if you have those filters
- 7:30which goes down as a part of this uh
- 7:33predicate push down that helps right um
- 7:36since we use do uh so people who use
- 7:39dynamic tables in snowflake we use do
- 7:42also so you use reflections and
- 7:44reflections actually is very helpful
- 7:46because you can cache and get the query
- 7:48be much much faster. Now in the table
- 7:50maintenance we struggled a bit initially
- 7:53because table maintenance is not easy.
- 7:54Even though you partition it correctly
- 7:56because in the engineering space we deal
- 7:59with pabytes of data people don't always
- 8:02kind of have a thought process in their
- 8:04partitioning strategy but they don't go
- 8:06into so we need to do constant I would
- 8:09say the hygiene of the system right we
- 8:11need to maintain like table properties
- 8:13and like you know optimization the
- 8:15vacuuming technique and then one thing
- 8:18which I realized that iceberg is not
- 8:20good at time series but in a world where
- 8:23we have lot of you know facts lot of
- 8:26fact table are time time series based
- 8:28table you know you need to have all
- 8:29these kind of data so we run our job to
- 8:33make it from fine grain to coarse grain
- 8:35but my my request for V4 and above is
- 8:38maybe we should look into make it also
- 8:40like a time series it's time series
- 8:42aware data or if you have some ideas I
- 8:45would love to learn from the from from
- 8:47the people who who knows it right
- 8:50now now the uh story begins okay this
- 8:53all kind of getting a stage. So now you
- 8:56have all the data in one platform. What
- 8:59are you going to do with it? Yes, you
- 9:01can give it stewardship, you know, based
- 9:02on the data stewardship, you give it to
- 9:04different people, they run the query
- 9:05engine, right? But then the idea is that
- 9:08if you have all this data, why not you
- 9:10use it to build a data intelligence
- 9:12platform? Okay, so here is the highlevel
- 9:16architecture of the data intelligence
- 9:18platform. At the lowest we call it PCH,
- 9:21plumbing, curating and harvesting. And
- 9:23what plumbing here means in the
- 9:25infrastructure where you bring in all
- 9:27the data lakes, create the data lake,
- 9:29bring everything in, push uh, you know,
- 9:32push the curation, you know, harmonize
- 9:34it. Those those are the part which we
- 9:36call as the plumbing. But on the top on
- 9:38the next layer is the layer of basically
- 9:42building that intelligence because think
- 9:46of let's say you have 10 tables okay now
- 9:49you want a data from one table it is
- 9:51easy right you know it but if you have
- 9:55thousands of tables and you want some
- 9:57data some AI now think about his
- 9:59question about the AI part of it if AI
- 10:02needs to learn it's not a magic right
- 10:04you put an expense as a profit or a
- 10:06frequency as Okay, column name what does
- 10:09that means that might be existing in 10
- 10:11different tables. How do you correlate
- 10:13that data? So you need to build a
- 10:15metadata. So we actually build for every
- 10:18table we have every table every column
- 10:20the schema description what this purpose
- 10:22is what this column means that we tell
- 10:25and not in a human language it is in a
- 10:27more like an LLM language because it
- 10:30understand the keywords you it doesn't
- 10:32need the whole sentence. So you don't
- 10:34have to make a very big metadata for
- 10:37each table. So in our metadata table,
- 10:39each row is a uh like it points to a
- 10:43table and its information. I'll show
- 10:46you. And at the top is the kind of uh
- 10:49you know u and I'll talk about the
- 10:51knowledge graph also in the next slide.
- 10:53But you know at the top is the
- 10:54harvesting. harvesting could be as
- 10:56simple as your dashboards which you
- 10:59create a PowerBI dashboards or you have
- 11:02a chat bots or your agentic framework
- 11:04where the agents can talk to it. That's
- 11:07all kind of in the harvesting layer. So
- 11:10I want to spend a little bit time on
- 11:13this slide. I think this creates the
- 11:15intelligence right. So with an example
- 11:19okay so let's say there are three table
- 11:22u one is the device table one is the
- 11:24application and one is the user let's
- 11:27say a user uses a device and the user
- 11:31also uses an application and then the
- 11:34application is running on that device
- 11:36let's say this is the kind of three data
- 11:38and you have gotten lot of metrics or
- 11:40data points out of it right now if you
- 11:44you build a layer on top of of it to
- 11:48explain explain it explain it to build a
- 11:51relationship. Now you are done with all
- 11:53this. You have an MCPL let's say and you
- 11:56query the uh query it to get an answer.
- 11:59Practically I tell you it works with 10
- 12:0220 30 100 tables. When it goes to
- 12:04thousands of tables you will it is very
- 12:07hard to create those relationship. you
- 12:09will join join up to two three is okay
- 12:12but what if if the join needs 10
- 12:14different tables right the context
- 12:16becomes so much confusing that it is
- 12:18very hard for an LLM to build a query
- 12:21even with the help of MCP to get a very
- 12:24realistic answer now why we are doing
- 12:26all this because
- 12:29you know LLMs are very smart why because
- 12:32they have the data of the internet they
- 12:33are trained in the internet data where
- 12:35is the next fuel of data rise it's on
- 12:39all the enterprise inside the enterprise
- 12:41we have all this data earlier they were
- 12:44silos now they are consolidated into one
- 12:46place right let's say you can to some
- 12:48extent like for example in our team IT
- 12:51we have already done it we consolidated
- 12:53all IT data into one single and all the
- 12:56pipelines thousands of pipelines are
- 12:57coming in and pushing data every you
- 13:00know based on the cadence of the you
- 13:02know the the data source right now if
- 13:05you have all this data can you build an
- 13:07intelligence on top of
- 13:09So in order to build an intelligence you
- 13:10need to be very context and ground
- 13:12aware. If you ask people can ask a lot
- 13:15of question but those questions need to
- 13:17be very grounded. How do you ground it?
- 13:19In order to ground it based on the you
- 13:21know this metadata you create a graph
- 13:24layer where the graphs comes into
- 13:26picture you know. So what happens is in
- 13:29this case like as you can see uh the
- 13:32relationship the user the device and the
- 13:35app are the nodes and the relationships
- 13:38are the edges. This is how we connect uh
- 13:41connect the create and create the whole
- 13:43graph. Now when you ask LLM to a you
- 13:47know to a very specific question then it
- 13:50is much more context grounded. So it
- 13:53gives you a very perfect answer and I
- 13:55can you can ask my team like now it
- 13:58works perfectly. There was never ever a
- 14:01single situation where we are not able
- 14:03to answer any question. So we have a
- 14:05chatbot today right that chatbot you can
- 14:08ask anything based on your permission by
- 14:10the way you know it it honors your
- 14:12permission because in the graph we have
- 14:15these properties like you know graph
- 14:17node has properties so you can define
- 14:19the properties to match whatever is in
- 14:22the data lake now there's another
- 14:24problem though how do you load the data
- 14:26in the in the graph many tools you get
- 14:29they cannot load the data in the graph
- 14:30and our our data in the icebuck table
- 14:33are coming at at a very high frequency.
- 14:35So that's why you know the graph has to
- 14:38be adaptive so that the graph can
- 14:41directly point and we don't have to have
- 14:43a different pipeline to load the load it
- 14:45in the graph and once this is done your
- 14:48LLM will work and if you want to share
- 14:50it because it's not only chatbots you
- 14:54need agentic framework your agent wants
- 14:56to talk to and make make decisions and I
- 14:58will I have a slide to cover that that
- 15:00will come out of the MCP layer. So
- 15:04um the zero ATL graph database which we
- 15:08use in this case in our case it is a
- 15:10puppy graph and I'm I'm sure uh you you
- 15:13saw it um I see some of the folks are
- 15:15here. So the reason which we chose is
- 15:18there's a zero copy you don't have to
- 15:20copy anything it is adaptive actually it
- 15:22has two mode one is the adaptive mode
- 15:24and another is the cache mode. If you
- 15:26have a sub millisecond kind of
- 15:28requirement to very fast query um of
- 15:31course you have to bound it with how
- 15:32many hops it should go. So you can
- 15:34control all that and you do it right is
- 15:37a full cache mode and it has an you know
- 15:39kind of a you know you know schema
- 15:42evaluation if your if your underlying uh
- 15:45you know table is schema evolving it
- 15:47understand it and it it load it as long
- 15:50as it has the information in the
- 15:52metadata and also it is uh you know it
- 15:55can be deployed in a k environment. So
- 15:57like in our case we have the leader node
- 15:59and the execution node. So we can have
- 16:01very high performance. If you need more
- 16:03and more performance, you can load it
- 16:05all in memory. But if you don't have
- 16:06much memory, you can move in the
- 16:08adaptive mode. So then as when the query
- 16:10comes, it will take some time initially
- 16:12to load it and and do it. But it's it's
- 16:15not that that slow. I'm just saying that
- 16:16depending upon your use case, you you
- 16:18can do it, right? Uh so this is the kind
- 16:21of tech stack uh which we do. So at the
- 16:25bottom we have minio because it's an
- 16:27on-prem implementation we are now doing
- 16:29in the cloud as well uh because there
- 16:31are some use cases there on top of it we
- 16:33have this iceberg with the park data
- 16:35file then we use nessi risk catalog and
- 16:39spark and do both actually we you know
- 16:42we are also working with snowflake to
- 16:43make it enable and then you know um the
- 16:47dio query engine so draio works also
- 16:49they have a feature I think a lot of
- 16:51engines have feature called copy into or
- 16:53merge into you know lot of these data
- 16:55like CSVs and JSON you need a SQS kind
- 16:59of hook and as soon as the data lands it
- 17:02automatically load into the table right
- 17:04and once the table is loaded the graph
- 17:06will take it and any query comes from
- 17:09the agentic side whether it's your
- 17:12agentic framework is LAN graph autogeni
- 17:14whatever you use right um you know uh or
- 17:17copilot studio so it will kind of query
- 17:21the uh you know graph using um you you
- 17:24know if you need to give it to other
- 17:25agents you can put MCP and then we have
- 17:28a a method called agent and critic. So
- 17:32we use claude for asking the it creates
- 17:35the query and then the critic always
- 17:38look into to before answering it it
- 17:41looks into have you understood the
- 17:42question because people can in in when
- 17:44they have a natural language people can
- 17:46ask anything. So you need to converge
- 17:48and make sure the data what what he's
- 17:51asking and what the what the agent is
- 17:53answering is correct. So this critic
- 17:56agents solve it. And if the critic agent
- 17:58and the and the you know the agent
- 18:00doesn't agree then after one try it goes
- 18:03back to the user say have you asked this
- 18:06question I'm not clear can you ask very
- 18:08clearly give me more context so
- 18:11otherwise if the critic says yes what
- 18:13you what you have done is correct it
- 18:15just goes and give the answer so
- 18:20uh I will stop here is there any
- 18:21question um because it is important yes
- 18:26>> access
- 18:27for your chat. have to deal with PII or
- 18:32personally identifiable information
- 18:34that
- 18:35>> yeah very good question so um I mean in
- 18:38order to have a PII we have to have a
- 18:41very every company has a very strict
- 18:44rule right in this case in our because
- 18:46we deal with lot of engineering data we
- 18:48haven't added PII but think of it if the
- 18:52if the you know your um you know
- 18:55security is built in at the at the
- 18:58bottom layer at the layer at which that
- 19:00each row of data it should percolate
- 19:02above you are not controlling anything
- 19:04above the in the graph also you are just
- 19:07it it is just follows through what is
- 19:09there in the in the data itself so
- 19:11that's why we feel like it is much more
- 19:13secure but you know and uh but we we
- 19:17particularly on our use case we don't
- 19:18deal with PI data
- 19:21>> yes
- 19:22>> so for the representation between like
- 19:24that step and a puppy graph step is that
- 19:27like
- 19:28>> is that oh there we go So for the
- 19:30representation in iceberg I guess is in
- 19:33that do step is that like already
- 19:35converted into like edges and and and
- 19:37nodes or is that going to be like is
- 19:40that what that do step is doing?
- 19:42>> Yes. So in in the in the in the graph in
- 19:44the puppy graph you have to set those
- 19:47adjacent nodes based on the metadata
- 19:49right. We are trying to develop an
- 19:51engine which we can uh you know which
- 19:53will read the metadata and continuously
- 19:55evolve. But once your schema is there in
- 19:58the graph, you don't have to do any more
- 20:00loading. [clears throat]
- 20:01>> Schema has to be built. But I think
- 20:03there is a room here for us for future
- 20:06to build the schema automatically based
- 20:08on because we already have the metadata.
- 20:10We know the what data is coming in.
- 20:12>> Awesome.
- 20:16>> Okay. So I you know um so in in summary
- 20:21you know what we get right we get a
- 20:23business impact. I'll start with the
- 20:25kind of kind of uh like the success
- 20:27criteria right so uh since we are it our
- 20:32goal is to reduce the you know um you
- 20:35know MTR right improve our quality of
- 20:38service how do you get that deflection
- 20:40of reduction of the cases one thing I
- 20:42told you that many a time we start with
- 20:45like think about AI ops but AI ops is
- 20:48like anomaly detection right anomaly
- 20:51detection is too noisy that's why it has
- 20:53not got into it is a 10, 15, 20 years
- 20:56old technology but it has not got into
- 20:58success. Now with all this engine what
- 21:02you can do is your agent can
- 21:04continuously look into various data set.
- 21:06Let's say you are getting from multiple
- 21:09data sets of your IT. You look into the
- 21:12anomaly. You create a baseline. You look
- 21:14into the anomalies. You don't just shout
- 21:16those anomalies or send those anomalies.
- 21:18That will be too much of noise. But you
- 21:20start learning it and create a graph
- 21:22again. It's called a dependency graph.
- 21:25And now once your dependency graph is
- 21:27built, you validate it in next
- 21:30iteration. In six to eight months if the
- 21:34events are happening correctly or it may
- 21:35take a little bit more time your your
- 21:38graph will adapt. It is not a human
- 21:39which is building it is built built by
- 21:42the system itself by learning the
- 21:44relationship that if this goes down oh I
- 21:46see when there is anomaly here there is
- 21:48an anomaly there there's anomaly in
- 21:50other places which means they are
- 21:52correlated let's first understand that
- 21:53these are correlated next step hey which
- 21:56was created at first where you got the
- 21:58the thing as first right and slowly it
- 22:01will learn that and then we are actually
- 22:03doing a you know a paper on that to kind
- 22:07of learn and and develop So if the paper
- 22:09published I will share with you but you
- 22:11know that paper talks about how those
- 22:13dependency graphs are created and then
- 22:16how do you do the anomaly detection uh
- 22:18based on that you know based on that
- 22:21graph right similarly you know um for uh
- 22:25standard operating procedure let's say
- 22:26you do certain work which is very
- 22:29mundane self-healing right for example
- 22:32if the machine kernel just hunks then
- 22:35the machine needs to be rebooted you
- 22:36cannot do anything So you know these are
- 22:39very standard procedure today you can
- 22:41make an agent to do all that how they
- 22:43will take the decision they can take
- 22:45decision but if they need more
- 22:46information to to triage they will go to
- 22:49that platform query it and get it right.
- 22:52So those are the kind of uh you know
- 22:55success criteria we we are we are
- 22:56thinking and maybe because of this also
- 22:59we are making faster decision you know
- 23:01in a in a leadership wants to take like
- 23:03a view of something let's say how much
- 23:05is my cost of doing this business or
- 23:08this particular service
- 23:10earlier they have to go to a dashboard
- 23:12no I don't want this I want something
- 23:13else they have to recreate or you know
- 23:16tell someone hey get me this data it
- 23:18takes a lot of iteration he can go and
- 23:21now do it a query and you know work with
- 23:24the query engine and and get all the
- 23:25data. That's the value of what we get in
- 23:28the you know with the AI. It's not just
- 23:31a simple data you know it's not a
- 23:33semantic it's just a semantic layer
- 23:35creation. It is more about getting the
- 23:37intelligence in the hand in the
- 23:39fingertip. We call it like a crystal
- 23:40balling with your data right and so what
- 23:44is the next step right so next step you
- 23:46know in in our is like there is an I
- 23:50think there's a talk about multimodel t
- 23:52so one one of the main challenge of is
- 23:54the multimodel t because text is fine
- 23:57but there are a lot of images we need to
- 23:59deal with for example your invoices
- 24:01let's say in conquer a lot of people put
- 24:03expense those are you know images how do
- 24:06you take those images build an AI
- 24:09pipeline which can extract the
- 24:10properties keep it in in iceberg table
- 24:14load it in the graph but you know today
- 24:16I have to store it in a vector database
- 24:19but maybe in future we don't have to do
- 24:22that so that's that's one of the you
- 24:24know area and in terms of uh the whole
- 24:27agentic ecosystem these are the four
- 24:29when we work with I already talked about
- 24:31AI ops digital twins what is digital
- 24:33twin means whatever you do let's say
- 24:36somebody else is doing for you for
- 24:38example Example, let's say you train
- 24:41somebody, what do you look for? For
- 24:43example, a presentation. Uh let's say
- 24:45you get a presentation, what do you look
- 24:46for? I look for only three things. What
- 24:49is the, you know, when my team does it,
- 24:51right? What is this? You know, what is
- 24:54the problem statement? What you're
- 24:55trying to solve? What is the solution
- 24:57provided and what they need from me? If
- 24:59these three things is not clear, I
- 25:02should send an email saying that hey,
- 25:04can you highlight these three things in
- 25:06your presentation? I don't have to look
- 25:08into my presentation or I just give an
- 25:10idea and like cloud is making those
- 25:12images of my uh thought process it will
- 25:15also create presentation and send it
- 25:17based on what I I prompted right so you
- 25:20know like a detail or in your simple
- 25:22laptop right if it is like your
- 25:24powerpoint or something is hung it'll
- 25:26tell you hey your generally you take
- 25:28this much of memory today it is taking
- 25:30more memory can I restart it can you
- 25:33save your data or at your evening when
- 25:35you are sleeping keep your machine on. I
- 25:37will I will do those uh you know hygiene
- 25:40stuff in your laptop right so those kind
- 25:43of things is there self-healing I
- 25:45already talked about and then the
- 25:47agentic workflow we believe that soon
- 25:50processes will be gone why we create the
- 25:52process because there is a guideline uh
- 25:55you want to so that humans can focus on
- 25:57it you want to give access to it for
- 25:59that there's approval there are 10 10 n
- 26:02yards to complete if this can be taught
- 26:05to an agent it can do it itself right
- 26:08you don't have to have the process so
- 26:09these are the kind of areas of focus for
- 26:12for us in AMD that's all um open for
- 26:16more question if time permits
- 26:21you're prompting
- 26:24you are giving a prompting for that
- 26:26workflow right uh the last option
- 26:29>> yeah yeah yeah so so workflow is like
- 26:33let's say um so thank you so workflow is
- 26:35like uh let's say you want um a simple
- 26:39workflow could be an NDA when you when I
- 26:42let's say went with puppy graph so we
- 26:45have an NDA you need to share so that we
- 26:47can start the journey right
- 26:49>> I have to put all these details which he
- 26:52has sent put it in a system somebody
- 26:54will look into it they have an approval
- 26:56process goes through it and then there
- 26:59is not much in the process it's a very
- 27:01standard process but it still takes two
- 27:03days to complete Right.
- 27:05>> Yeah.
- 27:05>> So you know it can be done by an agent.
- 27:08So that's why I'm saying that those
- 27:10workflows which is like you do it's very
- 27:13like five things five steps you do and
- 27:15those five testes there's nothing um I
- 27:18mean I mean great you know great about
- 27:20it. It is just that you are following
- 27:21that process or not. If you can do it by
- 27:24an agent then you don't need an human
- 27:26for it.
- 27:27>> So autohealing is is a debugging purpose
- 27:30or also operation purpose.
- 27:32>> It's it's both right. So um okay like
- 27:36let's say the example um so you have a
- 27:39server let's say a a graph server right
- 27:43and that server is not behaving properly
- 27:46okay because what are the signals the
- 27:48latency is increasing number of errors
- 27:50coming is increasing right it will
- 27:52traditionally it will open a ticket a a
- 27:56person will look into those tickets
- 27:58right and then they will debug it and
- 28:00then they will fix it right But let's
- 28:04say this in the in in the data
- 28:06intelligence platform we have all the
- 28:08tickets. We know what was done 10 times
- 28:1020 times before. It's nothing new. So it
- 28:13will get that knowledge do all the
- 28:16checks and fix it automatically.
- 28:18>> Okay,
- 28:18>> that's that's the auto healing. It
- 28:20connects to the GitHub repository and
- 28:22the previously deployed means what I
- 28:25rise this question is uh I think two
- 28:28three days before I saw the Google Uli
- 28:31uh something they're giving the that's
- 28:33the service for debugging and all those
- 28:36>> yeah in our case the data intelligence
- 28:38platform has all the we are continuously
- 28:41getting all the service now tickets and
- 28:42every every other tickets Jira tickets
- 28:44and everything so it has all the
- 28:46intelligence it know you can search that
- 28:47if there is a black screen on a ETX node
- 28:50what are the probable problems in AMD
- 28:53right it will go search all that and say
- 28:56most likely this is the problem if not
- 28:58this is the problem so you you as a
- 29:00human also do that that that okay this
- 29:02problem let me check it whether this
- 29:04this is the problem or not and there are
- 29:06certain uh like uh you know um signals
- 29:10or trends through which you debug that
- 29:13knowledge is stored here so you can so
- 29:15that agents can do
- 29:16>> so I think you build the knowledge graph
- 29:18>> yes Yes. Yes. I guess.
- 29:22>> Thank you.
- 29:24>> Thank you. [applause]
About this transcript
This page contains the full transcript of Graph RAG with Iceberg by Apache Iceberg™ Meetup, generated from the public captions YouTube serves with the video. The transcript has 5,409 words across 718 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.