AIOps: Leveraging AI for Incident Root Cause Analysis - Sathish Kumar — Transcript
Full transcript
- 0:13I'm Satish and I'm a senior principal
- 0:15engineer in Jira service management at
- 0:17Atlassian.
- 0:18Uh we work on a lot of products around
- 0:21support services, alert management,
- 0:23incident management, and AI ops. So,
- 0:24today I'll be talking about how we do AI
- 0:27ops in in one of our sub products. And
- 0:29and how we leverage AI effectively for
- 0:31uh
- 0:32automated alert management, automated
- 0:34incident management, right?
- 0:36The agenda of the talk will roughly
- 0:38cover the following topics.
- 0:40Like uh
- 0:41Like most of us will have some idea of
- 0:43what is a software instance, what is a
- 0:44root cause analysis, but I'll still
- 0:46start with a brief intro before going
- 0:48into the details. So, we'll talk about
- 0:50what are software incidents, what is RCA
- 0:53or root cause analysis, and the role
- 0:55that AI can play in in a root cause
- 0:57analysis both both in terms of like
- 1:00product in terms of in terms of
- 1:01operations, right? Like I'll be coming
- 1:03from the perspective of how we are
- 1:04building this into our products, but uh
- 1:06it's it's applicable even if you're like
- 1:08running production systems and you want
- 1:09to use it outside of the product, right?
- 1:11And uh I'll talk about the building
- 1:13blocks of uh how we've approached an AI
- 1:16SRE agent and the challenges that uh we
- 1:19have encountered along the way.
- 1:22So, what are software incidents?
- 1:25Software incidents are usually like
- 1:27unplanned interruptions and uh
- 1:29degradations or any kind of abnormal
- 1:30behavior in your software services and
- 1:33applications. It usually manifests in
- 1:35the form of some reliability or
- 1:37performance or functional issues to your
- 1:39customers.
- 1:41And a software incident usually goes
- 1:43through these following phases.
- 1:46So, you have a problem that is detected.
- 1:48Most of the time the problem is detected
- 1:49by your monitoring systems, but if your
- 1:51alerting or monitoring is weak,
- 1:53it can also be detected by let's say
- 1:55customer issues or support desk or even
- 1:57like if your product is like super
- 1:59popular, it can even be discovered on
- 2:01Twitter by your users and complain,
- 2:02right?
- 2:03And
- 2:04it results in some kind of customer
- 2:06impact and the the most important thing
- 2:09in any software incident is to
- 2:12reduce the customer impact or address
- 2:13the bleeding.
- 2:15Most of these incidents usually
- 2:16translate into something negative for
- 2:18the business. Like it can be revenue
- 2:19loss, it can be reputation loss, or some
- 2:21kind of like in worst cases even like
- 2:23data loss, right?
- 2:25And any software company or tech company
- 2:28usually goes through the following
- 2:29phases in order to effectively manage an
- 2:31incident. Like it it varies from
- 2:34organization to organization. The the
- 2:36rigor varies, the process varies, the
- 2:37tools vary, right? But still
- 2:40some of the common things are
- 2:41there's an investigation step. There is
- 2:44a effort to identify the potential root
- 2:46causes.
- 2:47Then we try to mitigate the issue. Like
- 2:49short-term mitigation is usually done
- 2:52with a goal of
- 2:53arresting the customer impact without
- 2:56like even if you don't find the actual
- 2:57root cause, right?
- 2:59And um
- 3:00once that is done, we want to update our
- 3:02stakeholders and we can spend the next
- 3:05few hours or few days finding out the
- 3:08actual root cause.
- 3:10I'll I'll go into details of what is
- 3:12root cause analysis and why it's
- 3:14important in this context. So, I'll be
- 3:16using the word root cause analysis to
- 3:17mean two things.
- 3:21Yeah.
- 3:22So, I'll be using the word root cause
- 3:23analysis in two different contexts. One
- 3:25is from the perspective of during an
- 3:27incident, how do you find out the
- 3:29problematic
- 3:31code, the problematic infrastructure, or
- 3:33the error to
- 3:34address the issue. The second is what we
- 3:36do as 5Y RCA at the end of the incident.
- 3:39Like we spend couple of days or couple
- 3:40of weeks to come up with a rigorous 5Y
- 3:42RCA, right? I'll be interchangeably
- 3:43using it in both, but it means both,
- 3:46yeah.
- 3:47So, what is root cause analysis? An RCA
- 3:50is a structured process to identify the
- 3:52reasons of why a problem has occurred.
- 3:55Um the the intent of an RCA is to
- 3:57understand the what, why, and how of a
- 4:00of an outage or an incident. And the
- 4:03most important purpose of doing a root
- 4:04cause analysis is to go beyond the
- 4:06surface level symptoms into the most
- 4:08fundamental or the underlying root
- 4:10cause. And the objective that we're
- 4:13trying to achieve by doing this doing
- 4:14this is that you can prevent similar
- 4:16problems from occurring in the future.
- 4:19And 5Y is a very popular technique in
- 4:21terms of doing an RCA, right? So So the
- 4:23basic idea behind 5Y is that you
- 4:25repeatedly ask the question, "Why did X
- 4:27happen?" And until you find the root
- 4:29cause. So there are a lot of uh
- 4:31conventions and good practices in terms
- 4:33of how you write a 5Y. It can't just be
- 4:35like some random 5Ys about five
- 4:37different aspects of the incident. Um so
- 4:40each Y has to ask a question which is
- 4:42trying to answer the previous Y. So that
- 4:45it's it's sort of like peeling the onion
- 4:46when you're trying to uncover the root
- 4:48cause of the incident. And most
- 4:51companies try to follow a principle of
- 4:52blameless postmortem because unless you
- 4:54follow blameless postmortem, you're not
- 4:56going to identify the fundamental root
- 4:58cause and prevent the issue from
- 5:00happening again.
- 5:03And the the better your RCAs are, the
- 5:05better you get a holistic understanding
- 5:07of your system. Like even people who
- 5:08have joined newly, like leadership who
- 5:10have joined newly, get a much more
- 5:11better understanding of the overall
- 5:13systems and the failure points in the
- 5:15overall system. Once you start doing
- 5:17sitting in more and more 5Y RCAs, right?
- 5:19So So eventually it helps you in from
- 5:22coming up with your architecture plans
- 5:24and
- 5:25system resiliency plans based on what is
- 5:27what is a hotspot or what is frequently
- 5:29failing in your systems.
- 5:32And the the uber goal of any kind of RCA
- 5:34system, be it a manual RCA system or a
- 5:37automated RCA system, right? Is is to
- 5:40attack this metric called MTTR, which is
- 5:42mean time to resolve. And the less time
- 5:45you take to resolve these incidents, the
- 5:47less time you
- 5:48your you
- 5:49you
- 5:50your customers are impacted. Like um
- 5:52what we're seeing with GitHub and Cloud
- 5:54these days, right?
- 5:55So, there are like different breakdowns
- 5:57in terms of uh MTTR.
- 6:00It is based on the phases that an
- 6:02incident goes through, right? The As
- 6:03soon as the incident starts, the first
- 6:05thing that we measure is what is called
- 6:07as
- 6:07MTTD or mean time to detect. Like did
- 6:10your monitoring systems detect the
- 6:11incident? Did a human user or a
- 6:14customer
- 6:15a user report the incident, right?
- 6:18And MTTD is usually good only if your
- 6:20alerting is good.
- 6:22The The next step is uh
- 6:25Did the on-call or did the incident
- 6:27responder acknowledge the alert?
- 6:29Like most of the time the the time spent
- 6:32in incidents is about getting the right
- 6:34people in the room, and
- 6:36the right people, the right owner
- 6:37services, and right stake right uh
- 6:41subject matter experts in order to solve
- 6:42an incident, right? This is what we call
- 6:43as MTTE or mean time to engage. So, if
- 6:46you optimize your mean time to engage,
- 6:48it's like the work is half done. Like
- 6:49you know which service to engage, which
- 6:52uh expert to engage, and which metric to
- 6:54debug, right? The next part is
- 6:56mitigation. Like mitigation is most
- 6:57important from a business perspective,
- 6:59because that is where that is the
- 7:01duration of your outage. You could You
- 7:02could do whatever hack it takes to fix
- 7:04the system,
- 7:05and resolution is about architecture
- 7:07resiliency, right? Like what are the
- 7:09long-term action items that we're taking
- 7:11after discovering the underlying root
- 7:13cause of the incident. And And all of
- 7:15these are important. MTTR is the uh
- 7:18uber metric, which which kind of
- 7:19represents all of these subparts.
- 7:23So, now that we know what is
- 7:25uh
- 7:25software incident, what is root cause
- 7:27analysis, let's look at what role that
- 7:29AI can play in in root cause analysis
- 7:32and software incidents.
- 7:34The Like when you're doing a root cause
- 7:36analysis of a complex microservice
- 7:38environment, like uh typically large
- 7:40enterprises deal with uh hundreds or
- 7:42thousands of microservices and their
- 7:44dependencies like databases, cache, and
- 7:47message queues, right? You you're
- 7:49dealing with a complex environment where
- 7:50you have to understand the dependency of
- 7:52multiple changes. So, there is a
- 7:54challenge in sifting through change logs
- 7:56in the form of deployments, pull
- 7:57requests, commits.
- 7:59And And what is shown here is an
- 8:00approach used at Facebook for divide and
- 8:02conquer of changes. Like every time
- 8:04there is a faulty broken build or faulty
- 8:07release that happens, they do a process
- 8:09like divide and conquer to
- 8:12to batch through the changes and find
- 8:13out the fundamental root cause.
- 8:15And how can AI help us irrespective of
- 8:18whether we are a small
- 8:19small enterprise or medium enterprise or
- 8:21large enterprise, right?
- 8:23LLMs have very good code understanding
- 8:25and they can look at a PR or a commit or
- 8:28code code snippet like a diff and try to
- 8:31conceptually understand what it means.
- 8:32So, using this, they can semantically
- 8:34relate how this commit or code change
- 8:37could have led to a certain incident.
- 8:41I'll go into details of this later. This
- 8:42is just to motivate on the different
- 8:44signals that that are important from a
- 8:46AI root cause analysis.
- 8:48The The second most complex part of root
- 8:51cause analysis is
- 8:52triangulation. Like how do you
- 8:54triangulate an incident with a code
- 8:56change with an observability signal that
- 8:58indicates that something is wrong. So,
- 9:01the the challenge here is that we're
- 9:02dealing with lots and lots of
- 9:04observability data. Depending on the
- 9:05scale of the company, this is this can
- 9:07literally be like terabytes of logs and
- 9:10uh terabytes of metrics, right? So, the
- 9:12different kinds of signals here are
- 9:14collectively called as melt, which
- 9:15stands for metrics, error events, logs,
- 9:18and traces.
- 9:19Like alerts, sentry errors are are a
- 9:21kind of error events.
- 9:23And the role of AI in helping us
- 9:26manage this voluminous data is that it
- 9:28can help us in pattern matching, anomaly
- 9:30detection. And most of these systems do
- 9:32not have very straightforward queries
- 9:34like your analytics database. You cannot
- 9:35just run a SQL query on a Prometheus
- 9:37database, right? So, it helps you in
- 9:40understanding the problem, trying to
- 9:42formulate like what is the right query
- 9:43that you need to run,
- 9:44>> [snorts]
- 9:45>> and and run those queries. And there is
- 9:47also correlation causation. Just because
- 9:48there are like 10 5xx errors at a point
- 9:50in time doesn't mean they're related.
- 9:52You have to know which error led to
- 9:54which error. Like the kind of output
- 9:56that you you you would have at the end
- 9:58of a human RCA of
- 9:595Y, that's the sort of thing that we're
- 10:01trying to do in minutes. Like by by
- 10:03relating those 10 5xx errors.
- 10:08And the other big challenge is that
- 10:10there are complex microservice
- 10:11architectures. Like most enterprises
- 10:13have hundreds or thousands of
- 10:15microservices, and and they are
- 10:16connected in complex ways. They have
- 10:18sync flows, async flows using SQS and
- 10:20Kafka. Each service like like we might
- 10:24in our head we might know that a service
- 10:26depends on let's say five services. But
- 10:28once you start looking at the network
- 10:29graphs, looking at the traces, you start
- 10:30realizing that you depend on like 20 or
- 10:3330 services which you did not even know
- 10:34in the first place.
- 10:36And the the goal of RCA is to come up
- 10:38with what is called as a fault
- 10:39propagation graph, right? Uh you start
- 10:42with an outage, like a business outage,
- 10:43like uh
- 10:44checkout is not working for bank XYZ.
- 10:47Like checkout is failing for bank XYZ
- 10:48can be a business error.
- 10:50How does a fault propagate all the way
- 10:52from a customer-facing cart cart service
- 10:54or checkout service all the way to the
- 10:56underlying service? The The purpose of
- 10:58this fault propagation graph is to show
- 10:59you in your system how this is actually
- 11:02happening.
- 11:04The The role of AI here is that it it
- 11:06kind of grounds your investigation. This
- 11:08is sort of like uh
- 11:10the context graph for AI to investigate
- 11:12in the right places instead of getting
- 11:13confused. And there is also unstructured
- 11:16understanding in the form of what does a
- 11:17service actually represent. Like a
- 11:19promise engine service can be different
- 11:21from what is a procurement service,
- 11:22right? What do the What are the concepts
- 11:24that they represent when you're facing a
- 11:25checkout error? So, those kind of
- 11:27relationships are unstructured
- 11:29relationships also have to come from the
- 11:33complex dependencies that you're seeing
- 11:34in production. And that's not all like
- 11:37just sorting these does not end the
- 11:39story, right? There can be like various
- 11:42other changes like you can have static
- 11:44feature flags, you can have
- 11:45infrastructure as code changes in
- 11:47deployment.yml, Terraform, or
- 11:50like something being merged does not
- 11:52mean it gets deployed, right? There are
- 11:54complex deployment topologies like
- 11:55canary deployments and progressive
- 11:56deployments which make the whole
- 11:59debugging story much more difficult.
- 12:00Like some of the examples that we have
- 12:02seen is there could be a deployment that
- 12:04happened like
- 12:051 month ago and the feature flag got
- 12:07rolled out 3 days ago and that caused
- 12:09the outage. So, it it doesn't mean that
- 12:12once you deployed a feature it's it
- 12:14starts immediately impacting and
- 12:15timeline correlation is is what it
- 12:16takes, right?
- 12:18So, let's let's look at what are the
- 12:21building blocks of like if you have to
- 12:23solve all these challenges, what are the
- 12:24building blocks of a AI on-call or a SRE
- 12:27agent that that needs to solve this?
- 12:29Like I'm I'm using the word SRE here a
- 12:31bit loosely. An SRE does a lot more
- 12:33things. An SRE writes code. An SRE
- 12:35improves systems as they go. I'm just
- 12:37taking a very narrow use case of
- 12:40during an incident what does a on-call
- 12:42person do in order to resolve the
- 12:44incident or what does a SRE do to
- 12:46resolve the incident, right? It doesn't
- 12:47talk about any other aspects of an SRE.
- 12:51So, the
- 12:52the principles are there is first
- 12:53principles thinking. Like humans have a
- 12:55runbook. Like you you can have a
- 12:57specialized runbook in your confluence
- 12:58or documentation which says this is how
- 13:01we solve issues of type X in the system.
- 13:03Like when there is a rate limit related
- 13:05error, this is how we resolve it and it
- 13:07can be codified into a runbook. But when
- 13:09you're giving it to an AI, it
- 13:10it can come across novel incidents. It
- 13:13can come across recurring incidents.
- 13:14Your runbooks are going to be completely
- 13:16useless when you when you come across a
- 13:18novel incident, right? So, that is where
- 13:20some generic runbooks will help. So,
- 13:22there'll be a typical runbook which
- 13:23helps you how to deal with code change
- 13:25related incidents. Like incidents where
- 13:27code change is the root cause, incidents
- 13:29where feature flags are the root cause,
- 13:31and incidents which can be explained by
- 13:33metrics, logs, and traces will have a
- 13:34typical runbook, right?
- 13:36So, uh the insights that we've gotten is
- 13:39that you need some kind of runbook. A
- 13:41runbook is like a
- 13:42like a cloud code to-do list here, but
- 13:44except that it's it's trying to debug in
- 13:46production.
- 13:47And uh you also need to correlate with
- 13:50these runbooks, right? Like, based on
- 13:52the incident, you need to have a dynamic
- 13:53plan.
- 13:54And you're trying to correlate between a
- 13:565XX error in your Splunk with a
- 13:59connection pool error in your metric
- 14:00with some code change that happened like
- 14:023 days ago in in GitHub PRs. So, this
- 14:04kind of correlation and triangulation is
- 14:06happening all the time.
- 14:08And and the biggest challenge with doing
- 14:09this in AI is that a human expert will
- 14:12have tribal knowledge or subject matter
- 14:14expertise, whereas the AI is always
- 14:16starting from scratch, and it has to
- 14:18somehow find the shortest path from the
- 14:21symptom to the underlying symptom. Like,
- 14:22the way the what you're effectively
- 14:24doing in a 5-way RCA is causal chaining
- 14:27of symptoms, and you're expecting AI to
- 14:29do that from scratch every time. So, do
- 14:31doing this
- 14:33triangulation becomes a challenge with
- 14:35AI.
- 14:36I'll I'll show a typical architecture.
- 14:37This is like a representative
- 14:38architecture. I've not used any real
- 14:40systems here, but just very high-level
- 14:41systems, right? The three building
- 14:43blocks of a AI SRE agent is One is the
- 14:46agentic interface. The second is the
- 14:48context engineering part, and the third
- 14:51is the LLM orchestration piece. I'll
- 14:53I'll zoom in on this. Most probably,
- 14:55this is not visible here.
- 14:56So, the first part is the interface,
- 14:58right? Like, um
- 15:00an on-call engineer or an SRE engineer
- 15:02receives alerts and incident
- 15:04notifications from an alert management
- 15:06system. In Atlassian, we have Jira
- 15:08Service Management for alerts and Jira
- 15:09Service Management for incidents. You
- 15:11can plug and play that with any any
- 15:13tool.
- 15:14And these alerts are actually coming
- 15:16because a threshold got breached in an
- 15:18observability system, or a a says
- 15:20something is broken, or even employees
- 15:23find out that something is broken and
- 15:24raise an incident.
- 15:26Once a notification comes to the on-call
- 15:28engineer or SRE engineer, they have like
- 15:30multiple interfaces to start working on
- 15:32the problem. Like we have RCA as a
- 15:34product, like root cause analysis as a
- 15:36product, and there is a product UI that
- 15:39on-call engineers can go to.
- 15:41We also have an agentic chat called
- 15:44Rover chat, which is a which is similar
- 15:46to like ChatGPT for enterprises,
- 15:48and uh
- 15:49this can this also has custom agent
- 15:51custom tools where
- 15:53the incident investigation can start.
- 15:55The third one is we have an equivalent
- 15:56of cloud code called Rover dev, which is
- 15:58like a CLI-based interface. So, you can
- 16:00just substitute this with anything,
- 16:01right? Like cloud code or cursor.
- 16:03And this is sort of like the agentic
- 16:05interface for someone like a on-call or
- 16:07SRE to start debugging an incident.
- 16:10The second part is context engineering.
- 16:13Um
- 16:14one of the new ones in AI is that
- 16:16in traditional ML, the more data you
- 16:18give, the better the system performs. In
- 16:20normal in LLM-based AI, the more data
- 16:22you give, the worse it performs, right?
- 16:24So, the the challenge is all about
- 16:26finding the right context and feeding it
- 16:27into the AI. It can easily get
- 16:30distracted if you give it the wrong
- 16:31context. So, we we have three ways of
- 16:34giving context to the AI.
- 16:36First is what is called as service
- 16:38catalog. So, think of service catalog as
- 16:40uh
- 16:42a what a one-stop service dashboard for
- 16:44your company where information about any
- 16:46service, metadata, services, owners,
- 16:49on-call can be found.
- 16:50Like a canonical open source example of
- 16:52this is backstage, and we have a product
- 16:56here called assets which does this.
- 16:58The second one is teamwork graph. Like
- 17:00teamwork graph, you can you can think of
- 17:01it as a context graph that that we
- 17:03built.
- 17:04And teamwork graph gives you context of
- 17:06what is the work that has happened in
- 17:07Jira, what is what are the
- 17:10RFCs, tech documents that got written in
- 17:12Confluence, what are the PRs that got
- 17:14raised in Bitbucket or GitHub, and what
- 17:17are the alerts that this service is
- 17:18currently receiving? So that is a one
- 17:20first-party context that we have, but we
- 17:22don't just stop there, right? Like
- 17:23Teamwork Graph is trying to build a
- 17:26like a context graph of
- 17:28of entire work, like
- 17:30So it it's not restricted to Atlassian.
- 17:32So we go beyond that into Google Drive,
- 17:34Salesforce um
- 17:36Sorry, Salesforce is a bad example here.
- 17:38Like GitHub pull requests or
- 17:40or even your Dropbox documents and
- 17:41whatnot, right? So the integration
- 17:43service pulls all of this and and builds
- 17:45a hundreds of billions of object graph,
- 17:47which which is like multi-tenanted per
- 17:49customer. And this graph is accessible
- 17:51to you whenever you want to
- 17:53get context. Like let's say I'm dealing
- 17:55with a checkout error. I can find out
- 17:57what are the recent payment
- 17:58gateway-related changes that happened in
- 18:00Confluence, discussed in
- 18:02Slack, and
- 18:03where have a have a PR in GitHub, right?
- 18:06So all of this information is something
- 18:08that I can retrieve in in a couple of
- 18:09seconds using Teamwork Graph. Like
- 18:12uh I can also mean that the agent can
- 18:14retrieve it on demand using the Teamwork
- 18:16Graph.
- 18:17The third part of context is the
- 18:19orchestrator. You're not going to be
- 18:21able to get all data you want upfront.
- 18:23Like there is always going to be some
- 18:25data that you want to pull on demand. So
- 18:27that is where tools like MCP CLIs and
- 18:29API calls will come in. And this is
- 18:31especially useful for us when we're
- 18:33dealing with large-scale data like
- 18:34metrics, logs, and and
- 18:38observability data.
- 18:40The the third important part of the
- 18:42architecture is is the LLM orchestration
- 18:44itself, which is once the RCA back-end
- 18:46service has received a received a
- 18:49question, the question can be about how
- 18:51do you uh
- 18:54like what is the root cause of this
- 18:55incident, which is a end-to-end
- 18:57question, or it can be something very
- 18:58basic like which service has a high
- 18:59number of alerts right now, which is a
- 19:01query-based question, right? So these
- 19:03queries are answered by the RCA agent.
- 19:05The RCA agent is made up of multiple sub
- 19:07agents and multiple skills. An example
- 19:09of a sub agent is something like a
- 19:11hypothesis or reasoning agent or a
- 19:13planner agent. An example of a skill can
- 19:15be something like change analysis,
- 19:17metrics analysis, log analysis. We We
- 19:19use a lot of custom ML models here so
- 19:21that it's it's not just like MCP
- 19:23integrations and tool calling.
- 19:25And
- 19:26the agent is connected to some kind of
- 19:29gateway in order to
- 19:31in order to decide the next tool to
- 19:33invoke or
- 19:34the orchestration to do.
- 19:38And
- 19:39good data often leads to good machine
- 19:41learning and AI. And Teamwork Graph is
- 19:43at the center of how we do good data,
- 19:45right? So, you can think of Teamwork
- 19:47Graph here as the connected data
- 19:48ecosystem for AI. Like the way rag is
- 19:51used in AI, we use a graph rag for this.
- 19:54And at a very high level, it is
- 19:57it is made up of nouns and
- 19:58relationships. Any first-party entity or
- 20:00third-party entity is converted into
- 20:01nouns. And the relationship between
- 20:04those entities are converted into what
- 20:06is called as graph relationships. And
- 20:08And we have an interconnected
- 20:11system of nouns for work. Like we call
- 20:13this as our system of work. Like all the
- 20:15Atlassian products are somehow connected
- 20:17to this whole Teamwork Graph and system
- 20:19of work.
- 20:20And in the context of RCA, what this
- 20:22gives us is a subgraph of deployment
- 20:25entities, pull requests, commits, Jira
- 20:26issues, repository, service. Like the
- 20:29catalog contains a repository and the
- 20:31service. And it helps us in grounding
- 20:34our investigations in facts.
- 20:37The The other part that that we don't
- 20:38actually own, but it's available in most
- 20:40of the observability systems and network
- 20:42monitoring systems is the service graph.
- 20:45So, a service catalog like Backstage or
- 20:47a service graph like a New Relic service
- 20:48graph is the is the system of record for
- 20:50this.
- 20:51And it gives you information about what
- 20:53are the upstream dependencies of a
- 20:54service, what are the downstream
- 20:56dependencies. Let's say you had an
- 20:57outage where
- 20:59checkout is not working for bank X and
- 21:01the underlying cause turns out to be
- 21:03something in
- 21:06a dependency service like promise
- 21:07service, right? So, these systems will
- 21:10trace the dependency from your
- 21:12user-facing checkout service all the way
- 21:13to the problematic service. And we
- 21:17get this information from SOR, but we
- 21:18also try to store this in our teamwork
- 21:20graph, so that at the time of incident,
- 21:21we have a holistic view of things.
- 21:24And the
- 21:26the end goal of this data is that we're
- 21:28trying to construct a sub graph which
- 21:30the AI can use. This is actually like a
- 21:32pretty big sub graph.
- 21:34And
- 21:35it it connects the incident with the
- 21:37service. It connects a service with all
- 21:39the changes like which repository is the
- 21:41service hosted in, what are the PRs,
- 21:43commits, and deployments that happen on
- 21:45this repository. It connects a service
- 21:47with a service. And and all of this is
- 21:48queryable in a
- 21:50uh like you can think of teamwork graph
- 21:51as a Neo4j like graph which the AI can
- 21:53use anytime it wants.
- 21:56And and the last important part of
- 21:58building block of an AI SRE system is
- 22:00the agent orchestration itself, right?
- 22:02I'll I'll start with a logical view.
- 22:04So, in the logical view of an RC agent,
- 22:07it had it is connected to like three
- 22:09different sources of data. So, there is
- 22:11change data, there is system knowledge,
- 22:12and then there is observability
- 22:13knowledge.
- 22:14And majority of this context comes from
- 22:17our teamwork graph context. And whatever
- 22:20is not available in teamwork graph, that
- 22:21is where we start using MCPs and CLIs to
- 22:23get
- 22:25to to augment it with the additional
- 22:26context.
- 22:27So, an example of change data can be
- 22:29GitHub or Bitbucket deployments, PRs,
- 22:31and commits. And static feature flags
- 22:34and Jira issues. So, the teamwork graph
- 22:36connects all of this and and gives a
- 22:37view of
- 22:39how does change influence the current
- 22:41problem.
- 22:43The second part is a system knowledge,
- 22:44which is there is structured knowledge
- 22:46and unstructured knowledge of a system.
- 22:48An example of a structured knowledge is
- 22:51your backstage system backstage software
- 22:53services and software service
- 22:54dependencies.
- 22:55But anything that is config driven is
- 22:57not going to going to reflect the real
- 22:59world. So, this is also connected to
- 23:00your
- 23:01active systems like neural network
- 23:03service graphs or uh
- 23:05distributed tracing based service
- 23:06graphs, right? So, it it reflects the
- 23:08real world and it's not some stale JSON
- 23:10that that is present.
- 23:12An example of unstructured data is
- 23:14architecture documents and
- 23:16RFCs and PRDs and confluence which tell
- 23:19like what each subsystem is all about.
- 23:21And the
- 23:22the teamwork graph is obviously per
- 23:24tenant, so
- 23:25it is grounded in the unstructured and
- 23:27structured knowledge of that specific
- 23:28tenant without using world knowledge.
- 23:30Like
- 23:31in addition to world knowledge, right?
- 23:33The third part is whatever data is not
- 23:35available to us, we we depend on MCPs
- 23:38and other systems to integrate that
- 23:40data. Like things like uh
- 23:43Prometheus metrics or Splunk logs are
- 23:45something that come through
- 23:47on demand on demand queries, right? The
- 23:49end result of the RCA agent is that it
- 23:51tries to come up with multiple
- 23:52hypothesis, like a ranked list of
- 23:54hypothesis.
- 23:55It's It's not just integrations and data
- 23:57here. There is also like obviously some
- 23:59amount of intelligence supplied at each
- 24:00place to know how to traverse all these
- 24:03dependencies, right?
- 24:04So, the structure of this is that it's
- 24:06made up of multiple subagents. It's made
- 24:08up of multiple skills.
- 24:10And for data that is not available
- 24:12through the context graph, it depends on
- 24:13multiple MCP servers.
- 24:15And there is also some uh
- 24:18internal intelligence like in the form
- 24:19of ML models and uh
- 24:22ML models for changes, ML models for
- 24:24anomaly detection, ML models for entity
- 24:26understanding, and so on.
- 24:29This is like a typical stack that we use
- 24:31for uh
- 24:32subagents and skills. It's It's just a
- 24:34layered architecture, which means that
- 24:35there is a agent harness or runtime.
- 24:38You can think of the agent as
- 24:40uh as a chef here. The analogy is a
- 24:41chef. And the agent harness depends upon
- 24:44multiple skills and multiple
- 24:46tools.
- 24:47You can think of the skills as some kind
- 24:49of recipe in prompts and some kind of
- 24:51recipe even in code, right? Like prompts
- 24:54cannot express 100% of the scenarios
- 24:55that a complex system is trying to
- 24:57solve. An example of skills here is
- 24:59again how do you do feature flag
- 25:01analysis? How do you do uh pull request
- 25:03analysis? How do you do metric analysis?
- 25:05Log analysis?
- 25:07It can't just be expressed all in
- 25:08English. So, it also needs a connect
- 25:10collection of MCP servers which
- 25:12translate that into
- 25:14uh API calls. And you also need to
- 25:16connect to your backend services which
- 25:18are acting as the intelligence layer,
- 25:19right? So, the interfaces like MCP CLI
- 25:22or function tools and even the backend
- 25:24services act as like the like the
- 25:26analogy to a chef here would be that
- 25:29these are the ingredients and appliances
- 25:30which are actually used to get the job
- 25:32done.
- 25:33And this can contain the intelligence
- 25:35that you're using in order to do an
- 25:36effective RCA, right? And obviously
- 25:38there is evals and stuff like that which
- 25:40I've not shown here. So, every problem
- 25:41will require some kind of evals data set
- 25:44that that is being used to solve solve
- 25:45the problem.
- 25:47And that concludes the building blocks.
- 25:49So, to summarize the building blocks of
- 25:51a AI SRE agent are
- 25:53one is the agent orchestration, the
- 25:55second is the data that you're using it
- 25:58up for grounding and context, and the
- 26:00architecture elements like context
- 26:02engineering and uh
- 26:05LVM orchestration and the agentic
- 26:06interface. And the other part is how do
- 26:08you actually represent runbooks? How do
- 26:10you uh
- 26:12do an investigation?
- 26:14I'll briefly talk about some of the
- 26:15challenges we've encountered before
- 26:17going into questions.
- 26:18So, uh
- 26:20we've encountered a lot of challenges,
- 26:21right? Like uh
- 26:23some of the significant challenges in
- 26:24change-based analysis would be
- 26:27like these there are causal
- 26:28relationships. Like these are uh
- 26:30unstructured relationships. Like your
- 26:32problem is in a different domain
- 26:33language, your features are in a
- 26:35different domain language, right? So,
- 26:36you're trying to interlink a problem
- 26:38domain with a feature domain. The second
- 26:40kind of problem is that a service graph
- 26:43is not enough. Like a service graph does
- 26:44not tell you which API calls which API,
- 26:46right? It it just tells you that service
- 26:48A calls service B.
- 26:49And the third is that uh like because
- 26:52it's a product like the the third part
- 26:55will vary between a product solution and
- 26:57a platform like an internal solution.
- 27:00Because it's a product we we deal with
- 27:02different levels of code understanding.
- 27:03Like we don't get access to the entire
- 27:04GitHub code of a customer. We get access
- 27:06to maybe the PR description and the
- 27:08commit description. So what you can do
- 27:10with
- 27:11with code understanding and without code
- 27:12understanding are widely different. Like
- 27:14some customers give us code
- 27:15understanding code access, some
- 27:16customers don't. So deep code
- 27:18understanding makes your RCA way better.
- 27:20But the challenge is also that you have
- 27:22to do it in a repository. Like what we
- 27:24are trying to do is a RCA without
- 27:27like it's not like we are using cloud
- 27:29code to login into repository and do an
- 27:31RCA, right? We have thousands of
- 27:32services and maybe like thousands of
- 27:34repos and we are trying to do an RCA
- 27:35across these thousands of repos.
- 27:39Observability has its own set of
- 27:41challenges. Like first is the scale the
- 27:43sheer scale of observability data.
- 27:46And second is that most of the MCP tools
- 27:48for observability are very low level in
- 27:50nature. Like it's it's mostly about uh
- 27:53how do you get metadata? How do you get
- 27:54raw data? How do you construct a query?
- 27:57It's it's that that is not enough for
- 27:59you to translate from a problem to a
- 28:01query.
- 28:02And the third part is you need to come
- 28:03up with way lots and lots of dynamic
- 28:06plans depending on the data sources,
- 28:07right?
- 28:08Yeah.
- 28:09Yeah, these are the list of challenges
- 28:11that we've encountered in this. Yeah.
- 28:14Yeah, thank you.
- 28:20>> [music]
- 28:35[music]
About this transcript
This page contains the full transcript of AIOps: Leveraging AI for Incident Root Cause Analysis - Sathish Kumar by Developer Summit, generated from the public captions YouTube serves with the video. The transcript has 5,304 words across 868 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.