YouTube2Text

AIOps: Leveraging AI for Incident Root Cause Analysis - Sathish Kumar — Transcript

by Developer Summit · 5,304 words · 868 segments · language en · Watch on YouTube

Full transcript

  1. 0:13I'm Satish and I'm a senior principal
  2. 0:15engineer in Jira service management at
  3. 0:17Atlassian.
  4. 0:18Uh we work on a lot of products around
  5. 0:21support services, alert management,
  6. 0:23incident management, and AI ops. So,
  7. 0:24today I'll be talking about how we do AI
  8. 0:27ops in in one of our sub products. And
  9. 0:29and how we leverage AI effectively for
  10. 0:31uh
  11. 0:32automated alert management, automated
  12. 0:34incident management, right?
  13. 0:36The agenda of the talk will roughly
  14. 0:38cover the following topics.
  15. 0:40Like uh
  16. 0:41Like most of us will have some idea of
  17. 0:43what is a software instance, what is a
  18. 0:44root cause analysis, but I'll still
  19. 0:46start with a brief intro before going
  20. 0:48into the details. So, we'll talk about
  21. 0:50what are software incidents, what is RCA
  22. 0:53or root cause analysis, and the role
  23. 0:55that AI can play in in a root cause
  24. 0:57analysis both both in terms of like
  25. 1:00product in terms of in terms of
  26. 1:01operations, right? Like I'll be coming
  27. 1:03from the perspective of how we are
  28. 1:04building this into our products, but uh
  29. 1:06it's it's applicable even if you're like
  30. 1:08running production systems and you want
  31. 1:09to use it outside of the product, right?
  32. 1:11And uh I'll talk about the building
  33. 1:13blocks of uh how we've approached an AI
  34. 1:16SRE agent and the challenges that uh we
  35. 1:19have encountered along the way.
  36. 1:22So, what are software incidents?
  37. 1:25Software incidents are usually like
  38. 1:27unplanned interruptions and uh
  39. 1:29degradations or any kind of abnormal
  40. 1:30behavior in your software services and
  41. 1:33applications. It usually manifests in
  42. 1:35the form of some reliability or
  43. 1:37performance or functional issues to your
  44. 1:39customers.
  45. 1:41And a software incident usually goes
  46. 1:43through these following phases.
  47. 1:46So, you have a problem that is detected.
  48. 1:48Most of the time the problem is detected
  49. 1:49by your monitoring systems, but if your
  50. 1:51alerting or monitoring is weak,
  51. 1:53it can also be detected by let's say
  52. 1:55customer issues or support desk or even
  53. 1:57like if your product is like super
  54. 1:59popular, it can even be discovered on
  55. 2:01Twitter by your users and complain,
  56. 2:02right?
  57. 2:03And
  58. 2:04it results in some kind of customer
  59. 2:06impact and the the most important thing
  60. 2:09in any software incident is to
  61. 2:12reduce the customer impact or address
  62. 2:13the bleeding.
  63. 2:15Most of these incidents usually
  64. 2:16translate into something negative for
  65. 2:18the business. Like it can be revenue
  66. 2:19loss, it can be reputation loss, or some
  67. 2:21kind of like in worst cases even like
  68. 2:23data loss, right?
  69. 2:25And any software company or tech company
  70. 2:28usually goes through the following
  71. 2:29phases in order to effectively manage an
  72. 2:31incident. Like it it varies from
  73. 2:34organization to organization. The the
  74. 2:36rigor varies, the process varies, the
  75. 2:37tools vary, right? But still
  76. 2:40some of the common things are
  77. 2:41there's an investigation step. There is
  78. 2:44a effort to identify the potential root
  79. 2:46causes.
  80. 2:47Then we try to mitigate the issue. Like
  81. 2:49short-term mitigation is usually done
  82. 2:52with a goal of
  83. 2:53arresting the customer impact without
  84. 2:56like even if you don't find the actual
  85. 2:57root cause, right?
  86. 2:59And um
  87. 3:00once that is done, we want to update our
  88. 3:02stakeholders and we can spend the next
  89. 3:05few hours or few days finding out the
  90. 3:08actual root cause.
  91. 3:10I'll I'll go into details of what is
  92. 3:12root cause analysis and why it's
  93. 3:14important in this context. So, I'll be
  94. 3:16using the word root cause analysis to
  95. 3:17mean two things.
  96. 3:21Yeah.
  97. 3:22So, I'll be using the word root cause
  98. 3:23analysis in two different contexts. One
  99. 3:25is from the perspective of during an
  100. 3:27incident, how do you find out the
  101. 3:29problematic
  102. 3:31code, the problematic infrastructure, or
  103. 3:33the error to
  104. 3:34address the issue. The second is what we
  105. 3:36do as 5Y RCA at the end of the incident.
  106. 3:39Like we spend couple of days or couple
  107. 3:40of weeks to come up with a rigorous 5Y
  108. 3:42RCA, right? I'll be interchangeably
  109. 3:43using it in both, but it means both,
  110. 3:46yeah.
  111. 3:47So, what is root cause analysis? An RCA
  112. 3:50is a structured process to identify the
  113. 3:52reasons of why a problem has occurred.
  114. 3:55Um the the intent of an RCA is to
  115. 3:57understand the what, why, and how of a
  116. 4:00of an outage or an incident. And the
  117. 4:03most important purpose of doing a root
  118. 4:04cause analysis is to go beyond the
  119. 4:06surface level symptoms into the most
  120. 4:08fundamental or the underlying root
  121. 4:10cause. And the objective that we're
  122. 4:13trying to achieve by doing this doing
  123. 4:14this is that you can prevent similar
  124. 4:16problems from occurring in the future.
  125. 4:19And 5Y is a very popular technique in
  126. 4:21terms of doing an RCA, right? So So the
  127. 4:23basic idea behind 5Y is that you
  128. 4:25repeatedly ask the question, "Why did X
  129. 4:27happen?" And until you find the root
  130. 4:29cause. So there are a lot of uh
  131. 4:31conventions and good practices in terms
  132. 4:33of how you write a 5Y. It can't just be
  133. 4:35like some random 5Ys about five
  134. 4:37different aspects of the incident. Um so
  135. 4:40each Y has to ask a question which is
  136. 4:42trying to answer the previous Y. So that
  137. 4:45it's it's sort of like peeling the onion
  138. 4:46when you're trying to uncover the root
  139. 4:48cause of the incident. And most
  140. 4:51companies try to follow a principle of
  141. 4:52blameless postmortem because unless you
  142. 4:54follow blameless postmortem, you're not
  143. 4:56going to identify the fundamental root
  144. 4:58cause and prevent the issue from
  145. 5:00happening again.
  146. 5:03And the the better your RCAs are, the
  147. 5:05better you get a holistic understanding
  148. 5:07of your system. Like even people who
  149. 5:08have joined newly, like leadership who
  150. 5:10have joined newly, get a much more
  151. 5:11better understanding of the overall
  152. 5:13systems and the failure points in the
  153. 5:15overall system. Once you start doing
  154. 5:17sitting in more and more 5Y RCAs, right?
  155. 5:19So So eventually it helps you in from
  156. 5:22coming up with your architecture plans
  157. 5:24and
  158. 5:25system resiliency plans based on what is
  159. 5:27what is a hotspot or what is frequently
  160. 5:29failing in your systems.
  161. 5:32And the the uber goal of any kind of RCA
  162. 5:34system, be it a manual RCA system or a
  163. 5:37automated RCA system, right? Is is to
  164. 5:40attack this metric called MTTR, which is
  165. 5:42mean time to resolve. And the less time
  166. 5:45you take to resolve these incidents, the
  167. 5:47less time you
  168. 5:48your you
  169. 5:49you
  170. 5:50your customers are impacted. Like um
  171. 5:52what we're seeing with GitHub and Cloud
  172. 5:54these days, right?
  173. 5:55So, there are like different breakdowns
  174. 5:57in terms of uh MTTR.
  175. 6:00It is based on the phases that an
  176. 6:02incident goes through, right? The As
  177. 6:03soon as the incident starts, the first
  178. 6:05thing that we measure is what is called
  179. 6:07as
  180. 6:07MTTD or mean time to detect. Like did
  181. 6:10your monitoring systems detect the
  182. 6:11incident? Did a human user or a
  183. 6:14customer
  184. 6:15a user report the incident, right?
  185. 6:18And MTTD is usually good only if your
  186. 6:20alerting is good.
  187. 6:22The The next step is uh
  188. 6:25Did the on-call or did the incident
  189. 6:27responder acknowledge the alert?
  190. 6:29Like most of the time the the time spent
  191. 6:32in incidents is about getting the right
  192. 6:34people in the room, and
  193. 6:36the right people, the right owner
  194. 6:37services, and right stake right uh
  195. 6:41subject matter experts in order to solve
  196. 6:42an incident, right? This is what we call
  197. 6:43as MTTE or mean time to engage. So, if
  198. 6:46you optimize your mean time to engage,
  199. 6:48it's like the work is half done. Like
  200. 6:49you know which service to engage, which
  201. 6:52uh expert to engage, and which metric to
  202. 6:54debug, right? The next part is
  203. 6:56mitigation. Like mitigation is most
  204. 6:57important from a business perspective,
  205. 6:59because that is where that is the
  206. 7:01duration of your outage. You could You
  207. 7:02could do whatever hack it takes to fix
  208. 7:04the system,
  209. 7:05and resolution is about architecture
  210. 7:07resiliency, right? Like what are the
  211. 7:09long-term action items that we're taking
  212. 7:11after discovering the underlying root
  213. 7:13cause of the incident. And And all of
  214. 7:15these are important. MTTR is the uh
  215. 7:18uber metric, which which kind of
  216. 7:19represents all of these subparts.
  217. 7:23So, now that we know what is
  218. 7:25uh
  219. 7:25software incident, what is root cause
  220. 7:27analysis, let's look at what role that
  221. 7:29AI can play in in root cause analysis
  222. 7:32and software incidents.
  223. 7:34The Like when you're doing a root cause
  224. 7:36analysis of a complex microservice
  225. 7:38environment, like uh typically large
  226. 7:40enterprises deal with uh hundreds or
  227. 7:42thousands of microservices and their
  228. 7:44dependencies like databases, cache, and
  229. 7:47message queues, right? You you're
  230. 7:49dealing with a complex environment where
  231. 7:50you have to understand the dependency of
  232. 7:52multiple changes. So, there is a
  233. 7:54challenge in sifting through change logs
  234. 7:56in the form of deployments, pull
  235. 7:57requests, commits.
  236. 7:59And And what is shown here is an
  237. 8:00approach used at Facebook for divide and
  238. 8:02conquer of changes. Like every time
  239. 8:04there is a faulty broken build or faulty
  240. 8:07release that happens, they do a process
  241. 8:09like divide and conquer to
  242. 8:12to batch through the changes and find
  243. 8:13out the fundamental root cause.
  244. 8:15And how can AI help us irrespective of
  245. 8:18whether we are a small
  246. 8:19small enterprise or medium enterprise or
  247. 8:21large enterprise, right?
  248. 8:23LLMs have very good code understanding
  249. 8:25and they can look at a PR or a commit or
  250. 8:28code code snippet like a diff and try to
  251. 8:31conceptually understand what it means.
  252. 8:32So, using this, they can semantically
  253. 8:34relate how this commit or code change
  254. 8:37could have led to a certain incident.
  255. 8:41I'll go into details of this later. This
  256. 8:42is just to motivate on the different
  257. 8:44signals that that are important from a
  258. 8:46AI root cause analysis.
  259. 8:48The The second most complex part of root
  260. 8:51cause analysis is
  261. 8:52triangulation. Like how do you
  262. 8:54triangulate an incident with a code
  263. 8:56change with an observability signal that
  264. 8:58indicates that something is wrong. So,
  265. 9:01the the challenge here is that we're
  266. 9:02dealing with lots and lots of
  267. 9:04observability data. Depending on the
  268. 9:05scale of the company, this is this can
  269. 9:07literally be like terabytes of logs and
  270. 9:10uh terabytes of metrics, right? So, the
  271. 9:12different kinds of signals here are
  272. 9:14collectively called as melt, which
  273. 9:15stands for metrics, error events, logs,
  274. 9:18and traces.
  275. 9:19Like alerts, sentry errors are are a
  276. 9:21kind of error events.
  277. 9:23And the role of AI in helping us
  278. 9:26manage this voluminous data is that it
  279. 9:28can help us in pattern matching, anomaly
  280. 9:30detection. And most of these systems do
  281. 9:32not have very straightforward queries
  282. 9:34like your analytics database. You cannot
  283. 9:35just run a SQL query on a Prometheus
  284. 9:37database, right? So, it helps you in
  285. 9:40understanding the problem, trying to
  286. 9:42formulate like what is the right query
  287. 9:43that you need to run,
  288. 9:44>> [snorts]
  289. 9:45>> and and run those queries. And there is
  290. 9:47also correlation causation. Just because
  291. 9:48there are like 10 5xx errors at a point
  292. 9:50in time doesn't mean they're related.
  293. 9:52You have to know which error led to
  294. 9:54which error. Like the kind of output
  295. 9:56that you you you would have at the end
  296. 9:58of a human RCA of
  297. 9:595Y, that's the sort of thing that we're
  298. 10:01trying to do in minutes. Like by by
  299. 10:03relating those 10 5xx errors.
  300. 10:08And the other big challenge is that
  301. 10:10there are complex microservice
  302. 10:11architectures. Like most enterprises
  303. 10:13have hundreds or thousands of
  304. 10:15microservices, and and they are
  305. 10:16connected in complex ways. They have
  306. 10:18sync flows, async flows using SQS and
  307. 10:20Kafka. Each service like like we might
  308. 10:24in our head we might know that a service
  309. 10:26depends on let's say five services. But
  310. 10:28once you start looking at the network
  311. 10:29graphs, looking at the traces, you start
  312. 10:30realizing that you depend on like 20 or
  313. 10:3330 services which you did not even know
  314. 10:34in the first place.
  315. 10:36And the the goal of RCA is to come up
  316. 10:38with what is called as a fault
  317. 10:39propagation graph, right? Uh you start
  318. 10:42with an outage, like a business outage,
  319. 10:43like uh
  320. 10:44checkout is not working for bank XYZ.
  321. 10:47Like checkout is failing for bank XYZ
  322. 10:48can be a business error.
  323. 10:50How does a fault propagate all the way
  324. 10:52from a customer-facing cart cart service
  325. 10:54or checkout service all the way to the
  326. 10:56underlying service? The The purpose of
  327. 10:58this fault propagation graph is to show
  328. 10:59you in your system how this is actually
  329. 11:02happening.
  330. 11:04The The role of AI here is that it it
  331. 11:06kind of grounds your investigation. This
  332. 11:08is sort of like uh
  333. 11:10the context graph for AI to investigate
  334. 11:12in the right places instead of getting
  335. 11:13confused. And there is also unstructured
  336. 11:16understanding in the form of what does a
  337. 11:17service actually represent. Like a
  338. 11:19promise engine service can be different
  339. 11:21from what is a procurement service,
  340. 11:22right? What do the What are the concepts
  341. 11:24that they represent when you're facing a
  342. 11:25checkout error? So, those kind of
  343. 11:27relationships are unstructured
  344. 11:29relationships also have to come from the
  345. 11:33complex dependencies that you're seeing
  346. 11:34in production. And that's not all like
  347. 11:37just sorting these does not end the
  348. 11:39story, right? There can be like various
  349. 11:42other changes like you can have static
  350. 11:44feature flags, you can have
  351. 11:45infrastructure as code changes in
  352. 11:47deployment.yml, Terraform, or
  353. 11:50like something being merged does not
  354. 11:52mean it gets deployed, right? There are
  355. 11:54complex deployment topologies like
  356. 11:55canary deployments and progressive
  357. 11:56deployments which make the whole
  358. 11:59debugging story much more difficult.
  359. 12:00Like some of the examples that we have
  360. 12:02seen is there could be a deployment that
  361. 12:04happened like
  362. 12:051 month ago and the feature flag got
  363. 12:07rolled out 3 days ago and that caused
  364. 12:09the outage. So, it it doesn't mean that
  365. 12:12once you deployed a feature it's it
  366. 12:14starts immediately impacting and
  367. 12:15timeline correlation is is what it
  368. 12:16takes, right?
  369. 12:18So, let's let's look at what are the
  370. 12:21building blocks of like if you have to
  371. 12:23solve all these challenges, what are the
  372. 12:24building blocks of a AI on-call or a SRE
  373. 12:27agent that that needs to solve this?
  374. 12:29Like I'm I'm using the word SRE here a
  375. 12:31bit loosely. An SRE does a lot more
  376. 12:33things. An SRE writes code. An SRE
  377. 12:35improves systems as they go. I'm just
  378. 12:37taking a very narrow use case of
  379. 12:40during an incident what does a on-call
  380. 12:42person do in order to resolve the
  381. 12:44incident or what does a SRE do to
  382. 12:46resolve the incident, right? It doesn't
  383. 12:47talk about any other aspects of an SRE.
  384. 12:51So, the
  385. 12:52the principles are there is first
  386. 12:53principles thinking. Like humans have a
  387. 12:55runbook. Like you you can have a
  388. 12:57specialized runbook in your confluence
  389. 12:58or documentation which says this is how
  390. 13:01we solve issues of type X in the system.
  391. 13:03Like when there is a rate limit related
  392. 13:05error, this is how we resolve it and it
  393. 13:07can be codified into a runbook. But when
  394. 13:09you're giving it to an AI, it
  395. 13:10it can come across novel incidents. It
  396. 13:13can come across recurring incidents.
  397. 13:14Your runbooks are going to be completely
  398. 13:16useless when you when you come across a
  399. 13:18novel incident, right? So, that is where
  400. 13:20some generic runbooks will help. So,
  401. 13:22there'll be a typical runbook which
  402. 13:23helps you how to deal with code change
  403. 13:25related incidents. Like incidents where
  404. 13:27code change is the root cause, incidents
  405. 13:29where feature flags are the root cause,
  406. 13:31and incidents which can be explained by
  407. 13:33metrics, logs, and traces will have a
  408. 13:34typical runbook, right?
  409. 13:36So, uh the insights that we've gotten is
  410. 13:39that you need some kind of runbook. A
  411. 13:41runbook is like a
  412. 13:42like a cloud code to-do list here, but
  413. 13:44except that it's it's trying to debug in
  414. 13:46production.
  415. 13:47And uh you also need to correlate with
  416. 13:50these runbooks, right? Like, based on
  417. 13:52the incident, you need to have a dynamic
  418. 13:53plan.
  419. 13:54And you're trying to correlate between a
  420. 13:565XX error in your Splunk with a
  421. 13:59connection pool error in your metric
  422. 14:00with some code change that happened like
  423. 14:023 days ago in in GitHub PRs. So, this
  424. 14:04kind of correlation and triangulation is
  425. 14:06happening all the time.
  426. 14:08And and the biggest challenge with doing
  427. 14:09this in AI is that a human expert will
  428. 14:12have tribal knowledge or subject matter
  429. 14:14expertise, whereas the AI is always
  430. 14:16starting from scratch, and it has to
  431. 14:18somehow find the shortest path from the
  432. 14:21symptom to the underlying symptom. Like,
  433. 14:22the way the what you're effectively
  434. 14:24doing in a 5-way RCA is causal chaining
  435. 14:27of symptoms, and you're expecting AI to
  436. 14:29do that from scratch every time. So, do
  437. 14:31doing this
  438. 14:33triangulation becomes a challenge with
  439. 14:35AI.
  440. 14:36I'll I'll show a typical architecture.
  441. 14:37This is like a representative
  442. 14:38architecture. I've not used any real
  443. 14:40systems here, but just very high-level
  444. 14:41systems, right? The three building
  445. 14:43blocks of a AI SRE agent is One is the
  446. 14:46agentic interface. The second is the
  447. 14:48context engineering part, and the third
  448. 14:51is the LLM orchestration piece. I'll
  449. 14:53I'll zoom in on this. Most probably,
  450. 14:55this is not visible here.
  451. 14:56So, the first part is the interface,
  452. 14:58right? Like, um
  453. 15:00an on-call engineer or an SRE engineer
  454. 15:02receives alerts and incident
  455. 15:04notifications from an alert management
  456. 15:06system. In Atlassian, we have Jira
  457. 15:08Service Management for alerts and Jira
  458. 15:09Service Management for incidents. You
  459. 15:11can plug and play that with any any
  460. 15:13tool.
  461. 15:14And these alerts are actually coming
  462. 15:16because a threshold got breached in an
  463. 15:18observability system, or a a says
  464. 15:20something is broken, or even employees
  465. 15:23find out that something is broken and
  466. 15:24raise an incident.
  467. 15:26Once a notification comes to the on-call
  468. 15:28engineer or SRE engineer, they have like
  469. 15:30multiple interfaces to start working on
  470. 15:32the problem. Like we have RCA as a
  471. 15:34product, like root cause analysis as a
  472. 15:36product, and there is a product UI that
  473. 15:39on-call engineers can go to.
  474. 15:41We also have an agentic chat called
  475. 15:44Rover chat, which is a which is similar
  476. 15:46to like ChatGPT for enterprises,
  477. 15:48and uh
  478. 15:49this can this also has custom agent
  479. 15:51custom tools where
  480. 15:53the incident investigation can start.
  481. 15:55The third one is we have an equivalent
  482. 15:56of cloud code called Rover dev, which is
  483. 15:58like a CLI-based interface. So, you can
  484. 16:00just substitute this with anything,
  485. 16:01right? Like cloud code or cursor.
  486. 16:03And this is sort of like the agentic
  487. 16:05interface for someone like a on-call or
  488. 16:07SRE to start debugging an incident.
  489. 16:10The second part is context engineering.
  490. 16:13Um
  491. 16:14one of the new ones in AI is that
  492. 16:16in traditional ML, the more data you
  493. 16:18give, the better the system performs. In
  494. 16:20normal in LLM-based AI, the more data
  495. 16:22you give, the worse it performs, right?
  496. 16:24So, the the challenge is all about
  497. 16:26finding the right context and feeding it
  498. 16:27into the AI. It can easily get
  499. 16:30distracted if you give it the wrong
  500. 16:31context. So, we we have three ways of
  501. 16:34giving context to the AI.
  502. 16:36First is what is called as service
  503. 16:38catalog. So, think of service catalog as
  504. 16:40uh
  505. 16:42a what a one-stop service dashboard for
  506. 16:44your company where information about any
  507. 16:46service, metadata, services, owners,
  508. 16:49on-call can be found.
  509. 16:50Like a canonical open source example of
  510. 16:52this is backstage, and we have a product
  511. 16:56here called assets which does this.
  512. 16:58The second one is teamwork graph. Like
  513. 17:00teamwork graph, you can you can think of
  514. 17:01it as a context graph that that we
  515. 17:03built.
  516. 17:04And teamwork graph gives you context of
  517. 17:06what is the work that has happened in
  518. 17:07Jira, what is what are the
  519. 17:10RFCs, tech documents that got written in
  520. 17:12Confluence, what are the PRs that got
  521. 17:14raised in Bitbucket or GitHub, and what
  522. 17:17are the alerts that this service is
  523. 17:18currently receiving? So that is a one
  524. 17:20first-party context that we have, but we
  525. 17:22don't just stop there, right? Like
  526. 17:23Teamwork Graph is trying to build a
  527. 17:26like a context graph of
  528. 17:28of entire work, like
  529. 17:30So it it's not restricted to Atlassian.
  530. 17:32So we go beyond that into Google Drive,
  531. 17:34Salesforce um
  532. 17:36Sorry, Salesforce is a bad example here.
  533. 17:38Like GitHub pull requests or
  534. 17:40or even your Dropbox documents and
  535. 17:41whatnot, right? So the integration
  536. 17:43service pulls all of this and and builds
  537. 17:45a hundreds of billions of object graph,
  538. 17:47which which is like multi-tenanted per
  539. 17:49customer. And this graph is accessible
  540. 17:51to you whenever you want to
  541. 17:53get context. Like let's say I'm dealing
  542. 17:55with a checkout error. I can find out
  543. 17:57what are the recent payment
  544. 17:58gateway-related changes that happened in
  545. 18:00Confluence, discussed in
  546. 18:02Slack, and
  547. 18:03where have a have a PR in GitHub, right?
  548. 18:06So all of this information is something
  549. 18:08that I can retrieve in in a couple of
  550. 18:09seconds using Teamwork Graph. Like
  551. 18:12uh I can also mean that the agent can
  552. 18:14retrieve it on demand using the Teamwork
  553. 18:16Graph.
  554. 18:17The third part of context is the
  555. 18:19orchestrator. You're not going to be
  556. 18:21able to get all data you want upfront.
  557. 18:23Like there is always going to be some
  558. 18:25data that you want to pull on demand. So
  559. 18:27that is where tools like MCP CLIs and
  560. 18:29API calls will come in. And this is
  561. 18:31especially useful for us when we're
  562. 18:33dealing with large-scale data like
  563. 18:34metrics, logs, and and
  564. 18:38observability data.
  565. 18:40The the third important part of the
  566. 18:42architecture is is the LLM orchestration
  567. 18:44itself, which is once the RCA back-end
  568. 18:46service has received a received a
  569. 18:49question, the question can be about how
  570. 18:51do you uh
  571. 18:54like what is the root cause of this
  572. 18:55incident, which is a end-to-end
  573. 18:57question, or it can be something very
  574. 18:58basic like which service has a high
  575. 18:59number of alerts right now, which is a
  576. 19:01query-based question, right? So these
  577. 19:03queries are answered by the RCA agent.
  578. 19:05The RCA agent is made up of multiple sub
  579. 19:07agents and multiple skills. An example
  580. 19:09of a sub agent is something like a
  581. 19:11hypothesis or reasoning agent or a
  582. 19:13planner agent. An example of a skill can
  583. 19:15be something like change analysis,
  584. 19:17metrics analysis, log analysis. We We
  585. 19:19use a lot of custom ML models here so
  586. 19:21that it's it's not just like MCP
  587. 19:23integrations and tool calling.
  588. 19:25And
  589. 19:26the agent is connected to some kind of
  590. 19:29gateway in order to
  591. 19:31in order to decide the next tool to
  592. 19:33invoke or
  593. 19:34the orchestration to do.
  594. 19:38And
  595. 19:39good data often leads to good machine
  596. 19:41learning and AI. And Teamwork Graph is
  597. 19:43at the center of how we do good data,
  598. 19:45right? So, you can think of Teamwork
  599. 19:47Graph here as the connected data
  600. 19:48ecosystem for AI. Like the way rag is
  601. 19:51used in AI, we use a graph rag for this.
  602. 19:54And at a very high level, it is
  603. 19:57it is made up of nouns and
  604. 19:58relationships. Any first-party entity or
  605. 20:00third-party entity is converted into
  606. 20:01nouns. And the relationship between
  607. 20:04those entities are converted into what
  608. 20:06is called as graph relationships. And
  609. 20:08And we have an interconnected
  610. 20:11system of nouns for work. Like we call
  611. 20:13this as our system of work. Like all the
  612. 20:15Atlassian products are somehow connected
  613. 20:17to this whole Teamwork Graph and system
  614. 20:19of work.
  615. 20:20And in the context of RCA, what this
  616. 20:22gives us is a subgraph of deployment
  617. 20:25entities, pull requests, commits, Jira
  618. 20:26issues, repository, service. Like the
  619. 20:29catalog contains a repository and the
  620. 20:31service. And it helps us in grounding
  621. 20:34our investigations in facts.
  622. 20:37The The other part that that we don't
  623. 20:38actually own, but it's available in most
  624. 20:40of the observability systems and network
  625. 20:42monitoring systems is the service graph.
  626. 20:45So, a service catalog like Backstage or
  627. 20:47a service graph like a New Relic service
  628. 20:48graph is the is the system of record for
  629. 20:50this.
  630. 20:51And it gives you information about what
  631. 20:53are the upstream dependencies of a
  632. 20:54service, what are the downstream
  633. 20:56dependencies. Let's say you had an
  634. 20:57outage where
  635. 20:59checkout is not working for bank X and
  636. 21:01the underlying cause turns out to be
  637. 21:03something in
  638. 21:06a dependency service like promise
  639. 21:07service, right? So, these systems will
  640. 21:10trace the dependency from your
  641. 21:12user-facing checkout service all the way
  642. 21:13to the problematic service. And we
  643. 21:17get this information from SOR, but we
  644. 21:18also try to store this in our teamwork
  645. 21:20graph, so that at the time of incident,
  646. 21:21we have a holistic view of things.
  647. 21:24And the
  648. 21:26the end goal of this data is that we're
  649. 21:28trying to construct a sub graph which
  650. 21:30the AI can use. This is actually like a
  651. 21:32pretty big sub graph.
  652. 21:34And
  653. 21:35it it connects the incident with the
  654. 21:37service. It connects a service with all
  655. 21:39the changes like which repository is the
  656. 21:41service hosted in, what are the PRs,
  657. 21:43commits, and deployments that happen on
  658. 21:45this repository. It connects a service
  659. 21:47with a service. And and all of this is
  660. 21:48queryable in a
  661. 21:50uh like you can think of teamwork graph
  662. 21:51as a Neo4j like graph which the AI can
  663. 21:53use anytime it wants.
  664. 21:56And and the last important part of
  665. 21:58building block of an AI SRE system is
  666. 22:00the agent orchestration itself, right?
  667. 22:02I'll I'll start with a logical view.
  668. 22:04So, in the logical view of an RC agent,
  669. 22:07it had it is connected to like three
  670. 22:09different sources of data. So, there is
  671. 22:11change data, there is system knowledge,
  672. 22:12and then there is observability
  673. 22:13knowledge.
  674. 22:14And majority of this context comes from
  675. 22:17our teamwork graph context. And whatever
  676. 22:20is not available in teamwork graph, that
  677. 22:21is where we start using MCPs and CLIs to
  678. 22:23get
  679. 22:25to to augment it with the additional
  680. 22:26context.
  681. 22:27So, an example of change data can be
  682. 22:29GitHub or Bitbucket deployments, PRs,
  683. 22:31and commits. And static feature flags
  684. 22:34and Jira issues. So, the teamwork graph
  685. 22:36connects all of this and and gives a
  686. 22:37view of
  687. 22:39how does change influence the current
  688. 22:41problem.
  689. 22:43The second part is a system knowledge,
  690. 22:44which is there is structured knowledge
  691. 22:46and unstructured knowledge of a system.
  692. 22:48An example of a structured knowledge is
  693. 22:51your backstage system backstage software
  694. 22:53services and software service
  695. 22:54dependencies.
  696. 22:55But anything that is config driven is
  697. 22:57not going to going to reflect the real
  698. 22:59world. So, this is also connected to
  699. 23:00your
  700. 23:01active systems like neural network
  701. 23:03service graphs or uh
  702. 23:05distributed tracing based service
  703. 23:06graphs, right? So, it it reflects the
  704. 23:08real world and it's not some stale JSON
  705. 23:10that that is present.
  706. 23:12An example of unstructured data is
  707. 23:14architecture documents and
  708. 23:16RFCs and PRDs and confluence which tell
  709. 23:19like what each subsystem is all about.
  710. 23:21And the
  711. 23:22the teamwork graph is obviously per
  712. 23:24tenant, so
  713. 23:25it is grounded in the unstructured and
  714. 23:27structured knowledge of that specific
  715. 23:28tenant without using world knowledge.
  716. 23:30Like
  717. 23:31in addition to world knowledge, right?
  718. 23:33The third part is whatever data is not
  719. 23:35available to us, we we depend on MCPs
  720. 23:38and other systems to integrate that
  721. 23:40data. Like things like uh
  722. 23:43Prometheus metrics or Splunk logs are
  723. 23:45something that come through
  724. 23:47on demand on demand queries, right? The
  725. 23:49end result of the RCA agent is that it
  726. 23:51tries to come up with multiple
  727. 23:52hypothesis, like a ranked list of
  728. 23:54hypothesis.
  729. 23:55It's It's not just integrations and data
  730. 23:57here. There is also like obviously some
  731. 23:59amount of intelligence supplied at each
  732. 24:00place to know how to traverse all these
  733. 24:03dependencies, right?
  734. 24:04So, the structure of this is that it's
  735. 24:06made up of multiple subagents. It's made
  736. 24:08up of multiple skills.
  737. 24:10And for data that is not available
  738. 24:12through the context graph, it depends on
  739. 24:13multiple MCP servers.
  740. 24:15And there is also some uh
  741. 24:18internal intelligence like in the form
  742. 24:19of ML models and uh
  743. 24:22ML models for changes, ML models for
  744. 24:24anomaly detection, ML models for entity
  745. 24:26understanding, and so on.
  746. 24:29This is like a typical stack that we use
  747. 24:31for uh
  748. 24:32subagents and skills. It's It's just a
  749. 24:34layered architecture, which means that
  750. 24:35there is a agent harness or runtime.
  751. 24:38You can think of the agent as
  752. 24:40uh as a chef here. The analogy is a
  753. 24:41chef. And the agent harness depends upon
  754. 24:44multiple skills and multiple
  755. 24:46tools.
  756. 24:47You can think of the skills as some kind
  757. 24:49of recipe in prompts and some kind of
  758. 24:51recipe even in code, right? Like prompts
  759. 24:54cannot express 100% of the scenarios
  760. 24:55that a complex system is trying to
  761. 24:57solve. An example of skills here is
  762. 24:59again how do you do feature flag
  763. 25:01analysis? How do you do uh pull request
  764. 25:03analysis? How do you do metric analysis?
  765. 25:05Log analysis?
  766. 25:07It can't just be expressed all in
  767. 25:08English. So, it also needs a connect
  768. 25:10collection of MCP servers which
  769. 25:12translate that into
  770. 25:14uh API calls. And you also need to
  771. 25:16connect to your backend services which
  772. 25:18are acting as the intelligence layer,
  773. 25:19right? So, the interfaces like MCP CLI
  774. 25:22or function tools and even the backend
  775. 25:24services act as like the like the
  776. 25:26analogy to a chef here would be that
  777. 25:29these are the ingredients and appliances
  778. 25:30which are actually used to get the job
  779. 25:32done.
  780. 25:33And this can contain the intelligence
  781. 25:35that you're using in order to do an
  782. 25:36effective RCA, right? And obviously
  783. 25:38there is evals and stuff like that which
  784. 25:40I've not shown here. So, every problem
  785. 25:41will require some kind of evals data set
  786. 25:44that that is being used to solve solve
  787. 25:45the problem.
  788. 25:47And that concludes the building blocks.
  789. 25:49So, to summarize the building blocks of
  790. 25:51a AI SRE agent are
  791. 25:53one is the agent orchestration, the
  792. 25:55second is the data that you're using it
  793. 25:58up for grounding and context, and the
  794. 26:00architecture elements like context
  795. 26:02engineering and uh
  796. 26:05LVM orchestration and the agentic
  797. 26:06interface. And the other part is how do
  798. 26:08you actually represent runbooks? How do
  799. 26:10you uh
  800. 26:12do an investigation?
  801. 26:14I'll briefly talk about some of the
  802. 26:15challenges we've encountered before
  803. 26:17going into questions.
  804. 26:18So, uh
  805. 26:20we've encountered a lot of challenges,
  806. 26:21right? Like uh
  807. 26:23some of the significant challenges in
  808. 26:24change-based analysis would be
  809. 26:27like these there are causal
  810. 26:28relationships. Like these are uh
  811. 26:30unstructured relationships. Like your
  812. 26:32problem is in a different domain
  813. 26:33language, your features are in a
  814. 26:35different domain language, right? So,
  815. 26:36you're trying to interlink a problem
  816. 26:38domain with a feature domain. The second
  817. 26:40kind of problem is that a service graph
  818. 26:43is not enough. Like a service graph does
  819. 26:44not tell you which API calls which API,
  820. 26:46right? It it just tells you that service
  821. 26:48A calls service B.
  822. 26:49And the third is that uh like because
  823. 26:52it's a product like the the third part
  824. 26:55will vary between a product solution and
  825. 26:57a platform like an internal solution.
  826. 27:00Because it's a product we we deal with
  827. 27:02different levels of code understanding.
  828. 27:03Like we don't get access to the entire
  829. 27:04GitHub code of a customer. We get access
  830. 27:06to maybe the PR description and the
  831. 27:08commit description. So what you can do
  832. 27:10with
  833. 27:11with code understanding and without code
  834. 27:12understanding are widely different. Like
  835. 27:14some customers give us code
  836. 27:15understanding code access, some
  837. 27:16customers don't. So deep code
  838. 27:18understanding makes your RCA way better.
  839. 27:20But the challenge is also that you have
  840. 27:22to do it in a repository. Like what we
  841. 27:24are trying to do is a RCA without
  842. 27:27like it's not like we are using cloud
  843. 27:29code to login into repository and do an
  844. 27:31RCA, right? We have thousands of
  845. 27:32services and maybe like thousands of
  846. 27:34repos and we are trying to do an RCA
  847. 27:35across these thousands of repos.
  848. 27:39Observability has its own set of
  849. 27:41challenges. Like first is the scale the
  850. 27:43sheer scale of observability data.
  851. 27:46And second is that most of the MCP tools
  852. 27:48for observability are very low level in
  853. 27:50nature. Like it's it's mostly about uh
  854. 27:53how do you get metadata? How do you get
  855. 27:54raw data? How do you construct a query?
  856. 27:57It's it's that that is not enough for
  857. 27:59you to translate from a problem to a
  858. 28:01query.
  859. 28:02And the third part is you need to come
  860. 28:03up with way lots and lots of dynamic
  861. 28:06plans depending on the data sources,
  862. 28:07right?
  863. 28:08Yeah.
  864. 28:09Yeah, these are the list of challenges
  865. 28:11that we've encountered in this. Yeah.
  866. 28:14Yeah, thank you.
  867. 28:20>> [music]
  868. 28:35[music]

About this transcript

This page contains the full transcript of AIOps: Leveraging AI for Incident Root Cause Analysis - Sathish Kumar by Developer Summit, generated from the public captions YouTube serves with the video. The transcript has 5,304 words across 868 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.