About this talk
This talk discusses the role of Site Reliability Engineering (SRE) in ensuring the reliability of complex systems through the use of AI agents. The speaker explains the challenges faced by SRE teams, particularly cognitive overload due to the multitude of tools engineers must navigate during incidents. The solution presented is a copilot system powered by specialized AI agents that assist in incident response by analyzing data, correlating alerts, and providing root cause analysis without overwhelming the engineer. The talk also highlights the collaborative capabilities of the agents, enabling them to share context and knowledge seamlessly, enhancing overall efficiency. A live demo showcases the functionality of the copilot and the effective integration of the AI agents in action.
Full transcript
Hello everyone. First I would like to thank our kind ghost for the nice introduction and to thank all of you who joined our session today. During the session we are going to show how series ensure reliability in complex ever evolving environment. and also how AI agents help SRES to work smarter and stay resilient. At the end, we'll wrap up with a live demo to bring up everything
to life. Let's begin with a brief overview of uh our global site reability engineering presence and the operating model. Site reability engineering combines software engineering and system engineering practices in order to ensure reliability, performance and scalability of production systems. It enables us to move quickly without compromising the stability and reliability of our services. Our SRE organization is operating across five strate strategically distributed locations and in total
we are 12 teams covering 51 life services through a 24x7 operating model. This global distribution is intentional. It enables us to adopt the follow the sun support model ensuring continuous coverage and enables empowers our teams for sustainable on call practices. This structure not only enhances our response times but also um enables us to have a resilient approach towards achieving our broader objective to deliver high quality and
reliable services. Hello from me as well as all of you now know what stands for. Let's deep dive in one of our challenges which we face in our day-to-day operations and that's the cognitive overload. Imagine that you're an SR engineer who is on call and around at 3:00 the night uncritical word comes in. You need to figure out what's going wrong and you need to do it
fast. But here is the catch. You're not working only with one or two different tools. You're actually working with five or even 10. Look at the slides where the S3 teams are connected to different set of tools. For example, dino trace elk thousand times. These tools are powerful. Yes. But when the engineer is is in the incident at three other night, he's context switching all the time,
jumping between dashboard to dashboard, trying to correlate data manually and to filter out all the false positive awards. We asked ourself uh the question, what do we do to help the engineer focus on what it actually matters instead of drowning in a sea of dashboards? And here is the answer. That is where actually the idea of a copilot powered by multiple agents comes in. And that's what
we are going to explore next. As you can see, the chaos from before is gone. The complex mess mesh of uh tools connected directly to engineers has been replaced by intelligent layer. At the core of the copilot is a system of specialized agents. Each one designed to emulate the expertise and workflows of a real s team. They are built to reflect real decision making and responsibilities but
operate with speed, precision and scale. Instead of having the engineer context switching at three night between tools and systems, these agents integrate directly with the entire ecosystem, analyzing alerts, correlating data and provide possible root cause analysis. With this, the role of the engineer evolves. no longer overwhelmed with the previously mentioned cognitive overload. They are now focused on rapid incident response, root cause validation and recovery during outage
situations. Let's take a quick step back to the slide with the teams and the tools. We actually don't work only with tools. We work with each other as well because sometimes the actual root cause could lie down in a dependent service covered by by another SE team and this hand of getting them into the context sharing walks them to deep dive introduces delays. Just adding one additional
additional SR team to the to the call can take up to 15 minutes which in an outage situation this is a lifetime. Now you can imagine when we need to add two or three teams the time adds up so does the risk. And here again the same question. What do we do to help this? As Peter already said, our agents, they're awesome, but that's not all of
it. They're able to share context knowledge while working together in a cooperative way while the outage transfs just like the real teams but without any delays. And to understand how it really look how how it really works, let's look at what makes up an agent. Each one is powered by three core components. Let's start with the context. This is the intelligent profile which defines the agent's identity
and its operational boundaries. Each agent knows what domain or services it's responsible for, what tools it has access to, what types of incidents it typically handles, and even its own personality and operating style, which is shaped by the practices and characteristics of the SRE team it represents. This contextual awareness allows the agents to act with precision. They know when to take the lead. They know when to
support other agent and how to align with the larger incident response effort. The second core component is the tool set. This is the operational interface that gives each agent the ability to observe, analyze and act. This includes of course direct integration with previously with the previously mentioned systems like Dino trace, elkstack, thousand ties and many many more agents can sorry can query realtime telemetry, inspect service health
and trace requests across dependencies. beyond this static analysis of data, the tool set also lets the agents to take action. When needed, an agent can suggest a remediation step or fix like restarting a service, draining the traffic from a bad node, or even rolling back a deployment. Every action happens within a human in the loop process. So the interventions stay safe, visible and under control. This keeps
a balance between autonomy and oversight, fast response with full accountability. And finally the knowledge base a rich blend of architectural documents past incident data dependency graphs and procedural runbooks that is from where actually the agents draws from in order to make informed decisions. accessories. We actually always had knowledge, tons of in fact different runbooks, documents, everything. But if you remember the SW with the teams and the
tools, again, it was a mess. So our next focus with Sopilot to was to turn this mess into something actionable. So we're introducing a separated component so-called data order, which is another component separated from the actual agent. that he has only one job to combine all the documents and provide them as a single source of true to see the whole picture. uh for the one who you
are familiar with a you maybe see that we are stepping up a little bit the things from the using the normal vector databases introducing a knowledge graph that's simple as the vector database couldn't handle the cognitive overload which we had in our documents and how does it work we have a simple injection pipeline which create which creates all the embeddings reankings it does everything for the documents
and then in the actual graph we have the embeddings inside the graph node where all the edges are semantically connected. So the agent could just query the the information in a natural language. Here I can also share you one story from my years in university. It was like I was sitting in math lecture and it was like 5:00 like just now and I was not really paying
attention. I was thinking when I'm going to need this why I'm going to need this. Well, but to the graph came with me at Avengers and that is just a small part of the graph which we had to handle in the S3 copilot all right let's see all of this in action in the demo which we prepared for you our AI agents jump into the into the
driver's seat they will spot issues coordinate responses and keeping everything running smoothly without even a single SR service reliability engineer um lift a finger. So let's jump directly to the demo now. So this is the interface. Give me a second. Maybe I not sharing the right screen. Okay. So this is the interface of the copilot itself. As we already mentioned um we are our target is to
cover all of our teams represented with the uh specified agent. So here are are all the agents which we currently have which are named according to the according to the name of the team. So um I will start with this one partic one this one particular agent. So I would like first to check whether it's uh okay as a size of the screen is it visible and
it's is it okay for all of you okay so I'll proceed. So what I'm going to do right now is to trigger a simple prompt to this and um what I'm going to check with it is what's the current state of healthiness of one of the most important services in SAP which is the which is a web- based control plane for all of the services part of
the portfolio of the company. So I'm going to simply ask please check the current state of the CF cockpit. So as you can see one important detail when I triggered the prompt a process kick kicked in and actually executed an anonymization of my message which actually ensures us that no user uh will be in a situation sharing sens sensitive data and this data to uh end up
in the large language model because at the end at the core of the system we are using LLMs. So right now the agent already responded but um I would like also to point out that at any time we are able to track the process through which the agent actually fulfills my request. So here this is actually the chain of thought of our agent. we are able to
see for example um in what language actually I asked him to uh execute a particular um particular process or action. Then we can see that uh how the which tool actually the agent decided to use in order to fulfill my request and of course the actual two execution. the result of this execution and the final response here. And at the end we receive this message which luckily
our one of the one of my our um most important services is up and running and there is no issues as I can see a lot of my colleagues are here and they are calm. So um we during the uh presentation we mentioned that the agents are able to collaborate and hand off tasks based on the context and their um their uh profile uh and of course
the services for which they are responsible. I'm aware uh that uh another agent which is namely this one Mina s agent it's responsible for another pretty important service part of our portfolio which is called web IDE and it is self-explanatory it's simply a web- based integrated integrated development environment so from this particular agent which we currently used I'm going going to ask the same question but for
this service and let's see what will happen. Can you now check the status of the web ID? Again at the beginning it's it's almost the same but the most interesting part right now will be exactly in the chain of thought and where we we will see something pretty interesting. As you can see, the agent already decided to use a particular tool which is called handoff. And the
idea of this tool is to provide the agents the ability to hand off tasks which for which they are aware they are simply not capable to fulfill but they know which of the other agents can help them. And that's exactly happening what's happening here. Similar to what Lubir explained, if we are not capable as a team to handle a particular issue, we are calling our colleagues and
they help us to do that. So you see here this handoff happened to the right agent and here we can even see what's happening under the hood. So uh handoff has been give me a second to find it because I'm asking for the availability status of web IDE and actually uh the agent which is capable to do that is exactly this one and it's already executing that
you see the other one already um started working on my requests and already actually provided the state and again I'm really happy that everything is up and running as expected. um unfortunately for the purposes of the demo of course we are not able to break our productive systems because right now what we see here is our productive systems which are located on 40 data centers. Actually in
our organization we are handling all that and uh for the purpose of the demo we introduced a dedicated agent demo agent simple name which is um and uh and together with that we introduced a simple application. This application is a python app which actually simply queries a database posgrsql database. So um at at first I will just trigger the same request to this agent. I would like
to know what's the current state of my demo application I'm expecting that it will be uh the application to be up and running as expected but let's see the agent is doing its stuff thinking and as expected the application and the message also confirms that. So right now I'm going to use this control plane which explained previously during our uh during the session with the previous agent
and it's the uh cockpit. What I'm going to do here is to simulate an issue and it will be simply uh related to that at here is the uh the demo application. I'm simply going to navigate to the instance of the database and we'll delete the key which is required for the application to be able to interact with the database. Without this key, the application will simply
not be able to authenticate and use the database. So I'm going to delete it. Straightforward. So the deletion is already in progress. I'll wait a few seconds more. It's gone. And now I will ask the agent to refresh the state to see whether the agent is aware that there is an issue and the same check will be executed but the expectation is that the result will be
completely different. And as expected it is the application right now is down because it's not able to connect to database and execute queries and operations related to it. Let's wait with the agent to finish. And of course the message confirms what I just explained. And imagine that I'm in a situation of out of an outage of one of our productive systems. I will be put in a
call and I need to act fast in order to resolve this issue. What's the first thing which I'm going to do? At least in my mind, the first thing which pops up is to check the locks of the application and see what's happening. For that purpose, I will trigger a command which will provide me an interface to extract the application locks. But not only extract and view
them but on top of that my agent is going to analyze the locks and pinpoint the most prominent issues which actually be which were observed. So I'm going to directly trigger the collection and the logs were already collected. I can see them in a row format here at the sidebar of the copilot interface and right now the agent is executing on that data and I would like
to see what's the issue because right now I'm simulating but in a real environment I cannot be sure that I need to go and recreate the key in order to fix the issue with the database. This can be anything and actually that's what I would like to uh show you Let's see. The analysis takes some time, but at the end we'll need to see the result and
pinpoint the exact tissue. as expected, the first the first point which our agent identified is a widespread database authentication failures. We can see more details in the drop-down. But this will guide me that there is something unusual related to how my application is authenticated to the database and that's why it's not able to work with it. So instead of going and recreating the key manually imagine uh
again imagine I am not aware about the issue and all the details around it. I am just aware that there is an authentication issue with the database. So what I'm going to check with the agent is can you fix this issue for me? Let's see what will happen here. What will be proposed to me at least if not fixed entirely. And again correlating back to the presentation
you see here the agent proposed to me to execute a recommended action. This recommended action actually is related set of steps which needs to be executed in order my application to be up and running again. And uh as I said correlating back to the presentation, the agent does not trigger it automatically. It waits for me to either confirm or to not confirm the execution of this recommend
detection. And of course, I want my demo application to be up and running again. That's why I'm going to confirm it right now. And the agent proceeds with the it's executing the recommend detection. We can observe it of course again with the um with the chain of thought and uh as soon as the recommend detection is started it's also checking the status whether it will be finished
successfully or we have an issue even with the recommend detection. Let's see what will be the result at the end. So a lot of checks here just pulling the state right now. It's just pulling the state of the execution of the particular recommended action and we are waiting for the final response. So here is the final response from our demo agent. The issue has been fixed. The
demo application is now up. And here we have also reference to the procedure which were were executed in order this to happen. So I will check one more time in order to show it visually here in the but it's already visible. Okay. Uh it it already do that for me. Sorry. You see we have the down state and right now the application is back up and running
as expected able to work with the database. One more thing I would like to show you also that the key were actually recreated for me following the procedures for which agent the the agent actually were aware because it's able to check its knowledge base what actually is needed to be executed in order this issue to be resolved. And here you see the key is available and at
the end while I'm in the call the wall threat which we had with our agent let's imagine that Lubu is together with me in the call I would need to provide transparency what actually I'm doing here how I'm actually working on resolving this issue I can quickly share the threat with him and he will receive everything which I see here in the interface on his machine and
he will have the awareness I just don't need to update each and every minute I'm doing this or I'm doing that. So at the end that was our demo and uh as I see that we are already out of time. I would like to thank thank you for your time and your attention and uh we'll be really happy uh to if there are some questions or points
for discussion stepping on this topic to approach me or Luir after that in order to discuss these are small this the thing the things which we demonstrated are small piece of the capabilities which our copilot has. So thank you very much. Thank you. All right. Thank you guys again. Uh as you said, we don't have time now. So you can reach out to to Lubameir and Peter
after that. Uh now this was the final session actually from this track. I will highly recommend you to join our last closing panel in the main track upstairs. It will be very interesting with interesting guests. So highly recommend to join us. Thank you guys.