From Logs to Decisions: Autonomous AI Agents for Real-Time Kubernetes Threat R... Willem Berroubache
About this talk
This talk focuses on the implementation of AI agent security in Kubernetes environments to enhance protection against cyber threats. The speaker, William Burbage, who serves as a security architect at Orange, discusses the importance of integrating AI features into security monitoring solutions for detecting and mitigating vulnerabilities. He highlights the gaps often found in legacy security stacks and emphasizes the necessity of real-time threat monitoring alongside traditional security audits. Utilizing tools like Falco and machine learning models, the speaker demonstrates how data can be leveraged to identify abnormal behavior and facilitate prompt remediation. The session also addresses the need for human oversight in security processes, ensuring that operational teams can effectively respond to alerts and implement appropriate measures against potential attacks.
Full transcript
In this session we're going to speak about AI agent security Kubernetes of course and how we can use all the solution against threats hacker and so on to protect our environment cluster and customers. Is there some people here who works on security topics and agent topics? No? Okay. Okay, great. So just to quickly introduce myself. I'm William Burbage. I'm working for Orange in Paris. Today I'm an
architect. I'm designing some security solution for security monitoring detect threats attackers hackers to protect our customers infrastructure I'm trying to add some AI security feature to protect and improve the security monitoring this architecture in our environment. In the past I I worked in 5G information system for network edge production Kubernetes cluster and I'm also the trainer for Orange Group one of the different trainer for Orange Group
on Kubernetes and cloud native solution. So just a quick disclaimer. All the things you will see during this session is unfortunately if I can say it like that based on different true story and it's not a fairy tale. So let's start. There is a big gap between what we expect about security from our partners what we can have in our different we all know this case where
some people said yeah we are secure. Yeah, we are cloud native and so on. But if you check if you are looking what you can have behind all this environment solution and so on sometimes you can have legacy stack all and if you are going to check and audit what you can have you can for example have a lot of CVs. Okay? you are going to check
exactly how they can deploy the application the animal file the helm chart and so on and trust us sometimes it's not very beautiful. with all these problems some breaches in our infrastructure and this a direct threat for customer trust user data and Orange environment. the audit is great. Okay? But runtime monitoring is non-negotiable because the audit say okay you're secure with this release in this version. But
you need to prove you're real you're really secure in production. okay. During security audit you can check our bug CVs and so on. Okay, but the hackers can try to exploit your application to find some exploit. Okay? And for example destroy your application. you will check the application the manifest the helm chart all the declarative information to deploy this application but it's just the map. Runtime events
are the real world general. we have full stack observability of course in our different environment. It could see the pod bridge the security issue but we need to understand to see how we can improve this automatic detection and have immediate remediation. So the threat is life. Okay? And the defense should to be too. So sometimes security and not only security operation could breaks at scale we have
a lot of logs we have of events lot of metrics alarm and so on we can have some fatigue with because everything could fire. Okay? No signal matters and people could mute the alarm. We have some event which are visible but the threats could be stay invisible. Okay? So that's why we're trying to improve the context to better to have a best understanding with this different don't
be blocked to take some decision. Okay? And avoid remediation lag with manual fixes slow response and all these elements raising the risk for our environment and customer. So the idea to use AI and to silence the noise understand all the signals and automate the response. AI is not just a tool. It could be the brain of our different was missing. Today we have a lot of a
lot of tools a new protocol like A2A MCP chat GPT and so on. So we need to try the best product the best solution to improve our operation our security operation and detect the threats. So the idea here is to use different agent different tools to improve the detection. So we collect a lot of data with logs metrics alarms from elastic from a few tools and so
on. You know all the observability stack on the CNCF landscape but we need to use all these data to detect some security events. So we have Falco for example. Okay? We can get all this message all the event try to correct to correlate sorry with application logs with some machine learning model isolation forest and so on to detect if it's an abnormal behavior and then send them
to our different agent. Here we have the coordinator agent. This agent will talk to the threat and the list agent with A2A. This agent will collect different information different event to say and try to give you this information if it's a real threat to give you some information about matter attack matter of fact if you work on 5G environment and so on and then pass this information
give this information to this remediation agent. This agent deploy in different cluster will help you to understand which tools you have in your environment offer you the best remediation possible. And then will send this information this coordinator will send this information to the notify agent. You can use matter most slack teams and so on. There is a lot now of tools in the market to send this
information to operational people. Okay? To help you to give you the remediation and know what to do against this threat this attacks. Okay? And to approve this remediation. Okay? To keep human in the loop. Okay? To have a feedback loop on improve this detection. Okay? And it allows teams to validate and to take the right choice. Okay? Finally there is three big axes feedback loop training pipeline.
Okay? To improve detection work signals and so on. Okay? And real time inference to have more context more understanding against against the threats and take the right here it could be an example what we can have in an alarm. We have a a little demonstration just after. So you can have the pod the name of the pod. Okay? What happened? Justification impact of if it could be
a false positive. So totally straight side we have developed a web application based on real vulnerability. We continue to fight in unfortunately. So this application is like a shop where you can buy some swag and so on. You can to bypass some security feature and so on. Okay, so here this is your orders. So you can use this web interface maybe to inject some command. Here nothing.
No order found. But if you check the response of the server you can list the file in this container. And yes in some environment we have this not in production of course, but some product could have this security issue. And we don't want to be to to have these vulnerabilities in production. So okay, let's continue. Let's try to list the files, get some secret rights I can
have the file in ETC directory. get all the file paste W paste W D shadow and so on. Okay. we have some security issue. We have an alarm. Okay. A shell spawned in a container with a a command injection. And we can have what an attacker could be do with is this exploit. We have what happens, justification and so on. Okay. So now let's do something else.
So we know we can now inject some command. So let's send in this container a crypto miner to be rich. Why not? now we will have a new alarm because we are continuing to injecting some commands in our container. Okay. And we are downloading a crypto miner. So now we have all the path all the chain to understand what the hacker have done in our application. Okay.
It allows us it allows us to have a deep understanding of what this guy is doing in our environment. Okay. So great. Now new alarm. Okay. With XML. Okay. Crypto miner. With the previous alarm that we can have here and what happened, first positive present and so on to have a deep understanding of this alarm. Right. here crypto miner as you can see and the So what
to do now to protect our environment against this attack. So here you will have a more detail explanation about this remediation, how to implement it. Okay. And you can have the choice to approve it or not. Okay. Because we know I can give you not very beautiful remediation. So keep human in the loop and approve it. So we have the attack the path, what could be the
impact. Okay. So now apply Our agent is working to implement all this remediation. So now okay everything seems to be okay. let's check in our cluster what we have. So we have label now threat detected here. Let's stop. Okay. Great. now we know in this container in this pod we have a threat. Okay. And let's check the other remediation implemented here in this container. Cilium network policy
to isolate and block the ingress traffic. So now the crypto miner could not send the information. Okay. What else? We have a vulnerable image. Okay. So now we have a policy with key word now to block now this image to prevent new deployment of this vulnerable Some practical outcomes about that say it's allows teams operational teams to have a contextual fidelity with cluster set cluster set tool
set. we can eliminate some operation noise to be focused on real attacks, real incident. And it will assist the teams during the defense because you can have fast commitment block the application. Okay. And keep human in the loop to validate It allows for one sec teams and help them to have this even correlation history for deep investigation. today we are continuing to finalizing this robustness the solution.
And we will test it in more sensitive environment. we know I I can have some risks. Okay. Hallucination, cascading errors model drift, overconfidence effect and we have human safeguard. So we have ethical judgment. Okay. Context awareness of your company environment, human oversight and validation. So that's why today is still very important to keep human in the loop. agentic AI is a new operating paradigm. we can offer
scale and speed to the different teams, the different project, the different environment. Humans have great judgment. AI could help you and process all these alarms elements classical automation execute some rules based on trigger, environment, message and so on. Okay. And agentic AI could help you to understand the goal. Okay. Choose the tools and adapt to your context. It's like a new primitive stack because agents are ones
plugins. Okay. It's like a new layer. So containers change deployment. But agents could help you to change your operation be more efficient. So agentic AI is a new paradigm and the real challenge is how we adapt it. Thank you for this session. If you have some question if you want to follow this topic, don't hesitate to send me emails to reach me to send me message and
LinkedIn and so because I think we have some great things to do with this topic, to share some code, some resources some resources and feel free to reach me. Thank you everyone.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32