AI Agents & Platform Engineering: Efficiency Boost... Hasith K, Vincent C, Sara Q, Idit L & Carlos S
About this talk
In this session, Carlos Santana leads a panel discussion on the integration of AI agents in platform engineering, exploring both their potential benefits and the challenges they present. The panelists, including experts from AWS, Red Hat, Solo, and Cisco, share insights into how AI is currently being employed within organizations, focusing on areas like automated incident management and operational efficiencies. They discuss the importance of trust, particularly in regulated environments, and the need for a structured approach to manage the interaction between humans and AI agents. Additionally, the talk highlights the significance of developing clear use cases and establishing safeguards to ensure the safety and reliability of AI systems in production.
Full transcript
Okay. Uh, welcome to the last day of CubeCon. Who's sad that this thing is ending? >> But don't worry, we'll have another talk about talking about AI. I was excited Cool. Um, so welcome everyone. Uh this is a session on AI agents and platform engineering efficiency boost or new source of trouble. Um my name is Carlos Santana. I'm a senior solutions um architect at AWS working primarily
with EKS customers, people running Kubernetes on on AWS. I'm building platform engineering um uh platforms. Uh let's have um a few seconds so everyone can introduce yourself and pass the mic. >> Good morning everyone. Vincent. I'm the chief technology officer for Redat in Asia Pacific. >> Hi everyone. I'm Sara. I'm a platform engineer working for entity data. >> Hi, I'm indeed Lavin. I'm the founder and CEO
of Solo. >> Hi, I'm Hasset, the director of platform engineering and CISO for Cisco's incubation unit. So, so to kick this off, uh we're going to have um an open conversation in this panel around um different areas just to I'm going to be the host and MC and also collaborate uh with my input around where are we today with in terms of AI um tools moving into
production platforms and what platform teams are doing about it. uh today what are the top challenges when you want to introduce um aentic workflows or maybe maybe support agentic workloads into your platform and the next one we're going to talk about um trust and trust is something that can go from trusting the model trusting the context uh but I think one of the most important ones and
and is going to agree with me is trusting the security who's actually making the call. Are you making the call on behalf or you calling it uh with uh permissions that anyone can do anything? So, and the last one is um how do we get started? Because I think um this is probably maybe the second CubeCon that the Linux Foundation and CNCF have uh started to pay
more attention into agentic um workloads. uh in the colloccated event uh on day zero we had a session we have a uh a track um in here and it was packed uh who went to those sessions so we're able to get in um I went by and there was a big line outside so I think there's a lot of curiosity for folks that want to learn about
this so let's jump in where are you um actually seeing AI agents using um in platform engineering today so some of the agentic uh use cases um what are uh areas that are being used today because there's areas that people are not putting in production today. So let's start with Vincent on from Red Hat. What are the current things that people are doing today in production that
are are working? >> So I think two two areas to look at is internally and externally for us. We are a software company. So of course you know agentic DLC I think is the type of use case that comes to mind first. Everyone is obviously very interested in generating code. But the reality is Agentic SDLC is a lot more than this, right? You're looking at how does
it impact your quality engineering? How does it turn your documentation into something that is more proactive and obviously allows you to be more consistent and probably better in terms of generating outputs. So there's a lot of opportunity there, but it's still very early and there's obviously a huge challenge in how this will integrate with open source projects in particular. So do you accept a contribution from an
agent? What are our quality standards? How you know when do we trigger a review? So that to me is big work in progress. The second biggest opportunity I feel is for the user. So platform engineering perspective. I think we've gone through the past couple of years or so in a huge consolidation of of telemetry with open telemetry standard. And now if you look at it from a
data science perspective that opens an avenue to feed this telemetry data into AI ops capability. So again we highly anticipate this will be done by agent with some degree of autonomy but there's a lot of opportunity here to really uh automate a lot of actions that are currently performed by human but in a much more proactive and efficient manner. >> Thank you. And you said autonomous and
we'll get into that in a in the the next session about talking about trust. But Sarah tell us like what is what is actually in the ground happening in the S sur organizations versus maybe hype that people are trying to discern what is hype versus actually being done in an organization because a lot of people um I'm guessing we have stakeholders here from the BP CISO CTO's
they make decisions on how successfully other organizations are and then they they kind of like manage risk through that. So um what is actually working today with the with the level of tools that are available? >> So um today what we see most in production are uh Asians operating more in the uh read path more than right path. um Asians are uh we don't see today in
productions um uh autonomous AI running uh end to end uh platforms but more uh AI uh assisting humans and not replacing them uh which means that for example uh log summarization uh use cases like uh um uh quering uh like obser observity data etc etc. So they are more uh successful in this uh read path where they can uh bring more uh summarized information to humans that
they can act on it to make uh act act and uh and make decisions uh and execute. >> And indeed uh you have an offering um what what are people like putting in production actually know if people are using in production now or today right you have an offering or what are people putting in production today versus maybe not not yet. >> Yeah. Yeah. So maybe I
I'm adding I'm actually going to wear right now three hats, right? The first one is that we in solo, right? We're using a lot of AI. So I can share about what we're doing there. We're relatively very innovative company. So there is a lot of motivation for the people to kind of like do it. Then there is the open source project that we created and that's all
of you hopefully there. And there's you know a lot of the innovators happening there a lot of the stuff and then the last one is all our customers. Right? So I can actually touch you all but I will argue that they are so different. Right? So I don't know do you want me to focus on one specifically or >> Yeah. What what what in terms of a
um platform engineering are they doing uh CI/CD in production? They're more doing monitoring or they actually working on like generating >> YAML files or Helm values. >> So to me if you're looking at it AI in the organization at least that's what we saw in the beginning. It started with the AI team right basically they got all the money. We were the cool people before but then
they become the cool people right they got their money they're making a decision most of them is Python engineers never done infrastructure in their life right but they made the decision which gateway to run which you know platform runtime to run that was their decision that's changing right now and honestly I predicted that it have to be changed that why a year ago we created all those
project because in the nutshells it's all going to come back to us and we will need to put stuff in production so in purpose because I knew that it's not the Yet when I started with K agent which is the runtime you always need to start with the runtime I actually created it for us. So the use cases when we put there we basically created the runtime
and we put their agent like you know how do you know I don't know agent all this right basically and a cilium agent and the reason is because I wanted to attract you guys right because I feel that we as a is it called a community we're so solid need to own that that's going to be owned by us so that's a good start for us right
let's play with this until they're giving us the platon to basically do what we want so that so that's where we started So what we saw is that in the community a lot of people running it right now in production. So they help. Now the good thing about all of this and I know there is a lot of fear. We will not be replaced ever guys like
I mean we are tier zero application. No one will trust agent for that in probably five to 10 years. So we're in actually not the people that that they replace in my opinion. I don't know like for instance I'm just hiring all the time. So I don't think that I think we are in a very good shape but um in the nut people doing so they are
using a lot of the debugability of the the clusters first things that everybody using. Um so so kind of like as you said like s like you know observability understand what's wrong it's everybody I guess doing it that's the first one and besides that in our customer you know they planning right now and this is a movement that we saw which is very interesting now we talking
to customer that basically said no we will run everything on kubernetes every agent in the organization will run in kubernetes so this is awesome right awesome for us and and last and not least that I would say is that um in solo itself I mean oh my god I You can I can give you a glimpse of what's going on. All our GM GM G go to
market is automated by AI. All everything that we're doing you know leads you know opportunity everything everything in marketing is automated and the amount of of of automation that we have everywhere in engineering is insane and by the way we also using our own product so like for instance agent get like you know we don't want to pay a lot on AI we need to make sure
that some something is rate limiting so everything is going through agent gateway it scales and so on so again I don't know it's a lot of stuff >> um I I can add uh what we've seen in AWS we we're both Ra was saying we are both consumers using AI internally and offering uh services that are uh offering this AI uh agentic composition of agents. So internally
uh we use it for operations. Um one um maybe that I was surprised from the service team was that we had the uh support uh engineering organization using things like uh we have an a tool called kuro which is cow code and traffic agents uh responding and finding things that the service team was actually surprised that they were able to to find it like they were so
so much surprised that internally we say hey we should hire that guy when we look through uh the logs and the things that he did actually he's very good at using the tools an agentic he was able to correlate the logs with the timing with the pipeline the previous version the next version and the source code but it wasn't the person it was using the right tools
with the right context with the right access uh to do that and uh externally um customers are using in production they're they're trying uh to use it for maybe read only uh operations in terms of like uh we have an offering called devops a devops agent that you can uh correlate uh has access to to logs to traces the topology of things in AWS. You have um
uh things that are related. You have the the servers with the networking with the microservices with the different aspects the topology awareness of these agents and having root cause analysis uh just when an event happens. So the the the person comes in and the agent have done all the work to find the correlations of all these systems and actually finding the root cause. So it's a matter
of and also helping with the with with that. Uh so in AWS we see a lot of customers using it for operations in terms of read read only and incident incident management uh and internally uh for um for operations right we also use it in in that in that in that use case. Um so talking about that hid um for you uh what have been you been
surprised um because I think you offer this agentic platform and you can talk a little bit to other organizations in your uh company. What have surprised you that they're using it for today? >> Uh I mean it's like slow adoption initially. Uh I mean main interesting thing is uh getting started because I wear two hats both security and platform engineering falls under me at the incubation unit.
So you can go quite fast uh and one of the challenges is usually from the you know security organization when it's separate how do you implement these uh capabilities properly and uh uh actually make them available in a large scale. So at the incubation we don't have that problem because I own both hats. Uh and we have been doing um agentic AI for platform engineering everything in
CI/CD both uh you know read and write and taking actions uh since end of 2024. uh that was very successful internally and that's why we decided to open source that uh project creating um uh cape uh which we donated to uh cloud native um sandbox um last week uh so stands for community AI platform engineering uh the whole um point is uh there's a lot that you
can do but it's very important to have the platform layers and uh my focus uh this year is to really uh you know prevent uh everybody creating you know different uh agentic systems whether you know not not just uh software development but your marketing team your you know business intelligence team your your product managers now I mean we are in a uh time where your product managers
would build something and do a pock and even uh have uh early early customers there right so how do we make sure that we do this in a way that uh can really uh is sustainable right that's kind of uh the surprises and challenges. But >> but that that will move us to to the top challenges. Um I included a few few ideas for for folks to
know what are the top challenges. But you brought one of like um internally everyone now that it's so easy apparently to create a platform. People are creating their own platform or people that are not done platform before they start creating their own tools or creating their own services. So you're trying to uh consolidate that into like let's do everything in a in a in a in a
standard way inside your company. Can you talk about what other challenges do you have in terms of like you said right access you're giving right access to to agents? Yeah, I mean we so today you can create you know LLM keys development machine Kubernetes clusters you know any of that is uh you know from creating a new application all the way to production we have agentic flows
involved however based on the risk and what you're being done you'll pull u uh humans uh into the loop uh now now the interesting thing there is uh you know for platform teams we have always been a bottleneck to streamline line by you know it's sort of a deliberate design right uh this really allows us to uh take all the toil and delegate that to AI and
improve our systems and and keep things uh you know competitive because if we are not competitive AI tooling and humans plus AI tooling are very good at finding loopholes and uh working around things and you you're going to get surprised in so many ways >> so it uh about challenges um You are offering a platform right to create agents and and build agents. Any challenges with skills
like are the users are because the performance of that agent depends on writing the right skills for the organization knowing what what NCP is what A2A A2A is all these lingo new things on AI. What do you see about people in those organizations of saying now you have a way to create agents we just need to learn how to write good agents or what it takes to
write a good agent. >> Yeah. So what we see in organization today and again I'm talking more about customers maybe not in the community well actually the community >> everybody getting like a you know a charter that charter is MCP right now I think this is the easy one and the first one that moved to us and the reason is because this is something that you know
stand not it's not something that uh that the II want to deal with. So what we see is that part is starting to move to the to the platform engineering. Um so so what do you need when you have you know what is the challenge when you're doing this as you said like what they see is they they worry about shadow IT they worry about security they
worry about identity spool because you have now an MCP spool that's kind of like showing you identity SP because in your organization maybe you hunting Andra or octa but then when you're going to Figma or to GitHub or there is their own identity and how to work it if you didn't see his talk you really need to watch back Christian Posta talk about this because they did
an amazing job there and explain what we're doing with all our our customers but that's a big things that you know this is like almost like 100% of the people that's what they're working on right now um then I we see in the platform area right we see two things we see a lot some organization just saying let's just go with agent core or you know a
little bit vertex we see sometime AI foundary because it's like they you know they don't want to decide excited. This is why you know to me it's very important we eventually will come to Kubernetes in my opinion and this is why if you saw I don't know if you guys saw what we created with the registry is basically exactly kind of like this abstraction that will allow
you to go to all those platform and so that's that in terms of how to learn the technology I mean skill is going to be the future probably even more in my opinion more than MCP uh I think that if you're looking at the the the vision of anthrop topic I know I think in my opinion and this is totally my opinion um all those companies that
doing today you know a workflow right so basically I don't know into N8A or even in my opinion they will be in trouble in the future because the thing the way I see the future and I think a tropics see the future and others is that eventually and and even we talked about A2 you mentioned we don't see that used a lot I will be honest and
the reason is because multi in my opinion multi-agent it will be an interesting subject but I think that the majority of the things will have be a general purpose agent a shell right if you saw what Nvidia did with open shell that that's that's the idea and it will get its functionality with MC with tools and skills that will be attached to it right and if you
think about what will happen the agent itself is going to be the orchestrator there's not going to be a state machine the the skills is going to be the note right the one that is actually executed and By the way, those skill might being actually written by an agent which basically mean that it need to run in a in sandbox which will be a very big subject
in my opinion in the next year. We're working on this like crazy right now. So it's bringing a if this is the future that there is a lot of interesting stuff to learn and a lot of problem to solve. For instance, this orchestration that the LLM will do is not deterministic. How do you taking care of this one? So that's why we announced eval yesterday, right? Because
that would be very very important because you're not controlled. It's not deterministic. So, so I mean I don't know I can take hours about the challenges. Oh my. >> Let's move to the next. >> Um so the the next and you said that they're not deterministic. So Sarah um what you've seen about like the levels of straws I have some here like human in the loop. Uh
you mentioned that yesterday we were talking about that you always need human in the loop or you need gut rails in terms of autonomy supervised full full automation. Do you see internally in your organization uh team members or organization asking you like this is so good to uh do incident management or go through logs? Why not have an autonomous agent that runs all the time and actually
fixes the problem? Have you seen that that request and um it goes with trust? Like are you going to trust or like that's that's the the conversation that you're having? Are you any requests like that in your organization? >> Uh yeah, sure. I think uh everyone here uh experimenting with Asians had this uh kind of questions. Uh the problem is that we are we spend a lot
of like a few years building reliable testable and observable systems and we are trying to bring nondetestics components as you said. Uh the problem is that this uh breaks the reliability model we we built. Uh so uh the the the big question here is that how to contain this nondeterminism in uh the deterministic systems um and how to safely run and operate uh agents. uh the uh
agents are running based on nondeterministics models uh which means that we cannot model their failure and this is the big question that we are trying as platform engineers and and SRES to to to to solve and this is the in my opinion this is the big challenge we are going to have in like few coming years uh in different organizations. >> Yeah. So, uh, Vincent, for you,
um, in terms of trust, um, in your in your role in Red Hat as a CTO and you talk to a lot of enterprises and I guess like enterprises, I'm assuming that enterprises that have regulatory aspects and even from from Europe have more strict regulations um, and also like important enterprises that run the world, right? Utilities, water, uh, electricity, um, air airlines, right? um government documents and
health. What are these organizations saying about like the platform team or the IT team saying like we can make this faster or better if we use AI, let us use AI. What is what is the contention there of like these are regulated environments? How do you introduce AI into a regulated environment um and trusting the AI? >> That's a great question. And I I wear actually a
double hat because I'm also this year a visiting scholar at Columbia University and I'm working with a financial safety institute lab within Colombia. And so it's quite interesting because what we've decided is is very hard to talk about AI safety in isolation of the use case. So we decided to use financial as a heavily regulated environment to really start to do some fundamental research around how do
we look at AI system safety. This year we are fully focused on agentic because this is where obviously the industry is moving and the reality is we are very far from production grade run and and environment. You know we we are still at the stage of defining how do we actually evaluate the readiness of an agentic system for different use case. So evaluation I totally agree is
probably the the topic I'm working the most on right now and uh I I actually share that with the Nvidia folks you know last week at NVIDIA GTC you know if you are interested into the topic look at what they did on the deep research agent. So they were sharing with me that to build a deep research agent essentially took a huge team of Nvidia scientists six
months to iteratively tune the performance of the agent and I was asking them the question what was a single decision that you took that actually improve the performance and the answer was none. It took us hundreds of different decisions, tuning, experimentation to actually get to a stable performance and the only way they realized that it's actually doable is by uh you know building extremely strong evaluation harness
and being able to iterate and run those harness to test every single decision from an outcome perspective. So to me was a good learning and that comes back to what you were explaining. You know I I think the the conclusion of this is reliability for agentic system is statistical. It's not deterministic. We can only look at it in a statistical manner. I will leave you though with
a thought, right? Because I have this discussion a lot with businesses. People are a bit freaked out. It's like well you know I'm not going to be able to run uh my test and get 100% confidence that the agent will actually deliver. Let me just take an example in banking right? All of you probably at some point had to apply for a banking loan. What happens when
you get a banking loan? There's a lot of system involved. They are also people. People have a 250 page document that they are reading which tell them this is what you have to do to approve a loan or not. Most of the time they give you predictable outcome for extremely simple decision. When you have more complex decision where it's kind of well it could be yes or
no. two human can actually get a different outcome. Are we actually going in a different world with agents? We're not. We actually reproducing a lot of existing business process just making them faster and actually more transparent. So my personal view is you shouldn't think of those system as getting 100% confidence. You should evaluate the system against their ability of doing better than the humans. Otherwise, you're kind
of approaching this whole topic wrong. I agree. I agree. Uh in terms of of trust, um one one thing that I maybe disagree to edit, right? Any every panel needs to disagree on something. Uh it's like cloud cloud skills um skills, right? which is context. Do everything through context these markdown files or text versus maybe programmatic programmatically adapters or uh gateways I would think I would think
right so I would take for examples in platform engineering we have to create environments that are in dev and they have a different configuration that in prod but the code still continues being the same. So one of the challenges and also trusting is like if you write a cloud skill is just differently in dev in staging and in production. Usually in Kubernetes we just switch environment environment
with config maps and as predictable we can switch which is the host name that is going to be accessed in dev versus production. So one aspect of the cloud skills that uh people are looking into platform engineering is like how do we make cloud skills that are adaptable and also configurable. So you want to use most of the cloud skill but when it comes to which access
where's the host name to access in this environment versus that environment you want to have the same cloud skill using different environments. So I think that's where MCP and AI gateways and the host names and environment variables we have to think about like not just giving access to the people using our platform that just give me a skill and I will load it. That skill needs to
be trusted right what is the level of trust that you're going to give that to the skill. in aspect right they always need a a gateway or adapter pattern uh to configure that >> answer that because you say I'm not saying that they don't need that all I'm saying is that they don't need the state machine it's very very different and state machine in my opinion not
going to be needed because it will be replaced by the LM but I 100% agree with you we're doing this we actually working dayto-day with skill we have a lot of you know we put skills in registry we put evaluator in the registry I mean trust me like that we 100% agree and get you know don't pitch me on getway I Okay. Uh, another thing is uh
you need different types of sandboxing. So, especially when you think about like a reasoning loop uh uh you give you know an agent u uh some context and some tools and hey go go and do execute this but don't do anything else. So that type of uh capability will be very useful. uh another thing to think about is uh now your CI/CD pipelines uh let's say you
have a baseline today with the level of uh um you know agentic SDLC going on and with agents doing other things this is going to increase 100x fold so think about uh how does your current uh infrastructure will be able to keep up with this and uh scale with the demand that that it's going to pose. So looking at the time, let's go into how do people
get started and and sandbox is something that u the CNCF is looking at it. You see it in the keynotes. If we're going to allow this autonomous agent, right, to write code on the fly versus where we used to write code on the development phase, we compile and that's the code that runs. That's alo change, right? This autonomous assistant, we'll figure out how to write code to
solve a problem. So sandboxing is becoming popular because we want to protect that container of not escaping the kernel looking things into uni kernel g visor kada containers that type of sandboxing also networking. So um because of time let's start with Vincent um how people can get get started or like what is the thing if they're already doing like aic and they have good experiments and like
you said uh some of the things are working today how can they uh get started maybe a a projects to look into uh open source projects to look into or maybe learning resources to look into so I think you know in fact anything you get started never start with the technology Start with what's your use case and how do you define value generation for your use case.
Second thing you have to look at again it's still not the technology is the data. What data do you have about your current process? How do you how are you going to integrate this data into your flow? So this is important because that will drive some of your choice from a technology perspective. Now you know there's two fundamental problem that you'll face as an enterprise. There's the
safety of AI and the AI safety. Those are two different problem. Safety of AI we can solve. We will solve it in 2026. Agent identity, how to be zero trust with agent. Look, we have freaking good engineer here. We have 13,000 engineer in CNCF working on those questions. We will solve it by the end of the year. I have absolutely no fear of that. What is a
lot more difficult is the the safety of the system itself. It's use case dependent. I mean we've we've discussed you know to me uh like take an example the the skill will not uh relate to the same uh to different model. So when you think your skills is good what does it mean? Let me tell you you swap from a quen model to a kim 2.5 you
will find that suddenly your skill doesn't work anymore. You've got to retest everything. So today the problem is we are really looking at systematic architecture of those systems which is a good news. It means engineer like ourselves still have a lot of work to do, still a lot of opportunity you know to get jobs. uh and at the same time you know those are the real complexity
and and so to finish though you know there is a number of technical element and dependency that will give you a lot of help to go into your direction to me I've I've also been saying similarly that AI gateway is a critical piece in your architecture if you don't have it you are going to spend so much time running around integrating pieces trying to understand what's happening
in your system your observability will not be easily available able you want to understand what's happening. So that's probably the first piece you need to define before you even start to architect around because it's basically your you know the orchestration of your entire system right it's it's like the brain >> can go ahead >> okay no I just wanted to say because the time is out I
think you think how to start I mean how did we start with kubernetes together right we are the community that's why we open on source all those projects so look at agent gateway in this sense very important piece and is in the Linux foundation look at k agent If you want to run those agent look at registry we just donated it u which I think will give
you every usually that's the next thing that you want and the biggest thing in my opinion will be two things number one is eval right so look at this and the other thing you will be limited by token so cost efficiency it's a huge thing look never mind I forgot about the the fund matter instead of the continuous they're progressing discloses look at these guys and yeah
good luck I don't But start with us, start with the community. >> Yeah. Sarah, um, advice for folks that are like looking into like do I do I add AI to the platform engineering stack or or maybe that's something to do it later or what how they can get started. So from a platform engineer perspective I think the starting point is to think stop thinking of AI
as a like a intelligent teammate but uh put it under like the same sale guard um uh guard rails as any other uh untrusted automation and from there try to build some uh sandboxing and uh uh limit the airbag and the u uh the uh the the execution context that we give to this uh Asians and from there start to build trust and uh and uh and
uh and expand the the trust contrast uh w with time. So for me like controlling what the AI uh in term of observability on term of permissions airbach etc and also um uh enabling policies to um override the uh behavior of the agents when they try when the when some like uh safety thresholds are are are crossed is the key points from a a platform >> Good.
So um we're out of time but I the key takeaways that I got was do system engineering end to end because the key skills and the models might change but you have a system it will work. Um security start with read only and as you progress you can give it more access uh check out projects like K K agent uh to get started easily in Kubernetes using
Kubernetes APIs and CRDs what things that you're already good at then you will get on top of it and uh we have se uh donating uh your project so if you have projects that are successfully right donated to the CNCF or join uh the project of of cape on how to make uh platform engineering even better uh using these type of tools. Um, so thank you so
much to the panel for today. Thank you for everyone that attended today. Hopefully this was a useful session. Thank you.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32