About this talk
This talk, presented by Anastasios Zafeiropoulos from the National Technical University of Athens, focuses on intent life cycle management within orchestration technologies applicable to computer science. It discusses the integration of Internet of Things, Edge, Cloud, and emerging 5G and 6G networks. The speaker explores how intent, defined as operational goals and outcomes, can be managed through application graphs and a knowledge graph that supports various deployment techniques. The presentation outlines a simulation kit developed for experimenting with different resource deployments and introduces an agentic version that incorporates generative AI and machine learning for improved intent management. Highlighting the importance of continuous operation data, the speaker demonstrates how to optimize resource usage effectively while ensuring service performance and energy efficiency.
Full transcript
[music] >> Maybe maybe we can we can start. Um okay, it's a transition to to an open source solution, but must more more heavy, I would say, in terms of of technology. So, probably the audience goes coming from different sciences. In this in this topic, we're going to speak on technologies applied to definitely to computer science and to mostly orchestration technologies related to Internet of Things, Edge,
Cloud, but also let's say 5G and 6G networks. So, my name is Anastasios Zafeiropoulos. I'm coming from the National Technical University of Athens, specifically the Network Management and Optimal Design Laboratory. And we're going to speak about intent life cycle management. So, coming to the domain of orchestration, whether we have compute, but also network resources, we want to deploy services and applications, and we want to manage their
life cycle um in uh terms of provisioning and resource So, for this, in many cases, we have a request. So, we have an intent that has to be satisfied and um um processed in order to uh manage the application in a way that the provider will be happy, but also the one that provides the resources will optimize the way that the resources are used. This may be
related to cost, to energy, to optimal performance, whatever. So, for this, we have adopted the intent definition coming from uh IETF RFC 9316. So, we have intent defined as a set of operational goals and outcomes described in a in a declarative way. Uh so, we have this definition, but then the way that you implement and you manage the intent is dependent on your implementation. Uh so, we
can consider that we have a high-level expression of objectives and constraints that have to be satisfied uh in in the way that we manage this deployment. And for this, we consider critical to uh semantics, actually knowledge graph, that we can um represent the information to manage the intent. Okay, in this figure, in this diagram, we try to see how we can model such info. I will try
to go shortly through it just to give some highlights. Okay, we have graphs. We call them application graphs, we can call them service graphs. It's a graph with a set of components interconnected that we want to have different type of deployments for these graphs. And these graphs have components. So, we interconnect components, and these components are connected through relationships. And for each uh type of them, we
make host information related to their type, if they are classified as computational heavy, memory intensive, network intensive, and it goes on. And this type of applications, they're somehow classified. Uh whenever they are deployed, we have an application graph instance. So, this is running. It has to be provided over resources. So, then we speak about deployed components and deployed relationships. And for each case, in the deployment, we
have an intent. I want to deploy a media streaming service achieve high performance or achieve low energy consumption as much as possible. the combination of the application graph instance with the intent is translated to requirements for the components and for the relationships. Both computational requirements and network requirements. And this has to be deployed over infrastructure. Infrastructure is coming on the right part of the figure from infrastructure
providers. So, we have cluster nodes and network links that interconnect these cluster nodes. And this goes on. Application is deployed, we collect lot of data, we process them, we check if we have uh good intent status or we have violations, and we can proceed to corrective actions. So, the important thing here is that information from different type of deployments is mapped to this graph, this knowledge graph,
and then we have it available for future deployments, for runtime actions. So, we are learning from the continuous operation of the uh net- network and computational infrastructure. So, we consider this uh very very important for for the scenario. And then we move on to this figure, where we have uh three control loops in the life cycle management of the first control loop is I get a request.
I have some high-level objectives, constraints, and the application graph, and I want to translate it. So, this can be also done, let's say, with generative AI tools, to a formal intent description. And I can also do kind of conflict resolution. If you ask me things that are conflicting, I can notify you that, come on, this you have to improve in your request, so to have a good
intent description. So, we this first control loop does this syntactic and semantic validation. Then, having this, we have a valid intent, we move on to deployment. So, we infrastructure coming from multi-providers. So, control loop two actually supports solving of op- optimization problem. What is the best way to serve this intent based on the offerings of multi-providers? uh it gives a solution, and it also supports runtime management.
Okay, I see that my intent is not well satisfied, so I want to do customizations. So, control loop two does all this thing in runtime. And control loop three actually provides long-term feedback. So, uh intent is processed, I get intent reports, I can see what went well, what went wrong, and I can do customizations. And in all control loops, we share information in the knowledge graph already
mentioned. So, we can take advantage of reasoning, recommendations coming from the knowledge graph for this purpose. Now, for this, we have developed a an a simulation kit, okay, an open source kit that we try to support different type of deployments uh with dynamic creation of uh a multi-cluster infrastructure, dynamic creation of application graphs and application workloads, so as to be able to experiment uh let's say, a
lot with different type of graphs and infrastructure. It would be really hard to do this in on real infrastructure because we need lot of resources, and also playing with different type of graphs and microservices was really hard. So, we decided to develop a simulation kit and maintain it for this And based on the deployments, we assess three three main metrics related to intent violations, um acceptance ratio
of the request, how many of the requests are properly served, and uh metrics related to costs in terms of CPU usage and also energy consumption. moving one step further, for this simulation kit, we are working on making an agentic version for this. here, we said, "Okay, we have a complex setup, we have different control loops. Let's try to train agents to support us in the different parts
of the process." again, we have an intent definition. We have generative AI tool supporting the first control loop with uh intent translation, and we have different agents supporting deployment and scheduling, adaptation on runtime, and continuous monitoring and reporting. And for all of them, we have, as said, the knowledge graph storing knowledge, learning, uh making suggestions. So, reasoning is supported, and all these agentic workflows are are exploiting
the reasoning part uh by the knowledge graph. And for sure, for resources, we speak up above the overall computing continuum, okay, considering also radio resource units and 6G technologies and moving to the edge and to the core part of the of the And um uh this is, let's say, a high-level figure of the agents. Okay, we have a supervisor agent managing the requests and having different tools
for the initial intent management. The intent agent supports conflict resolution, semantic and syntactic validations, and proper uh creation of an intent manifest. We have the deployment agent that actually is the one managing the deployments uh with the runtime adaptations. Uh we have provider agents that are actually uh providing resources and managing the part of the providers, and the knowledge graph. And in each of them, we have
different control loops activated, as said in in our previous work. And in the agents, we combine uh large language models driven decisions, but also rule-based logic. And in some cases, we we work also with reinforcement learning. So, agents are of different types and combine uh different type of decisions to satisfy the the intent. Okay and coming to just to show some some of the results. Prior to
the agentic approach with a simulation kit we work with assessing intent violations during the time slots by activating the different control loops. So the green the green light is what what happens when we have no intent control loop activated. So let's say baseline and then by activating different loops we can see what is the impact in terms of reduction of the intent violations. And we have done
this for different objectives. In one case my intent was give me this service with the highest performance possible and in the second case give me considering the lowest energy Okay and the lowest line in both cases is the one with the all control loops activated. So we see a major improvement in the reduction of Then having this we also checked how the deployed components stay change over
time. So in one case we asked to have high performance in the other energy efficiency. But the providers are classified in three categories. I have a provider that internally wants to support intents targeted to high performance. Another provider wants to to have a moderate cost and another provider has to wants to be energy efficient. So in the first case we saw that most of our components went
to the performance oriented providers and when this is let's say full full with request then we have request going to the others and the opposite in the other case. And now having this having assessed the what is important is also moving to the agentic approach. So here we have three lines. The blue is the non-agentic the one with the simulation kit. The orange is the agentic approach
without the use of the knowledge graph with an empty knowledge graph so it's agentic empty. And the green is the agentic approach with populated knowledge graph. So we have knowledge that we take advantage in our decisions. And what we saw actually it's clear that non-agentic approach we had this type of violations. This percentage was reduced a lot while moving to the agentic approach and much better while
moving to the agentic with knowledge graphs. So we saw very high percentage of improvement and we assessed also different type of metrics on the behavior of the agents in terms of reliability automation capacity for adaptive management and energy efficiency but also continuous learning because we we assessed it by just populating the knowledge graph with three four deployments. It's not that we populated the knowledge graph with a
mass history of data. So results can be further improved if we consider an an operation running in larger period in time. And for sure okay there is a lot of ongoing work. We want to to add more metrics rather than CPU usage that we mostly consider [snorts] at the moment. We want the agent to support to support more orchestration actions such to support also scaling and other
type of We work on a game theory approach to support negotiations and auctions between the intent agent and multi-providers. So we have offerings and we have negotiations in trying to let's say simulate a real ecosystem of multi-providers that offer resources and the agents take place how they manage them and also development extension of the knowledge graph to to what is called a context graph that is also
support tracing in the decision making. Why the agents took this decision and if it was successful. So I can learn also from the decisions of the of the And that's all on on my side. Okay this work is done as part of the 6G 6G SNS project but is applicable both to 6G resources but also in general computing continuum. We'll be glad to discuss for any question
or or feedback from from your Thank you. >> Thank you. >> [applause] >> I don't know if we have question. No. So we can move on. >> [laughter and gasps] >> Don't give work to him. So slide 10 please. Okay. Slide 10 can you show slide 10? I'm sorry slide 10. Yes yeah. So what is very interesting is this this thing here. Okay so using agents using
agents with knowledge graph. Okay. So here you only have prediction here have a prediction plus deterministic determinism because the knowledge graph are providing the deterministic thing. So this is important. So actually we need to promote this as an architecture pattern. Sorry for that. But I need to work with you on this because this is really and you have to know that we might be about to launch
in SC42 a project on multi-agent framework. I I I thought I asked for to call it multi-agent reference architecture but we'll see. It's launched by Korea. It was this morning at 2:00 in the morning. So and I asked to be created and this is an example. So so Antonio many thanks for for the feedback. We can definitely discuss as an example of multi-agent approach. Many many thanks.
Any other question? If not Thank you very much. >> [music]