About this talk
This talk introduces the One AI Ops Framework, an experimental open-source solution developed by OpenNebula for enhancing observability and automating cloud infrastructure management. The speaker explains how observability is the ability to analyze system operations through relevant data, which is crucial for detecting anomalies and improving decision-making. They explore the integration of AI into observability, highlighting its advantages such as automated anomaly detection and predictive analysis. The framework combines OpenNebula's cloud management capabilities with Prometheus for monitoring and Grafana for visualization, along with machine learning algorithms for intelligent resource allocation. The speaker also discusses the challenges associated with third-party solutions and the benefits of open-source transparency. Future developments for the framework include enhanced anomaly detection and automated operational decisions.
Full transcript
n [Music] hello everyone I'm Victor Palma CL engineer at open nebula thank you very much for being here and join us it's really a pleasure to be at this year dep Bobs Pro Edition sharing this presentation with all of you this session we introduce a new experimentals AI Ops Frameworks developed in open neula and its promis interation for the evaluation of AI algorithms in order to provide
intelligent workload for casting and infrastructure orchestrations capabilities to automate and optimize the provisioning and deployment of edge Cloud notes so let us start with the presentation so first let me introduce myself again I'm Victor Palma of open nebula I come from Madrid and I've been working for open nebula for almost three years developing and innovating in the cloud Edge world well well let's move on from the
introduction and let's start with some context so first what is observability well I guess it's not a strange concept for most of you it's nothing more than the ability to understand and analyze the inner workings of a system by collecting and analyzing relevant information okay so or in other words it's just transforming a data put into information into something that can be useful for us so we
can have a lot of numbers of or data but if we are not able to give the them a meaning they are of no use to us so once we have the information understanding and analyzing it's fundamental this is observability thanks to observability we can detect anomalies patterns or potential risk we can identify performance improvements or observability can we help us in a decision making as the
say saying goes information is power so we need to um we need to add the the observability isue to to our to our infrastructure so well let's let's talk now about AI in the end that is what we want to address in this session is AI useful for observability without a doubt yes CH it's true that nowadays we want to use a in everything well you know
marketing guys are the responsible for that but well observability it's a totally natural process for for thei so thanks to AI and data processing algorithm we get many advantages such as in Hing data analysis automated anomaly detection Dynamic scaling or predictive analysis it opens up a war of possibilities however while this all sounds good there are a number of challenges and concerns that we need to be
aware of first ER there are concerns with third party Solutions so many organizations currently entrust sensitive data to external observability providers raising concerns about data ownership and privacy the information and the dates on the provider servers so this in certain uh situations can be a risk so H I propose a solution the h following an open source um philosophy so providing transparency allowing organizations to scr design
code address customization needs and maintain control over the data thanks to open source we can have the control of the code and the data so that allowed to ask to avoid the vendor loing risks vendor loing is just the look that we have we we when we um depend on on a certain provider so organization might find themselves TI a specific vendor limiting flexibility and potentially increasing
cost it's no easy migrating between a providers for example if you want to move all your workload from AWS to gcloud or to another Cloud Prov providers each provider has has different data models and API so it's a very difficult thing so now what how can we address the challenges we have lines the solution for this is the one AI Ops framework the open source solution for
a driving observability the 1 AI framework combines open neula a cloud infrastructure management platform promethus and grafana from infrastructure information monitoring and visualization tools and a set of machure learning and artificial intelligence algorithms to provide these eight driving observability Frameworks so let's take a step by step approach first what is open nebula open nebula is a simple open source solution to build and manage Enterprise clouds that
combines existing digitalization Technologies with Advanced features for multi-tenancy automatic provision and elasticity in order to offer on demand virtualized service and applications so open neula use the concept of hybrid Cloud that means the combination of on premise Data Centers Public cloud and even notes on the edge in order to operate with the different notes that we have so all of them managed from the same platform that
is open nebula here we can see some of the possibilities we have with open nebula thanks to open nebula you can use the same interface to control old Network and storage resources shared between different hypervisors such as bware KDM lxu or even firecracker opena has several interventions that that facilitate the creation of automatic workload CL and applications such as terapon kubernetes anable Docker or cou apis open
Neola also has its own web portal that we called Sunstone which with which you can interact with the opena core in a simple and combinate way other important feature that to highlight will be multi it so different users with can access access to different resources Self Service where every user can deploy a VM for itself elasticity for multi-m services the possibility to create multi-tier apps High ability
to BMS and open nebula instance the option to create a fate solution the ability to provision resources on the edge automatically multicloud management and the possibility to combine BMS and container workload within the same platform speaking of multicloud as I said before open nebula allows you to manage any infrastructure with automatic provision of resources from cloud providers such equines awbs or gcloud all of them with and
uniform management with the homogeneous layer for user and clo administr trators over this management platform we can deploy any applications such bmss multi PM Services containers or even kubernetes clusters on a shared environment managing all of them from the same portal very handy now that we have seen what top nebula is let's move on to the next important part of one AI Ops Prometheus promethus is also
an open-source platform used for event monitoring and alerting it records metrics in a Time series database built using HTTP pool model with flexible cues and real time alerting open nebula has a specific intervation for prus that allows to intergrate the metrics of viral machines and containers in Prometheus through its own exporters for this opena implements ins size the host the the notes own exported offer offered by
promethus to export basic metrics of resource usage on the physical host open nebula's own exporter to export metrics related to the virtual machines using the L beard library and finally open neula own monitoring system that export the data to theola monitor located the front this turn export all the collected data together with General usage information from the opola instance to promethus this way from promethus we have
a huge amount of information about the state of our club but what can we do with all information so it's time to talk about the third and last part of one AI Ops which is obviously artificial intelligence by adding AI to the formula we can predict a lot of relevant information we can predict through machine learning the CPU memory and network traffic usage we may have in
the future based on all past metrics and then thanks to the impl mentation of decision algorithms we can allocate resources based on the current context of our Cloud also using the prediction calculated through the machine learning algorithms as a guide for the deployment so as a result we obtain the architecture that you can see in this slide one AI Ops is based as we have said before
on artificial intelligence algorithms implemented on the P architecture using promethus for data extraction so it's an architecture that it's currently still on the development so it's still subject to to changes however we can already see here some of the main components that we say before us the current open architecture with the monitoring system the infrasture manager the physical infrastructure manager all the BMS containers um all the
type of resources that we can make of neula like on premise resources public resources or resources on top of of that we have the new one AI Ops architecture layer with a database with all theorical traces onls of from our on cloud a prediction and anomaly detection module that is on charart of analyze all the informations the only information that we have in the database and some
H models related to sorry to the elasticity managers the V infrastructure orchestrator and the physical infrastructure orchestrator so thanks to this parts of the of the schema we can um allocate and optimize resources in our Cloud h of course we also have a reporting and aler modled in order to create alert based on certain rules is is based on the promethus system so it's it's more or
less is going to be the same for for nebula so thanks to when AI Ops and the architecture s above we have CPU usage prediction for individual VM CPU operation per hour the general CPU usage and we also have accuracy that means the accuracy between the last day real usage and the prediction for the last day generated by the tool we can also optimize with one AI
Ops the allocation Consulting suggestion per each BM based on some algorithm such BL balancing to balance the workload in our Cloud reducing the migrations in order to reduce well the migration in the optimization process or the resource contention algorithm in order to join all the BMS in a few host and safe resources very useful for on premise ER infrastructures so here we can see the main dashboard
in grafana with all the metrics supported by promethus uh we have a lot of interesting information such a grafts with the memory CPU and storage usage of the host an overview of the status of our cloud with all the resources deploy it as well as the status of or build our host one AI Ops also has other dashboards like the following one where we can see all
the results of the predictions made we can consult the CPU usage pred for each host which allows us to get an idea about the use of our infra in the hours so when iops also offers an average usage over or entire Cloud as well as an accuracy percentage that allow us to to check the veracity of of the prediction so in this case we can see that
it's a very good value and Below we can see the migrations that the tools suggest for a core optimization algorithm in this case this way we can take better advantage of our Cloud resources as I said before very handy to on premise infrastructures so here finally H we can see a prototype of the one AI Ops implementation we we use the open neula and Li bit sport
to provide General information about the status of the open neula and BMS running on each host then we use the one AI Ops component in order to optimize the placement of rbn based usage on the usage and the data collected and then once premisus has all the metrics collected we can check all the information using the grafana dashboards so since is a prototype no automatic migration action
are performed here this is only limited to suggestions for the cloud administrator however according to the results of our Labs the recommendation are are quite good and help greatly to optimize the cloud so what are the next tips first of all we will implement the Pio operations in order to apply the automatically since as I said before currently 1 AI Ops only shows migration suggestions since it's
still in development stage once this step is done H The Next Step will be include AI Ops as part of the vula software distribution installing it within the same package finally we also hope to SP the functionality to support anomaly detections allocation based on memory prediction allocation based on network traffic and alerts and warnings based on these metrics and detections one AI Ops is um a open
source so you can contribute to the project or check the source code in this repository that you here and speaking of contributing I would also like to mention the opena Forum where you can collaborate discuss and help other opena users to get the best out of their clouds I really recommend it you will also have the opportunity to learn a lot about the platform here apart of
course from the official opula documentation so as a closing slide I would like to commend that this project is being found by the European Union named cognit a cognitive server framework for the cloud Edge coner one AI Ops together with OP nebula will be used as fundamental Pilar within this project for the deployment and optimization of resources in the cloud continue so oh sorry so well that's
all I hope this presentation has been of in up to you I hope to be able to to share more news about Wang aops in the future so stay tuned to the official op nebulas information channels like the Forum Twitter and so on for updates so thank you very much for your attention I hope to see you the next time so I think that now we can
let's move on to to the Q&A section so in case you have any questions we can see now so no questions
More from this event
See all 58 talks →
Halil Ibrahim Kalkan: Building a Kubernetes Integrated Local Development Environment
45:20
Paco Orozco: Growing at the Edge: Doubling Traffic While Changing the API Gateway
45:03
Viktor Vedmich: Ideal Blueprint Versus Reality for CI/CD Pipelines
46:03
Koray Oksay: Continuous Deployment: The GitOps, The Pipelines, and The Ugly
43:03