About this talk
In this presentation, Aldin Osmanagich discusses the evolution and capabilities of Zabbix, an open-source monitoring tool that has been developed over the past 20 years. He highlights its features for anomaly detection and trend prediction, emphasizing how they enable users to anticipate issues before they cause downtime. Osmanagich explains the advantages of using Zabbix, such as its universal application for monitoring a variety of systems and its commitment to being fully open-source without hidden costs. The session also covers statistical methods and machine learning approaches for detecting anomalies, providing insights into when to use each method. The talk includes practical examples of how Zabbix can help in operational efficiency and reducing costs through proactive monitoring.
Full transcript
Today uh you'll be hearing about a lecture a lecture about anomaly detection and prediction in zabics. Uh our host will be Aldin Osmanagich who will pro who will try to convince you that open source is more than just good enough. Please welcome to the stage. [applause] Hello. Can you hear me? Okay. Uh so today I want to talk about a cool open source project that started 20
years ago and now has matured enough to rival any monitoring tools out there. In part this will be success story about Zabix and how their decision to be 100% free opensource without hidden costs paid off for them and but the main focus of this presentation is uh to show you how you can predict uh problems before they happen and for that we will use trend prediction and
anomaly detection. uh functionality. Okay, this is our agenda for today uh that I just introduced you. I just forgot to mention the near the end of the presentation I will talk when to use machine learning and when to use statistical approach but few words first thing first few words about me I'm Alden uh Osmanagich I'm I'm I came from Croatia I live near near Zagri so this
is my first time in Sophia and I I have been playing with for since college almost 20 years and most of of my time since 2011 I have been playing with Zavis also I work as a system engineer at Telink business services and in Telink we decided even though we have a long history with open source we decided to aggressively invest in open source so let's start
So uh first question, how uh how many of you have ever used Zabix? Can you just raise hands to I have two two potential Zabix files. Zabix is difficult to to describe because he's constantly evolving because the Zabix is trying to cover all domains of monitoring but I will not be wrong if I said that Zabix is universal all-in-one open-source platform for fault and performance management and
this definition will change if you go on Zabix website they are always changing because they are adding more more things to Zabics. but three characteristics of Zabics have stickked with them since the beginning and that one is they are universal, they are open source and they are all in one tool. Universal means that Zabix can monitor anything. When I when I go to customer site in a
new environment, I have a courage because I know whatever is out there, I will be able to monitor it. If it is a legacy device system or it is a cutting edge technology, I I can manage it and somehow monitor with Zavix and they worked in the last maybe seven years uh on integration. So you can just go on official Zavix website and and you can see
here just part of integration and you can download what you need. If if you cannot find integration for your application system middleware you can go always go to uh community uh repository and find your integration and Zavix is all is trying to be allinone. 20 years ago, maybe some of you will remember, you didn't have dedicated tool for fault and performance. You didn't have one tool for
for performance. You had dedicated for fault and then for performance. And Zabix was the of the one of the first one that was able to merge those into one. I mean there was lots of other that but they did that efficiently because it makes sense if you get alarm CPU utilization is 80% it will be nice to know what was the average for last seven weeks if
it was 5% then that's a big jump it was if if it was 75 then the something happened but not drastically something bad as as you can see here uh you have uh some features that I just mentioned synthetics monitoring you have g maps dashboards graphana look alike and uh automatic action trend prediction baseline anomaly detection that is something that we talk about today etc the third
characteristic Why Zabic succeeded? Because from the start they decided to be open source 100% free without corporate tradition without limits. So now you may ask yes >> I was just going to ask is it kind of does it is it put together? Yes, it's standalone the only I mean it's uh built on the open source like posgre MySQL engine apache but their uh their program agents are
written for example they're using PHP for uh web and for uh for P collector they're using it's written in C so it's very efficient they have agent uh they have agent C and in Golang So they have their own thing since the beginning. And like I said, uh why I'm saying 100% free open source. When you say open source, most people think of course it's free. It's
GPL license. I can use it. But let me show you something. This is uh I took the other day from Google Trends. They're showing you basically popularity of some monitoring tools and you can see that Zabix monitoring tools is only rising. Everything else is going down and this is not cherrypicked data. I started working in telco I don't know maybe 17 15 16 years ago and I
was working network operating center and back then nagus was standard for monitoring from 2000 to till 2010 everyone talk about nagus and everyone used nagus but something happened nagus decided that they will uh commercialize some cool features if you want to if you uh want to I don't know connect to active develop directory or you have want some custom reports you need to pay so they have
the core nagus as open source but also they have uh stuff that are hidden under a price and the community didn't like that and because of that the community splitted they started in Singa and in Singa said okay we will stay 100% free and I also put here Zenos because back then I was testing Zenos Nagus Kaki um and later on I even tested it uh to
see what what what is good and Zabis at that time decided we will be 100% free forever We will not have hidden costs, special modules etc. Everything is free, we will give community to use it as is and if you need education, if you need services, pay me and that's what's open source is about you must give uh to to the world great software. I mean give
the software and you don't owe anyone to maintain they need to pay you for that and that's okay and Zabix build business only on that on education they providing certification program and with services and ina for example continued open source but they just didn't manage to be to go on to the top to compete with the best monitoring solutions. And the best part is you have open
source free and you have freedom to choose with who you will work because there are 300 partners all around the world. You don't need to work with Zabix. you need to you can choose someone from I don't know Croatia where uh where we have support for Zabics all here in Bulgaria but you can choose with who you want to enough of the history let's let's go in
more technical stuff uh I want to show you today I want you to leave this room and get an idea what you can do with Zabics. So it's a bit more technical but not too much. So what are trend prediction? Trend prediction enables in Zabix to predict when some problem will happen. For example, when the disk utilization of of some server will be 100%. And why that
matters? Because if you detect issues before they cause downtime, it's everyone is happy. No one seees that downtime. Uh you you save time. You also get early warning system. If you have alarm that says your disc will full 30 days, you don't need to rush and change and and change the disc. You have time. You know that you have time. Also uh one guy that I know
is planning planning his network with trend prediction when he received alarm that network link between city will be full for two weeks or four weeks I don't know what he set as a threshold he contacts local telco and ask for link upgrade and when you sum of sum all of those benefits you have reduced operation cost and something to mention that most of people forget you can
do automation between prediction. So if you have VMware hypervisors and and you can set pred uh to predict when the CPU or memory usage will be full and then you automatically provision more without doing anything. Zabix will do all the stuff. He will send where to provision this and that and you will get a report that says what was what was happen and you can always go
uh check that later on and remove if if it is not correct but most of the time it is correct because those are most of the time linear uh usage. Okay, let's let's see some alarms from monitoring tools. These are the standard alarms or on most infrastructure monitoring tools. You have disk usage is 96%. Another one service SD CRM some application has high memory usage and you
have chassis one in bad state which is on main load balancer. Now when you look at this you know you have problem and the next step is you need to troubleshoot that that and that takes time. You need you you need to see okay if this usage is 96% when it will be full if this memory used utilization is high will it go down in 5 minutes
10 or through two days and if chassis fun is out of order okay chassis has five fans maybe can without that fan if you implement thread prediction you can get something like this so first alarm tells you this usage is 96 6% you know there are seven uh 100 megabytes left of 20 giga disk and zabix tell you that in 9 hour it will be full so
you now know how much time you have another example where you have in-house application this is our inhouse in Croatia in company application uh is using uh application is using 3 GBs of memory and at this rate it will exhaust memory Mor in 30 days. So developers pushed new version of application with bug that causes memory leak and now you know that you have 30 days time
to tell them that and chess is fine one you know it's it's don't it's not working. tells you okay immediately you know the order what and when you need to do so how that that works in Zabix so let's go through high overview zabix is saving data he has of course row of data that if CPU is reporting every one minute uh value he will save that
but he has consolidated data that is called trend data and you can and you can have one two three years of trend data and you can use that for statistical uh calculations so first if you choose okay use 30 days of data and predict something in the future so he uses that trend data then he's trying to fit the data to formulas most of the time I
use a linear formula but you can use others formula as well as as well. So he's trying to fit that to find the pattern and then when he finds he extrapolates into future and says your problem will appear for two weeks and it's doing that with two function you have threshold to reach function that predicts how much time you have until specified value is reached. Here's a
here's a good example. Predict when free space reaches five gigs uh reaches five gigs. Use one day of history data. So he's using trend date of of one day to find how behaves to define the trend and then he he predicts that value five gigs will be I don't know in two days. Another uh function is reverse where you where you give time and then he gives
you when that value be. So in in the example below predict CPU usage value one hour ahead two hours of trend data. So you put okay use two hour and then predict me what that value be in one hour in the future. Uh here's a link for more advanced example. Uh this is link for to my blog post where I uh decided to show community how can
how can how can they resolve one problem with trend prediction and that problem is imagine you have you configure trend prediction on your hyper on your data stores or or disk and you are looking 30 days of data to predict when disk will be full. So what if someone in the last hour copies terabytes of data and that that's he will fill the disk in 20 minutes
or one hour whatever it will he will rapidly fill the disk but trend prediction is looking three days into the past and then this this spike just uh just fits into the average bunch of low values and cannot see it because the average was smoothed and he don't sees that then he says this will be full in one week and not one hour and to to work
around around that issue you need to put multiple multiples uh time periods. So in this example I configured okay first look 1 hour then look 4 hour then 8 then day and then look week and month and choose the worst case in that way you're still looking 30 days but if someone copies something in the last hour you will also get information about that okay now let's
uh analyzes analyzes history data to predict abnormalities, what is not normal, what deviates from the average. And it's most when there's lots of normal data of course and luckily for us data is mostly on on every metric because if 30 days 5% goes 200 that is anomaly but if it goes every day 200 that becomes the normal and trend prediction uses that trend data that I mentioned
earlier. earlier some understand better sudden spike CPU usage that I just told unusual memory consumption memory leak packet lost or bandit anomalies in uh on the network query response time anomalies in database uh sudden increase of fail login someone is trying to log in to brute force you will receive alarm and temperature or power consumption spikes in several room. Someone leaves the door then suddenly drops or
spikes uh if someone turned the air conditioning uh more than it has to be. And in this graph you can basically see there is uh and human can easily detect anomalies but the problem is you have thousands of these sometimes even in millions of graphs like this in big environments and no human can check all of that. So you can use anomaly detection to tell you when
something like this happens. So how you configure anomaly detection in Zavix? We have a one wizard but we have problem with terminology. First I need to what ST algorithm is and then what uh seasons and deviation. Okay STL is very easy to understand. Look at graphs. If you look the top graph, this is the normal data. It can be it can be disk data. It can be
u CPU data, Docker containers performance. It's data that you see normal in monitoring tools. The second graph is how algorith algorithm works. He first expects trends. So you have normal data then algorithm extracts trends and then algorithm extracts seasonality the repeating pattern those are not anomalies and then whatever is left on the bottom graph are the noise and anomaly if the data is it will shown as
anomaly and that's basically how algorithm works uh to explain deviation are very easy to understand. They just indicates how far value is from the average from the mean and season are easier to understand if so seasons are something that is we have the we have daily pattern 24 hours. So if website traffic increases during the day and drops at night that's one cycle that's and that will
repeat the other day. So you choose 24-hour season weekly pattern when network load is higher on the week days and lower on weekends. That's a pattern one cycle shift the base. If there are factory where they are working every eight hour and at the start of each shift load goes up that's a season and monthly part pattern is when some report goes first of the month every
month or some backup and you have that spike so you can use that season. Okay. Now, now when you know terminology, we can uh item CPU utilization is basically we choose what metrics to use in Zabix metrics are called item. In the next part we say okay we will use STL function to predict anomalies and then we say evolution period look 28 days to the past to
define what is the line period shift is not important. You can shift time to look not from now but from yesterday for example. So for now thection 20 of day 20 eight days of data to predict anomalies in last 24 hours and you def it will be seven days because it it's uh uh it's that is quiet over the weekend. uh deviation default three. You can choose
what you want. Meaning if data points is three deviation from the average then uh classify that as anomaly and we have a M is most robust for to for for avoiding noise. So I use that most of the time. This is the default. If you don't enter anything it will use that and result. Now you need to specify if there are two anomalous in 24 hours fire
alarm and how that looks. Uh CPU usage outside expect uh expected range. There is two anomalies detected in the last two hours. okay so now you know that something is happened. There are anomalies in the on the in the data and you go can go on the graph and the and see for yourself what is happening happening. We we have a topic baseline anomalies. Baseline anomalies gives
you a option to use flexible threshold. Let me show you example so you will understand better. This is the standard alarm. Alarm average CPU utilization is uh 70% for the last eight hours. And in Zabix you write that as in in alert alert expression like this. You call trend average function. You find eight eight hours look trend data for eight hours. Then if it is more than
70% trigger alarm. But what if you have a CPU utilization uh 5% and then suddenly goes to 65%. It's not 70 you don't get alarm but it is a indicator of potential problem and you now you can also uh detect that here we can upgrade this alarm to say every CP utilization for light last hour exceeds the normal four week baseline by two and and you only
need to remove the 70 the fixed number and add baseline v WMA function where you define okay look for the last 28 days define baseline and uh then whatever it is trigger alarm so this is great because you know not all the not all the services are using uh CPU utilization in the same way they're not using memory in the same way and if you If you're
using baseline, you can have lots of indicators that are telling you something and for that he's using baseline WMA function. Uh they are calculating baseline looking to periods inside of seasons. Here's the example. If CPU usage at 10:00 on the last three Monday were uh 30% 35 40 is last value Zabix will give give that last value I mean that algorithm will give that last value more
weight and then he will check in uh three Mondays values and you will get 36.7%. Uh here's uh uh the image that describes that. So you have here uh four periods uh inside four se seasons. So, Zabix is looking and then checking find baseline and when you understand seasons periods it's to configure period one day seasons peak number of season four and you will get the value
instead you writing your own fixed threshold that is not always correct. effect. let's talk about when to use machine approach. Uh first I need to say that Zabix uh does not have out ofbox deep learning neural networks but with open source tools and you have stepbystep almost uh tutorials that can you explain you how you can do it you can implement neural networks for zab so when
is the time to use deep learning and when is okay to stick with statistical approach with statistical method like we saw earlier earlier it's simple to imple implement easy to understand works well with limited data so it's okay to have I 30 days of data, one year, two years, whatever statistical will swallow that and you will get some results. Low of computational requirements. You don't need lots
of lots of hardware and it's reliable for baseline and trend based detection like we uh saw earlier And machine learning can do something of course that statistic can't. He can view multiple metrics. He you can feed to machine learning CPU, memory, disk, network, everything you know about uh server and he will be know the context of uh of the data and because of that he can detect
anomalies on multiple H of also he can learn system specific behavior and based deni and based of that he can create dynamic thresholds like we saw with baseline with deep learning you can do that but it is it is it needs lots of data and lots of tweaking to positives I I noticed something interesting on the last uh Zabix conference. Last year uh uh I mean last
year Zabis conference uh not this year. Uh there was one guy that says okay we have implemented deep learning uh with Zavix. You can scan this uh link to see his lecture. You have you even have on YouTube his lecture we have implemented. It works great for us. They I think they even selling that to uh customers of their own and uh and they you they use
the deep neural network specifically LCTM autoenccoders. What is interesting to me that uh on the same conference few hours later there was another guy I mean he's from Zavic so he may be biased but uh that said uh deep neural network are not enough matured to to to beat statistical approach and he quoted some papers he quoted papers from 2022 saying we found a deep learning learning
approach just are not yet competitive despite their higher processing effort on training. And then from 2023 our experiments show that the classical machine learning methods outperform the deep learning methods and 2024 where basically they say uh they are creating illusion of progress. So this was interesting to me. Uh one guy saying statistical approach is the best. Another one saying that deep learning works well. And I will
help you to make your own conclusion with this. I don't know how many of you has have heard for Will Smith spaghetti test. Sorry. I was just going to say that uh I don't think that it is very informative to categorize all the uh solutions into either statistical or machine learning solutions because very obviously there will be both useful and useless machine learning uh analysis and there
will be useless and useful statistical analysis. So it kind of doesn't make very much sense to me to compare these categories as a whole. >> No, I'm not comparing what is better. I'm comparing when to use one, when to use another one. And this was more an anecdotal uh story that I noticed that there are there are two guys saying one contradictory things. Yeah. And and it's
kind of weird that they >> it's in the answer is in the middle. >> Yeah. Yeah. Yeah. Yeah. >> I understand you. Yes. I agree with you. And and of uh my team also played with we are trying to uh almost the same thing that uh this guy did with deep learning. We are trying to predict on time series data when something will uh and we we
we couldn't get 100 of course no no one gets 100% accuracy I think we get 70% accuracy accuracy so uh the conclusion but that was back in 2022 so I I will leave you with this let me continue and you will know the answer so there is a will spaghetti test if no one heard you can check that on wiki uh pedia where it's informal benchmark of
uh generative videos models uh to see how much progress they have in generating human activities uh spatial expression etc. And five months ago this is the latest result. So looks pretty realistic to me. I'm looking at I I have a feeling it's a AI because of the perfect color grading, but I can easily be fooled and say this is the real thing. This is a Will Smith
eating spaghetti. what is interesting is just how much generative video models evolved in only two years because they started with this two years ago. they they the top uh quality video that is generated by AI was something like this. So this is mini horror movie and only in two years we have almost perfect uh video of Will Smith eating spaghetti. So if AI uh is not evolved
enough to be uh very precise, it will be eventually. In the meantime, you can play with statistical approach or you can give a chance to AI because maybe they already have better results. So that's that's all for me. Thank you. [applause] >> Hello. >> Okay. Okay. So I see we have some time left. So for next five minutes we can have a Q&A section where you can
ask Audin any questions you have about the >> also I will be here this they can always approach me if something is uh >> yes after the lecture you can always approach him so if anyone has a question please raise your hand I'll come with the microphone >> Yes. >> What do you what do you think about the useless uh usefulness of time series databases uh as
a specialized uh format for storing uh metrics data? Uh because Zabix uses kind of ordinary SQL databases. Do you find any benefit in using specialized time series databases? >> Uh yes. for we right now the best performance that we got was using uh posgre with time scale database and but zabix is uh uh planning to implement click house integration and we will give a chance. Click house
is promising but more for like I when I saw the um the benchmark they was they were using uh logs and something like that but I'm not sure how well it work with uh numbers with time series data. So right now the best that you can get performance is using time scale database uh with posgrade uh and we tried almost everything and that eventually worked the best
for for us. Okay, thank you for the question. >> Okay, questions. There's also a microphone in the middle if anyone wants to go up. Okay, in that case, big big round of applause for Audin, please.