TestCon Europe 2025

Peter Sabev: Measuring Performance with Functional Tests: A Low-Effort Approach for High Impact

42:15 · 21 Oct 2025 – 24 Oct 2025 · YouTube

About this talk

In this talk, the speaker discusses various performance issues observed across different systems and technologies, emphasizing how minor delays can lead to significant problems. He shares impactful examples from various sectors, including healthcare, finance, and transportation, illustrating how slow software can jeopardize operations, user experiences, and even lives. The speaker transitions to performance testing, highlighting the common misconceptions, challenges, and underutilization of such testing in development teams. He proposes the idea that functional tests can be adapted to serve as performance tests, sharing key strategies for effective performance measurement, including establishing baselines, repeated testing, and effective data visualization. He concludes with a call to action for software engineers, particularly QA professionals, to take ownership and responsibility for software performance in their organizations.

Full transcript

All right. So, first thing first, few words about me. Actually, it has been much more than 20 years because uh I'm coming from Bulgaria, Eastern Europe, the land of yogurt, roses, and reverse engineered Apple 2 computers. I don't know how many of you know but we have uh had a factory in a small city of praitz where they made Apple 2 computers. So when I was 7

years old I actually saw that computer for first time and since then this is my profession and everything I do. Uh currently I'm software development manager at IBM. It took uh quite a long career path going through small startups of that were two people to big uh company names. I also have PhD in informatics and uh also I'm speaker traveler and uh also co-founder of two of

the biggest conferences on the Balkans QA challenge accepted and dev challenge accepted. Uh, and I have to admit something to you. Actually, this year I was promoted to a manager and they took the most precious thing from me. They took me the QA team. So, I'm no longer QA. Uh, but uh, as you know uh, as there is no ex policeman, there is also no ex QA.

On the personal side, uh, actually I'm married. I'm proud father of two brilliant kids and one spoiled cat. I'm also traveler, photographer and junior motorcyclist. This is my fourth time in V news and uh who knows maybe next time I come with my motorbike. So, okay, one important disclaimer so I can say whatever I want from now on. So the views and the opinions I'm going to

share are not related to my employer or any particular organization and may not much cool time for stories right and I actually will start a story with my wife. Two weeks ago she slipped. She fell badly. Her leg hurt. So, we went went to the emergency and made an X-ray image of her leg. Then we went to the doctor waiting and the doctor came out and said

the following thing. He said, "You are lucky. There is nothing broken, nothing twisted. Still, you need to wait for 30 minutes for the results." And I was, "Why? Wait, you know that there is and I still need to wait for 30 more minutes. How do you know that there is nothing broken? Because uh doctor said actually the X-ray is ready, the analysis is ready, but we have

a very old CD burner that is just one speed and we have to wait for half an hour to burn you the CD with the results. So you have and I was really surprised how something that small can cause that big problem, that big delay. Everything was ready within 3 minutes, the X-ray, the scan, and we had to wait for 27 more for the CD to burn.

And actually in our wives it is often that something that is really tiny delays entire system really a what uh how many of you have been to the previous keynote the Chris one he had some stories okay most of you so thank you for being here I also have some stories for you and they're uh related to performance I'll start with uh the year 1992 Wondon. They

are having a really nice uh ambulance uh coordination center with a really cutting catch software for the date. The thing is the software get overloaded, get swole and actually the chaos that it caused lasted for three days and thousands of patients were reportedly dead because they didn't receive their medical help on time. We continue to 2013 NASDAQ on August 22nd. It was freezing roughly for 3 hours

because of overworld and complete slowdown of the of their security information processor system. The result is millions if not billions of losses. Another sad story for from 2018 Yuber self-driving test car that uh was going uh to recognize a pedestrian. So the system was really slow. The pedestrian was uh actually pushing their bike and the system was saying uh this is a pedestrian. No, it is a

cyclist. No, it And reacted pushing the brakes only 1.3 seconds before the crash. that was another case when software is slow and this costs wives. fail swallow in data centers. There is a really interesting study. Basically, all the big companies in all the big uh data centers are having uh servers, routters, network cards, hard disk drives and all of these gradually slow down. Some of these devices

slow down to the uh moment where they are completely frozen. However, uh as they are still working, the system is not detecting them. And just imagine your servers starting to work uh 100 times slower because of a network car year 2020 COVID. It takes one minute because of SW SW software to access the medical profile. You can imagine what it was for the doctor. thousands of patients

needing immediate care and you are every second is very crucial and you wait for the medical profile to load. 2022 DJI drones they released a firmware update that made their drones react only 200 milliseconds slower. This was obviously enough for the drones to crush in different poles, fences, trees, and some of these were reported. Of course, DJI were forced to issue a corrective patch. And uh something

that is another interesting thing, one of the big UK banks had a delay, technical issue according to them. But according to newly wet coupleo who had reserved the restaurant actually this ruined their wedding. The venue did not receive the money. They cancelled the reservation. So ended they ended up uh waiting at the park at the parking quad instead of the restaurant because of that. And uh yeah,

let's let's go 2025. Actually in June there was a major slowdown but uh to be honest I experienced that yesterday as well and actually slow down in chat GPT didn't cause only chat GP to freeze. These are services like uh chat GPT, open AI API, Microsoft Copilot, BIN, Snapchat, Dualingo, Notion, Khan Academy, Slack, GPT and many more. So basically in complex system it acts like a domino

effect. Failure may not always start with a crash. It may start with a delay that quietly sets a chain of events no one sees until it is too late. Cool. Now I want to raise your hand if first you have done any performance testing in the past two months. All right, I see about 15%, hold your hands, please hold your hands. Keep them up if your team

actually has some time or budget that is dedicated for performance testing. Uh others. Okay, I see three, four, five, six, seven hands. And keep it up if your tests automatically measure performance in your CI/CD pipeline. All right, three, four, five hands. Okay, five people in the entire room. Why is that? Why we never do it? Okay, it is expensive. When people hear about performance testing, they hear

about uh a lot of time and investment uh powerful machines what that need to simulate lot of users, lots of requests and uh this is complex and often postponed. The other pink is many QA engineers simply love the idea but cannot do the setup. They don't understand the things. there is another reason. This is test rail. Look at that. We have passed failed blocked. But do you

see any execution times? No. And the same is actually happening to practice test to X-ray to mercury to Q test. All of these are just having passed failed and no times. this is also a problem. Well, what happens when I go to my management and start talking about performance testing? Usually, this opens actually a lot of doors and those doors are the way out. Yeah. Why? because

we need to be cost effective and usually there is no budget, no money, no time for that releases are coming. So what I usually hear is just see what you can do and trust me I have been working for so many companies in so many countries and everywhere it is absolutely the same. Just see what you can do. That's it. All right, let's see what we can

do. What if I tell you that every functional test could also be a performance test? Most of us are [snorts] having automation that is up and running. It produces results and uh all we need to do is to measure those tests and see if we can do something about that. the idea is really really simple. We have some test execution. We put start and end timestamps. We

have the duration and basically that's all. We know how much it takes and we have a performance test. So that's it. Thank All right, no worries. Just kidding. [laughter] But uh as a conference organizer myself, I can tell you that it is always the case. Uh when people are hearing to some boring guy speaking some stuff and then they hear applause, that's it. Oh, what's happening? So

Europe was for the technical crew of Tescon Europe. It is a hard job to do. And happy birthday test gone as well. 10 years. Wow. [snorts] So, let's continue. Is this idea a real performance testing? What do you think? Yes. Oh, well, my answer is big. No. It is not the real performance testing and here I have to say it is absolutely not a replacement for the

normal performance testing. Basically if you imagine your performance testing like an athlete with nice body what I'm proposing here is uh more like AliExpress performance testing. So it might not be the best but considering the circumstances, considering your time and budget, it might still do the job. actually most of the tools we have, they have already some kind of field that catches time and execution duration. So

for some of the tools it's code time for some is duration for some is execution time for some is execution underscore time but basically the idea is absolutely the same for all the tools. So uh this talk is not related to any particular tool. You can do it with anyone. What I'm going to give you until the rest of my talk is actually 10 tips on how

to make this idea actually working. first thing you need to do is pick your right tests. Those could be all of the tests, especially if you don't have such a big uh suit. It could be a specific very important for you scenario or it could be some kind of combination smoke test. And uh the tricky part is you can measure everything. You you may be actually already

measuring it considering uh the previous slide. But you have to determine what you are going to track and inspect every day because small measurements every day will beat one big crisis test before the release. So next thing is you need to add really lots of measurements. The same way you are putting asserts in your unit tests, you have to put measurements in your performance test. Let me

give you an example. So you have uh just let's say login test which is you are entering the login details, pressing the login button, uh the dashboard loads and then you get some stats. If you have stopwatches set only in the beginning and the end of this entire test, you can compare that it has been just 12.2% slower compared to yesterday. Good, but not enough. If you

put much more stopwatches, you have much better granularity. And as you can see in this case we can clearly see that the much slower part would be the main dashboard loading. So this is where we have to concentrate our investigation. Another thing that we should consider is we have to have only valid performance results. Only the passing functional tests should be used for performance data and actually

failing tests they will tweak they will distort your data and that could be a problem in future. Basically why I put the flat tire the same the same way uh if you have your car with a flat tire and you are measuring it maximum speed probably it will not reach it. So make sure everything is working okay and then you can consider those as valid results at

least in most of the cases. I will give you more examples for that. So next important thing is to set a baseline. Baseline is basically your reference speed or the first table known good result. What does this mean? This means how [snorts] fast your software works when everything is all right and uh uh you are getting the baseline. The same way if you don't know what is

the water level without waves, it is hard to measure the wave height. So you need to know how fast your software works. For instance, if you are having the login API, you start five runs. For instance, we get uh those at 2.2, 2.3, 2.2, 2.2, 2.1. Okay, the average is 2.2 seconds. So, this is the normal performance. From there on, every time you run your test, you

compare it to that baseline. Simple as that. or almost. You need to keep that baseline alive because let's face it, you will get different security updates uh some optimizations, some third priority libraries that get updated and uh this will shift your timing permanently eventually. So frequently you should do the so-called baseline recalibration and uh set it accordingly to what the new performance is doing. Something important everything

I'm going to say from now on every tip I'm going to give is valid for both setting the baseline and the individual performance This is important. All right. You might also want uh to get rid of the noise. This is optional step but what I mean imagine you have to send a request of uh e-commerce site where you have uh you don't you want to know the

top 100 people who have most orders and uh you receive uh a response. Let's say it's a complicated system and it it takes 17.4 4 seconds. But if you don't want to measure the time for connection DB queries, you are just interested in the processing time, you can do the so-cal bracketing, which means you try to send a new test uh empty test, something like that. In

this this case, instead of 100, we say give me top zero people with most orders. and uh you will eventually receive a response that is on empty list. If this takes 2.3 uh before the execution and 1.9 we get the average of those two we get also uh the 100 people query and we get the real for most of the cases that works pretty So another thing

you need to have for your baseline is performance score. Of course you can have many like this but uh usually when you are reporting your team and your managers they don't care too much about particular feature that has slowed down. Although some of you as engineers will have to investigate that. But what your manager usually cares about this is the total performance. So you should think uh

if you have let's say as here test A and test B you should have this performance score that can give you what you need. A tricky part here is that if you have many tests and [snorts] you have really significant change, you might uh have a problem because it will be insignificant change in your performance score. And uh the way to do that is the most important

tests you put higher weights. But yeah, the idea I have put this oranges and tapos on purpose. The idea is to compare oranges to oranges and apples to apples. So try to find some kind of index of performance score so you know uh where you are. Now let's see this example. We have 100% then we have one slower test and it became 136%. Is this good? Is

this acceptable? Is this really bad and unacceptable? Here is why you need to put the so-called thresholds. Uh talking about the water level again, we can see on that bridge that uh we have uh level of the water that is in the green zone in the yellow zone and then we have a dangerous red zone. So we need to determine when is SW really when we will

say okay we have a problem and here [snorts] is the tricky part for you as engineers because it really depends on your project. If we have 20% [snorts] slower than the baseline and it is 2.2 seconds we just calculate our thresholds. Okay it is 2.64. So if it's 2.5 [snorts] it's fine, 2.6 is borderline, 2.7 is warning. But where to put that threshold? I can give you

nothing but advice from my own practice. For the critical features, put a [snorts] small threshold. For the non-critical, you can be one way one idea more relaxed. So you can start with lower values but if you are getting too many false positives in the time you can adjust that threshold. But one more time it really depends on your tests on your environment on your product everything I

mentioned environment. So the next thing is always to run your tests in the same environment because let's face it, your local developer machine will probably uh not be equal to the production environment as number of records as uh fast uh as speed of processing. Uh and what we need to do is to measure one consistent place and [snorts] trend the graph that will become really modern arc

if uh the if we are measuring on several different environments. Time of the day also matters. Trust me having software in the day and having software in the night works completely different. Especially if it is let's say bank software during the day they have customers they are processing something during the night they have different batches that are archiving and processing stuff. So you this software will probably

work completely completely different in terms of uh latency and uh speed of processing. So consider to execute on the same environment, same time of the day and also consider the time zones as well. I can give you an example. In one of my previous companies, uh we were working in Bulgaria and there had to be a backup that is uh chron job scheduled at specific time every

day and our DevOps team they considered midnight GMT time. So when in London it is midnight actually uh the system started the backup everything slowed down no one cared until we understood that we have very very important customer in Japan because in Tokyo at the same time it is 8:00 and you know uh how punchual the people in Japan are 7:59 they are getting off the train

8:00 they're in the office 8001 they're walking in their system and they're swole deadly sw because of the backup. So that is something that you should uh really consider. Another frequent problem is the phantom spark. So you have no code changes usually but uh suddenly the system works slower or gives you slower results. Usually in almost all of the cases it is some external dependency maintenance window

some jobs running or downtime of your network or uh very high CPU or this quote. So it could be any of these. Advice number seven, repeat, repeat, repeat. You cannot trust a single run. You should always measure those tests more than uh if you want uh to be reliable, if you want good results [snorts] that can really give you some information. That is why you will need

to run several iteration. Imagine the following. You have a very very quick small test that takes usually uh 0.001 second. So 1,000th of the second. Then if this test for any reason and it will happen almost always takes 0.0012 this is actually a 20% difference. So usually the threshold will give you uh warning and [snorts] you will have a problem. So micro tests like this you will

have to repeat them many times in order to get some reliable And here is another thing that you need to determine as SQA engineers the balance between more repetitions and less repetitions because more repetitions will give you more reliability of your test but will also mean slower execution. So depending on how much time you have to do those executions, you can decide how many time you would

like to repeat those tests. But what is more important than that is to compare only what is consistent. What do I mean? Look at that guy early in the morning. That lead. He's sleepy. He hasn't drank his coffee. He hasn't warmed up. Do you think he's going to make his first Wap his best W as well? Probably not. So same is with software more or less. Sometimes

uh software uh most of many uh hard drives might be hibernating or sleeping. Some connections may not be established. So usually you need to few to run a few very basic call them no tests if you want just to uh make the system up and running again. warm it and have all the caches, compilers and connections uh working. Exception of that would be if you are testing

exactly the caching functionality but that's another story. What is the bottom line here is you need to avoid cold starts in Another thing that happens often is the following things. you get suddenly your test running like 20 times faster. So let's say the wagon test it runs 2.2 seconds 2.2 2.2 and suddenly it takes you you haven't made any changes but it is much much faster. Well

usually 2,000% faster means 100% wrong. Usually this is uh some database that uh is down and disconnecting or you have really caching or you have some system that is uh returning uh strange uh or empty responses. So be very careful and check Another important thing for people who love mathematics use median not mean average. What do I mean? Look at those tests I have taken from my

previous example. So you have 2.1 2.3 2.2 2.2 2.2. In that case everything is relatively consistent and we see that both are the same. Both are fine. The mean and the median. What happens however if we have one spike. In that case, if we are still [snorts] measuring the mean average, our average time would be 3.18, which is not good for a baseline, right? To form your

baseline, it is much better if you get the middle value, the median. So consider that when you are forming your baselines mostly. It makes a big difference. Next thing is okay you have that data but you still need to make it visual and uh here I'm not going to limit you with uh use that and that or this and this too you can do whatever suits your

team you can have uh a simple excel sheet and uh you can go through all of your CI/CD tools uh you can use Jenkins GitLab up. You can use test reporting tools like X-ray test trail or lure or whatever suits you. You can even use observability tools uh graphana, y promote. So the most important thing here is when something goes wrong you to have easy way to

see what has happened in the past, what is happening now, which tests uh have taken more time and this is really really important. And uh another thing was a mistake that I made in the beginning. We said okay if the test is slow we should not allow the build at all. This didn't work well. There were many false positives. So what I would suggest to you is

use warnings instead. And uh actually the result of uh our effort was notification in Slack like the one you are seeing on your right. So basically every day you're getting uh something that is uh for each of the tests better performance, same performers or so performance. Of course you can exclude tests that are not working or having temporary problems. But in general, if you receive that on

Slack, let's say every morning, you just have a very quick look and you see, okay, something is wrong with payment gateway and order history. And basically that is my summary. So if you have to remember one slide from my entire talk, that's the slide. Start measuring duration for your existing tests as granular as possible. Pick the right tests. Form a single performance score from your baseline. That

should be mediumbased. Define a threshold for each test or the entire suit. Repeat those execution daily on the same environment, same time of the day and consider only reliable passing tests. [snorts] And don't be ignorant. If you hear if you see something is wrong, just investigate I will finish with uh the photo of that car. Works nice, right? So if we don't measure the performance, it is

basically saying, okay, I'm fine with that with that cart as soon as it takes me from point A to point B, no matter that its maximum speed is 20 kilometers per hour. Basically, we are doing this and this is simply not enough in today's world. uh for every 100 milliseconds of added page world latency, Amazon reports losses of 1% in sales. You know how much is that

and how much Amazon sells. Walmart one second one second improvement in their page time means 2% increase in conversions according to their stats. 1 millisecond advantage in NASDAQ actually can generate millions of dollars per year. And let's talk about human wives. 2C delay in the RAN refresh is already critical for safety. 0.1 second software reaction time in self-driving cars doubles the risk of a crash. So I

started with such examples and uh I want to finish with something very important something that uh Chris put the fundamental of and uh I would back him up as well with my presentation. Such problems happen and speed quality. It is responsibility of our entire engineering team. But QAS QAS are their last resort. They are the heroes, the quiet heroes who can prevent the next performance disaster. So,

it's up to you if you want to be that hero or you want to be the reason for the next uh headline in the press that everyone will read about. Thank you very much. [applause] >> And uh yeah, if you can just leave the PR code again uh the QR code uh just to say something more from my presentation. Yeah. uh you can download the slides from

this link and there uh I also have PhD and there is also scientific paper on that topic if you are more interested in the paper. Thanks one more time. >> Thank you Petra for uh amazing talk. We have some questions in Slido. I want to give you a moment to upvote the ones uh that most interest you. And um uh before I do that, we have an

announcement. Someone has lost a Samsung phone in an Otterbox uh phone case. So if that's you, please uh go ahead and find the volunteers uh and hopefully you'll you'll get it back if you're able to identify it. All right, back to our Q&A. Our most upvoted question is, how would you account for the test framework's impact on the performance? because that also takes takes a toll I

guess. >> Uh that is a very good question actually. Test frameworks could affect the performance. That is why I said it is really good to have one empty or new test first to uh ignore the warming up initialization of the framework and also you can use the bracketing technique I mentioned. So you use that zero test uh to decrease the time of the remaining tests. If you

do it enough times, it would be good. I think it will work. Next question. >> Um all right. Uh the next question is how many iterations are enough for a performance test? Should a performance test also be a stress test? I think you partially answered that question, but maybe you could elaborate. >> Yeah. Uh my practice shows three to five is usually enough. It really depends on

on the product. Uh however uh should a performance test also be a stress test? Well, as long as it does not break the system, that could happen. But if it breaks the system uh then we are uh reaching the flat tire case where we are measuring something that is not working. So that's my answer. >> Yeah. If you make the the environment fail totally then your performance

tests are pretty much over, right? >> Yep. [laughter] >> All right. Uh so let's take it a step further. So once we discover those performance issues, what are your suggestions on solving them? Uh someone is asking uh we experience uh things like this uh we experience like this is the most difficult thing when it comes to executing them. >> Yeah. Well uh finding the reason is maybe

a good topic for entirely new talk uh about that. But >> part two coming next year. >> Yeah. But I will tell you one word uh on the top of the regular investigation do profiling. Google for profiling tools and see what can profiling do. >> All right. We got some feedbacks on your talks. Uh some uh are very positive that it is very clear and good presentation.

Thank you. Someone tried a code it didn't work for them. So maybe you can find uh find some time to uh debug that with the with the person. come come talk to Peter. Um all right so someone is sharing that they have a struggle uh that they don't uh certainly know if the test CI/CD runner performance or the software performance any tips to u investigate any tips

to identify which >> this pretty much matches the testing framework one whether it is the testing framework or the CI/CD2 I would say the problems and the solutions to that are pretty much the same. So if you have any particular uh framework or tool in mind do not hesitate [snorts] to reach me and we can get deeper into

From event

TestCon Europe 2025

21 Oct 2025 – 24 Oct 2025

All event videos
Back to Watch