About this talk
This talk explores the impact of artificial intelligence on quality assurance (QA) within software development. The speaker discusses how their company, ZenFlow, integrates AI tools to enhance QA processes, such as test writing, bug tracking, and root cause analysis. They provide examples of AI-driven bots that assist in generating test scenarios based on code changes and analyzing feedback from team communications. The speaker also addresses the limitations of AI in managing complex systems, emphasizing that while AI can automate routine tasks, the demand for skilled QA professionals continues to grow due to increasing software complexity and quality expectations. Overall, the session highlights the transformative potential of AI in QA, demonstrating how it can improve efficiency while underscoring the necessity for human oversight.
Full transcript
All right. Hello everyone. Welcome to the very first very last talk of the conference. Hope you can hear me. Yeah, okay, great. Yeah, it still feels a bit weird. >> [laughter] >> To talk with you know, with all the background noises. Yeah, as I mentioned this is the last talk of the conference and thank you all for choosing the this talk over the whatever networking mixer. We
literally like on this meme right now. That's us in the corner. Yeah, thank you all for coming before we dive into AI and QA and all that. I want to learn a bit more about you. I want to understand in those survivors of the conference who are QAs, who are not and if you are using AI or not using AI. So I know that those answers might
feel a bit tricky to read so please read carefully and yeah, if you can just scan the QR code and open that URL. And let's see how many QAs we have on the QA talk. Okay. So far QAs are the minority here somehow. Yeah, I still see people voting for the company not allowing to use AI. Which feels kind of weird in 2026 but uh yeah, I
feel feel sorry for you, I guess. All right, but yeah, so far majority of the audience is not QA. Well, I guess hopefully you will learn something new about QA today and how AI is transforming QA. Now the second question is about MCP and agent skills. I'm curious if you are following the latest trends, the latest technological trends when it comes to AI and different tools for
AI and in particular MCP and skills. Okay, looks like uh most of you are familiar with at least one of them. That's good. Great. So, yeah, I don't think I will talk about MCP and skills in this uh in this talk. Sorry for Sorry to all uh 16% of the people who voted for uh the awesome CP and skills. Um but also won't really be diving deep
into MCP, so you know, it's uh it's uh you won't uh you know, feel uh stranded here. Now, a few fun facts about me. Uh sorry for people who attended my talk uh just before. Uh this will be a repeated slide. So, I've been in IT for more than 10 years, almost in uh ML, uh back end, DevOps, all that. Uh been uh one of the founding
members of the company called Zencoder. We are an AI coding assistant. And I've been, well, as I said, I've been the founding member, so I've been here from day zero. Uh uh for the last year and a half, I've been mostly DevRel. So, this is my event number 57, which resulted in quite a lot of flights. That's from last year. So, I pretty much spend most time
on the on the plane rather than at home. Uh I also, as you know, as DevRel supposed to, I started my YouTube channel. Uh not a huge subscriber uh subscribe following, but you know, feel free to uh help me fix that. Uh and I'm split between Portugal and Singapore. So, on the Portugal, I have uh my forecast, and in Singapore, I have my girlfriend, so I'm sort
of oscillating between those two countries uh between half of the world. Actually, I arrived here uh from Singapore just yesterday. All right. And I will be flying back to Singapore tomorrow. So, it's uh yeah. interesting interesting experience. But uh throughout these, you know, uh throughout those uh couple years I've been an encoder. It also allowed me to follow the progress, the rapid progress or rapid evolution of
AI. And um I built for myself a timeline of model releases. And I as you can see I stopped at February uh because it became it became quite tedious to follow everything. And I didn't add the interactive slide of the question for you to vote, but just I will just give you a couple seconds to think about which model, which Sonnet version we had uh 1 year
ago on this same date. All right, so just as a reminder, right now we have uh Sonnet 4.6 and Opus 4.7. So uh just think for a second which Sonnet version we had just a year ago. Was it 3.5, 3.7, 4.0, or something else? And the correct answer is uh 3.7. So th- was released on 22nd May 1 year ago. So exactly 1 year ago we had
Sonnet 3.7. And we've been through roughly five generations. And of course this rapid evolution um affected well, not just the models themselves. They are getting better, they're getting bigger, they're getting more capable. But also it affected uh a lot of a lot of people, a lot of roles, right? So you've probably uh seen a lot of uh headlines like this, right? So in this case Anthropic CEO
uh suggested AI will replace software engineers in general. No wonder what the 6000 funds. Uh but also you might have heard a lot of news um more about QAs, right? So for example this tweet from uh what month ago where apparently uh company CEO fired the whole QA team and then they suffered quite a um unexpected uh loss of $6 million due due to some um hallucination.
So you know, um those things are kind of I guess can become a bit um scary, right? When you especially when you hear that all that every time all the day every day. So in this talk I wanted to show first how we as ZenFlow how QA team as ZenFlow uses AI how we are adopting different AI technologies, different AI skills and transforming how how QA is
working how QA is approaching their work and their role and how they change. So before we do that I need to briefly talk about ZenFlow. Not as a promotion although of course you know feel free to try but just so that we are on the same page and you know what our QA team is testing what is there you know what they need to basically work with.
So ZenFlow is a standalone application. It's a application for managing a fleet of agents. So I have here on the right on the left different projects, different repositories. I have tasks. So each task is a separate agent running or doing something. Like here is adding dark theme to our website. Here is adding some templates. We have quite a few features built in which QA of course needs
to test. I think we have built-in browser. We have some automations. We have some terminal and all that. So essentially there are a lot of surfaces which uh um QA need to verify need to test every time we our engineering team or product products our products actually do code uh using ZenFlow as well. Uh so every time a new feature a new release is coming up of
course QAs need to test a lot of things, right? So a lot of services. Now it helps with that in our application the front end is essentially an electron app, right? So that kind of helps. Uh but of course um AI is not limited to just testing electron apps. You can use AI to test the mobile applications uh finish standalone application and so on. But in our
in our case, it is an Electron app, so all the cases I will show uh are essentially related to that. Now, let me close that. Let me switch to my slides. So, first and the most obvious case, right, is uh writing tests, right? So, uh you know, AI can write tests for you. AI can AI can write code for you. And, of course, it's nothing new, right?
You people have been using AI. Hopefully, you've been using AI for that for some time now, right? Uh but, we can go beyond just beyond just that beyond just writing tests beyond beyond just writing unit tests or end-to-end tests and so on. Uh we can do, for example, uh things like that. So, uh this is our uh one of our bots, which uh for each pull request
uh analyzes the end-to-end test coverage and give us gives us a report gives uh us the report. So, in this case, uh this is a pull request for from me for uh and it is a simple pull request, which was adding some templates. So, not much of the code change. It's just a bunch of JSONs and so on. Uh but, in this case, suggested a few uh
testing scenarios, uh which uh were not were not covered uh just simply on the based on the changes, which I've done uh in the repository. And then, of course, I I further used um the agent to add them. But, here essentially it uh don't just simply analyze the code. It also adds another layer. It uh um intellectually analyzes what is what's changed. It analyzes uh which tests
you already have, and which which things which uh scenarios you might want to cover uh on top of that with with uh new uh with new tests. And, it also ranks them. So, uh you can see here, there is P1, P2, and so on. So, essentially, it's just uh levels of Yeah. Um levels of uh how how important those those tests would be. So, that's the most
obvious scenario, Next, of course, QAs in our team are adjusting feedback, and I'm probably one of the main sources of this feedback. I spam a lot in in our Slack, so that's 42,000 messages from me. So, I do send a lot of them. And of course, a lot well, not a a lot I'm not sure what the ratio is, but a lot of them are bugs or
feedback, and of course, our QAs trash them, need to analyze if there if there is already a ticket for that, if there is if that's a new bug, and so on. So, we have a bot for that, right? So, I can actually go item by item as well. So, in this case, for example, there are some relevant Jira tickets, so maybe this bug already was reported, and
there are a few tickets related to that. Or if if, you know, if that's a new bug, um the bot can actually analyze it, confirm it, and create a ticket for us. And this also kind of gets us into the sort of use case three, which is the root cause root cause analysis, right? So, I already showed this uh back and front end tracked. And uh you
can connect AI to a lot of different tools, right? You can connect it to Sentry, you can connect it to any other data log what wherever your logs and errors are stored. And through that, it can detect errors, it can analyze the code, and help you help your QA team, uh help your developers essentially not just log the uh log the bug, but also try to analyze
the root cause, uh find what what needs to be changed, and if there are any uh you know, potentially modifications. But then, of course, why stop on just root cause analysis when we can actually verify, reproduce, and fix, right? And uh let me actually switch to video. This is a screenshot from the video. So, um we have internally a bunch of skills, which essentially instruct the AI
how to reproduce a bug, how to um report the reproduction, uh how to fix the bug, and so on. Essentially, um as part of that, as part of the reproduction, uh AI is supposed to create artifacts, right? To for people to review, for people to verify that there is indeed a bug, and then that the fix is actually working. So, uh this is a video produced by
uh AI, and let's see if you I know I'm not sure if you can hear the sound, because it's also annotated by this kind of robotic voice. Let's see. yeah, I'm not sure if you can hear it from through the my microphone, but yeah, basically, uh it automatically starts the task. So, in this case, the panel uh the panel sizes, it when you open the to-do or
files or terminal, uh were different of different size. So, it created a sort of jumping effect when you switch between tabs. So, the width of the uh to-do is uh one is this one, but then if you go to files, it's wider, which mean which makes it um me makes it kind of jump. So, essentially, uh AI on its own, based on description of the of the
bug, was able to run the product, run the application, uh start the task, and click click around, verify that the bug is actually uh working. And then suggest a suggest a fix. Uh yeah, I'm not sure why it why it keeps adding those constant animations, to be honest, but uh yeah. Uh so, that's a reproduction. But then, of course, after that, it can also fix, and it
produce artifact of that as well. So, let's uh see. Again, it's just starting the starting the task with the fix, post fix. So, in this case, if you look at the visual panels, to do and files will be the same width. And then files. Yeah, they are essentially the same width. So, you as a human, as a developer, as QA, you can just verify visually based on
those artifacts that the fix is applied. Fix is actually working. mute it. And yeah, as you can see, there's also this kind of kind of presentation style slides as well, which explains what's happening, which explains what the reason for the bug is, and what is the And yeah, sometimes it uh Sometimes AI is a bit stubborn, so when I was trying to create a video for that,
and it's actually our the whole prompt for this. And we have a bunch of skills, as I mentioned, and we have instructions for like the bug description, some reference recording, and so on. And some reason in this case Opus 4.7 was a bit lazy. So, it didn't want to produce a video until I actually asked it to do to do the video, and then didn't create a
narrated video, but uh No. Sometimes AI is being AI. But you know, through through AI, through the use of AI, you you don't just you don't just verify the bug, you can actually even fix it. And of course, uh with that, it kind of let me get back to the slideshow. Uh with that, you know, the kind of kind of begs the question, well, let's just fix
all the bugs, replace all QAs now with AI and maybe all the engineers and everyone with AI, right? Well, um the problem with that is um Well, of course, on the surface, AI is, you know, great with reproducing, with fixing, at least the simple bugs. Uh there are a few sort of hurdles or blockers for that. And first one is, of course, the cost, right? And here
um I have a screenshot of the AI solving a task for me. It's not related to bug fixes in this particular case. It's just adding a dark theme to the website. Uh but this particular um feature implementation uh would cost uh two almost two and a half dollars in the API cost for um on GPT-4. So, of course, if you have, I don't know, hundreds or maybe
thousands of uh potential bugs or issues which you need to fix, it can add up add up really quickly. And especially when you want AI to do a lot of things like reproduce, analyze, fix, triage, and so on, it can uh quickly add up to tens or maybe hundreds of dollars in API cost. And not every company uh can afford that. Uh not every company can just,
you know, run all the all the bugs through AI, uh run all the features through AI, and so on. Um So, that's uh one one limitation. And then, at the same time, um the other limitation is is that AI is, at least for now, is not as good at at prioritizing things. So, if you uh for example, again, have a list of bugs which you want to
potentially fix, uh you probably still want to prioritize them. And uh if you have some sort of metrics on which kind, you know, automatically collected for you for these issues, and maybe you prioritize using that, sure, then you can uh potentially rank them. Uh but in reality, in my experience, it's not as simple as as that. So, you might want to talk with your you know with
your product team, with your users, prioritizing features or bugs becoming much more sophisticated, and AI is still not at that level to uh help you automatically with that. Uh but then the second hurdle or blocker is complexity. Complexity of the systems and requirements. So, in 2024, uh which in AI years like 2 years in in real life in real world is like an eternity, uh Gartner uh
mentioned that uh despite all the AI uh evol- evolution and so on, and despite that despite the fact that AI will automate a lot of a lot, uh still demand will actually increase uh for QAs and for software developers in general because the complexity of software and the quality expectations are rising, right? So, if you for example uh think about Tesla, and they have the Tesla stand.
Uh here in 2020, uh their um so their cars were much simpler than now in 2026, right? And at the same time, all the different requirements, all the different safety requirements, all the different software requirements which are involved in producing those cars and developing those cars are increasing. So, let's say um they don't just um they don't just you know test drive the cars, right? They have
some simulators. Uh I'm not sure what was going on behind me. Uh they they have some simulators which they need to also test. They need to verify the physics of those simulators, and so on. So, essentially, of the new software, of the new um products being developed, being created is creating potentially new sort of either uh branches of QA or even potentially completely new roles, right? So,
I know maybe 10 years ago there was not um a demand for um people who would verify and test how uh simulators for solar cars would work. And now it well, we have a lot of companies which are developing solar cars. And of course the demand for for those roles for those people is well, first of all was created and then it is rising rising, right? So
AI automation is more manual or routine tasks, still the demand will rise because the complexity will will rise. Yeah, so a lot of people say that, you people won't be replaced by AI, they will be replaced by people who are using AI. And to some extent it is it's true, right? Because uh again, if we look at the example of Tesla, I think I've seen some of
these statistics somewhere that uh I think in 2020 Tesla uh QA team had 200 manual QA engineers and in 2025 it became I think 50 manual QA engineers. So the amount of manual QA engineers dropped because they invested heavily in the automation, but at the same time the size of the QA team actually increased. Uh the total size uh of the QA team. Uh which means that
again, as I said, uh people uh roles some roles are getting shrinked, yeah, but some some roles are getting replaced. Uh but at the same time uh we are uh through AI can um automate a lot. We can um create more software, which also would require uh more uh more testing, more QAs. Yeah, so that I won't take you uh much longer so that you can maybe
still have time to uh mingle in that uh party that's that's happening behind me, behind us. So, uh just a few QR codes for you. Uh this one is for uh you to try some flow if you are interested. And this one over here is uh to connect with me. It has all my socials. So, if you have any questions, you're interested in, you know, uh exploring
AI and QA in general, feel free to reach out. And with that, uh I will uh let you all uh go and finish this conference uh for today and for the whole conference. Thank you. >> Woo!