DEVWorld 2026

Maxim Scheplin - The Formula for Faster Outage Recovery

23:42 · 07 May 2026 – 08 May 2026 · YouTube

About this talk

In this talk, Maxim Skepelin discusses the importance of reliable incident response in software systems to minimize revenue loss during outages. He outlines a structured approach consisting of three components: time to detect, time to acknowledge, and time to repair. The speaker emphasizes the significance of defining service level objectives (SLOs) to establish clear expectations for system performance. He also highlights the need for a structured on-call process, effective communication, and team training to improve incident response times. Furthermore, he advocates for investing in observability, creating training programs, and preparing for common service failures to enhance the repair process. Ultimately, the talk provides guidance on refining incident management strategies to help organizations better respond to production issues.

Full transcript

Hello everyone. Ready? Hi folks. >> [sighs] >> Um I want to open with a question. How much does it cost when soft software you are responsible for doesn't work? In my experience, for a global company, an outage can cost millions in lost revenue. And with so much to lose, reliability and efficiency of incident response becomes critical. And that's what we're going to talk about today. My name

is Maxim Skepelin. I work at booking.com. In the In the past 10 years, I worked with many engineering teams to improve incident management, how to handle production outages, and how to respond faster. figuring out best practices of incident management is costly and stressful. With this talk, I want to give a structured approach how you build your incident response or tweak your existing ones to save time in

in and stress. >> And to do that, I want to think about an outage as a simple formula. Outage duration equals to time to detect plus time to acknowledge plus time to repair. Time to Time to detect is from the moment something bad happens to your system to the moment when you know it. And that time is not zero. Time to acknowledge is the time from knowing

there is a problem to someone actually jumping in to handle the issue. And the last part, the repair, time to repair is from acknowledgement to complete resolution of the issue. And those three parts cover the all time spent in the incident response process. And if you want your outages to be short, you need to work out how to make each of them shorter. And how to maximize

efficiency of that And I will talk one by one of the of which part of the of the formula and give some ideas how you can your existing incident response or tailor the the one you're building for the team. Let's time Let's talk Let's talk about time to detect. And again, from the moment you know there is a problem, from the moment something bad happens to the

moment when you know it. I've seen I've seen that part is often ignored because the common argument that well, kind of obvious that what an outage is, you see a problem when when there is a big one. Except, it's not that exactly true. For example, I work on a mobile app. I push back end changes. App is working, but older versions are broken. Is it an outage

or not? Hard to tell. It depends on the on the impact. Similarly, more crazy even more crazy my website doesn't work with certain browser extension installed. Like, how crazy that could be? And yet, it could be considered as an outage if the impact is sufficiently big. And the point is, if you can't draw the line, if you can't if you don't have a tailored definition for your

specific system what an outage is and what it is not, you have no control over how long it might be because you don't you might have one and don't even know about that. And the common solution here is to have to define service level objectives for your systems. Service level objectives is a very simple framework popularized by Google SREs, uh but invented way before that. Where you

basically describe expectations from your system as statements, X must be true Y percentage of time, where X is the metric that you care about, and Y is the threshold, usually expressed as a number of nines for that metric. For example, you might say 99% of page loads must have must complete under 300 milliseconds, or 99% 99.9% of API requests must be processed successfully, if rely API reliability

and page load speed is something you care about. And for each of your service level objectives, you define those those things you care about, you start measuring them. And you have a graph something like that for every for each service level objective, where it shows you right now how you how how it is true. What is the load speed percentage of percentile of load speed under page

load speed under 300 millisecond, what how what is the API reliability right now? Basically, red line shows you how you want your system to perform. The green line shows how it is actually performing. And that gives you binary answer whether you have a problem or not. As long as green line stays above the red line, you don't have a problem. If any of your service level objective

violated, means it's down, you have a problem. And you're not spending time debating whether it's an outage or not, whether it's a bug or a feature, you can tell by looking at the graph. But, there is a catch. How do you define your SLOs? if you go to an engineering team and ask them, or work with an engineering team to define service level objectives, you are very

likely to end up with something deeply technical, database uptime, error rate, latency, something something something, which is all well and good, but none of it tells you whether your your users actually getting what they need from your system. And that's an important part which I want to emphasize in this talk. your software does something useful, whatever it is, like sending marketing emails, transferring money, selling tickets, generating

reports, doesn't matter. Like there is a reason your users interact with your software. Find a metric that represent that and and try to tailor your definition of your service level objectives close to 99% of marketing emails must be sent within 15 minutes from from the scheduled moment. That sounds like a good good thing to to measure. And in the same tech stack and the same technology, but

you're measuring the business outcomes your software to create. And it's actually a good exercise to try to find what what describes what it is your software is supposed to do. And a colleague of mine wrote a very nice article. Its title is your system is fine, your users are not, and that goes in depth about how to define service level objectives in business terms, how to what

are the best practices in big companies and so on. There is a short link if you still know how to type things in the browser, and there's a QR code that leads to the same thing. Uh and that's how you get time to detect under control. If you can't say whether you have an outage or not right now, you have no control of how long your incident

response will be because you have no way to detect an outage. And that's the first part of the The second part is time to acknowledge. And whenever I talk about acknowledgement, I often hear like naturally alerting pops up. Topic of alerting, sending more alerts, making better alerts, more alerts. Alerts. It's a kind of solution to that. That's true. Alerting is important part. And it's but it is

a small part of the of that thing. And it's an easy part because when you have meaningful service level objectives, alerting is easy. difficult part is what happens after. Okay, alert fires, but who responds to it? Like what if it is at 3:00 in the night? Do people even know they supposed to react? Do you know how to reach those people? Basically, slow part is the on-call

process, not firing an alert. That takes seconds. And in order to improve how you your time to acknowledge, you need to transition from chaotic response, alert fires, and then everyone crazy running searching for people who can solve the issue to a structured, predictable on-call schedule. And for that, you need to start far, far ahead. And the very first step is to actually to build your on-call process,

is to identify critical services, review your tech estate. Again, whatever your your specific area or the whole company, and find business critical systems, like we spoke about the impact of software a minute ago. Not everything is equally important, right? Service name, components, system, whatever terminology you use, what it does, and when it's down, why why it's bad. And focus on the immediate impact. Again, giving an example,

if I run a web shop, my index page is broken, even worse, I pay ads for for ads to Google, users come to my website and see an error page. How great it is. It's bad. On the other hand, if order confirmation email's delayed by 2 hour, that's not a big deal. My point is that even in business impact, not everything is equally important. So, find the

critical ones with immediate business impact on company reputation, compliance, revenue, like something tangible and bad happens immediately when that thing is down. And that's the definition of business critical services, I think. step number two, map ownership. Like, you Who are the people who supposed to take care of that component? And you would be surprised how often I've seen a business critical service that nobody owns or only

one person in the whole company knows how to deal with that. If that's your situation, your incident response will not be great. You need sufficiently big group of people, like at least five, better eight or 10, to take care about uh of each system, of each business critical system. It takes effort to run critical systems in production. Like, you need people to do that. And if you

don't know the owners, have those conversations. Escalate those the issues. Talk to managers because you have impact of the downtime. You you identified that. Like, if you don't find owners, you're basically accepting that it's going to happen for indefinite time for unpre- at any frequency, which is probably bad. Like, that's why you need impact of the downtime to have those conversation about ownership. And the first step

is actually easy. Build a schedule for on-call, which is like we have a bunch of tools of like PagerDuty, OpsGenie, Spike, Alert, probably dozens of others, where you map to time blocks so there is always someone available. And and the tool actually takes care of scheduling, overrides, escalations, notifications to send to responders like it your phone goes crazy when you have an when you have an app.

The selection of the tool is really not important because I worked with dozens of them. More or less they do all all the same things. The tricky part is again people part. If you're building on call and on call process and if you're building it 24 hours 7 days a week coverage, you need to give people motivation to do really. Like I never seen forcing people to

take on call shifts work well. You don't want to hold hostages in your on Like either people people engineers need to make that choice. Like yes I I want to be there. I understand why I made that decision. Either because you pay them extra, you gave time off, you managed to find people who deeply care about success of your company or or give some form of recognition.

Uh like you need buying from engineers. It can't be forced really. Like it never works. The the on call never works when it is forced. And one important thing check your employment contracts. Talk to HR. Talk to legal if you actually how your contract structured with engineers whether you're allowed to do that or not because you might have an a bigger problem if you can't uh on

that side. So and that's time to acknowledge. Again, to make time to acknowledge shorter from to finding the person who can act on on the solution, you need three steps. Identify business critical systems. Not everything is equally important. Uh find the one that immediately impacts revenue, reputation, compliance, something. Map ownership. Find people Find people will take care of those systems. Like It not always happens naturally. Sometimes

for like you might end up with business critical services not owned by them. And actually sometimes definition of a critical service is weak. Like spreadsheet might be a business critical service in the end. And then you you build on-call schedule. Like there is a bunch of tools for that. They all very much similar. Pricing and UI changes, but that's matter of taste and budget. acknowledge under control.

the time to repair, you never know what's your your next outage to be is going to be, so and yet you can prepare for the unknown. And one thing, we spoke about time to detect, we spoke about time to acknowledge, and those things the interventions that I touch, you do them once and they work. You define service level objectives and they work for weeks, months, and even

years. You You might tweak them here and there, but time to acknowledge, building on-call also you do it once. People join and leave, but the process stays more the same. The people get It becomes a practice. Time to repair is a tricky thing because it drastically different from one outage to another because the root causes are different. And yet there are a number of things you can

do. >> Thing number one. Invest in observability. Really, in 2026 I shouldn't be telling from the stage why it's a good idea to have metrics. But yet, you need granular metrics to understand what's going on inside your systems. everything from incoming traffic to CPU and and memory utilization on the hardware, even if you don't own that hardware, there is still hardware, and everything in between. Write metrics,

build dashboards. Uh you really need to to have telemetry data to correlate events, to spot anomalies, to know that because otherwise, even if all parts works perfectly and you have no idea why it's happening, that that's the thing that will help you answer that question faster. Uh yeah. Second thing. Training. this section of the time to repair is the most obvious things that I'm talking from the

stage. Training from the engineers. Have you ever been in a situation when someone in your team or not you, someone you know, pushes a changes to production, it breaks, and then that person spends 30 minutes to figuring out how to do get revert because they never done that. That's the thing I'm talking about. Like and there are more stuff like that. So depends on the cost of

your outage, depends on depend on depending on how much you lose, like you don't want to pay for that, Uh what There are tools and scripts and whatever web interfaces that you used to operate your systems in production or someone using to operate them in production. How to scale pods in your Kubernetes clusters, how to disable availability zone in AWS and whatever you use. How to revert

changes, how to I don't know, skip uh steps in CI pipeline if you need and so People need to know how to do that like because when you when you get an alert, when you page at night, you are not at your best to learn something new like no knowledge should be on the fingertips. And to do that, simulate alerts. Like make those situation in a safe

environment, in a test environment where people can practice. yeah, basically make sure that everyone in in the team knows how to do that. second third thing is to plan ahead. Like certain things if you have so if services in production and they live long enough in production meaning your company is successful you will get next thing. You will get sudden traffic spikes. The load you didn't expect.

You will get DDoS attack. your network will be slow or partitioned entirely. Your dependencies will fail or like those things are very likely to happen in the long term. Don't force people to be creative at the moment on figuring out when it's already Spend some time in advance looking at your architecture identifying those weak spots like what if what if that happens. And then write down mitigation

plans. Like and that's how you create playbooks. Okay, when you get scale workloads, activate our firewall rules, enable another zone. Like but to write it down like don't force people to be creative because again at 3:00 in the night you are not at your best to come up with great ideas. And keep that collection of runbooks. Yes, for you can even link it to specific alerts and

give tools to triage. Do that. And that's how you get better at repairing systems. Once you know that it's something broken when there is someone when someone acting on the it invest in observability, metrics. We have open telemetry standard. We have bunch of open source tools, bunch of vendors who can help you to that for relatively reasonable price. Training. Really, like how many people know how in

your team know how to do Git revert on on the master branch or or how you name it now. And write playbooks. Like anticipate changes. Like I anticipate failure modes. Like there are there are certain things that will happen. Like prepare that. Write docs. Wri- Write scenarios. Don't write novels. Write like step-by-step things. Go here. Check this dashboard. Confirm the problem. Run this tool. Verify the solution

here and so on and so on. Like it's it's actionable down-to-earth >> [gasps] >> And that's the summary of all three components. Again, you spend time in three stages in incident response. Time to detect, time to acknowledge, and time to repair. And that thing can be indefinite. Like you like to prevent that, define service level Time to acknowledge from the moment you know there is a problem

to the moment there is a person acting to resolve it. Build on-call process, build on-call schedule, use tools for that, but remember the people part. You need to give people motivation. and check that you legally can do that. it's all invest observability, train engineers, and plan common failures for common failures. Certain things will happen, certain things will go wrong. You just just prepare for that. get better

at responding to incidents. That's how you build incident response if you don't have one yet, the process for your team. And the thing is, it won't going to be perfect from the first time. Learn from every outage. When something happens and you resolve the issue, get people in the room who were involved in the in the resolution process, and run a postmortem. discuss what happened, what was

the impact, how severe it was. How the problem was was detected. Was it a human who found an issue or was it an automated alert? Were there too many alerts? Were they too late, too noisy? Like you can you can discuss those things and tailor uh tweak your systems to get better. How efficient was the incident response overall? Did you have tools? Did your playbooks worked? Did

people know how to how to how to use the tools and so on and so on. And most importantly, can this issue happen again? And if answer is yes, then you have some work to do. Like you never know what's your next outage will be, but you can prepare you can start preparing for how efficiently you will respond. Good luck. And now it's a minute for shameless

self-promotion. Uh my colleague and I we wrote a book for engineering leaders about building efficient process processes running planning for the team, defining strategy, and so on and so on. If you're interested, check it out. Like on-call is one of the things where actually managers and engineering leaders can make a big difference building the efficiency about that. It's one of the topics, but there are many others.

Thank you. If you have questions, catch me in the hallway.

From event

DEVWorld 2026

07 May 2026 – 08 May 2026

All event videos
Back to Watch