1000 Services, 1 Year, 0 Downtime: Airbnb’s Zonal Cluster Migration - Sunny Beatteay, Airbnb
About this talk
This talk presents a detailed case study of Airbnb's significant zonal cluster migration, which involved migrating 1,000 services over a year without downtime. The speaker, Sunny B, discusses the challenges faced while transitioning from a regional to a zonal architecture, emphasizing the complexity involved with managing multiple control planes and workloads. The talk highlights the technical and organizational hurdles that were navigated, including the refactoring of deployment systems to accommodate a new abstraction layer called 'cells' that isolates workloads and improves incident recovery. Additionally, the speaker outlines the strategies employed to integrate the migration seamlessly into developer workflows, such as automating portions of the migration process and aligning it with existing deployment practices. The successful completion of this migration not only improved operational resilience but also set the groundwork for future multi-regional service deployments.
Full transcript
Um, so thank you everyone for joining me. As you can see from my top slide, my name is Sunny B, and today I'll be presenting to you 1,000 services, 1 year, zero downtime, Airbnb's zonal cluster migration. So, this is going to be a case study going over Airbnb's largest zonal cluster migration to date. I can say it's the largest cuz it's the only one we've done. Hopefully
the only one we'll ever do. while I'm your presenter, I'm just one piece of a larger team that did this effort. So, I just want to give them a quick shout out cuz as you'll see through this talk, we went through quite the ordeal together. Um, so since the KubeCon schedule went live, I've had a handful of people communicate to me saying they've had similar discussions in
their teams around should they adopt a regional or a zonal architecture. In every case, although ultimately they picked going with a regional, which is more than fair cuz I honestly if you were to ask me, I'd probably say 99 times out of 100, you should do the same. Uh, regional clusters are much easier to just reason about. You have for infra structure engineers, you have the one
control plane spread across multiple AZs, giving you a high availability high availability setup. And same with the app developers, very similar. You have one deployments usually with possibly HPA. Um, you just scale up and scale down. So, you in both cases you have a singleton workload but still highly available. So, it's pretty good deal. And reason why AWS itself recommends this kind of architecture. Contrast that with
a zonal where you have instead of having one cluster running across the entire region, you have a separate clusters per AZ. So, for infra engineers, you're dealing with multiple control plane APIs, and then essentially with your app developers, you're also dealing with essentially sharded workloads. So, instead of having one deployment image, you now have multiple. And so, you're adding a lot more complexity to your system. Which
is why when Airbnb itself decided to and have had discussions, we also decided to go with a regional architecture. And so, we essentially we had our setup kind of typical way, uh, control plane APIs spread across multiple AZs. For our deployments or our services, normally it was a deployment with a single horizontal pod auto auto scaler allowing to scale up and scale down, as well as a
pod topology spread to allow pods to balance across the multiple AZs. And from about 2018 to 2023, the service very well. Uh, we scaled up, our customer base grew, and the scale architecture scaled up with it. And so, around uh, 2023 when we had about 1,000 services running on this architecture. Uh, to equate that into Kubernetes terms, 1,000 services at Airbnb equates to around 5,000 namespaces, 1,200
jobs and cron jobs, 3,500 deployments, and around 36,000 pods. So, we had a fair amount of stuff running on our architecture. And as our customer customer base grew, our system became more and more complex, and so did the financial cost of instance. So, in response to this, we invested a lot into instant prevention. So, at Airbnb, developers get a workspace environment called Airdev where they can test
and develop their code in essentially containerized spaces that can mimic staging and production. So, it allows them to kind of test in a environment similar to what they see in production before they even merge into the codebase. And then even after they merge their code, they also we've also invested into canary frameworks that will allow them as you deploy your code into production, you can test for
regressions in a canary environment and you're basically seeing a slice of the production traffic. If you detect any regressions, we can halt it from going to production and kind of roll back. So, this has definitely helped us in many cases from preventing instances. So, why did we still choose to migrate to zonal? And the all it comes down to essentially instant recovery. Cuz while we had invested
a lot into instant prevention, there's only so much you can do to prevent instant instances, and there's only so much that regional clusters can do to prevent instances. Cuz while regional clusters do protect protect against AZ failures cuz you do have the redundancy in the other AZs, it doesn't protect us from ourselves. So, app deployments and infra infra updates can still potentially cause risk at the cluster
level. And since the cluster essentially the blast the blast radius for any changes could potentially be the entire region. So, kind of put this in concrete terms, we can kind of go over an instant scenario comparing regional and zonal. Uh, to emphasize this is a hypothetical scenario. Whether it was based on true events or not, who's to say? so let's say that we were doing a typical
cleanup operation. Say we had some stale objects that we wanted to clean up to give more disk space to our SE cluster. Um, but accidentally our wrong the wrong handful of resources were marked for deletion. And these resources were very critical for our SE deployment as service mesh. So, this resulted in essentially cascade failures, and hypothetically this could take out essentially all of your traffic to your
cluster. Not a very good thing. Uh, you might imagine though that okay, you just deleted some of your resources, you can just redeploy them and hopefully and so it'll be recovered. Um, but what if I had said that the service mesh is also a critical dependency for your CD platform? So, now you can't deploy that fix. So, now you basically have to root cause and find a
hot patch for this. And so, while you're doing that and while the world's on fire, uh, since there's no traffic going to your cluster, all of your workloads because they're not being utilized has scaled down. And so, finally you do figure out a hot patch, you get your service mesh back online. Now traffic is flowing back to your cluster, but now since all your workloads have scaled
down, now you're essentially having all minimum pods trying to handle millions of traffic coming in. And so, you experience a tremendous amount of back pressure. And there's not really much you can do, you just kind of have to wait for the system to recover and allow for the pods to kind of scale up. And since you're only running one HPA for deployments, you're essentially having one HPA
to scale up for the entire region for each workload. So, it can take a little while, and obviously there can be a big financial cost to this. Now, contrast that with zonal. So, with zonal, let's say the same thing happens. You take out your SEO, uh, service mesh, but in this case it's only isolated to 1/3 of the region. Still not great, but it's better than before,
but you still have about 33% of error rate coming in. But the great part is that now you have the ability to do a zonal traffic failover. So, even before you root cause, even before you figure out what the issue is, you can traffic shift away from the unhealthy cluster, giving your app developers and your infrastructure engineers time to root cause it, figure out the issue without
the world being on fire. So, at this point you've essentially already recovered from the incident. And so, once you figure out the issue, you patched it, the cluster comes back online, you can slowly traffic shift back to the healthy cluster, and essentially have recovered from the incident. So, this was the big gap that we wanted to fill in our own instant recovery story. We wanted to reduce
the blast radius of changes, and we wanted to have faster time recovery without having to do root causing and investigation first. We basically wanted to have a break glass in case of emergency button that allows us to traffic shift away from any of these large incidents. Um, but this is still, you know, it's nice to say that you have that ability. So, why don't you just kind
of why why doesn't everyone do this? Well, because obviously it's a lot more engineering, a lot more complexity to your system. And to migrate entire fleet from regional to zonal is a big lift. So, before we even thought about doing this, we wanted to have essentially a proof of concept. We wanted to see if it was even feasible to do in the first place. So, we had
a subset of engineers essentially do a small scale proof of concept where they would just create the zonal clusters and find some volunteer service owners to allow us to migrate their services over to the zonal clusters. And it was a success. Was it a success? Uh, well, kind of. We migrated two services in about 3 months. Uh, to put that in perspective, two migrations 3 months equals
1,000 migrations in 125 years. Uh, so if we wanted to finish this migration within our lifetimes, we knew we'd kind of need to scale this process up. But the nice thing about the proof of concept is it did give us or it did expose the friction points in this process. Um, so there was kind of two main categories of friction points. First, we had the technical challenges.
Um, and to understand the technical challenges, we have to understand the deployment system at Airbnb. And so, we have a system called OneTouch at Airbnb. Now, OneTouch is kind of its own thing to talk about, it could warrant its own talk. But to summarize, OneTouch is essentially our abstraction layer over Kubernetes for app developers. They we give them essentially a DSL and a minimized config language to
uh, configure their workloads and say what we want to run in Kubernetes. We then take those service workloads into our CD platform, hydrate them using our CubeGen tooling into the actual Kubernetes manifests, and then deploy them onto the infrastructure. Uh, but the big thing about this was that ever since its creation, there was one core assumption was that every workload would run on one regional cluster. So,
that assumption was baked all throughout the system. We even had hardcoded contexts saying which cluster each service work on both in the uh, CubeGen the Kubernetes manifests as well as our CD pipeline. So, that integration was baked throughout. So, we knew that we'd have we'd have to address that if we're going to go to zonal. Uh, the next was just the service migration itself. Cuz during the
proof of concept, it was a very bare-bones process. Uh, step one was just a handful of Cube cuddle commands. Step two was just a silent prayer that things didn't break. Uh, so we knew that we'd have to bench better operate opera uh, better make this better for us going up to 1,000 cluster services. And the second domain of challenges was just the organizational challenges cuz we were
essentially going to shift to a whole new paradigm, not just for infrastructure, but for app developers. And we knew that this would be a fundamental shift to how the product teams both deployed and both deployed and, um, managed their workloads. And so, basically it was a us telling them, "Trust us. This We believe this is the right way to go. We know it's going to fundamentally change
how you do your work, but we think this is the right way to go." So, we knew that if we were going to do this migration, we'd have to keep in mind that we want to best integrate this migration into the app developer workflows and not disrupt them while we're doing this cuz all of us, as infrastructure engineers, we do migrations all the time. It's normal for
us, but for product teams, we don't want to like bog them down with migrations. We want to best As we kind of came up our our migration strategy, we want to um, fit into their workflow as best as possible. So, that's when we started the essentially the project summarization. This is the nickname we gave for our intra migration from regional to zonal. The reason why we called
it summarization was because we introduced two new abstractions to our infrastructure layer. We had the cell. So, you can think of a cell essentially a Docker container before clusters. It was basically a complete self-contained unit with all the infrastructure dependencies you would need and you can package it up and make it reproducible across multiple clusters and different AZs. So, each cell would have its own Kubernetes cluster,
its own service mesh, its own AWS account, its own rate limits. Basically, the whole point was to isolate each cell from each other. So, that way if one cell went down, the other ones would be completely isolated. And then we had, uh, the other abstraction which was a cell set. So, basically it's a grouping of cells in a region. So, the fundamental change was that no longer
would workloads run on a regional cluster, they would now run in And so, we now had to basically take the abstraction and bake it throughout all of our one touch that was previously assuming a context. So, it took about a good amount of work, about, you know, just the large refactoring of our process, um, and just create a way to essentially untangle all the assumptions we had
baked in over the years into our deployment system. And so, we went from now assuming regional clusters to now each would run in a cell set. So, instead of having a regular deploy job, we had a cell deploy job and we were deploying to multiple clusters. And this also, um, made its way into the config language. And so, instead of now instead of having a hardcoded context,
we just had a cell set field. And the nice thing about this is that it allowed us on the infrastructure side to better abstract away. So, no longer did app developers know exactly which cluster they'd be running on. So, it gave us the freedom to kind of migrate them if we needed to in the future. And then because we we knew this was going to be a
big change app developers now they're basically managing multiple deployments. We want to make it as seamless as possible for them to manage these workloads as if it was a single unit. So, we updated our tooling to now incorporate cell, uh, nomenclature in there. And now the app app developers had the ability to actually do more granular operations. So, they could scale up, scale down by cell. They
could get logs for a certain cell, deployment in a certain cell. So, it gave them both the ability to operate at a whole unit and on a cell by cell basis. And then we also updated their dashboards as well. So, that way they were able to do metrics and alerting by cell and by AZ as opposed to the entire So, that handled our deployment system, but we
still had the task of how we were going to automate all of these migrations to zonal clusters. So, we we knew since we knew we wanted to bake this into their normal workflow, we figured the best time to migrate them would be to during the actual deployment of their service. So, we integrated our migration tooling into our CD platform. Essentially, when we migrate a service, we would
inject uh, special stages. So, that way we, um, allowed them to actually migrate. And the special stages were uh, essentially we broke down the migration into three phases. So, we had the prepare phase where we would scale up the, uh, regional deployment, the service in the regional in the regional cluster. The reason we did this is because we wanted to account for any bursty traffic or, uh,
increase in traffic that might happen during the migration. We just want to make sure that the migration is as stable as possible. So, we would scale up first the regional deployments. Then we would deploy the services to the cell clusters. As a finalized stage, we would incrementally decrease the, uh, scale down the regional deployments and allow the, uh, cell deployments to naturally scale up with increased traffic
until we got to zero. And then after that, we would just, uh, clamp and then, boom, migrated. That was the happy case at least. >> [snorts] >> Um, so we had at least the technical challenges for the most part figured out. We still knew we had the organizational challenges of how are we going to spread the word, how are we going to get this message out there,
how are we going to get the buy-in from the rest of the company to adapt to this new paradigm. And so, essentially we went on a road show in a way of just telling everybody, "Hey, this is coming. Our summarization is coming. Here's why we're doing it. Here's why it's very important. How And here's also how it's going to impact you." That was the most important part.
And so, we tried We tried to tell them that this was going to be as seamless as possible for them cuz with a product team, you're normally doing PR reviews and you're normally deploying your services. So, these are things that they already do. And so, we itself, when we were updating your configs, it's just going to come in the form of a PR review. And then when
you actually migrate your service, it's just going to come in the form of a of a deployment. So, as hands-off we took the onus on ourselves and try to make it as hands-off for our app developers as possible. And so, we estimated around a half a day's work for a service to get migrated. And then obviously this led to a lot of questions, a lot of concerns
of what's going to happen. So, we tried to document everything as possible. We had guidebooks, tips and tricks, uh, self-service migration guide, FAQs, all that stuff we could think of to best, uh, accommodate any questions app developers would have. And we even gave them a schedule. So, before we even started the mass migration, we kind of gave them a schedule of like, "Here's when you can expect
your services to get migrated." So, we kind of allowed them to best prepare for it. And so, with all of that said and done, the next thing was actually execute migration. So, we started the migration in the beginning of 2025. Uh, at that point we already had about 19 services migrated just from trial or trialing and beta testing our our work our migration workflows. And we had
about 30% of the total compute running on cells already. Um, but we knew we still had a good way to go. So, we worked in batches. The best way to streamline this was each developer or each engineer on the summarization team that did the migration work in batches. So, anyone who was working on more critical services, uh, were batched in about 10 to 20, uh, services and
migrating in parallel. And anyone working on my less critical ones would be 30 to 40. And then, uh, for So, in order to do this, we also created the automated PR review cuz we weren't going to be manually creating these PRs. So, we created a batch job that basically we would take the batch of services that we were migrating, pass it to a batch job. It would
for each service get the source code, clone it down. We had created a config migration tooling that would basically automatically migrate their configs from the hardcoded context to now the cell sets. And then we'd create the PR. We would then get the information from our service catalog for uh, team that owned the service, the on-call engineer for that service. And we would create a Slack thread, uh,
just going basically notifying the team, "Hey, you're next up for migration." Um, and we're paging the uh, the on-call engineer cuz usually at Airbnb at least allow the on-call people who are in charge of doing support tickets. So, we would have tagged the on-call engineer. And but the critical thing about this kind of Slack thread was that, or I guess the communication was that we didn't want
to migrate each service in one go because there was no guarantee that a one to a service would automatically adapt well to a zonal cluster environments. So, we wanted to do it in phases. So, the best way to test that nothing would happen beforehand. So, a lot of services at Airbnb run with one environments. You have your staging environment, testing, load testing, kind of anything you can
think of for a service. And so, anything that was non-prod, we would migrate those first and that way we would catch the bulk of the issues that would happen. So, if if you had like sticky sessions or things that were basically your hardcoded in your application logic to be running on a uh, a single cluster, we would kind of identify those services, uh, identify those issues first
and allow teams kind of fix that. And then we'd go to partial prod which is any environment that had a partial amount of production traffic. This is mainly canaries. And then only after those two stages were successful would we even think about migrating their production environments. This is how we were able to kind of catch a lot of the issues before they even get migrated to And
then part of this, uh, Slack thread was that we also tagged ourselves in this so we can kind of keep track of the ongoing migrations. And we severely underestimated how much support that would require. Um, but we knew that it was important because for a lot of teams, you know, going through migrations, we want to make sure they they didn't feel we weren't paying attention to them
or if they didn't if they had questions and concerns, we want to make sure that we were there to answer it and they felt that, you know, we weren't just kind of throwing this over the wall to them. They could have back and forth with us. And so, in the beginning, it was definitely a bit slow, kind of new learning curve, um, getting used to this new
workflow. And so, we were able to but with time, it did get quicker and quicker. And so, by at the end of the first month, we were able to migrate up 10%. And then as the months went on it was really cool to see we had a chart or a dashboard showing the converging of regional clusters to cell clusters and so it's my favorite thing that favorite
pastime at being every week to see what would the new chart look like and so finally in June is when we hit the 50% mark and that was kind of a cool market to like okay we're actually doing this this is happening we're on track we're halfway through the year and we're already at 50% and we got our process pretty down pat where even people thought we
were using AI for the process but no it was still a lot of human intervention maybe now if we're doing it we definitely would probably leverage more AI but back wasn't quite ready for it but that wasn't to say that the process was completely seamless so we did reserve a subset of projects for the very end of the year and these we called essentially the omega ones
cuz these are the ones that we didn't want to take lightly to give some perspective for the average service we require very little coordination we could probably complete the migration in about one or two weeks with very little coordination with the service owners for these subset of services require a lot more coordination a lot more white glove treatment because a lot of these services had a lot
more stringent requirements for example a lot of these services want to run on specialized node pools away from the kind of the general services so we had to update our web hooks to allow for them to label their services and then with that label we can allow them to migrate to sh get routed to a special node pool they also wanted to make sure that the migration
didn't happen during a time that would disrupt both their traffic and their development so there's a lot of Friday night deploys cuz that was the best time we could find so I don't know how you feel about that but still had to do it and then we even had to upgrade our Etsy cluster from 8 gigabytes to 16 to just these are very very resource intensive services
so we want to make sure we had the capacity to actually migrate them to our cell clusters but with time still even with these services that had all these demands we were able to kind of just following our process of three phases caught any issues and allowed them to with time they still were able to get knocked down until finally we had the final omega migration so
we actually had a chart that showed we gave this mostly for developers but I like to watch it too cuz you could see the migration happen real time so the green in this case is them running on regional and as they migrated you could show the percentage of the pods more like going to cells so this is with the final omega migration this was very cool to
watch and once the final one happened all we could do left was to celebrate so like I said within a year we migrated from the bottom all the way to the top and with it required a lot of PRs around 3000 migration PRs in total to migrate all the services but as of today we have about 95% of our compute running in cells and then we've also
as I said migrated over 1000 services all with and we even as a cool little accent mark got to delete quite a few of our regional clusters so this is the dashboard showing and so you can see when the metric stopped collecting is when it got deleted so it was really cool to actually get the chance to delete our regional clusters not all of them but a
good chunk of them and then I mentioned at the beginning of this talk that the whole point of this was to allow us to have kind of the ability to do traffic shifting and that and so we actually were able to complete our first full zonal traffic migration shift in September 2015 in less than 10 minutes without any instance and when I mean like we actually did
go to our live clusters shift them over and make sure that there weren't any issues so we were able to complete and having completed so that was also a very very cool milestone >> but with all that said there is still much more work to do so since cuz as the as part of the cell division team I can tell you we interacted with a lot of
service owners we got first hand account from them the feedback like I wish this was better I wish that was better all of the kind of things we heard from the pain points they have with using our tooling to migrate to manage their services so since cell division has ended we've been at work trying to improve the experience for those developers we've been creating better debugging tooling
to kind of help visualize the cells and we're actually using headlamp for this to kind of visualize and debug in real time we're actually experimenting with AI as well to help simplify the configs and reduce the amount of boiler plate and just kind of cruft throughout the all of our service configs for the entire organization so you can see we've done a good amount of work reducing
the total YAML in our code base and probably important thing for me is that we're actually now working on workload placement controllers that will allow for PR-less migrations so no longer we have to go to service owner and get PRs we can do it all behind the scenes and it's music to my ears thank you very much for listening >> [applause] >> I think yeah we do
have time for Q&A so in case anyone wants to ask a question I'm not sure do you there's like a microphone there I'm not sure if you want to go to the microphone or you can just talk or if you talk say it to me I'll I'll repeat it out hello just I want to ask all your application and services you have only one region or more
or how it works that's a good point so right now we're we are experimenting as part of this we that was this this is also essentially a first step towards going multi-regional so originally we're doing one region but as part of this now that we have the ability to run on multiple clusters we're going to expand that to now running in multiple regions and we are in
the process of doing that we do have some services already running in multiple regions and this was kind of the first step towards going to that all right well no other questions thank you very much and then feel free to come talk to me afterwards if you see me if you have a question afterwards for desk thank you very much everyone
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32