Navigating Multi-Cloud High Availability – Architecture and Operations with Kubernetes, Service Mesh
About this talk
This talk discusses the significance of multi-cloud architectures in enhancing availability for enterprise applications, presented by a member of the SAP Cloud Identity Services team. The speaker highlights the critical role of availability, using an incident involving Shell's SkyPad application as a case study to illustrate the impact of outages. He defines what a multi-cloud strategy entails, emphasizing interconnected environments that prevent vendor lock-in while promoting flexibility and resilience. The session also covers the technical approaches employed, such as leveraging Kubernetes and service meshes to facilitate seamless interaction between diverse cloud environments. Additionally, the speaker outlines various availability strategies and the challenges they present, ultimately advocating for a well-planned multi-cloud adoption to mitigate risks and optimize performance across cloud providers.
Full transcript
Hello everyone. I would like to start with three questions for you. So first one, how many of you are running your applications on Kubernetes? You can raise your hands. Okay. Okay, that's expected. Um how many of you rely on multiple on multiszone high availability for your projects like to host your applications in a single data center uh between the different availability zones. Okay, some hands here and
there. Great. And the third one, how many of you have an environment which is spread between several cloud vendors and act as a single solution? Okay, not so many. that was expected. For those of you with your hands down, you're on the right place because after this session and during it, I'm going to share with you some real world examples of how um what problems the multiloud
can resolve for you and for the rest with the hands up. I'm interesting to speak with you after the session. Um but first I would like to share with you who are we and what we are doing and why the availability topic is important for us. I'm part of the SAP cloud identity services team, one of the most critical kernel services inside of the company. We are
managing the authentication, authorization and user provisioning for all the SAP C customers and CL and uh internal employees. Basically when you need to access some of the SAP applications when you buy one you we we have to allow it and as being in the central of the company this means that when we go down everything stops from internal services to external customer business scenario and I'll give
you only one example my favorite one with a company called Shell you should have heard of They have developed an application called Shell Skypad which is basically running on a tablet and they're using it to fuel the airplanes all over the world. And several years ago during one of our outages which are not so many, we managed to impact them and this caused a massive traffic jam
and delays on the airplanes all over the world. And this is only one of the examples for our customers. So this makes the availability topic important for us and we consider a 3minut outage critical. Thanks to the multicloud we are able to increase our availability and to make our product flexible on infrastructure changes. And now let's address the elephant in the room. What does the real multicloud
mean? Well, it means interconnected environments, making your application and product and platform cloud agnostic, meaning that it can be spread between single environment of yours or a single region. The statistics shows that 89% of the enterprise companies have a multi cloud strategy. But do you know what those companies understand of multiquote? They think that the multiquote is using different cloud products from different vendors together to build
your software on. This is not what this session is about. We are going on the next level here because it's simple to take the database from Azure and the block storage from Amazon and build your product on top of them. And it's totally okay by the way because you're using best of both worlds. But you just made your product vendor locked which is something we would like
to avoid. There are many articles in the internet IT podcasts videos stating that interconnecting your environments is something you should not do because it has it adds too much complexity. It has many moving parts and it needs more skilled people to operate. And this is true. Having a multi cloud is increasing your complexity and rising your cost. And the question is, does it worth it? To answer
this, you should first answer the how much question. How much will I have to spend in multiloud? How much efforts are needed? And if you take a closer look at those articles, you can see that they are dated several years old. Even the first one multi-quality set a trap was dated seven years ago. It was written there. So seven years ago we run our production on virtual
machines all over the place and we all know what the virtual machine instrumentation what skills are needed. They were different between the different cloud vendors. Nowadays we have the technology which allow us to abstract those differences away to interconnect our environments and to achieve multi cloud easier than before. Imagine that you could make your different environments talk to each other in a secure way over the internet.
This is now possible thanks to the service mesh mutuals which is even more common these days. It allows you to bridge your environments together. What if I tell you that you could instrument your environments and applications the same way independent on the cloud vendor they run on? This is now possible thanks to the Kubernetes which gives you the the common ground which is the same everywhere. At
this point some of you might think okay multi cloud is a bit possible a bit more achievable these days but what is it for me? Why should I care about that? And I'm going to give you four four real examples of the things which I value about this solution and which we're using. First, it allow you to mitigate the vendor walk-in. Imagine that you receive a better
price, better availability, or something better from a competitor. You can migrate without the fear of losing your service for a certain period of time. without the fear that you should invest and spend time in refactoring just to substitute one database with another. How to do this? Simple. Just run on Kubernetes. Put your database in. We're using that since many years and it it is successful. Second, you
can not only move between the cloud vendors, but you can do it zero downtime without impacting your We were in a situation where one of our cloud providers decided to sunset a data center we were running on. And thanks to the multiquote, we were able to spin up a new data center closer to the first one, binding, bridge them together with a service mesh, interconnected the databases,
sync the traffic, sync the database data, and then move our customer traffic to the second data center and stop the first one with zero downtime. No change management justifications, no customer impact whenever we wanted to do and this is something powerful. Third, multicloud helps also to fix latency and compliance issues. We are currently in the process of migrating to an external database and because we have 18
production regions for some of them the database is far away from our servers and for those thanks to the multi cloud we are able to just do this zero downtime migration move to a closer data center and resolve our latency issues otherwise we would impact our customers by having a longer response which is not acceptable. It also helps you to fix compliance For some customers, it's crucial
their data to be located in a certain part of the world. And for them, you could choose either to run in a single data center multiszone and accept the lower availability or to combine several data centers in the same region but from different cloud vendors and twice their availability. It depends on the business critical case for your customer. So you have options. And the fourth one, multi
cloud help you to avoid major cloud provider outages. And we all know that those things are not very common, but when they happen, the impact is usually massive and it's not resolved within several minutes. And if I want to give you an analogy, I would compare the multicloud to the new car. see them. They both are considered more expensive and more complex compared to the other options
out there. But they they both save time. The new car save time from repairs. The multi cloud save time from downtimes and actually prolonged outages. And it is up to you to decide whether your time is worth the cost or not. As I've gotten older, I've came to value my time much more than before because nowadays I can I simply cannot afford my family trip ruined because
a car breakdown, especially with a young child. The same principle is valid for the multi cloud adoption. See, the largest enterprise companies simply cannot afford downtime and impact for their customers during prolonged But [snorts] is the multiloud always the right the right option? Definitely not. There are other two which we're going to talk about multiszone and multi data center. But first to choose the right one for
you, you should address the availability topic. It's crucial during the initial design. So from this table I would like to emphasize that the more the more availability you aim the less outages and or the less downtime you would be able to afford which at the end will bring your complexity up and also your cost. If we take the three 9ths availability which is the more common the
the most common availability used for most of the services which are not considered business critical. This allow you to to have more than 8 and a half hours every year of outages. This means 43 minutes every month. This time is enough for you to wake up somebody and he or she to do some manual magic efforts to resolve your your service to revive it. It's enough time.
And for this profile, you could go perfect with the multiszone architecture. You can go with the multiszone even with the three and a half nights. Here you can no longer tolerate outages more than 22 minutes per per month. But if you invest some some efforts in manual in automatic tests and some other stuff to not bring down your system during an upgrade, you can survive with the
multiszone. If you need a bit higher availability here, the things are getting even more interesting. The multi data center is perfect for this one because here you cannot tolerate more than four minutes every month of outages. Here you should rise your complexity and invest in redundancy, self-reovery and other mechanisms which at the end will rise your cost. So what are the high availability options which we mentioned
already starting with the multiszone three and a half 9s it's okay most of the cloud are telling you that they will they will give you four 9s if you host your service inside of their data centers and spread your work your your applications between the different availability zones. So, but this is not feasible because now first they won't achieve it every time and and second you have
zero window for maintenance breakdown because those things happen. So the three and a half nights is the availability which you which you should aim with the multiszone which by the way is the cheapest solution. Rising the availability you should think outside of a single data center to the multiDC. MultiDC means same cloud provider, several data centers which act as one environment for you. The private network is
the same. And if you need to go even higher with the availability, then you should choose the multi cloud because multi cloud can help you to survive in global outages. In our team, we're using all three architectures depending on the customer case. So let's talk about about our architecture. We rely on the same thing deployed everywhere. Meaning that one of our regions independent QA, production, staging, whatever
is the same three landscapes below. You can think of a landscape as a set of servers enough for you to operate your product on. But we have three because the database is inside of our Kubernetes and because we need connectivity and because we need to survive with one landscape down during an outage. The rest two will continue serve the customer traffic and it will give us enough
time to resolve the problem with the first landscape without impacting our customers. Moreover, we have invested in DNS failover mechanism. Basically each of our regions have an URL and the IP below is under our control. Whether the first landscape will receive the traffic or the second one, those actions are automatically synchronized and done uh from this system which is external. We are using Kakami global traffic manager
but there are others many of them use uh are offering the same principle. We also have added Kubernetes everywhere and not only that but we are using is as a service provided by SAP team called gardener. Maybe some of you have heard So this allowed us to now move away from the OS patching and and um OS maintenance anymore and focus more on other DevOps stuff. We
also introduced the service mesh everywhere to all of our environments because we wanted the freedom to choose what to do with them, how to migrate them without without the impact. And by adding more things on the table, [snorts] we have to address the design and it should have be it should be common to reduce the complexity. We needed interchangeable components and centralization centralization on more places. So
now let's talk about how those components allow us to achieve our availability goal. Starting on our original concept, this is an example of a multiszone. And here we have a single data center and our landscapes are spread between the different availability zones. The private network is the same provided by the cloud vendor and the availability was three and a half 9s. Now here this is a multiC
basically it's the same picture but now instead of different zones we have different data centers provided by the same vendor. The communication can be private if you wish. By having the automated DNS failover mechanism on top, we are serving the traffic with this landscape and in case of failure the the automatic f the automatic the the traffic is forwarded automatically to the second data center between one
and three minutes worldwide which is powerful. we also since we added Kubernetes on the p to the picture now our landscapes each of them is a different Kubernetes cluster which allow us to survive with a gardener cluster downtime a Kubernetes cluster downtime which also give us some more higher availability and by adding the service machine to the table now we can run our different landscapes in different
cloud vendors and make them talk over the internet using mutual TLS. By having that we able to achieve even more than the 49 So by adding more components to the table as I mentioned already we have increased the complexity for our service and for that we needed a common design interchangeable components the same set of software deployed to all environments minimum differences centralized solutions for management which
at the end limited the complexity as I mentioned one word centralization we rely I on the central git repository which is hosting the code for the deployment of all of our environments. One repository before we had four. Also all of our Docker images are infrastructure agnostic. The same image deployed everywhere. The changes needed for the environment to function are deployed with with Kubernetes config maps and secrets.
those are different from the image and the image is then signed, packaged and uploaded to a central artifactory storage available to be consumed for all of our environments. We have we are using customize to structure our g repository to be able to achieve 95% of our code to be the same basically and the rest 5% is delta and it's added on top of the manifests thanks to
the kubernet to the customized overlays we are also invested in in central secret storage and for that we are using hash corp vault As a continuous integration system, we decided to go with the Jenkins because we were having a long history with it and it's the one responsible for our code build [snorts] image signing automated tested tests and even in it instruments our continuous delivery process and
for that we go with the Argo CD. Why? because we wanted to have a central place for Kubernetes orchestration. The the changes have to be visible for us whether the change hit our system in it was successful or not. So we also wanted to be in control of what is delivered to which systems and when. So for our most business critical scenarios, we are able to deploy
the changes on demand and also automatic if you want. And I would like to finish with a strong sentence. Multiquality is not a trap but with is a way forward. And I would like you all to think for your customer scenarios and for your product needs. And after this session, answer the following question. Can you benefit from multiquil or not? Thank you. [applause] >> Okay. So, we
have five minutes for Q&A. Someone >> the gentleman over there. >> Uh yeah, we have Hello. So my question is you said that you have for example a database which is hosted on Kubernetes directly right? So, how are you balancing? It seems to me like a trade-off. If it is not, you know, correct me. Um, if it is, how do you how do you balance the tradeoff
between having to manage this database and any other infrastructure yourself versus uh outsourcing it to the cloud providers managed say database um publish subscribe whatever. Mhm. Well, if we if you move our database to for example an an Amazon database, it will be located at some at some part of the world. And if you want to move our region to another one, we we should sync to
the same cloud vendor database, right? If we want to switch to cloud vendors, we are not able. So we we are vendor walking our database to a certain cloud vendor. I mentioned that we are now in a process of migrating toward towards an external database and it is SAP HANA. So SAP HANA some of you might know it but it's basically the same thing. Why we choose
SAP HANA? Because we have to choose SP HANA because we're SCP. So that's the reason the for me there is no trade-off if the database is inside inside your Kubernetes there is not no problem we were running in MongoDB since many years without any problems >> yeah we Yeah, the tradeoff is that we are a bit more database administrators at some point, but it's still manageable for
us. >> Yeah, >> I have a question on the slide with the infrastructure providers. There was Asia AWS and another one CC3. What is this one? >> CC3 is an internal it's converge quad. It's internal SAP cloud provider which we are using for maybe half of our production environments and the rest are Azure, AWS and GCP at this point. So we are running in four cloud providers
at >> Thanks. There is a question on the hand. >> Thank you. Uh if you can show again the slide with the percentages of the availabilities. Uh you mentioned that the multicloud is for going to from uh 3 9ths and a half to 49s or up. >> Multi data center is to go to the to the 49s. So multi uh multiszone was the first one. It was
able to with by using it you were able to >> the table with availability percentage. >> Oh, okay. Just a second. >> It was one of >> you mean this one? >> Exactly. >> So, multicloud is somewhere below the >> Yeah, it's below 49. >> Yeah. My question is by what margin does the cost increase for going in that step compared to the previous step? Well, the
the exact margin I'm not I have not done the calculations, but with the multi data center, you're basically using the same thing. You you are combining several data centers to speak together. By going to multi cloud, you need to invest in also service mesh which increases some of your resources and resource consumption of your clusters. But it's it's uh something like 5% or something like this. It's
not it's not >> the big jump is from >> it's from multiszone to multi DC. Yeah. >> Because here you're running in a single and here you should have three [snorts] >> which basically triples the cost. >> Okay. Thanks. But from those three two of them may serve your customer traffic 50/50. And the third one is for the database clustering. So it's with less resources. So if
you have to think of the the resources because in SAP we are we are thinking more over the availability topic but still you can you can lower the resources on those landscapes and at get some benefits about the availability as well. So thank you all.