Cloud Native Theater | Istio Day: The Good, The Ugly, and The Bad... Alfonso Ming and Jorge Turrado
About this talk
In this talk, Alonso and Jorge discuss their experience with adopting ambient mesh to improve microservices management. The session begins by introducing key concepts like service mesh, API gateway, and the sidecar model, outlining their respective roles in security, observability, and traffic management. They describe the transition from the sidecar to the ambient mesh model, highlighting the advantages such as reduced operational costs, better resource allocation, and simplified traffic management. They also address potential challenges, including the increased complexity of managing layer 4 and layer 7 policies, monitoring, and maintaining authorization rules. Ultimately, the speakers emphasize that while there are hurdles, the benefits of ambient mesh for their infrastructure significantly outweighed the drawbacks.
Full transcript
First of all, a quick introduction. >> So I am Alonso. I work in Spart. So I intend to improve the life of our teams. Happy to to stay here. >> And my name is Jorge. You can call me George if you prefer. I understand naming it's complex. Horge Turd to improve even more the spelling. I work as principal site reliability engineer at Esbar Digit. And there are
my achievements GADA CNCF Microsoft MVP and which first of all thanks to everybody for joining us today. There are a lot of super interesting meetings. I hope that this meeting will be at least useful. I'm not going to commit interesting. No, I I it's interesting where we are going to explain our adoption process of ambient mesh. Erh and luckily the topic that we have chosen to to
guide our session is quite aligned because I see that so many people here could request to be a risk group during the next flu campaign. So the idea is starting by a quick introduction just to ensure I'm pretty sure that all of you know what is what service me is and what gateway API is. Sorry if not or if you are there we are going to invest
five minutes just to put everybody of us in the same page to start with our uh lessons learned. So well as you know and service mesh is an uh software that works in the first layer. So can help us in order to manage our microservices and is built in this case is was built by three pillars security for the one hand observability and traffic management. So uh
as uh if we are talking about security uh it provide uh authentification and uh muts for all in the all the services and if we talk about observability we can uh see and point of our in network and we can detect uh detect errors so fast. So if we talk about traffic management uh it controls the flow of data and uh lot the balance balance the load
between the all the services. So >> yes sir >> yes sir. So well the first thing that we did when we introduce a service mesh in our cluster it decide about CD car model that is the standard for the last years about this or ambassar model works with a proxy every pot has a proxy and all the communications uh is between the proxies. The proxies handle all
the traffic and pass the traffic to the service A from the uh from the service A sorry to service B and all the communication and the authentification is provided by the proxies. So this is a model that work very fine but present at least two problems >> before before jumping to the next step everybody here knows what is this approach right? >> Yeah. Yeah. >> Perfect. >>
No someone that doesn't know it. Nice. We can continue. >> So what this CD car model presented at least two problems. The first one is the related about the overh commitment because uh if you have 1,00 running in your cluster you will have thousand proxies running too. So it mean that at the end of the month you will pay a lot of more money in your cloud
doesn't matter Google Asure and uh your product will be more uh expensive. So another problem related with the CD car model is about operational complexity because if you need to uh update or some feature about the core of the you will need to restart all the balance inside the the mess. of well imagine if you have in our case that we have a lot of cluster a
lot of pots running uh could be a high risk that you cannot suffer in a production environment. >> Yeah. And we ended with a drift of five different versions in our uh proxies because one is point 211 but the other the other one is 212 25 obviously supported versions. I mean we are not using unsupported version but you know what I mean the drift that you have
in the amount of versions that you have because critical warloads cannot be restarted at any point just because the proxy needs to be updated. So with this in mind uh yes next. >> Yes sir. >> With this in mind we started to think about uh work with ambient mess that is a different approach. In this case the proxies is in the top of the node. Uh when
you deploy ato in ambient mess you will deploy a demon center that is called c tunnel and in this case it means that you have a c tunnel par per node. So in this case the proxy is in the top of the node and not in the uh in the pot. So in this case the communication is between these tunnels and use uh for the traffic basically
we can say that in MBMS we can talk about mult observability way points and policies. So if we talk about boots uh cel provide the boots basically is their job when you configure and level n space for uh to be part of the mess uh is to configure a so listener in the bots in the in network space and sitan can intercept those traffic. So this is
in a high level how am IMS works for providing mult >> if you are interesting on this topic it's known on each bone http over ts you can look for it it's a really clever approach to be honest but it's how it works if you want to go deeper into the topics >> so if you need uh the power of layer 7 uh in this case you
will need a new element that is a waypoint The in this case the traffic is uh flow the traffic is around C tunnel C tunnel through H1 and H H1 send the traffic to way point and the waypoint to the C tunnel in the other node. So when you need a capabilities of LA 7 the waypoint will be will be applied the authorization policies and reg policies
in the level of waypoint and after that it forward the traffic to the pot or the services. So if we talk about observability too now we see Daniel we can we can see who is calling here who is calling who sorry and if we talk about policies now we have policies for the layer 4 and the policies for layer 7. >> Okay at this point have you
already work with ambient pets? Have you already evaluated it? >> Nobody here. Nice. the oh once the first one have stand up their hand now there are more volunteers so probably the first okay do you work with or this in general only three people I'm surprised >> well if you work with and this is just the last step of the introduction I promise the boring part will
end with this slide the if you plan to jump from from sidecar model to ambient model O has a clear announcement thereto ambient suggest to use API gateway. What API gateway stands for? Probably if you are not using or if you are using engineext not please remember that is deprecated engineext ingress controller. If you are using any other ingress controller of the market you could have discovered
that the amount of capabilities are a bit limited. It's an API that fit really well 10 years ago, but it hasn't grow with Kubernetes in terms of features, a lot of annotation, a lot of specific inte implementation configuration that is hard to maintain. API gateway comes to propose a solution or gate gateway sorry gateway API because it's an specification not an implementation. Gateway API try to solve
all the solution. First of all, splitting the responsibilities. The SRS are not the in charge of the endpoints that are exposed as well as the developers are not in charge of things like TLS or certificates or something. So it comes to solve why it's this relevant because it's another new API that you need to know to jump into ambient mess. Is it a problem? We will discover
it. But it's something that you need to be aware. It's not just transparently ch change from virtual services to uh ambient mess virtual services because the right approach is using gateway API. So from here have you seen the the good the bad and the ugly right the movie. Perfect. I think so. The the >> for now now is the the process that we pass for for moving
from cycle motor to ambient mess and the good things that is good. >> Yeah, good things >> obviously the good things with the cost reduction and now our bosses are happy because the product are much cheaper. So it's it's great. just to give you row numbers, root numbers because I don't have all the numbers in mind depending on the product but we have reduced the the not
the overall cost the proxy cost the the service mesh proxy cost around third between 30 and 60%. Because obviously the proxy is not something that you scale specifically for your workload. No default value 100 millores 1,000 pots with 100 millores consuming one millore. So the amount of the amount the the the reduction that we have measured is between 40 and 60% of the proxy resources allocation which
is huge in terms of h saving and also the other part the other good part is that we have reduced the drift between versions especially running on the same cluster because or across different clusters because of no this warload cannot be scale. Oh, but my API is running with two different Isto versions because the old ports are running with the previous version and the new port have
been injected with the new one. So there is a drift even there. Is it a problem? No, it's not. But it's hard to maintain because maybe the problem that you are facing is specific to the version that is not valid anymore in your setup. aligned with this cost reduction. Ah, another thing is that now our traffic rooting footprint and cost is predictable because it's not well depends
on the amount of the ports well depends on the amount of the proxies. We can estimate based on the nodes that we are running the prices that uh we will have in terms of service machine because we are going to have a single proxy in the cluster set tunnel will be there just proxying all the things. So yes, for sure. >> No worries. And the last the
last benefit is that first of all the metrics were really really high fragmented because you need to go proxy by proxy pulling the metrics that it's worth. It's super easy if you are using Prometheus because Prometheus does all the job on your behalf. But the size of your monitoring system in terms of a scrapping all the metrics is also bigger because it needs to create a pooling
job for each single proxy that you have. So we ended with Prometheus H scrappers with four CPUs. I don't remember the amount of gigabytes of RAM because obviously the fragmentation of the metrics was terrible amount the amount of proxies was terrible. So if we talk about the bad things. No. >> Now it's it's time to talk about the bat. >> This is the invited. Who is the
ugly? >> And I don't want I don't want to spoil who is the ugly but he is. >> This is this this is happen when I am the student, he's the teacher. So >> the picture will be nice. But >> so well okay. uh so in the past things obviously as I told you uh we have now the way points so now this is another element that
add more complexity in our infrastructure so the policies now are in level four in number in level seven in layer seven so it's mean that is more difficult that in the cycle model for example >> yeah there at this point is important it's something that we have we have said during the explanation Alonso has said that now set tunnel that component there the one the blue one
on each node asked as act as a note proxy let's call it in that way so that node is handling all the traffic but only at layer four I know that there are vendors that support layer 7 on set tunnel but that that's not part of the open source the open source at this moment is a stick is a attached to only layer 4. It's really nice,
but who only runs layer 4 applications on their clusters? Don't you have any single rest API or something? So, probably you would like to have layer something related with layer 7 at least metrics or something related. So using waypoint even is optional and luckily it's mandatory in a modern application because layer 7 is something that you cannot get rid of. It's part of our the main internet
are rest APIs and aligned with that the new complexity that we have introduced we have get rid of the operational complexity of updating proxies and we have increased the complexity of the topology that we are managing because currently in terms of not in terms of metrics that is part of the ugly is part of his uh picture. Now metrics are a bit more fragmented between layer four
and level and layer seven. But for us and for in in our case we have we are not high regulated environment but as we are processing payment methods and we are processing payment we have high authorization constraints in terms of uh controlling the traffic inside our clusters and with waypoints with sorry with proxies it's easy proxy A calling proxy B easy peasy it's authorized yes go ahead
no I add police come here because someone is h calling us. But now it's more complex. Why? Because the new authorization scheme depends on what are you authorizing. If you are going to authorize layer 7, you need to build an authorization policy for the service and attach it as a backend service to the service that you want to authorize. Why? Because it's the waypoint, the layer 7
component, the one who will apply and enforce the policy. So you need to apply it there. But if you only do that and someone overpass the waypoint for any reason, your world can receive the the traffic. So you need to apply to be even it's not mandatory. It's recommended to have both policies. One enforcing layer 7 at waypoint and one enforcing layer layer four at set tunnel
allowing only the traffic from the waypoint. So instead of doing a warload A, I want to call to warload B. Can I? Yes, you can. Nice. Thanks. You there. Here is your payload. Now it's no, can I call it? No, I'm not sure. Call to the other. It's more complicated. It's true that you can avoid the layer four rule, but in my arrogant opinion, you shouldn't because
it's a open door that you are keeping there. >> One thing more about the bad things. I think we have not slide but it's about the C tunnel now uh in the second model that's correct. So if we have a port that is crash doesn't matter all the networking is is work is working but now with sitan if the sitan files all the communication inside this node
are fall down so this is a problem in a high for our products in a high in in production environment uh well this is a could be a disaster uh we have to h work with this in mind this is a in this approach >> the good part is that set tunnel is a really resilient software. I mean we haven't had any crashing problem yet. >> H
with set tunnel >> and the ugly >> this is my >> he's like the wine years old man. >> He's like the wine improves over the years. >> And which are the ugly things? Set tunnel works as a transparent proxy. It cannot be excluded because if you have some weird X scenar scenarios for instance related with a random protocol that your database uses and you need to
exclude the port that's super easy with proxies because proxies were were just there modifying IP tables. But here the the proxying is doing transparently for the port. All the traffic is managed by the by this. What does it mean? For instance, are you using a strongly blocked or strong block policies in your So, uh if you are using some web hooks, some metric servers like KDA, not
because it's the project I maintain, trust me, it's just because it's another metric server. In those scenarios, the typ the the super easy part is excluded in the income port, the incoming the the inbound port of the web hook matrix server or whatever or whatever and that's all the policy has been automatically excluded. But in this case, you need to go directly to the policy and allow
the thanks time to speed up and allow explicitly saying okay ignore this port at policy level. So you need to be aware about that. Uh well if we talk about the observability point in this point uh we have to rewrite all our uh rules ours because uh it was different within the cical model. So now we have C tunnel for the loss for C tunnel and the
waypoint. So it's more complex uh in order to have a very good uh observability in this point we can lose if we compare >> telemetry is the next slide. H I don't know this is what happened when we don't prepare the the slide so continues you >> yeah now another ugly thing and ugly means that it's something just to take into account but not really really painful
because met um metrics have changed as you are introducing set tunnel there the metrics change a bit between the typical metric that we are a bit too because they change a bit the set tunnel metrics. Okay, I can live with that. The flow is a bit more complex because instead of being my API calling local host local host to another proxy and oh surprise working now there
are there are some magic proxy in there between my app the hbon pro the hbone protocol or or set tunnel through the the CNI injection hijacking the traffic and rooting the rooting there to other the flow it's in general more complex it's not rocket science I mean it's another step there, but it's something that you need to be there. And the most painful thing was modifying all
our dashboards, all our alerts to start to use the new the new metrics because things like blocking traffic at layer 4, it's a new metric that you have to take into account. It's not something that happens automatically. And the last point that we have I think >> so well we use a gateway API because it's needed but uh we have a different problem here because we use
gateway API for the part of the routine but we have to h use the AP of for applying the policies. So we have to work with this kind of appies and this is more complex for our teams. So >> and the the last ugly part that we had to face with do you use third manager for HTTP01 acme resolver or something? So okay third manager does support
it of course but you need to enable the experimental feature you need to enable to change your approach. Do you use Argo rollouts? Nice traffic. Do you use also with traffic management? is doing. >> That's correct. That's the point. Isto has native support, >> but is without CRD this >> aka Gateway API needs to use the gateway API plug-in for Argo rollout as well as for other
tooling because gateway API is a the new in my opinion the new standard for for ingress but it's not native in all the systems. You need to use the plugins already there. You don't need to code the plugins on your beh on your behalf, but you need to swap from the super easyto next layer whatever to installing the plugin, configure the plug-in. It's not a big deal,
but it's something that you need to deal with and that's the other ugly part that we have faced because we use search manager, we use argo rollout and we use another tooling that had to be updated. So at this point is ambient worth it? Of course it is but not without an small pain there. So in our specific case the savings has justified the things and also
the simplicity of the infrastructure for our developers because it's better making the things more complex for an small group than for the whole company. So the reduction of the operational complexity it's another topic and the pro tip there if you don't need levels layer seven don't use set tunnel v be a stick to or sorry don't don't use waypoint be a stick to set tunnel because your
life will be easier sadly that me is worldly recognized I guess that yes of course but it's not always possible Do you have any question? At this point, we are done. I can give you more details if you want, but more or less. Yeah. Do we have some microphone please? Thanks. Uh another ugly part I guess I read in the documentation the comparison between the sidecar and
the uh ambient and I see it mentions that multicluster installation is not supported. Is there any plan in the near future? >> I'm not part of maintainer panel so I cannot give any right answer but I know that they they are going to do it. >> Yeah. Uh I'm happen to be an STO maintainer who happened to work on STO multicluster in ambient. It's now in beta.
Uh you can try there some known uh corner cases. Uh but I I I I tested it myself. It works. >> Any other question? >> Perfect. Thanks. Appreciate it. >> Thank you so much. One question is about the is really feel that we we could have a good savings in related to cost optimization. For example, if I need a circuit breaker, I'm going to need a waypoint
proxy. Yeah. If I need um a detection, I'm going to need a waypoint >> But if any of my services I need a waypoint proxy, how many waypoint proxies I need? For example, how many pots minimum tree is that waypoint is a an enboy proxy configured by those lovely guys. Thanks for your effort. So it's a you can use as a s proxy or you can spin
up isolated proxies. You can choose which waypoint can handle it if you are using your own cluster for your own warload is an assert envoy. So the size of those waypoints is significantly smaller than the proxies because it's optimized there and it's shared. The good part of this approach is that resources are shared across multiple services. So can be better scale and also at as it's another
workload. You can autoscale it. You can deploy an HPA and say okay scale on demand as you need. I don't care during the night two instances during the peak of the middle of the day I don't know 200 doesn't matter because and as it's not a sidecar it can scale based on demand >> a way point is is like it's like no it's a deployment so it's
say you can scale on depending on your traffic on your >> but in that case the waypoint proxies are cluster scope or name space scope That's a really good point. They are as default namespace scope but you can specify cluster scope if you want. In our case as we are running h isolated things we run one waypoint per product. So we configure which name spaces should use
each waypoint and we can have four three x amount of way points >> understanding way points and waypoint for SSO waypoint for payments waypoint for product X. We can have n waypoint and the name spaces for the the SSO name spaces can use the SSO waypoint and I can choose which one I want to use. As default is name spaces but you can configure it globally cross
name space. >> Okay, thank you so much. It's a very good insight. >> Oh, I will do a lot of running today. Thank you very much. Uh maybe this is a stupid question but >> there isn't any stupid question. >> Yeah, but you mentioned that consumption ofite cars uh each 100 millpu why you just didn't reduce the requests because that in our experience it's not real. That's
really fair point and we did it for the most demanding warloads. But if you have 10,000 warload, it's real to to go workload by warload adjusting the commitment of the proxy. That was our problem. The problem is that we started doing okay, let's execute a load test to configure the proxy and we did it for two products. The other was like well the over cost doesn't justify
the effort that we are putting just deciding the right size for each warload. That's why we ended saying not 100 but 50 millord for every proxy and go ahead just because of operational cost. Go. Um, if you're working in an environment where you have very restrictive egress policies for every part, I noticed that it was no longer possible to just do that with side conotations. Is there
anything else you can do to keep that implicit deny on egress traffic while still allowing specific traffic? >> That's another really interesting point. I think that my colleague there has a better answer but I will say because I we don't have a a strict agric policies but based on docs it should be possible. >> Uh I cannot answer from the top of my head but uh they
still boot there. Uh reach out >> passing the bucket. >> Actually the documentation say that scikar egress policies are not it's not real security. It's fake security. So you should never use it actually from what I know and the recommendation is to use network policies. So that's the def facto standard based on the documentation. >> Well over there you can >> um do we have any other
questions? >> Oh well thank you guys. >> Thanks for attending.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32