KubeCon + CloudNativeCon Europe

Cloud Native Theater | Istio Day: Panel: Horrors and Successes of Running Istio in Production

39:38 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This panel discussion focuses on the experiences of various end users running Istio in production environments. The panelists, who come from different organizations, share their journeys in adopting Istio, including initial motivations, challenges faced, and lessons learned over time. The conversation covers significant features and capabilities of Istio, such as multicluster support, ingress and egress traffic management, and the use of authorization policies. They also delve into some of the horror stories associated with upgrades and deployment issues, as well as the positive impacts Istio has had on their operations, including advanced traffic routing and improved observability. The panelists highlight the importance of a strong community and ongoing development in their decision to use Istio as their service mesh solution.

Full transcript

Okay, cool. Everyone can hear me. Um, so welcome to our lovely panel, the tales from the mesh, horrors and successes from running ISTO in production. Um, I'm going to give a quick intro then introduce the panelists. Um, I'm Yophikova. I was the Kubernetes 133 release lead. I work at Solo and have done a lot of work integrating K gateway and agent gateway with ISTTO. Um, I also

last KubeCon Amsterdam I gave I was the moderator for a panel of a lot of ISTO developers. So I'm really excited to be back to moderate this panel that's filled with ISTTO end users. So with that I'd like all of our panelists to introduce say your name, where you work and the year you started using ISTO. So uh let's start with Augustine here. >> Okay. Hi, I'm

Augustine. I belong to Kaisante. Kaisaban Techch is a company inside the Kaisaban group and Kaisaban Group is a group of companies which belongs which has the biggest banking in Spain called Kaisa the biggest insurance companies also in Spain and one of the major banks also in Portugal called BPI. So I'm the communities lead architect for Kaiser Bank. I'm managing from the infrastructure to all the marvelous DN

CMCF products running on top of and well we've been working with histo22 main at the beginning as an English controller but also starting in 2022 like for a zero trust as a zero trust product I mean nobody can go out outside of his name space without a uh without submitting ordering I a rule for getting out of the nice space there's we have a self-service portal where

you ask where you want to go and in what in in case of what you're ordering some rules are approved automatically other rules are are not and then we have some magic we have charts argo CD and all stuff and that get automatically applied in the class >> awesome thanks um maybe Alex you want to go next >> yeah hi everyone my name is Alex Williams I'm

a principal software engineer at Skyscanner, which is a travel meta uh website. Hopefully, you've heard of it before. Uh we've been using STO in production since 2019 when we did a massive overhaul of our Kubernetes estate, moving from one or two clusters, which were PETs into um a salesbased architecture of 24 production clusters and many other around the periphery and yeah, using it mainly for our um

ingress gateways, our east west gateways, and of course, service to service communication. Um, so yeah, hopefully share some some stories from the good and bad days of uh of ISTTO throughout the last six or seven years. >> Awesome. And uh who's next uh on the list? Vladimir, do you want to introduce yourself? Uh where you work and what year you started using? >> Thanks. I'm Vladimir. I'm

working in TomTom as a software engineer basically dealing with Kubernetes. So we have a big challenge with uh few hundreds of Kubernetes clusters that has to be consolidated because you know 100 teams having and running their own clusters inevitably end up in uh quite a mess. So um we started consolidating these things and now kind of comes to be one of the main components that we're looking

to solve problems like multicluster service mesh multi-reional support. So it would be one environment which is easy to use for engineers and migrate workloads, easy to use for us as uh platform engineers who can you know swap workloads, move them from one cluster to another uh do whatever we want. Uh with these two we started working um well TomTom other teams uh who provide gateway centralized gateway

started long before uh my teams. I'm working probably for two years with these two before that for a decade and more working with engineix and basically doing all the magic with roing with engineix and yeah I I do have a couple horror stories that I can uh perhaps share later. >> Exciting. Um and then last but not least Roland do you want to introduce yourself where you

work and uh when you first >> Yeah so my name is uh Robert Cole. I work as a system engineer at B. I'm part of the cloud networking team. Uh we are responsible for operating ISTTO among other things. Um we started looking at ISTO in 2019 and start evaluating it and we brought it in production in 2020. >> Well, our first question is uh what motivated you

to adopt uh ISTO and what benefits keep you using it today? So maybe we can uh start with you Roland this time and uh give us a little backstory about why what brought you to >> Yeah, sure. Um in uh somewhere in 2018 we started our cloud journey. We set up our first Kubernetes cluster and within half a year or so we ran out of IP space

um because these clusters are not easily uh cannot be changed on the fly. So we had to create new clusters to accommodate our growth. Um in order to facilitate crosscluster communication we looked at a service mesh and yeah was was the one that supported multicluster from the start. Um so that was the primary driver for uh service mesh and another important component ofto is the eus gateway.

Um and that also solved a lot of issues with dealing with egress access based on host names instead of IP addresses. And uh and the last last uh item is uh virtual machine integration. So so uh making sure that the mesh spans both our cloud and our on-rem uh workloads. >> That was a pretty quick integration. You went from you know one year in uh starting to

adopt it. Um Alex, how about you? What was your you know introduction to STO? Yeah. So as said we went from having uh two clusters per region um called cluster and cluster two in very very individual names and having services that are stuck in those clusters. So like my service.cluster 2 skyscanner and then realizing that as one cluster broke it took out half of the services of

our microser architecture. So we started looking at um the kind of cells based architecture and having single a clusters um so two clusters per AM across four regions. So 24 clusters and trying to then get to a point and we really had a need of making the clients um completely agnostic to where the the service they were trying to reach is deployed. Um so we have a

sort of my service.skyscanner.io name which is fully backed by service entries and virtual services and that was the the main driver I think was the traffic routting traffic management and outlier detection so that when a cluster goes boom or there's a bad deployment of a service we don't have to drain an entire region. Um, especially when you multiply that by 600 microservices and I think as John

said earlier now it's 60 million requests per minute like trying to get that and make every service work at all times isn't realistic. Um, and having this abstracted away from our service owners means they can get on and make a travel website and not have to worry about YAML. >> Very cool. Uh, yeah, it seems like both of you are using SEO with multicluster. So that's that's

great to hear. Um, how about you Vladimir? How are you using like what first brought you to ISTO and why what makes you uh you keep using it >> Well, in TomTom as I said was used it for quite a while. It was a part of a group product a commercial product and in our team we went maybe slightly different way. We started looking in the more

community supported uh software versions because you can end up in in in costs quite quite deep, right? And um I think the main motivator was a pragmatism because you want to have a tool which have a strong community. You have a tool which has a strong kind of development progress and uh Easter at at least two years ago was showing all those benefits that we wanted to

grasp. So yeah, pragmatism of course we could do things with engineics or with other tools but in the end you end up with halfbaked solutions which you are not able to develop uh further because there is also no community who would be motivated to develop complex features. >> I like that pragmatism why adopt um >> Yeah. Well, obviously we use because it's one of the strongest product

for managing communication inside the cluster. We have almost 100 clusters and 30,000 cores running in those clusters. So we need some strong product for managing all these. No. So first thing in controller one of the main features we need a zero trust system. So that's the second reason why we use and then we need also for routting traffic through ingress rule ingress nodes eress nodes. So it

was easy for us to manage all those rules with this. So that that is >> very cool. I think uh uh the next question I have is uh kind of a reflection. So there are other service mesh options out there. So when you were first starting your STO journey, why uh did you ultimately choose over other options? And I think some of you touched on it like

Vladimir, you mentioned that the STO community was very strong. You like the road map. Um does anyone want to tell us your reason for using STO? >> Yeah, I can start. Um I think in the early days back in 2019 when we were looking at this salesbased architecture we actually dabbled with linkerd um maybe linker d1 looking at colleagues um and it wasn't as mature at the

time um for the things we wanted like outlier detection and what the power you can get from the destination rule CRD object I think for me was the the real draw towards this and having that abstraction and failover at the um the sidecar level completely like transparent to the service owners. I think that was the bit that was um the draw. Um and linkad then to go

on to be a CNCF sponsored product which is still there but um there were other options that available from vendors but having to avoid vendor lock in or maybe things that were slightly less mature and not open source because ultimately you're going to want to contribute upstream. You're going to want to have a look at the source code and backing that behind a vendor wasn't really an

option for us at the time. >> Anyone else uh want to give their story of why they they chose STO and stuck with it? Yeah, a little bit like the same like uh Alex said, we at the time 2019 there was not there was there were no real options beside Isto. There was linkad but linkad didn't have multicluster support. So there was only one option left for

us because multicluster was like the go the the must-h have feature. So yeah >> multicluster nice >> multicluster. Yeah. >> Um anyone else? Augustine did you look at other >> four seven capabilities routting capabilities and and you can do whatever you want with this team if not then you have envoy filters no so I mean this >> definitely a lot of flexibility um so we touched on

this briefly but I think it's still um useful to you know describe how you deploy today so now has a lot of different modes sidecar mode ambient mode ingress only eress gateways um multicluster as you mentioned Um so can you give us a little bit um you know of a short overview of how you deploy today maybe not how you started but how does yourto deployment look

today in your clusters um let's start with uh >> yeah now we're just using sidecar but we starting last year to do to do some PC's and for this year this will be something in production for multicluster for us also now it's super important multicluster we have to ensure resiliency not only in two in the two data centers in Barcelona we have to have more resilency in

other regions. So this year we moving to multicluster for is more important now than ambience capabilities. >> William do you want to add anything >> sir? >> Well we didn't go for sidecar because it was obvious for a DK though everyone was using sidecar and everyone was unhappy and we were looking obviously on uh at ambient mesh as a as a capability that would allow us to

get rid of this complexity. Um so yeah mainly we use it in ambient mode. We are still to get to the m multicluster setup. Um right now it's just a dedicated deployed in every cluster. Yeah there is some naming convention for gateways for endpoints and stuff like that. But uh the real hardcore stuff with the multicluster service mesh is still on the road and we are hoping

in the next quarter to actually crack that problem. >> Very cool. Um, and I know Alex, you you've been talking a lot about multicluster, but you also mentioned VMs, I think. >> Um, yes. So, in in our setup, um, I can't go into the full details of, uh, of what John's talk was about earlier, but it's, um, really important for us to have the outline detection at

the cluster level. So we have over the years had a few iterations of trying to have a single IP address that represents a cluster of pods rather than having true flat network multicluster um with multi- primary. So we are right now inside car mode with um an egress gateway to handle outbound traffic to neighboring clusters. Um, but I think if we weren't here this week, we'd also

be rolling out multicluster in ambient mode, which is pretty exciting, and we've been working closely with the community and with Solo to actually get that capability out there. Um, and it'll be saving a lot of money. Side cars are are great. Um, and there's other challenges with ambient which we can maybe go into later about observability and and even just moving an existing setup to it, but

it's um going to save a lot of money of CPUs of those 30,000 containers that are running. >> Yeah, definitely. And Roland, uh, how's your setup look today? >> Uh, it looks very similar to the setup that we started with uh, in 2020. So we have a um multicluster flat networking setup is with sidec cars of course. Uh the only major change that was done over the

years was we moved from uh remote primary setup like we had like a central control plane cluster and then a lots of remote clusters. We we changed that to a multi- primary setup. Um, and yeah, today I'm uh working on getting off side cars and into ambient mode, but that's a work in progress. >> Exciting. Okay. Well, um, our next question is, uh, the horror stories you've

had with ISTO. So, the reason everyone's here today, um, so if you think back through your ISTO career, uh, do you have any memorable horror stories that happened to you and how did you work to resolve it? Um, let's start with um, let's see. Any volunteers? >> Sure. >> You sure? Go for it. >> I'll uh, cast your mind back. So maybe mine's more of an operational

night nightmare rather than an actual instant, but there's been plenty of those for other reasons. Um, so if you don't know, way back in STO 1.4, there was a component called mixer um, that handled centralized telemetry. Um, the move to STO 1.5 removed that component and flattened things down into STOD. Um and there was a massive change for for the better, but those centralized uh telemetry metrics

that would now by default be emitted from sidecars would increase the um cardality of our metrics by 35,000. So I think our observability provider was happy looking at the the dollar signs, but we weren't too happy with it. Um and I think that the pain for us, it was maybe three years of running a very outofdate working like for a long time to even get to 1.5.

Um and in the end we went from directly for STO 1.4 to I think 116 in like one massive jump and >> big jump. >> By doing that we actually created an entire new set of 24 clusters. We did Valero backups of our production clusters. We did traffic drain with an internal component that we call the drain operator that modifies load balancer weights to drain clusters and

then re-edit those weights to send production traffic to the cloned Valero restored cluster in STO 116. That project probably took an entire team of people 12 months to upgrade a minor version of of so obviously from a company point of view like what are you doing? Why are you spending all this time doing this thing? But where we are at now means that we have um a

slightly novel setup of using open telemetry collectors and a span metrics pro processor to actually count spans and produce telemetry. And it works really well actually going into the future with ambient and waypoint proxies. We can still send 100% of those spans to those collectors, count them into metrics, and then optionally throw away the spans if we wanted to. >> Amazing. Um were your upgrades after that

more smooth hopefully? Yes. Yes, they were. Um, considering the jump we are at now, I think we're at 129 >> front row. No, more than I do. But yeah, 129. So, it's we've moved on. We didn't get stuck and but that first move um was painful from going Yeah. four years on one version. >> Yeah, I can imagine. Um, >> who wants to go next? August. I

can talk also about another painful upgrade for Fanas was from 118 to 124. It happened the same. You do that grade, you find a lot of bugs. You see bugs get resolved every week in new minor versions. But I'm not going to talk about that. I'm going to talk one about the last week >> last week. Oh no. >> Yeah. Yeah. Yeah. >> Yeah. Yeah. Barcia, I

will still look for some plot as well. This was I mean it's related with you, you know. It's not 100% but as I said before everything all the application security is managed by the internal helm where we map all the requirements to a specific objects network policy objects. So this helm that we use for every application in the in the bank for in each name space when

we we create a new version and we need to migrate to new to this new version. The problem the migration we didn't migrate to the latest one we to we migrated to the previous version and the problem is that we didn't have aligned the network policy with the objects such as service entries authorization policies and all these objects. So what happened in our onrem clusters which use

OBN and multus the problem with network policies is that some type of network policies they they have an explosion problem. I mean they they grow quadratically. So if you create lot of network policies the problem is that if you don't take you don't have care you can't create million of entries in your in your IP tables in the node IP table. So what happened as the script

was running to to all one 10,000 repos upgrading to this new version and create I started creating all these network policies all the everything start to die the nodes start to die master nodes start to die API server stop degrading so it was it was fun no because the problem is that we didn't have a roll back script so we we lost a precious time developing this

back backup script And and then the other problem was the the GitHubs approach as we were working with one,000 repos. It's not it's not very fast to upgrade 1,000 repos or to to execute a roll back against one 1,000 repos. No. So lesson learned, you have to have a although you work in a GitHub approach, you have to have a roll back scripts. And second lesson learned

is that you have to have a panic button. I like switch down Argo CD and start deleting things manually. Not this. >> That sounds very painful. Um, who wants to go next? Uh, Vlad. >> let me try one. Well, besides uh occasional shooting yourself in the foot, I think one of the problems stand out the most for us is um being able to understand whether to uh

uh got everything it needs to do the proper routing. For example, we had a couple problems with certificates which might not be ready on time. And uh the most maybe bizarre things that happened was uh we had a development environment set up and running. everything working smoothly. there is some roing flying around and uh all of a sudden next day when we recreate this development environment in

the new namespace on the same cluster uh it doesn't work and we sp we spent probably three hours trying to understand the problem and in the end we found out that a straight certificate with the same host name in it which wasn't deleted uh in the old name space was still used by Easter instead of the new certificate which also had the same uh host name. So

it turned out there was no any message about this certificate. the wrong certificate is used and I only uh um saw it when I made a curl request with verbose output and saw in the uh certificate information a subject alternate name which does not belong to the certificate which I was expecting to see and uh this kind of things around uh resources which is not really responsible

for but depends on visibility in the uh collaboration of these kind of components is probably the most invisible thing for us and just recently we also were thinking how we can improve visibility for our customers who use platform to understand that uh the routing is not working because of some problem that has to be kind of seen through very cumbersome ways and customers can't really understand those

things so we need to bring them to to the surface and this is where I believe where could do a little bit better by uh finding those edge cases where uh incomplete state of deployment has to be reported somehow to customer if is depends on on on those items which are actually mentioned in in the chain of resources. >> Do you use like uh any STO debugging

tools like um I I feel likel would not catch that but um like what's your usual flow to debug a similar scenario >> to debug? Well, h the problem was not let's say in a in in dimension where you would go to some network debugging tools logging to node and try to understand things. It was clearly something related to misconfiguration. Yeah. and basically trying to reconcile all

the Kubernetes manifest that there are all the proper things in place rolling back to the revision which was absolutely guaranteed working in in the previous day and trying that one comparing the output that was kind of a solution but uh as I say with multicluster service mesh there are a few things that can go wrong and we will be paying a lot of attention in to to

kind of having all the tools which show us the picture. Thanks. And then uh Roland, would you like to finish us off of the STO horror stories? >> Yeah. Yeah, sure. Um yeah, so we're running for about six year is production and overall it's been really stable. Of course, we did have had some weird issues every now and then, but nothing that really took down production until

about a month ago. So I guess maybe maybe this this panel session jinxed it for us. So what happened uh we did a um disaster recovery test in our data center and yeah we got a lot of disaster but we didn't get lots of the recoyree unfortunately. So what happened is what um we enabled DNS proxying in the data center uh earlier this year and it turns

out that there's a DNS >> there's an issue there's always a DNS issue right so there's an seems to be an issue in in envoy DNS proxy uh functionality that when you disconnect it from the network but you still send it calls it's going to queue them up and it's got and it keeps retrying buying them until forever basically. So we had a lot of we had

a lot of side cars with DNS proxies that could not connect to the DNS servers, but we had still a lot of uh software running on these machines that did DNS calls. So well, if if all those calls queue up in the in the DNS proxy for about an hour or so and then all of a sudden you reconnect the network, what do you think is going

to happen? everything is going to flood your DNS servers and then all the rest of your uh of the infrastructure that wasn't affected is also going to be affected because your DNS servers are down. So that was quite a a problem. Uh it it took quite some time to recover also also because if DNS doesn't work how how do you lock into your machines turned out to

be tricky. >> So how did you resolve it? Uh well I did and my colleagues did but uh yeah it took a lot of effort. Uh we had to we had to turn off proxies uh disabled DNS caching uh DNS proxying for now until we actually redefined the the root cause or some kind of configurability to make sure that it doesn't keep retrying forever if in case

of connectivity problems. So this was the the worst that we had in six years. Oh, >> all very bloody. Uh, >> so let's end on a more positive note before we open it up for audience questions. So, uh, you know, forget about the bad. Now, let's look at the wins. Um, so what's the most valuable thing that ITto has helped you with? Like, uh, in production. Uh,

I know we we talked about multicluster a lot. We talked about ambient. Um, but if you think back and you can like name only one thing, what's that one thing that has helped you with in your journey? Um maybe you can start with flamer. Yeah, >> I think for us it's uh level seven traffic management capabilities. level seven like any specific policy >> uh request rate limiting

circuit >> I mean everything where you need to let's say in in in previously running let's say engineix you would have to manage engineix configuration uh those capabilities exposed in kubernetes manifest and resources provided by is much more kind of convenient and useful for customers >> so if you're using level seven with ambient are you using the waypoint then >> yes though as As I said, we

still need to get to the point where those things are really heavily relied on because uh most of the workloads uh at TomTom are served still by the ETO in a glue uh product and uh on our Kubernetes as a platform. Uh we are gaining traction with customers. Yes, we have some situations where we have like 200,000s requests per minute and things like Um getting to a

point where some customers would have to uh carefully manage their workloads is kind of an important thing for us and having these capabilities well defined it uh and worked out in E2 is the most important >> It's a great one. Uh anyone else want to go next? Yeah, for us it's uh for us it's the uh the multicluster routing and the layer 7 authorization policies to make

sure that only specific services can't talk to other specific services. Um that was the biggest driver to adoptto and it is still providing the best value for today. So if you're using L7 authorization policy, what kind of policies are you writing? Like >> so for example, you allow a name space and then you allow on specific uh paths or specific methods. >> Uh but also uh for

users accessing services, we use um JWT. So we have policies that look at claims groups in particular that allowed them to access a certain pass on a server. Yeah. >> Very cool. Um who's next? Good. >> Yeah, I think I'd have to agree like traffic rooting is just so good. MTLS, zero trust by default, but to maybe take a different spin on it, the extensibility that you

have. So, a recent example in the last few weeks is that we wanted to add um an external processor service um to every public inbound request to Skyscanner. So, this is like a huge amount of traffic and it would just be a thing that'd be very very hard to do without. But instead like the extensibility of using an envoy filter CRD object to actually then modify the

STTO generated envoy configuration. In this case, it was maybe 40 lines of YAML to use a fairly new feature of Envoy, which is the external processor that points to a small Golang service. And it's doing header manipulation and session management, which is managed by a completely different team, but our platform team with one CRD deployment rolled out through the channels mechanism that we have means that that

that like identity team at Skyscanner can now inject headers backed by a service and store those session identities. And I I don't know how you would do it otherwise like maybe lambda edge which if you have lots of money or maybe an engineext gateway that has to sit in between and having another level of expense but as I said it's 40 lines of YAML we did it

in one week and it was pretty much like a support request over Slack rather than months of design. >> Oh that's amazing. Um and then >> that you can do whatever you want the L7 capabilities and on top of that I will say also observability. you know, you have ton of metrics of ingress, air res I mean, so that's also really nice the observability of the product.

>> Amazing. Um, well, now let's open it up for questions from the audience. Um, so if you'd like to come up and ask a question, um, we have a a mic. Great. Um, any questions can be about the horror stories, how you know. Oh, good. It's it's not necessarily a question, but I could talk about a horror story that we had that would also tie in pretty

nicely. Uh so, uh we also of course use STTO. We use also all of the features. Um especially for us, the fact that you can delegate certain parts to developers, but you don't have to require them to become top-notch is engineers or YAML wranglers. You can just say, well, here's your chart, couple of flags, there you go. That's really nice. One of the things that we do

is we abstracted traffic and L7 configuration in just a set of flags. So you can say I want public traffic and I might want internal traffic but the public traffic is only allowed to do safe methods. So that's get and head and the internal traffic is allowed to do other things. So simple they don't need to know about the policies but they can just turn it on

or off which is easy mode for developers. Now for external traffic we also want to do authentication and authorization but we had a problem and it was a couple of years ago we were migrating from one authentication system to a different one which means from non-standard tokens to JWTs JTs are easy you guys are using them you can just reference them you reference the claims easy mode

but the non-standard ones it's not supported so what we did and I think it's deprecated uh we had a Wom plugin and we did that because we wanted this to be as fast as possible. So it runs like in a WOM runtime VM inside envoy inside your STTO gateway um which is super fast. However, uh I didn't want to write it in C because I don't trust

myself writing C. So I wrote it in Go. But Go is a managed language which a garbage collector which is not supported in was so you run tiny Go which is meant for microcontrollers but it also kind of runs in Wom and I thought I did it right but I did it wrong. So our gateways kept crashing every now and then and you looked at the graphics

and you see this or the metrics and you see the sawtooth pattern where you're like wait a second if the memory goes up consistently and then suddenly goes back to zero because it crashed it means that there's a memory leak somewhere. So I had a memory leak in go in was in envoy in myto gateway uh which >> has been fun to debug. >> It's fun to

debug. Yeah because you at first like wait are we being attacked? this is like slow loris but for isto what what's happening and it took quite a while to figure that one out but we did find it out and of course as you do you rewrite it in rust uh and then you don't have that problem um but yeah it was definitely an an interesting thing and

you would also have because these are of course load balanced and not everyone would hit the same gateway at the same time so some requests for some customers would sometimes fail but with completely different parameters so there was no way to correlate any of this until of course we figured out that pattern smells like a memory leak, even though you would not look for that when you're

thinking about Go. Yeah, that was definitely an interesting problem. Uh, and something that took quite a while to fix and well find in the first place. >> Great horror story. Anyone else have any horror stories or or questions about >> uh for our end users? >> Maybe just a comment on that. Maybe >> if not, I Yeah, go ahead. >> Sorry. I have some backup ones that

don't. >> Uh, yeah. On the topic of shooting yourself in the foot, um, do you have ISO APIs in mind like CRDs or configs that you look at and like you use it and it causes a problem and you read documentation and says, "Oh yeah, well works as it like supposed to, but I really really don't like this API. It's very confusing. I don't understand it and

it causes problems for me." Do you have examples of such? Yeah, I I think authorization policy and multiple authorization policies when you put them together. So if you have one default policy, customers not always understand that they are kind of compounding. They are not or or and uh kind of giving even giving documentation is not always helping because you know you can read it once, twice, three

times but some people don't really get it because it's kind of a complex concept. >> Cool example. Yes. Anybody else? >> Maybe the >> envoy filters. >> The envoy filters. >> Which one? >> Envoy filters. They're like a black box. >> You can do anything. >> Yes, you can do it. Anything. Uh I imagine it was created to shoot yourself in the foot. >> Yeah. Are we

on the topic of uh of things that we dislike? Uh we have a drain operator controller which modifies the uh locality load balancing weights. So we can have a custom CRD that says drain cluster X and over the course of 20 minutes or something it will keep modifying those weights that need to add up to exactly 100 or everything blows up. Um you can have it at

the config map level of the STOD or you can have it per destination rule. So per destination rule sounds like a great option but now we have 600 services with 600 destination rules that all need to be modified across 24 clusters and it becomes a bit of a pain. So yeah, having having a way to have both failover and uh the distributed like load balancer locality priorities

would be would be great and yeah having to write custom tooling to to remove traffic is it's a bit of a pain but touch something would it hasn't broken so maybe I shouldn't have said it out loud on stage. >> We will see tomorrow. Uh one follow-up question. Did you file a bug for it? Yes, sure. >> Yeah. Don't check. >> Any other questions? Oh. Oh, we

have one here. >> So, anyone ever been bitten by TLS origination? So if you do egress filtering and you're like if your pod wants to connect to someone on the outside it you have to configure your service to not use TLS and it goes to a service entry and a destination rule and then isto had to add TLS before it goes to the outside. Well, you can

make it worse by using HTTP2 and gRPC which expect it to be TLS, but it can't be TLS. But you will actually add TLS between the side cars and then the last cycle will strip the TLS off of it. And if you don't align all of them, right? So that's four places. Then your TLS origination will essentially break you and have you running in circles for a

week before you find that. >> Yes, we've had something similar. I I do have an interesting story about it's TLS related but not not this but uh STO 129 it comes with a a small uh permission change. It's going to it it adds config map reader um permissions. If you're a multicluster, make sure that you roll that out before you do all the other stuff. Otherwise, the

STTO instance cannot connect and read the proper configuration from the remote clusters and you get incomplete spiffy addresses and it gives you lots of TLS errors. We call this in staging for fortunately it's going to be fixed in the next release, I think. So, be careful if you're going to 129.1. good public service announce announcement. Um I think on the TLS uh side, did anyone else want

to add to that? >> Yeah, I can maybe layer on. We do some some fun magic uh for stateful services that live in a different AWS account where we wrap TLS up in TLS. So the underlying protocol might be for a Postgress database or for Reddius and we tunnel that out through a VPC endpoint into a different Amazon account with a a cluster that only runs an

ingress gateway that unwraps the first level of TNS and then carries on the maybe or maybe not TLS encrypted underlying data thing. And when you have weird hang-ups of the outer connection, you have developers saying my database connection dropped. It's your fault. and you kind of have to say yes, it is our fault, but sorry, like someone tripped over a cable in Amazon. Um, so yeah, that

was a headache to get setting up. And again, one of those things that it just kind of works and hopefully we don't add any more use cases to it. >> Have I ever >> Yeah, debugging through Wireshark. Um, absolutely not. No. >> Um, I think we're running out of time. uh Antonio shows like that it's time for me to stop talking. >> Uh thank you everyone for

joining. Uh thank you our panelist and thank you Nina for uh moderating this discussion. that wraps up the ISO day. Thank you.