Cloud Native Theater | EnvoyCon: From ingress-nginx to Envoy Gateway at Zapier: The... Kalen Wessel
About this talk
This talk covers the migration process undertaken by Zapier from Ingress Engine X to Envoy Gateway. The speaker, a senior site reliability engineer, explains the reasons for the migration, including outdated technologies and accumulated technical debt. He details the approach taken, such as a full ingress audit, creating baselines for latency, and the use of weighted DNS for incremental traffic shifts. The speaker discusses unexpected challenges faced during the migration, such as handling large payloads and connection pooling issues, while emphasizing the importance of observability and metrics. He concludes by highlighting the benefits realized post-migration, including improved stability and a unified view of their edge infrastructure.
Full transcript
Hi everyone, my name is Kayn Wessle and I'm a senior site reliability engineer at Zapier. Today I'm going to be walking you through our migration from Ingress Engine X to Envoy Gateway. I'll touch on why we made the move, how we approached it, some of the unexpected issues we ran into, and the new features we're capitalizing on. So why we need to move? Well, for starters, the
writing was on the wall. I'm sure the EngineX end of life announcement is why many of you are here in the audience today. Along with the ingress API being marked as frozen for quite some time and the gateway API gaining traction, ingress's days just felt numbered. But not only that, years of tech debt. Like many companies, we didn't wake up and choose chaos. We accumulated it. Over
the years, we had built up a collection of differing load balances. uh various ingress ingress enginex controllers annotations started to vary by teams environments behaved inconsistently staging and production slowly diverged. This naturally created pain points. The fragmentation introduced drift which occasionally even resulted in the odd outage. And sure, we could have simply migrated to a different ingress controller, but this felt like the right opportunity uh and
right moment to rebuild the ingress layer the way we always wanted it. And this time with the Kubernetes API gateway at the center. So what do we all get with the Kubernetes API gateway? Well, for starters, a clear separation of responsibilities. Platform teams manage their gateways and gateway classes. application teams handle their HP routes and backend traffic policies. Strongly typed CRDs uh now let us finally drop
those delicate annotation strings which I'm sure have all bit us one time or another. And with proper schemas in place, CI linting can catch those silly syntax errors a lot earlier. Gateway API also gives us richer tech primitives uh sorry traffic primitives out of the box. uh traffic splitting, headerbased canary matching and more and all again without relying on those controller specific annotations. At this point we
were sold on Kubernetes API gateway uh but which implementation did we want to go with? Last March we evaluated uh multiple implementations such as stto glue and a few others. Envoy gateway uh stood out as the winner for us when compared to other options. its Kubernetes API alignment uh was the strongest. You have to remember this was already a year ago, so I'm sure some of the
others are catching up. It also provided clear extensibility path allowing you to build some of those more advanced use cases. The consistent release cadence and active community signal to us a really healthy and just trustworthy project, but most importantly, it's everything you need from day one without the feature you care about being stuck behind an enterprise license. Now, prior to any migrations, housekeeping was in order. We
took inventory of every ingress resource. Most of them were managed by GitOps, but uh we did still find a few handcrafted resources here and there. Uh this is also when we decided to roll out a kerno policy to make sure any new ingresses that might come up audit would get flagged. Then we establish baselines 4x 5x latency and all over a meaningful time window. Because if you
don't know what normal looks like today, how will you know if you broke something? Next, we conducted a full ingress audit mapping annotations to their envoy equivalents to ensure expected functionality was preserved post migration. You'd actually be surprised how many annotations you can probably just drop algether. Once we identified all the ingressives, we created extensive tracking sheets. Uh this easily allowed stakeholders to stay organized and up
to date uh on the status of the migration project. When it came time to finally start migrating, uh the services we decided to cut over first uh were some of the highest traffic ones. Uh and that might seem counterintuitive, but these have the most eyes on them. So that meant faster detection of edge cases and discovering most of the pain early on. Finding those sharp edges up
front uh directly informed our design decisions and made the rest of the roll out much smoother. This plan gave us wins to focus on. For example, adding header based canary support to our monolith uh deployment demonstrated early uh value and uh it strengthened our testing strategy. It also helped us maintain momentum for a project of this size because it's easy to uh yeah to slow down. Now
our migration had to be incremental, observable, and reversible. Each batch of services started in staging followed by a bacon time. And this could range from a few days to a few weeks all depending on the service tier. Once the service team was happy with what they saw in staging, uh we would get the green light to move on and proceed with production. Since in our case, every
service lives on its own domain, we used weighted DNS to for a more gradual traffic shift. Once you enable support for HP routes within external DNS, the new record would automatically register in DNS just like it did for ingress. Traffic shifting then became a matter of simply adjusting DNS weights that allowed us to move traffic incrementally 10% 50% and finally 100%. At each step we compared those
envoy metrics against the engineext baselines I was telling you about earlier. If we noticed any deviations, we could simply reweight DNS offered the control we needed to migrate safely and confidently. So, we've got our strategy, our baselines, and our rollout plan. But I think we can mostly agree almost no migration ever goes according to plan. These next few slides are the unexpected surprises we didn't see coming.
the stuff that doesn't show up in planning docs and can require sometimes days of troubleshooting. For over a week, we chased 503s on our busiest web hook service. The error rate wasn't massive, but it was consistently above those enginex baselines. At first, we went after the usual suspects. We tuned timeouts. We adjusted TCP keeps. We're running load tests. Nothing was reproducing in production. It wasn't until we
dug deeper into the envoy metrics uh that the story started to come together. We already knew retries were happening. That wasn't unexpected. But then we noticed this other metric retry or shadow abandoned. That was the missing piece for us. Retries were being abandoned before they even started. That's not a timeout issue. That's not an upstream connect failure. That's envoy straight up saying I can't even attempt this.
And the culprit buffer limits. Some of our web hook payloads were well over 1 megabyte. Envoy's default buffer is 1 megabyte. During a retry, Envoy needs to fully buffer the entire request body. And if it doesn't fit, the retry simply gets dropped. Once we understood that, the fix was straightforward. Increase the buffer limit inside the client traffic policy. Not to be confused with the buffer uh limit
that you can set inside the backend traffic policy. Those handle returning 413s for large payloads. After that change was applied, abandoned retries dropped to zero. Our metrics were back in line with our baselines. We're back in business. Reconciliation wos. Uh, our envoy gateway controller pods in our dev cluster started ooming. Our dashboard showed over a 100,000 reconciliation events in a cluster that should have been mostly idle.
After digging deeper and tracing events, we realized the reconcile loop wasn't random at all. Our vault operator was refreshing the vault clls secret in every nameace. Envoy gateway happens to watch TLS secrets. So every single update triggered a full across all monitor namespaces. The fix in this case was relatively simple. We dropped the CA secret distribution um because we're not using that functionality. It was some leftover
crud. But really the takeaway here is be aware of the operators running in your clusters. If something's churning resources your gateway is watching, your control plane will absolutely pay the price. Post migration, we saw some elevated HP latency for a few services. Not too many, just a couple. Nothing was obviously wrong, but our baseline showed a clear increase in response latency. Thank goodness for those back end
or those uh baselines. In our case, EngineX worker model inadvertently protected us by fragmenting connections across multiple worker pools. Envoy removed that accidental protection. It uses shared pools and aggressively reuses connections which is great for efficiency, but these are synchronous services. Requests on a single connection are effectively serialized. Envoy concentrated onto fewer upstream connections. So when one long request landed on a reuse connection, the short request
behind it had to wait, resulting in response time, an increase in response times. We fixed it by forcing one request per connection in the backend traffic policy circuit breaker settings. We traded some efficiency for isolation, but latency stabilized immediately after synchronous services, make sure your connection settings match your concurrency model. Uh here are a few differing defaults that caught us off guard. Ingress engine X will forward
encoded paths unchanged by default. Envoy Gateway on the other hand will unescape encoded slashes. So things like percent 2F turn into a normal forward slash before the request is passed onto the application. In our case, Argo CD and Vault rely on those encoded uh values for some of their paths. When Envoy normalized it, the path changed resulting in a 404. The fix was to explicitly set escape/action
to keep unchanged. Uh now there are some security considerations with disabling normalization. So that's something to evaluate carefully. In our case, this only affected internal services behind our VPN. So we were comfortable with the trade-off. Another default behavior we ran into with ingress engineext is it'll drop headers containing underscores. While Envoy core allows headers containing underscores by default, but envoy gateway rejects the request entirely. During our
migration, this led to 400s until we made the policy explicit by updating the client traffic policy header with underscores action drop header. This brought us back in line with ingress engine X and we were good to go. Uh the last behavior I want to share is how ingress engine X retries certain upstream failures automatically. Envoy Gateway won't unless you explicitly configure retry triggers. So to match the
original EngineX behavior, we include the following retry block within the backend traffic policy. If you migrate without matching this behavior, you may see an increase in 503s on some of your maybe unstable services. Now most of us in the room can agree that it can be difficult to stay up to date with system upgrades. What we really care about though are upgrades that resolve real operational pain.
In Envoy Gateway 1.5.0, our control plane memory hovered around 80%. And we hit frequent out of memory issues. After upgrading to simply 1.5.3, memory usage dropped to under 20% and we haven't seen a like an oom kill since. By seeing such drastic performance improvements, it really helped solidify that we made the right decision going with Envoy Gateway. So stability was step one. Step two was control. We
built a custom OZ service called Wavefinder. It integrates via the external OC endpoint communicating over gRPC and it will validate JWTs and enriches requests with customer context all while introducing only submillisecond overhead. Now routing rate limits and region selection can depend on verified identities all enforced before the request even reaches the application. And because the context flows into our observability pipelines, instance look different now too. We
can immediately see which accounts and customers are affected and prioritize accordingly. This is just one of the small examples of how we've started building on top of envoy gateway. I'll leave you with this. For the first time, we have a single unified view of our entire edge. No bespoke dashboards, no stitching metrics together, just one consistent picture. We migrated roughly 500 production ingresses in under 6 months
that handle 60,000 requests per second. And I think maybe only one incident came out of it. But the real scale, it's the consistency framework. That consistency reduces risk. It improves developer experience and it gives us the foundation we can confidently build on. Thanks for listening to my talk. Uh if you have any questions, feel free to grab me in the hall. And yeah, and I also have
a um a more verbose uh blog post you can check out. >> See if this Oh, the microphone's on. Thank you so much, Kayen. I'm so excited to see so many people coming to hear you speak. >> We do actually have a couple of minutes. If someone has a question for Kalin, you have to put your hand up and wave. David, keep on going. We're going. >>
Hi, thank you for the presentation. Uh I would like to ask why you chose Go uh for the off Z instead of Louis script which is embedded. Uh mainly just because the SRRES who wrote that Wavefinder application just preferred Go. That's that's why there was no real rhyme or reason. >> Okay. Thank you. >> Any more questions for Kalin? Sorry, I'm bumping into people here. Oh, there
we got hand. Fantastic. Yeah, thanks for the talk. Um what was the most interesting thing you found in the custom annotation configuration extensions? So, we all know you can do bad stuff there. But was what was the most ridiculous? >> Oh, man. There's the most ridiculous. >> Did something make you laugh? >> I'm not I'm not even sure. I I I have to think about it. I'd
have to get back to you on that one. And I can't think off the top of my head at the moment.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32