Project Lightning Talk: Envoy Today: What’s New In Managing Cloud-Native And AI Workl... Karol Szwaj
About this talk
This talk celebrates the 10th anniversary of the Envoy project, which has evolved from a proxy for a lift company into a key component of the cloud-native ecosystem. The speaker discusses the core features of the Envoy proxy, highlighting its role as a high-performance Layer 7 proxy utilized across service meshes, API gateways, and edge proxies. They introduce Envoy Gateway, a Kubernetes-native control plane that simplifies the adoption of Envoy proxies by utilizing Gateway API, and the new Envoy AI Gateway designed to manage AI traffic with its unique requirements. This talk also covers the roadmap for Envoy, detailing improvements in inference management and the integration of new telemetry and rate-limiting features. The speaker emphasizes the significant growth in the adoption of Envoy technology, with impressive metrics in both Envoy Gateway and the core proxy.
Full transcript
Hello everyone. Thanks for attending this talk. So yeah, this year is something very special. It's 10 years of envoy project and it all started 10 years ago as a from proxy at a lift company and now it become a one of the key component in the cloud networking cloud native ecosystem. So let's take a look how the envoy ecosystem looks now. Uh what we have is our
core envoy proxy which is of course high performance L7 proxy and it was been foundation for all these years. It been in every service mesh API gateway and edge edge proxy. Uh another sub project is envoy gateway. It's our Kubernetes control plane. It's Kubernetes native uh gateway and it makes uh adoption of envoy proxies much easier and it runs uh using gateway API and manages many infrastructure
envoy proxies and the latest addition is a sub project named envoy AI gateway. It's a data and control plane extension. uh key key key functionality is to manage AML traffic uh using X X processor and internal CRDs for in AI inference. So why AI traffic is so different? Uh first of all streaming inference calls is long live connections not traditional HTTP request response. Everything is based on
the token. So single request can widely cost different amounts depending on the model size and also in aentic workflows we have like one triggers triggers many chains of tools across multiple backends. So that cause that we need to have perod observability. So not just the endpoint is slow but we need to find out which model how many tokens it uses and at what cost. So if we
take a look at the road map across all the envoy projects we see that we are moving many features from AI traffic management into the core envoy proxy like uh native protocol support MCP and agent to agent. We have also inference improvements there like request buffering and cross service failover and in another gateway we also added many extensibility points and filter integration that supports uh our use
case a gateway and na gateway AI gateway is moving into the API graduation we moved to from alpha to beta and we are planning to reach the our general availability milestones. Uh also it supports open responses API for open AI standard agents. We provision like token budget and rate limit policies and have a richer telemetry for LM backends. for envoy gateway we also are focusing now on
gateway API version 1.5 conformance and that also we have a lot of new load balancing features and if we take a look at the adoption I think it's very impressive that uh recently like gateway reach out 2.1 million hamchart pools over 30-day window like this exponential growth shows also like uh recent events with the ninx ingress migration We have also 5 million image pools in a gateway
and also on the core envoy proxy we have over 27,000 stars and we also reaching uh over 300 unique uh gateway contributors. So if you are keen to learn more about the what we do and what end users are doing with Envoy, I prepared a Luma calendar that you can scan and see. It contains all Envoy related talks. uh I recommend you all of them. Uh there
are many end user specific performance uh details on on how they they using envoy and also we have on Tuesday the celebration of the envoy 10 years uh anniversary and yeah you can meet uh maintainers and contributors in project pavilion. So show up by the envoy kiosk uh from Tuesday to Thursday at P11A and also there is an invitation on envoy slag where all the feature and
uh user discussion are happening. So yeah thank you. I'm Kosi software engineer at Kubernatic and an envoy gateway uh maintainer. Thank you. >> All right, thanks Carol. Good work. [applause]
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32