10 Years of Cilium: Connecting, Securing, and Simplifying the... Bill M, Paul A, Marcelo M & Neha A
About this talk
This talk celebrates the ten-year anniversary of Cilium, a Cloud Native Networking Interface (CNI) that provides observability and security through eBPF technology. The speaker highlights Cilium's growth from its inception to now becoming a standard for cloud-native security and networking, with over 50,000 stars on GitHub and a community of more than a thousand contributors. Recent developments include advanced capabilities with Gateway API, multi-cluster connectivity, and enhanced support for IPv6. Notably, the integration with Zunnel enables mutual TLS for secure data communication without altering application behavior. The talk also covers the importance of network policies and monitoring in managing the complexities of modern infrastructure, particularly in large-scale environments. Various community contributions and use cases illustrate the transformative impact Cilium has had across industries.
Full transcript
Okay, Syllium maintainers track session. Thanks everyone for coming today 10 years. So if you aren't sure where you are right now, this is silly maintainers track. Psyllium is a CNI. It does observability. It does security with Tetreon and all of that is powered by EVPF. But before we kind of dive into the details of this session, I'm going to take you back a decade. So first of
all, Selium finally 10 years old. Pretty exciting. And when the project was founded 10 years ago, there was no AI. There's barely containers. There's barely Kubernetes. And so in the past 10 years, obviously, we've seen a massive growth growth in cloud native. And what's kind of happened? So one decade in where are we? Well from that first initial commit that you saw Selenium has grown to community
of over a thousand individ individual contributors and in the latest annual report uh the project as a total has almost 50,000 stars on GitHub. Even though even if those are a bit of a vanity metric it's pretty cool to see the community the people and all of you in the room here today too. But kind of besides those things, what are we actually doing? So in the
latest state of Kubernetes networking report, uh we asked people which CNI are you using in your environment and the top result uh with two-thirds of the votes was selium and isalent. I think that's a really good testament to how is really becoming the standard for cloudnative security observability um and networking. Everybody is standardizing on top of selium. And you'll hear later today from some of our end
users about that too. And in the same state of Kubernetes networking report, we asked what are you looking at in the next uh 12 months? What kind of issues are you looking at? And people are talking about gateway API multi multicluster connectivity ebpfbased capabilities. And when we looked at the selium user survey, you know, that's exactly what selium has. So choosing Psyllium as a platform for your
cloudnative environment is setting you up for today. It's solving your networking or network policy or network observability challenges, but it's also setting you up for the future too with gateway API multicluster and now MTLS with ingress which we'll also uh with um Zunnel which we'll also hear about later. Selium is also the only graduated project in the CNCF. Even a decade in, it's become the leader in
the space. And the innovation doesn't stop. 10 years in, the project continues to lead the way. And so in the latest 1.19 release, uh we had integration with Zunnel, which Neo will talk about later, um we added support for Gateway API 1.14. We're increasing support for IPv6 um and IPv6 dual stack or uh IPv6 only clusters. Uh moving forwards with BGP to integrate deeper into your network
and doing things like uh multi-pool IPAM. And then on the security and observability side, it's things around uh DNS host firewall, being able to track trace packets through your network, so you're not stuck uh trying to debug your network at 2 a.m. with no And I think that's actually the reason why in the past 10 years, Selium has really become the industry standard CNI. And you don't
have to just take it from me just standing here on the stage. If you go to the Psyllium website, there's almost a hundred case studies about different companies across every industry, every single challenge, and the things that they've been able to overcome and how they've able been able to improve their business with Psyllium. But don't just take it from me. Uh we actually have one of those
end users on stage with that today. So with that, I'll hand it over to Marcelo. >> Thank you, Bill. Thank you, Bill. Hello everyone. I got to say that the view from here is very different. Like I appreciate your present here. Thank you so much. Um so let's get to it. Uh my name is Marcelo and I'm platform engineer at Salonus. We are a group of 30
engineers and we're splitted um around like an office like in German and also in US. I have some colleagues here. So like thanks for you guys being here. Yes. Uh we are a small team but we do some uh infrastructure at scale. We run some infrastructure at Uh for people that are not familiar with Salonus, Salonus is the com the company that define process mining. So in
a nutshell basically what we do we connect to ARPs you know any software custom any APIs and we collect this data for customers. We massage this data we put a you know our process into and then we return this for customers as a value. So our customers use our platform to understand and visualize this process from a unique view. In order to do this, we have to
move a lot of data and networking is pretty crucial for us. To give you a perspective on our scale, we're not the biggest well yet, but we are growing every day. We run on a multi cloud capacity. Our presence majorly so far it's been like AWS, uh Azure and and GCP. uh our daily bandwidth rounds around like 3.5 terabytes a day and this is like running 24
by7 right so like all this connectors are uh collecting the data from our customers we call this as a data ingestion and then it gets processed by our pipelines and our product itself we render out like 360 million daily requests right and we are among like 160 clusters and probably like is different today in this diagram here we can see a little bit of very very very
much on a high level. Uh how is our infrastructure? So in our concept of fleet we have three different clusters and we have like the main cluster we have like the machineing learning cluster and we have the query engine cluster that all combine you know they provide the salon value to our what I want to bring for you guys here you know don't have much time um
but what I want to bring for you guys it's like our um environment and how we were like droning in complexity before right beyond the traditional cloud provider um host Kubernetes offerings like AKS AKS and G and GKE we are also running other flavors like we are running like a Garner we running cops and we're running open shift uh each environment pretty much like a brought it
on CNI and really put like a toll you know in the operation on a day-to-day basis like on top of that on top of each CNI we had different MTLS implementations so it was pretty complicated and the just to keeping the lights on was consuming a lot of energy from our team we really needed like a unifi scalable solution and this is why I'm talking to you
today and then this is where ceiling came in here I decided to just show one slot for you because like from the AKS strategy from the open shift we decided to install new environments and do the migration from scratch but from the EKS ecosystem we already like with a heavy footprint. So we really needed to create a solution with as close to a zero downtime and we
could not really afford to reinstall our Kubernetes clusters of EKS Kubernetes clusters. So when we started this, it started in 2024, right? And back in the day, AWS and EKS didn't have an option for install a cluster without a CNI. And we already had like like I said a very expressive uh fleet of clusters. So basically we created a strategy where we needed to run Celion and
get the workloads migrated to the Celon nodes and our requirement would be like to run a full EPF mode right and and get in the Q proxy replacement. So our strategy basically was getting the AWS CNI nodes and we went to the demon set. We labelled those demon sets and we labelled those nodes. So basically like only the only nodes like we're running AWS CNI. Then as
a next step we basically corded on all those workloads running on that nodes. And this was a strategy to force the audi scaler to whenever we were running or loading rollouting our workloads would force to land the workload on a celium managed node. So like I said there's a a step maybe I get ahead of myself. We installed the the celium in the hybrid mode and hybrid
in this concept is just basically like we have the two CNI. So Celium was doing its part for the manage nodes and then AWS and II was handing like the legacy to be migrated nodes. This draining process could have take like a few hours or in some cases like days and we basically like in some cases depending on how busy the cluster was. We were relying on
the alleyscaler itself just to get and roll out the pods you know landing on the new clusters and some other environments we manually force them to land in the new cluster in the new nodes I'm sorry. And then as a finally step uh we flipped Celium to run an exclusive mode as we call internally. And this is basically where you know we removed the AWS CNI we
removed the Q proxy and then we had the the CNI the fully managed by Celium running on the cube proxy free mode. Uh just a little bit of like our journey here you know and our troubleshootings. I want to you know if you go into similar process one of the lessons I want to say it's like that we had to really put a lot of effort and
a lot of like coaching you know with our engineering team to redesign the network policies that we had in place today our strategy is like each development team has its own you know owns it the network policies and we are there to support them but like the IP block was a big one and was heavily used in our infrastructure so we spend quite amount of time you
know on this particular topic. Another one it's like that we unfortunately learned the hard way on this was like really monitor your BPF maps because this will lead to problems and you know make sure that you have this stack you know integrated on your receivability um as well. We don't have much time we have I posted here like a QR code for you guys you know so
if you want to read more you know learn more about our experience I do invite you guys to to check this out. Thank you so much. And with that, I bring it to Paul. Thank you. >> Hello everyone. Uh my name is Paul. I work as a community builder for Isovalent. That's the company behind Celium and Tetragon. So today I'm going to be talking to you about
uh cloudnative uh runtime security with Tetragonon. Uh so what is Tetragonon? I'm pretty sure everyone here is probably familiar with Celium, but Tetragonon is sort of like more um the youngest uh sibling in the Celium family, right? So, um Tetragonon essentially does what we call runtime security. It's able to like um monitor like your tetragon sits in the kernel is able to monitor like what your applications
or what your workload are doing, right? And this is possible through eBPF that powers uh Celium. All right. Um so the security architecture of tetragonon is based on um four four signals. We like to call it the four golden signals of security observ observability which is a process execution um network observability um also um file access and layer 7 network identity. So with Tetragonon you're able to
monitor like the full process life cycle over the entire of everything that happens in the environment right you're able to see what from a file assess perspective you're able to see um which files are being assessed by which workload and all of this is tied back to the specific uh Kubernetes identity right so you can actually associate like for example you can say a specific port open
like a specific um a specific uh directory Uh so the journey so far uh uh the journey so far first the first commit uh for tetragonon was made in uh 2022 and tetragonon became like GA in uh 2023 and since then we've shipped like um a couple of features that that's making tetragon like setting tetragonon as the standard for runtime security right so we've been able to
ship Kubernetes identity aware uh policies uh we've been able to ship like reduction reduction filters This reduction filters essentially allow you to basically um when you don't want say for example sensitive like information like creeping into like your logs, you can actually write this as you can write regular expressions that actually filter away like things like passwords or argument. You don't want this sipping into your logs,
right? Uh we shipped we've also shipped u support for like two of the most popular runtime uh container runtimes which is cryo and uh containerd, right? We've also shipped persistent enforcement. So pesticon essentially um basically if for some reason inadvertently the tetragon agent dies right the the policies responsible the security policies responsible for enforcement still stay active right and we've also shipped like some other features that
make um the ergonomics of of writing of deploying tracing policies like really easy easy right so for example attribute resolution basically allow allows you um write tracing policies in an easy way. Um uh we've also shipped wider support for like user space hooks. You can actually hook into like user space functions and like attach EPF programs to this user space function. There's al we've also shipped support
for uh event throttling. So basically event throttling is um in tetragonon is basically uh when you've got like this event and you can limit like the per group you can limit the events per c and if the event exceed a a certain threshold like you actually stop this. So this particular is this is very useful for like actually helping like the security posture of tetragonon itself uh
in case like there's an overload of events uh you can actually um stay within the threshold. Uh where are we headed to with Tetragonon? Um Selium has been able to um give us like a standard for what uh container networking should look like. Kubernetes has given us a standard for what um for what application orchestration should look like, right? And runtime security is still relatively new, right?
Uh we've the the landscape is evolving very very quickly. uh we've got AI now and we've got like uh the threat landscape is evolving like very rapidly and with tetragonon what we're trying to achieve with tetragonon is uh set a standard for runtime security one that is cubernetes native that's built on the very powerful foundation of EBF and that's highly performant um thank you very much so
I'm going to introduce uh Neha to come take over >> thanks All this is such an amazing presentation so far. It's amazing to see this room full and I love to be part of it. I want to know who all love Celium. Woo! This is so good to see. Hey guys, I'm Niha Garval. I'm principal engineering at Microsoft and today we're going to be talking about the
Microsoft investments in Celium. So, Celium is a mature EVPF power CNI that industry has rallied behind and Microsoft is all in. From co-chairing the six scalability to shipping some foundational upstream contributions, we are not just the users of Celium, we are the builders of it. We all are sprinting into the AI era now with running training jobs, serving inferencing workloads or even hosting the agentic uh agentic
AI apps. Kubernetes is your de facto platform. Now let's talk about the challenges we are facing in running all these work type of workloads. Whether it comes to the scale we have to we have to offer scale at per with performance and the security which is non-negotiable. So in last cubecon we talked about the scale investments we are doing in the single cluster. How we can scale
a single cluster CL celium. How we can span how we how we can we can go beyond 5k nodes. How can how we can give up to 100k nodes in a single cluster. But now we have to think the scale we have to give a scalable and performance solution across clusters. So we have to span our workloads. We have capacity constraints. Our GPUs are spanned across different
workloads, different environments, different vendors. We have to fix the spuff like the single point of failure if a one cluster goes down. So all these efforts require the need of going beyond single cluster. So the investments we have been doing in the multicluster the two problems the one in the hybrid routing mode. So what it is basically you are your single clusters are performant we have we
have solved by by by giving the native routing but now when we are connecting the two clusters they they are either using their own overlay range or their IP addresses are are overlapped or they're flat but there is needs to have a connectivity crosscluster connectivity so what we have shipped is we have we have given hybrid routing mode so that you don't have a single routing either
it's all native or all encapsulation you can build the best of performance give the best of performance in your local cluster while encapsulate the crosscluster uh crosscluster the second aspect is the filtering I'm talking about the scale challenges now in the scale in this in the cluster mesh we are able to scale up to 250 clusters but within those clusters as we grow more and more workloads
we all know Celium is is the The baseline of Celium is exposing identities which is which are label based identities. So the more and more workloads we we add it adds a pressure onto the Celium agent by by managing those identities. Now in the multicluster scenario where your service is spanned you're not spanning every service you're not spanning your cube system workloads. Now you can with this
you can even you can filter it out like what services you care about spanning and only share those identities across the clusters. So with this we make the Celium cluster mesh really enterprise ready to hold the enterprise scale uh cluster. So I highly encourage you all to look into these CFPS and see how these this can be enhanced further for the enterprise workloads. Every enterprise hesitates to
run open claw, isn't it? And for a good reason. Security isn't a feature you guys we bolt in. It is a foundation you build Now, Celium already secures your network with it has extensive network policy support from L3 all all the way from L3 to L7. But as the AI era, as the AI workloads is pushing into the enterprise clusters, we all want to run the agentic
workloads. We all want to be performant. We are tightening it further. Here's how we're bringing the So we have Celium has identities. It has end to end. It has encryption but it doesn't have the complete end to end story. Now with that we bringing the MTLS and workload level identity which is natively part of the platform CNI Celium. So yes, it's a drum roll moment that Celium
has embraced the Z tunnel uh enhancement integration. It's a joint effort by Microsoft and isol. Now with this we have the complete end to-end story of MTLS encryption which is which is which gives you complete complete part-to-part encryption. So what for for those who doesn't know Z tunnel it's a CNCF graduated Rust level proxy. It it it basically introduce a transparent tunnel. So any traffic which is
going from your pod workload it it it it gets encapsulated in an edge tunnel and it is the enhancements we did in the Z tunnel itself. Z tunnel um does not have the integration with with the spire which is an again industry standard production grade CA authority CA. So we extended the Z tunnel to also be integrated with Aspire which gives us the full picture of um
end toend encryption with the with the full uh with the full C authority. There's a link to the blog. I highly encourage for you guys to to check that blog a little bit on the side on the architecture. Uh we I'm assuming we all be knowing the Z tunnel is already integrated with ISTO. Now Zunnel is integrated with Celium. So you have a Z tunnel proxy which
is a sidecarless based proxy running on every node and the Celium agent becomes the control pane of the Z tunnel and it it is it it is already watching the workloads it is aware of the workloads it is aware of the connections being being created. Now if your namespace is annotated for all those namespace pod workloads, it it projects projects the data into the Z tunnel creates
a transparent redirection proxy redirection rules from your pod workloads to the Z tunnel and then get the get the searchs from the integrated the spire uh server which is running there and establishes the tunnel establishes an encryption tunnel. I have a small uh demo video for you guys to see how we made this inception into into an action. >> Hey everyone, >> okay. >> The video has
to be connect. I can't >> I I need help to connect to the video >> Can you guys hear the audio? Okay. Okay. I think I Whose laptop is this? Bill is your laptop. I'm exposing some personal emails here. >> Oh, no. It's >> I think it's Hey everyone, in this video I'm going to show how Psyllium MTLS encryption works in Azure Kubernetes service and how it
enables workload level authentication and encryption without sidecars or application changes. At a high level, Psyllium MTLS combines three components. The Psyllium agent transparently intercepts pod traffic. Z tunnel enforces mutual TLS at layer 4 without pod sidecars and spire provides cryptographic workload identity based on Kubernetes namespaces and service accounts. Together this design enforces security below the application layer with no sidecars, no certificate management for developers and no
changes to application behavior. With that context, let's jump into the cluster and see this in action. Here I have an AKS cluster with psyllium MTLS enabled where I can see the psyllium agent Z tunnel and spy already deployed. I've also deployed a simple client and server application in a namespace called Z tunnel test enrolled. On the right I'm running two TCP dump sessions. The top one is
capturing CEX traffic on port 80. The bottom one is capturing traffic forwarded to the Z tunnel proxy on port 15008. First, let's look at the baseline behavior. Right now, MTLS is not enabled for this namespace. When the client sends traffic to the server, you can see the request payload clearly in plain text on port 80. There's no traffic going through Z tunnel yet. Everything is flowing through
the standards psyllium data path. Now, I'm going to enable MTLS by labeling the namespace. With this single label, all pods in the namespace are automatically enrolled in MTLS. When a namespace is enrolled, all service account identities in that namespace are automatically registered with Spire. This is what allows both sides of the connection to authenticate each other cryptographically. Now, let's generate the same traffic again. This time, you'll
notice that there's no Cleex traffic visible on port 80. Instead, all traffic is being forwarded to Z tunnel on port 15008 where it's encrypted and mutually authenticated. We can also confirm this by looking at the Z tunnel logs which show encrypted connections between the client and server workloads. And that's how Selium MTLS encrypts and authenticates workload traffic in AKS. >> Yeah, thank you guys. This one thing
I want to tell here is with this we do not need to have very heavy service mesh running. We do not need to maintain two control planes. We are bringing the foundational security building blocks into the platform CNI which is the Celium CNI which is the default de facto CNI for Kubernetes and I need you all I need support from everyone to contribute and make it even
b further make it even better enhance the security to bring the identity aware policies to even bring MTLS encryption into the cluster mesh and complete the end to-end story. So I'm looking forward for everyone to contribute and make it even better. Thank you. Thanks both. >> Thanks, Niha. So now to round out the session, I have just a couple updates from the psyllium communities. So the first
one, yesterday we had the dev summit. Nia was talking about contributions. We all got in a room, talked about what the future of the project is, and I'm pretty excited about what's going to be coming for next CubeCon. So make sure you come back to our maintainers track session. Then uh beyond that there's two new books about psyllium. Uh if you want the advanced version there's psyllium
up and running from O'Reilly written by a couple of my colleagues at Isovalent. And if you want the fun version you can have uh the illustrated children's guide to psyllium written by me. Um there will be his signings later at the Isovalent booth. Uh there's also a bunch of new case studies. A lot of these was written by my colleague Katie. Um, and you can learn all
about uh not only what Solonus is doing, but what lots of companies are doing with Selium. There's also two new white papers out from the EbFound talking about EB, how EVPF, kind of like the underlying technology behind Psyllium is enabling all these things around networking, observability, and security and what it's actually doing in production. So, I can recommend checking out those, too. And there's also a new
EBPF meetup program from the foundation too. Uh with that, um I'd like to close like thank you all for coming. Um we'd love to have you as a part of the community. We'd love to see us get from 1,000 contributors to 2,000 contributors. If you want to be part of the community, come by. There's the weekly developer meeting every Wednesday. Uh SIG scalability, SIG policy, and comm
and SIG community all meet once a month. or if you like Paul's part of the talk, make sure to take out check out the Tetragonon community meeting. And with that, uh, thank you for coming. I think we have one minute for questions or come see us afterwards.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32