KubeCon + CloudNativeCon Europe

Kubernetes Third Party Aud... Iain Smart, Amir Montazery, Rey Lejano, Tabitha Sable & Pietro Tirenna

32:55 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk provides an update on the third-party audit conducted by the Kubernetes SIG Security team, led by Ian and his colleagues. The discussions revolve around the importance of external audits for enhancing Kubernetes security, detailing previous audits from 2019 and 2022 conducted by vendors like Trail of Bits and NCC Group. The speaker highlights various security findings such as path traversals, weak default credentials, and architectural issues that emerged from these audits. They also address ongoing efforts to resolve open findings, the complexity of managing Kubernetes security, and the potential vulnerabilities introduced by integrating different operating systems and components. Insights into the audit's scope, including features prioritized for evaluation, emphasize the meticulous approach taken to ensure the security of Kubernetes projects. Throughout the talk, the speaker encourages audience engagement and collaboration to tackle existing security issues.

Full transcript

Good afternoon everyone. Thank you for still being here at a talk which is fairly late in the day. I hope you're all more awake than I'm currently feeling. This is the third party audit update from Kubernetes SIG security and friends. I'm Ian. I am one of the project leads for the third party project. I am joined by Ray Amir Tabitha and pro who will have more detailed

introductions as they talk. I've done who am I? So, as I said, I'm one of the leads of the third party audit sub project within Kubernetes SIG Security. We are a group of people who work with external vendors to try and get an external view of Kubernetes security. So, as you probably all know, we've got quite a lot of developers and contributors. And while we do encourage

them to pay attention as much as they can to security, that's a specialism in and of itself. So what we've found over the last few years is it's really useful to get dedicated external feedback as well. Where possible we try and work with some vendors to get that feedback. That's a QR code not a Rick roll as tempted as I was every time. Uh this is just

a link to our kind of organization docs for the sub project. In the past we have had two audits performed. uh they were performed the big audits were performed by atrisis and trail of bits in 2019 and then another one in 2022 which was performed by NCC group. Weirdly that was a job that I was actually on the uh project side of so I was delivering that

report. Now I've found myself on the other side which is still quite weird. We've had a few findings over the over the years. I don't know if anyone's read the reports. They're quite good bedtime reading. Uh depending on if you want to be interested or fall asleep probably do both. We've had quite a lot of findings though in quite a good range of topics. So we've had

things like path traversals, weak default credentials, uh cryptographic material that could benefit from improvement. And we've also had some really good architectural findings which have led to quite a lot of discussion. Again, if you're struggling to sleep, GitHub issues, they are all tracked. There's plenty of reading you can go and do. Uh my personal favorite finding just as a a little bit of flavor here. Um there's

a finding from the NCC audit where you can trick the API server into connecting to itself with sufficient permissions. So there's a public GitHub issue for this. Um, if you've got the ability to patch a node object in Kubernetes and rewrite what CD thinks is the node, you can actually trick the API server into talking to itself as admin, which is a really weird niche technically kind

of privilege escalation, even though it starts from a privileged position. If you want to hear more about that one, um, feel free to catch me later. I can talk about it till you're fed up of me. But moving forward, we had the two audits in the past and we've been trying to keep track of all of the findings that we had. Um, anything that was high-risisk or

high severity has already been fixed before the report went public. You'll be happy to hear. Uh, we do also have quite a number of findings which are still open from previous audits and we are aware that there's a bit of an accountability issue here. So, we're trying to keep track of these. If anybody has any time free and wants to help the security of Kubernetes, more than

happy to have a conversation about which of these findings we think are are worth looking at. The oldest one technically it's not a finding from the audit because the audit finding was a duplicate of I don't know if Rory McKun's here, but his favorite GitHub issue 18982 lack of Kubernetes certificate revocation for user identities which I believe has celebrated its 10th birthday and is still going strong.

So if anybody wants to weigh in on that discussion, I'm sure people would be delighted. We did also have some issues get marked as stale and rotten by that pesky automation that we have. So again, we've gone through and rescued those from the bin of automation saying this is a finding that we are not currently uh able to fix. There are about half of the p previous

audit findings um still open, still needing some love and attention. So, as the first call to arms of the day, if anybody wants to dive into some some findings that have been there for a while, uh, please absolutely do. If you're a member of one of the SIGs that those findings were kind of allocated to after the report was published, you're probably going to see someone drop

in on your triage meetings at some point because I believe the findings we reopened will need another round of triage. So, have fun with that one. That's just a little bit of flavor text for what we've done in the past. Those audits have been completed, but you're probably not here to talk about three years ago because that's ancient history in Kubernetes land. I'm going to hand you

over now to Ray to talk about what we've done more recently. Hey folks, my name is Ray Alano and I also I'm also co-lead along with Ian of the audit sub project. So one of the first steps of running a security audit for Kubernetes is defining scope. And so with 2.2 two million lines of code uh it is almost impossible to get the entire of Kubernetes into

into a single audit. So uh since the last two prior audits have been on the core components uh like cublet cube cube API server and more this audit is on non-core components. Uh so with that we take a look at how do we define scope for that non-core components. We looked at the 665 uh features that have graduated to stable since the last uh audit and until

when we started running this audit. So between 124 131 here. Uh then we looked and we also asked all the sigs is there any kind of area or component that you would like to be in the next external audit. So with that we took those and we also identified a bunch as well and we created an audit roadmap which is on GitHub as well. And so from

uh so from the sigs and from the graduated features as well and from our own uh uh our own work as well we identified uh 93 features and areas that we initially identified uh we labeled or we we identified uh 17 as high priority and there's a link over here uh for that for the RFP and within the this audit scope part of it includes cluster API

so you could create clusters uh pod security mission So which we play as pod security mission control I'm sorry pod security policy uh croup v2 since croup v1 is in maintenance mode uh and so at this point I'm not going to go throughout the whole scope of the audits but I'm going to hand off uh towards our next step here was to hand off to the open

source technology improvement fund uh for for vendor selection. >> Thank you Ray. I'm Amir Montazeri, managing director of OTF. That's the open-source technology Improvement Fund. We're a nonprofit that specializes in security for open-source projects and maintainers. And for the Kubernetes audit, we were tapped to help with this one and were what I call the third party audit champion. uh with champion being the operative word here because

even with a project as well supported and with a large community like Kubernetes, security audits are very complex and require a lot of work to even get to the stage to start an audit. And that is what we specialize in as a nonprofit. that we've been doing it for CNCF for a number of years now as well as a number of other organizations and for the Kubernetes

audit uh was no different but some of the main points being that yes the the audit itself was was quite a complex process working with Ry and Ian and the rest of the the communities and the sub uh communities on gathering that information for what to focus on how to scope the audit and had to sift through quite a lot of interest from uh the number a

various number of security firms who expressed interest in in doing this audit. Uh we wanted to consider um all of the previous uh vendors who who want who uh who put in their hat to work on the previous uh Kubernetes audits. So there was quite a lot of interest in in doing so and going through our process to get proposals, analyze them and eventually um assess proposals

and choose the winning one. So in short, OTIFF served as that independent neutral third-party body that really championed the audit from the very beginning to the very end and was able to be that third party and independent party to help with all of the administration and a lot of those touch points that go into making the bringing the work together. And uh with that I will hand

it off to Tabitha to tell us more about that. I'm Tabitha Sable and one of the ways in which I serve the Kubernetes community is as a member of the security response committee in Kubernetes governance committee is basically your license to do things in private. Most work because it's an open- source project done by SIGs must be done in public with participation of of everyone involved. But

there are certain aspects that really need to be handled in secret for the safety of the community. The code of conduct committee for example, security response committee. So that means that we do things like incident response for security incidents within the Kubernetes project infrastructure and we handle the triage and release of fixes for security vulnerabilities. One way that we help with the thirdparty audit process to improve

the safety of Kubernetes for everyone is by reviewing the audit findings before they get released to the public. So the SRC members have divided the findings from the most recent audit up amongst ourselves and we're currently working with the various maintainers across the Kubernetes project to ensure that all of the sensitive ones are able to get fixed before we release the full audit report to the public.

In order to help this talk be full of concrete examples, we've approved some of the findings for discussion in public, even though we haven't yet been able to release the full report. And so, Petro will walk us through some of the findings as examples of how the audit process works. Thank you. I'm Petro Denna and I work in Shielder, the company that performed the audit. We are

an infosex security boutique. We do um ethical hacking, penetration testing, security research, which means that on our day-to-day we encounter Kubernetes clusters. We find ourselves in nods. We find ourselves in targeting applications that run on nodes. So we are always always wondering with the question what can go wrong where where can the vulnerabilities be? So when you tackle a scope that is this big uh among like

hu high high priority low priority PRs full-fledged GitHub repositories you need a process you need a methodology of course because you can I mean you could like just throw a hundred of auditors but of course time is constrained people are constrained so you need to be smart with that what we did is uh let's say a hybrid divide and conquer approach so initially we do threat modeling

together for each and every project so many of you um who are not conf I may have heard this word threat modeling a lot what actually what I mean by threat modeling here there are lots of methodologies lots of different taxonomies I don't want to go into that but what for us threat modeling means here is asking are there in trusted untrusted user input vectors is this

untrusted input handled unsafeely anywhere and if it happens that untrusted input is used anywhere what assets are at stake here that is basically doing brainstorming on what can go wrong and how attackers may want to do things and then we tked each uh single component to a single auditor to provide findings and what I mean by findings here I think this is important item I mean actionable

findings because maintainers are already uh we we don't need to flood the backlog of maintainers anymore than already they have so it's important that we provide actual findings that can improve the security posture rather than just filling a 200 pages report with things that no one will ever read. So um the summary of the audit is in a few in a sentence securing the thing is hard.

What I mean by that um is that Kubernetes is a very highly distributed highly eerogenous ecosystem which is great for many reasons but in security it creates some challenges because you need to integrate those right. So first of all trust in supply chain. You have lots of components, lots of integrations and even a single and malicious uh trusted component will bring down your cluster will compromise your

cluster if you don't build defense in depth detection and mitigation. It's important to do fudging because many times you have complex parsers and you cannot help your auditors to understand all the possible edge cases that you might have there. So if you have a parser or a piece of code that uh is doing some complex um logic it's it's great to uh help ourselves with fudsing unsafe

defaults. I mean who who hasn't built some unsafe defaults in their project right? If you think about it, you could I mean you have an HTTP tool and a tool that exposes HTTP API right now you should always provide a TLS certificate mutual TLS authentication between the client and the server. Okay. Yeah, you should that you should do that. But of course you will leave a default

of TLS insecure so that anyone can connect to the class to the service and then you will write in in a very small font size in the documentation. Oh, be careful. You need to provide certificate certificates there. Please do that otherwise you will be pawned and uh um so g given that um it's also important that um not maybe not um many of us already play with

it but Windows in Kubernetes is a thing right now and uh it's native this support is native and assumptions that stand on an operating system do not necessarily stand on other operating systems. So that's another point that introduce bugs and problems. Now of course if you're here you want to know about some details. U I stand by the rule of proof of concept or GTFO. However uh

as as you was mentioned the full report is not public. So let's say that we are going to peak your interest and curiosity so that when the report is released you can go and read it as a bedtime story or as something to tell your children or parents. The first proof of concept that we're going to bring uh is about the image builder. Um the I'm just

going to give you a few um tell you a few words about the image builder. So um the idea of the image builder is that uh of course Kubernetes can be installed anywhere of and it's it's one of its great uh points but you shouldn't install Kubernetes everywhere. You should you want consistent way of having uh Kubernetes deployed on multiple infrastructure and multiple cloud providers. And the

idea of the image builder is to provide you with consistent virtual machine images that are identical among different infrastructure providers. And essentially the image builder orchestrates two things, two tools. Packer and an anible. Packer is a Hashi Corp tool that um with um code allows you to declaratively um declare yeah sorry um a virtual machine and then you can pass this code to the packer and you

can tell what provider you want on what cloud um infrastructure and you will create virtual machine images for that and then anible is the tool that with playbooks and again declarative um code can install Kubernetes or install whatever you need on those systems. This is great and it's very useful. Of course, we audited I hope the slide is quite read. I think it's readable uh hopefully. And

we audited the actual scripts that were packed with image builder and we found that one of them contained a very very uh common thing in security a hard-coded admin Windows administrative password. Now uh what happened is that every Windows um every Windows image built with Packer could potentially ship with a hardcoded password. Um luckily because of the infinite ways of computing many other providers were um for

many reasons replacing and and overriding that variable. But for a specific cloud provider, Nutanix, that was not the case. Which means that as soon as we discovered it, we realized that all of the Windows um nodes deployed on Nutanix via the image builder had a coded Windows administrator administrator password, which means that it's trivial for an attacker to get um to log the not into the node

and compromise the and compromised it. So this is the stuff for an ar urgent not. Uh we didn't wait for the full report. In this case, we immediately reported it and this was promptly triaged and fixed and a CV was released. Round of applause for the CV. No, I'm joking. Let's continue. Uh so um the second our second proof of concept uh tells us about two stories

pod security standards and Aparm. How do they how did they meet? Uh first of all about Aparm some of you may know it, some of you may not. Um, a parmer is basically a Linux security module that allows you to enrich the standard Unix permission model. The idea is that you can create policies which contain actions that can be allowed or action actions that are allowed and

actions that are denied. For instance, this policy can read files but it cannot write files in the /Home/ coupube directory. As soon as you create it, you apply it to a process and uh then the Linux security model of a parour will take care of isolating that process and forcing the policy. So far so good. Especially because Kubernetes natively supports Kubernet a pararmour which means that if

you if your node is on a Linux system that has a parour installed for instance on Ubuntu you can put in your security context uh the name of the profile the policy that you want to use and your container will be uh isolated with a parour. This is great because it can help us um increasing the security of our um of our cl of our containers. Now

let's put a pause a bit on what a parour is and let's think about the pod security standards instead. The pod security standard is a way to have kubernetes automatically apply a certain level of hardening into your to your name space. The idea is that you have you can have three different layers of privileges. Uh it can be privileged, it's unrestricted baseline, it's reasonably secure or restricted,

it's paranoid. So uh you the the pod security standards also touches a parour configuration. What does this mean? This means that if you enable pod security standards and pod security admission which take takes care of actually implementing it, you cannot anymore disable a parour for your containers. So once you have ps enabled with baseline, you only have two options for a parour. You either use the runtime

default policy which is a policy supplied by your container runtime or you can choose a profile that was installed on your local on your local so on the node by um one would expect the cluster administrator. However, this is not the case. In fact, what we discovered is that Aparm ships with a ton of pre-bundled profiles and many of these profiles are basically just plain alliases to

the unconfined profile. This means that if you use um this aparmour PS if you have a parour and you enable PSS any attacker could trivially create an unconfined container by just piggybacking on an existing profile that is pre-bundled so it exists that actually maps to unconfined. And third we're going to discuss a bit about the cluster API um and its supply chain with providers. So cluster API

is um basically it works very well in tandem with the image builder. The idea of the cluster API is to give users um and cluster administrators a way to create clusters declaratively in the way that we know about Kubernetes. So um you have a management cluster which is a central cluster that understand what a cluster resource is. You have providers you can imagine AWS, Asure, VMware and

uh and everything else. Um and the idea is that the the um Kubernetes admin will take a template of a cluster which is just a YAML file from one of the providers. Then it will apply this cluster definition to the management cluster and the management management cluster automatically creates the actual workload cluster for you and this is great. However, we should really focus on the fact that

a cluster template is just a YAML file that is applied to the to the Kubernetes control cluster which means that and this is the place where I will not go into too many details. Uh we have found multiple ways that attackers could use under certain conditions uh where the attacker can sneak, inject or control how the YAML template of the provider is installed in the management cluster

which means that the the attack scenario is that the attacker can then sneak for example an arbitrary privileged pod on the management cluster which means keys to the kingdom because the management cluster typically has control over every other cluster in the federation. So again um if we look if we um think about u the the use cases that we showed we see those problems in the trust

in the sub in the trust of the supply chain in insecure defaults but mostly the I think that the main take is that none of the issues that we found were due to um insecure coding or not following the best practices really like the code is very mature but the problem is that each developer of has the has visibility on their own project. It's very important to

have external audits because those are the people that can see the whole picture how the things connect together and it's at the connections that typically bugs arise. I guess this wraps up our content and think we'll be very glad to answer our questions now. Any questions? >> Entirely possible that you won't be able to speak to this. What was the most surprising thing that you found or

didn't find? >> Um surprising thing? Um yeah. Well, I guess uh I guess the >> this >> Yeah, I guess the more surprising one was like it fell in the category of assumption on one OS do not fall do not um stand for other oss like we have found a particularly specific vulnerability for one for for an operating system that um that let's let's just say that

it's weird to have this appearing in together with the Kubernetes hard I would say. Yeah. Well, and and from the point of view of being able to discuss publicly released vulnerabilities that have existed in the past, you know, if you look through the history of the Kubernetes CVES, you will see that there are a fair number of these sorts of problems in the history as well where

you know some bit of code that constructs a path that was written when Linux was the only platform for Kubernetes. Then when it's used to construct Windows paths, sometimes things can go wrong because Windows has certain fancy features that Linux doesn't have. Linux has certain fancy features that Windows doesn't have. And as a developer, you get used to how do you work around the sharp edges of

those fancy features. And so as you expand the set of fancy things you could poke yourself on, you do occasionally get poked. Thank you for the update on what's happening. Um the nutanics packer thing uh my company found that and do a lot of different things with cis prep and get rid of those credentials when I build the image in the experience of the team as Windows

is starting to become a bit more prevalent into the Kubernetes ecosystem. Do you think we're going to see a lot more what I will have called onprem Windows AD problems creeping into possibly impact Kubernetes security? Does anyone >> I wouldn't be surprised if we see more of that at some point purely because ultimately we've just got components running on it's it's just an operating system thing doing

operating system things. So we've got components running. I think as adoption grows in any new environment, we're going to just see more stuff happening. I also think we might, and this is this is more from my experience, pentesting systems than it is from the audit particularly, uh we might just see people following default credentials or default install paths and just installing components as entity system or domain

admin rather than necessarily tightly scoped credentials. partially because as things tick up, developers write stuff fast. They write the minimum documentation as quickly as they can. The security guidance historically has come a little bit later. So if you are developing for something now that you're hoping to release soon, please make sure the security docs are reasonably good. Uh and we are more than happy to help you

review those before you launch something if you want to. Uh feel free just to ping channel 6 security on Slack if you ever want to. But yeah, I think we we may well start to see things along that line. Yeah. Yeah. I I think that there's a a crosscultural exchange element of that too, you know, based on having spent a lot of time running on prem environments

with both, you know, Windows and Unix environments, those folks don't necessarily always talk. So then when you start mixing it in this sort of cloudnative mixture you know those pods may not be made by somebody who has been administering you know active directory since the NT days. >> Uh hello and thank you very interesting presentation. I wanted to ask with uh more and more AI engines around

do they tend to help you or make more trouble for you? Yeah. Well, um I think I think we need to be cautious with as with everything However, um for for ourselves like I can speak of course for the experience of of my company and um we tend I I think that what AI can definitely give us an edge right now is really understanding previously like unknown

for us code bases like right now it's it's great that we can poke on a new codebase and instead of spending weeks reading understanding how the entire components fall into pieces you can just like ask let's these smart questions. I where is authentification handled? Can can like a a question that of that I often um ask to the agent or I I generate questions out of it

is um is it possible for an input that comes from in from this source to go into this sync and this kind of source sync analysis which we call taint analysis is I think I think it's much much easier when you have this this tools. >> I think I'd second that. Yeah. Um I had a specific question about a weird edge case in Kubernetes that I've spoken

about before the other day and had never worked out where it came from. Uh one question to the right model just saying this is the thing that happens. Why does it happen? Even in a code base is relatively complex as core Kubernetes which is what did we say about two and a half million lines of Go something ridiculous like that. I'm not going to read two and

a half million lines of Go. Not a chance. um the models were able to get through them in I think I was about three minutes before I got a reply with the exact line of code causing the behavior. So with the right targeted question, yes, absolutely helpful. I don't think they're at the point where you could just go find me all the bugs in Kubernetes and it'll

go and find them all. Thankfully, but also kind of sadly. >> Are we at time? Uh >> we are at time. Um quick question or long question? Go ahead. >> Thank you very much. Thanks so much for your work. Um, also this great presentation. Um, one quick question. How do you deal with trust um, on such a sensitive project uh, with a global team because you don't

want to be skeptical with each other, but at the same time it's highly critical for nearly the whole IT world. And thanks again. Yeah, this is a great question, especially for an open source project where we like to uh work in the open. Um, and security is one one of those uh areas that uh that we might that we will take privacy as a priority. So unless

you wanted to say >> I mean I mean I can I can comment on this a little bit and say that uh practically speaking human trust and and the building of human relationships is is a big part of it. You know, like for example, on the security response committee, the membership in that committee is not decided by a vote across the project the way that, for example,

the steering committee is. And there would be there would be some advantages to that, but there would also be some clear drawbacks, you know, like in a game theory kind of perspective. And so instead there is a private nomination process and so on. And so, you know, that allows folks to build a reputation by having done work with others within the community. So, that you have some

sense, you know, this >> genuinely means well and is capable of handling private disclosure information without tweeting about it or or whatever. But, but yeah, it it can be a challenge and and I think part of it comes down to trying to structure work in ways so that there is the maximum amount of things that you can do in so that you minimize the scope of where

you have to have these worries about things that are private. I hope that helps. Well, thank you all so much for coming. I think we should put the mics down, but uh you can find us around. Several of us you can find at the SIG security booth in the project pavilion. I think we have one more shift tomorrow.