Audit-Ready Kubernetes: How Chase UK Leveraged Policy as Code for Co... Jim Bugwadia & Nischay Goyal
About this talk
This talk discusses how to prepare Kubernetes clusters for audits in regulated environments, emphasizing automation and compliance. The speaker explains the challenges of maintaining compliance while managing large-scale Kubernetes deployments, specifically in the financial sector as exemplified by Chase UK. They introduce Kyverno, a Kubernetes-native policy engine, as a key tool to enforce policies and automate compliance processes. The session highlights how Kyverno integrates reporting and exception management, helping platform engineers standardize their environments and ensure adherence to regulatory requirements. The speakers share their experiences and results from implementing Kyverno, including improvements in efficiency and audit success.
Full transcript
Today, we're going to talk about getting your Kubernetes clusters audit ready. And of course, as platform engineers, our focus is going to be on how you can automate this process as much as possible. So, in terms of the topics what we'll cover, first we'll go over some of the key challenges that you see in regulated environments, what it means to get audit ready, and also some of
the, you know, decision processes involved. Then we're going to talk about one of the key tools today, which will be Kyverno, a policy engine which we'll use for, you know, the audit and compliance process. And then finally, we'll go through the full stack and and talk a little bit about what other tools are involved, how you can, you know, do this end to end. So, quick introductions.
I'm Jim Bugwadia, co-founder CEO at Nirmata, and also a maintainer of Kyverno. And along with me, I have Nishant. Um thanks, Jim. Hi, guys. I'm Nishant Goel, working as a senior lead platform engineer at JP Morgan Chase, and specifically for Chase UK. Glad to be here. So, as Jim mentioned, we're going to talk about Kubernetes in a highly regulated environment, right? I think the first key thing
is when you think about Kubernetes, suddenly we start thinking about scalability, reliability, networking, all these challenges. But when we are in banking or in financial regulatory industry, the first thing is will this pass an audit? That's the first thing what we think about, right? With that, who are we? Chase UK, right? it was launched back in September 2021, and in less than 5 years, we have more
than 2.5 million customers, with 20 billion plus deposits as customers' money, and this is the scale what we are working upon. Now, along with that, think about thousands of nodes with multiple clusters, around 15 our environment running in different environments with different configurations. And this is where things gets interesting if you merge that scale of the customers with the technology scale we are running at. Now, if
you think about when you are in this financial regulated environment, right? There's always strong audit expectations, because you're actually working for the real customer money. You want to protect their money as well, right? And if you've been in this space, you know that speed and sort of where engineering teams want to ship fast independently because they want to go to market, versus with expectations of are they
secure, are they safe, are they compliant? So, we as platform engineers or platform operators is always trying to balance the speed versus the compliance, right? And as you know in this space, these two naturally don't get along very well. So, what was our challenges, right? We sort of realized back few years back that we didn't have a compliance system, we actually had a compliance process. Now, what
this compliance process means, Your security teams, your auditor teams, or some other teams are coming and asking you questions about how you're doing this in your environment. So, you start producing quarterly reports, maybe spreadsheets, maybe running some kubectl commands that, hey, I'm running these many pods with these many configurations with maybe mesh is enabled, it's not enabled, all sort of that, right? Which means you're giving reporting,
but inconsistent every now and then. And this is what we were doing at Chase UK as well, because we were growing so fast, all the requirements was coming sort of ad hoc, right? We didn't have a proper framework around it, and it was all manual at that time. So, what do we need from a audit ready perspective? What was our sort of a requirements, right? The first
thing was standardizing the whole environment in terms of each resource, right? We want to make sure that if a pod is running in our environment, no matter which pod you pick out of thousands pods, it looks exactly the same. You could say this is a baseline, right? We want to enforce guardrails, not documentations, not recommendations, because when you give recommendations, I think it gets ignored in a
real reality world. In idealistic world, no. I think everybody should follow that because recommendations are there to help you for your applications, right? Along with that, we want to uh we want to apply policy once and gets rolled out everywhere as part of the whole estate. We want to enable developers, not blocking the developers. And finally, as I mentioned about reporting, it's quite important that we wanted
to build a unified reporting, whether is it for developers in the back-end engineering team or front-end engineering teams, whether it's a reporting for us as platform operators, or it's a reporting for some team leads, or it's for it's a reporting for some stakeholders, and which includes security teams, auditor teams, or some regulatory teams who needs access to your environment to look whether are you doing the right
things as per the regulations or not. And looking at all that context, what are the solutions we sort of evaluated which we needed at that time? So, the first thing is we had a OPA, which is Open Policy Agent with Gatekeeper. We already had that deployed in our Uh I'm going to talk about that in a minute. That was the first thing. Second is, as natural engineers,
the first thing comes in comes to mind is can we build something in-house? Can we write something which will I'll maintain it, this is a cool product, and then I'll write about it, and I'll go to a conference and talk about what I have done there. As engineers, I think that's the first natural instinct. Third is Kyverno, a policy engine, and no surprise, I think you've already
guessed it by now which we picked up, which we'll be speaking about in a minute. And the four fourth was other policy engines back in that space, So, these were the solutions. As I mentioned, we already had a OPA uh in our environment, right? So, as Chase UK was launched in September 2021 in production for the live customers, right? It took almost 3 to 4 years to
develop and to go actually live. I think less than that. But given that within 4 years, we had four policies in production. That's the first thing that So, I'll take a timeline of 3 to 4 years. There was a very limited expertise of Rego Rego back in 2022 and 2023. This timeline is very important, because if you ask today, I think everybody's using AI here, and I
don't think expertise is a issue now. The team or the engineers who developed these policies or who initially developed the environment was no longer part of our team, which is a classic the domain or the technology is lost now, and we don't know how this thing is working or what we need to do here. And finally, because there was no reporting, when something breaks or there was
no ownership present for the resources, because we as platform, we provide a platform as a Kubernetes platform for engineers to deploy, but there was no standard ownership for those resources. And with that, I'll hand it over to Jim to actually talk about Kyverno, and please. Thank you, Nishant. Yeah, so Kyverno, as many of you know, is a policy engine which was designed for Kubernetes. But other than
just trying to simplify how to write policies, one thing which was very important to us from the beginning is how do you automate the entire process of governance and compliance. So, also making sure that the right other tools were integrated and part of the policy engine itself was extremely important. So, first of all, the approach we took was, you know, as Kubernetes focused engineers, we all have
to learn Kubernetes, how it works, what a KRM looks like, how do you operate Kubernetes resources. So, the first thought was, why not make policies, exceptions, reports, everything very native to Kubernetes itself. So, that was one core decision in Kyverno, and that philosophy stands today, too. As Kubernetes has evolved, Kyverno continues to evolve, much like in our newer, you know, you'll see in the newer releases, we've
adopted some of the newer features in Kubernetes. The other thing, you know, that was very different about Kyverno versus OPA or other policy tools, to us policy was not just about enforcing validation logic. So, not just about allow or deny decisions, but also being able to automate the full resource life cycle. So, you'll see in Kyverno, there's policies for every phase of a resource, and you can
do policy-based automation, which is extremely powerful for platform engineers. The other two things, you know, I mentioned briefly already, like say integrated reporting, and making sure reports were available not just for compliance, but also for observability for developers, for operators, and even security teams to consume. And we all know, wherever you have a policy, there will be exceptions. So, if you have policies, you need exception management.
So, we didn't want to make that an afterthought or something decoupled. It's integrated into Kyverno, works like just any other Kubernetes resource. You can use GitOps for it, any other best practice. So, those were some of the core things we have always built into Kyverno. One change in the project, you know, other than the new policy types I'll cover, is Kyverno started with the focus on Kubernetes,
but since then has evolved into operating anywhere in your infrastructure or application stack. So, you can use Kyverno for IAC, you can use Kyverno at the app layer. In fact, there was a session yesterday about using Kyverno for MCP authorization. So, lots of different use cases where Kyverno applies across the entire stack. So, in Kyverno, we have five validate um five different policy types. The first one
is to validate resources to make sure your configurations are consistent, compliant, follow the best practices that you want as an organization, or, you know, from your regulatory perspective. The next one is to change things, like mutate things on the fly. Um so, there are, you know, there's sometimes you'll get into discussions with people who want, you know, who are kind of about GitOps, and they'll say, "Well,
you shouldn't mutate things in cluster." That's not, you know, always like having a hard stance on that is not always best cuz if you think about what deployments do to pods or other Kubernetes controllers do, you are mutating, you are changing resources in cluster. And the tighter the control loop, the faster you can do that in cluster, the better. Um Kyverno also generates resources, and there are
very good use cases for that, too. So, for example, if you have a new space created, Kyverno can automatically create roles, role bindings, fine-grained permissions. You can do things based on label triggers, you can create network policies. Again, think of this as closed-loop automation, which you want to run in cluster cuz attackers don't go through your pipeline, they are in cluster already. Right? So, you want these
things to be generated in real time, just like any other Kubernetes controller would. And GitOps works very well with this because now with, you know, admission controls and by, you know, setting the right parameters either with Argo CD or Flux, they can work together. And then finally, we have a deleting policy cuz again, you want to clean up resources. So, based on either TTLs, time-to-live labels, or
based on other constraints, you can automatically clean up resources, even look for orphan resources, dangling resources in your cluster to make sure everything is tuned and operational. So, like I mentioned, you know, has expanded outside of Kubernetes, too, and here's just a quick example of what that looks like. The only thing I want to highlight here is if you go to our website, there's plenty more examples.
Is no matter where you're operating, a validating policy still looks like a validating policy. It's still very much like a Kubernetes resource format. So, everything you kind of know and love about, you know, Kubernetes resources, um but the only one change you'll see here is based on where you're operating, either it's JSON or Envoy or, you know, other kind of, you know, environments, versus Kubernetes, which is
the default. And then, of course, the body of the policy, the cell expressions will vary based on what payloads you're applying the policy to. I'm not going to go through all of the use cases, but just really want to highlight over here, Kyverno gets used for a lot of automation, optimization, as well as security, of course, use cases, right? So, you can start with the basics, like
pod security is a must-have, you need that. Uh so, start there, but then look at other automation. One, you know, very simple type of use case, which is often ignored, is even things like label inheritance, right? So, making sure not only namespaces, but every resource in the namespace gets the right label. Stuff like that is super easy with And then finally, just talking about how do you
go from enforcement, which is one part of governance, into that full compliance and governance life cycle, right? So, this is where, you know, reporting was always integrated into Kyverno. We've spun that out into a separate project, Open Reports. So, now that Kyverno has graduated, which, by the way, just got announced yesterday, um we are going to we're focusing a lot more on Open Reports and some of
the other projects around Kyverno. Um so, that gives you standardized reports for Kubernetes and beyond, and other projects have already started, you know, adopting this format. And then finally, exception management, like I mentioned, is very fine-grained. You can have granular exceptions, not only on resources, but on like things like containers within a resource or specific settings, which makes it easy to manage. With that, now Nisha is
going to go into more details about the full stack and solution. Thanks, Jim, for the overview of Kyverno. So, now let's look at you have already understood sort of the problem and the context and Kyverno, what it can do, right? So, we take both of them and go into what is why we selected sort of Kyverno, right? How it is helping Chase UK. So, first is everything
is Kubernetes native, right? Where there was no learning curve because everything is written in YAML, and I think each one of the person in this room can write YAMLs, and I think that's why we all here in the Kubernetes KubeCon conference, right? So, everybody knows how to write YAMLs. apart from that, reporting, which Jim mentioned was a built-in reporting was enabled and was there in term of
Open Reports as well. That was a massive win for us, right? Because while platform operators, we can look at resources by calling some APIs and some automations, right? But when you look from the other side of the stakeholders, right? They're always looking for some reports so that they can visually see it, right? Uh I'm going to share the reporting how we have done it Chase UK in
the next slides, but that was a massive key sort of decision point for us when we selected Kyverno. Right? The different um modes in the policies, like you could enforce a policy, which means that from here onwards, you won't allow if you're failing the policy, or a audit mode, or a warning mode, right? Uh so, the different modes according different use cases. And policy exception is a
first-class citizen. This is quite a key again because when any policies in general, not in technology, but in general, like, "Hey, this is a policy which everybody has to follow." Or the terms and conditions when we go to any products, right? Sorry. in the financial world, I think 100% compliance does not mean that I am meeting all the policies. I think from my perspective, the definition is
that I'm managing the 100% risk, right? Now, 99% workloads can be, yes, I'm meeting the policy and compliant. What about that tiny 1% use case, let's say, my it's a edge case and it cannot meet a policy, right? And at that time, you need exceptions as a first-class citizens. Cuz if you don't have that process defined, other you need to come up with some custom process to
let them allow into your cluster. Because in the end, they are products which needs to go live into the market, and you have security, and security is there to support the products, but not to actually block them, right? That's the whole idea behind it. This is what audit-ready Kubernetes stack for us or what we are actually using for Kyverno there. So, if you look, we have Kubernetes
on which we are running all the workloads. Uh and in each of the Kubernetes clusters we are running, Kyverno is sitting as a admission controller, so it is Chase UK, when you build a cluster, the definition of a base cluster is actually Kyverno has to be present in each one of the clusters, right? That gets there from the minute we build the clusters. And then we use
Prometheus and then probably Thanos in the end, like to for metrics and for all of that to flow, and then finally, Grafana for the visualization to actually give in the end the data real-time reporting dashboard for all the stakeholders. Now, if you look at some of the other took some screenshots from the um our real-time reporting, right? This Grafana report Uh so, if you look, this is
pod security standards. There's only one Grafana dashboard which we use. We don't have multiple dashboards, but there are a few panels. So, I'm going to go through a few of the panels. So, on this side, if you see, there are different filters on the top of the dashboard, which actually tells you that if you which if you're a security engineer, right? And you want to have a
look at a specific category, you select the category, which is pod security standards restricted. Uh I think back in 2020 three or 2024, when PSPs got deprecated, the pod security standards came into picture. That's one. You can look what is the state across all clusters. And because it's pod security standards, that's why there are no failures right now, even today, right? Everybody has to pass that. that's
one example. The second is it's all policies right now because if I don't show failures, I don't want to say that everything is working perfectly fine. I want to give you the real picture. If you see, I have selected all the categories, now all the policies. Here, if I'm a specific team lead, right? Who is looking for a specific namespace of across all of my workloads, you
can select a filter of a namespace, and it will show you that specific information. If I'm a platform engineer, I'm I'm rolling out a new policy, I will select that specific policy to see who is actually failing at that minute to actually set up a baseline. That tomorrow, there's a new requirement, I have to write a policy, right? That we should have this on the platform from
any XYZ requirement, right? I can actually set up a baseline right now, 95% or 80% of our workloads are actually passing that or not. It helps you to sort of project what are you going to do, whether this policy makes sense or not. What is the actual work is hard to be done by different teams. I remember back in when we were doing the deprecation that's when
we introduced one policy because where all the capabilities used to get added by PSPs by default as mutations, right? So we used to introduce a policy rather than going to all the developers in our in our company, can you add those which we actually introduced it by looking at the failure rate that 99% was not adding it by default. It was PSPs that was doing it. So
it gives you different ideas with that data. On the similar on the same dashboard violations, right? Now if you're an engineer and you know that your workload is failing a policy, what information is required by an engineer to fix your resource? So we sort of mention the cluster name which is from production here. We mention your the name space on which in which the resource is running.
The type of a resource can be deployment, any Kubernetes resource it can be that type. Gives you the name of the resource. The policy it is failing which in this case was a pod spread we from a reliability perspective we wanted it to be in a different AZs. That's what and the rule which is the specific rule of that policy. And then tells you the categories. So
all the information which is required is present on that single board so that engineers don't have to go at multiple places to actually connect all the dots together to work on it. Cuz usually it's been seen that fixing a problem takes less time compared to gathering all the information to fix the problem. So that's the whole idea behind the single dashboard which we have been using. Performance.
This is sort of I took a screenshot from our of latency of the We've been running for more than 3 years almost. So if you look at the scale what we've been running, this is a latency which is less than 20 milliseconds in our whole estate. On an average, right? And the best practice is latency should be less than 100 milliseconds or should be less than 100
milliseconds. The reason being is because if your admission gets slow, it has a knock-on effect on all of your workloads which will get into the cluster. It's like if there's only one door to come into the hall. And there's 1,000 people waiting there, you will start seeing a queue. Listen. So you need the in process you need it to be quick. What was our journey? As I
said, in 2021 when Chase UK was launched, there were sort of four policies which was OPA and Argo. 2023 is when we started we got a uh audit issue that we had to work upon. Right? And we started looking at that and that's when in 2023 because in yesterday's keynote I forgot his name he mentioned that it took him a month to go from this point to
production. And literally that was the effort as an engineer it took where I spent a month from a research on that problem and to actually taking it to production was 1 month of effort in a bank. This is very important to mention in a Right? We we managed to migrate all of those four policies running in production and we decommissioned the whole OPA gatekeeper from our environment
in Now if you look at the growth in 2024, 5, and 6, right? We had four policies and it's gone up to 46 policies. The key thing is we managed to crack a framework or a culture. It was never about policies, right? It was all about culture. Now when we as platform operators or even the engineers who want to build something and provide a service to their
on the back end our customers, right? They actually start thinking from the get-go how do I protect my product? How do I make sure my product is reliable? Be it about labels, be it about ownership, be about reliability. I think all the use cases what Jim mentioned in his slide, right? I think lot of those things come into these policies, but there are custom use cases which
we need as well, right? So I think that's the key here that we managed to crack that framework and now policies are not written by one single team. Uh one of the main main objective was can any developer or can any security engineer raise a PR for someone to review it rather than you come and raise a ticket with the platform engineering team it's a niche thing
only they can do it, right? We should raise a ticket, wait for them to write a policy and come back with their cells. Right? So we have different policies in different ways. Uh networking, multi-region what we are doing right now. So in that space. What was our results back in 2023? Because when you do something end of the end of the year you go back in reality.
What was the sort of success you achieved with that, right? Roughly 7 to 8 months of engineering effort was saved at that time. In terms of 25% automation because when you do like one of the example which I remember was when we when you do Kubernetes upgrades, right? We always use deprecated. We look at deprecated APIs. You read the release notes and look at who's been using
it, right? So we actually end up writing a policy which helped us to show how many applications actually using those APIs, right? That's just one example. Uh in terms of automation and cleanup as well because there was a lot of dangling resources which we cleaned up as well as part of the policies. I think the key thing by this time was our auditors were happy. And I
think that was one of the hardest SLA or the KPI to meet. that was the success. When you go through a journey there's always learnings, right? From our these are the few learnings which I had with One is when we wrote our first policies in Kyverno we wanted to ensure that the policies when you write they are being tested properly in the CI. It wasn't there before.
We were not testing the policies because one single policy can impact a whole cluster. Okay. I think they got to finish. Yeah. So always test your policies in the CI. Make sure you write your test for the policies. Policy enforcement. This is a learning we had on our journey. If you start a policy in audit it takes ages for you to go to enforcement mode. If there's
an opportunity or there's some way which you can find out that I can start with the enforcement mode you should do it. Because from that point you know that all the new stuff which is coming into cluster is always compliant. And now from that point the non-compliance are bound to go down. Now I'm calling it a non-compliance but it can be any resource configuration which you don't
want in the cluster. example being if you have a if you're writing a policy for a pod versus a deployment, right? Pod always gets come back into the cluster through admission whereas deployment doesn't come back into the admission every time until unless you deploy, right? So think about how you want to write a policy which will enable you to go into enforcement from day one. The last
one is interesting. so when you write a policy you could actually define the scope of a policy. the scope of a policy is pick up all the resources with the specific label as an Which means the policy will skip all the resources which are not in that scope. At the same time when you write exceptions exception also means skip. So this is what audit highlighted to us,
why you have so many exceptions in a cluster? And we're like we don't have exceptions, right? So it gets confusing where you have a exception but actually they are supposed to be skipped but not an exception. So I think that's where we had a confusion. We're still looking into it. How do we improve that sort of thing? These are actually exceptions and these are actually genuine skip
cases. So that was another thing during this With that I'll hand it over to Jim for the conclusion. Yeah, so as we've talked about policies can be used for a broad set of use cases, right? And today we focused on security compliance but there's operations, optimization, automation, and other things. But to get compliance right, you also need reporting, you need exception management, and all of this has
to be auditable and tracked. So super important as you're looking at the processes to make sure all of these can be automated. You by actually using GitOps, use it across the board for And then as you were, you know, kind of automating this also think about the audience, developers, security teams, all of them need visibility, which is why, you know, what's great about Kubernetes is you're standardizing
on a set of configurations. Now with Kyverno and other tools like that, you can extend that to other processes around your clusters and your infrastructure. And Kyverno as we've talked about it was designed for these use cases. So if you haven't tried it, you know, give it a try and if you have any questions, we do have a booth in the project pavilion. So stop by, we'd
love to chat. Thank you.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32