Alex Olivier: Layered Security: What Aviation Safety and Cheese Can Teach Us About Zero Trust
About this talk
This talk discusses the intersection of aviation safety and cyber security, particularly through the lens of zero trust security principles. The speaker, Alex Olivier, emphasizes that as cyberattacks grow more frequent and sophisticated, a layered security approach is vital to prevent breaches. He introduces the Swiss cheese model, which illustrates how multiple layers of security can mitigate risks, akin to the precautions taken in aviation. The discussion extends to how the same concepts apply to software architecture, where authorization must be externalized to ensure robust access control. Olivier also delves into tools and technologies like Kubernetes, OAuth 2, and Cerbos, a policy decision point for managing authorization efficiently and securely.
Full transcript
Thank you for joining me today. As it says on the screen, we're going to be talking about planes, cheese, and zero trust security, particularly one layer of cheese that I think is a bit underserved. Before we go into that, a bit of scene setting. A bit of a pop quiz, I guess, for this time of the morning. It's the morning, still morning, yeah. Um Any guesses how
frequently cyberattacks are happening right now? Every hour? 30 seconds? Every 37 seconds is the latest report that people are willing to put numbers against. That's on Cyber Security Ventures 2023. Percentage of businesses affected? 60, I heard. 68, that was Sophos. That is definitely much higher. Uh probably guess from the accent, I live in the UK. Just this last week, three of our major supermarket chains and their
supply chains have been hacked, and I went to the shops last week, and the shelves are empty. Um There's literally like no food on some of the shelves because cyberattacks happening at the moment, which is great fun. And the average cost of a data breach, um this number has rocketed since this stat came out, so I'll spare you that, but around 5 million back in 2022, this
number is increasing exponentially. So, that's the world we're living in right now. Um The landscape today, if you've not read this book, I highly recommend it. This is how they tell me the world ends. It is a terrifying read if you have anything to do with technology. Um but really lays out the challenges that we're facing. Attacks are getting more sophisticated, expanded attack surfaces across all of
our technology stacks. There's remote We're now in a kind of world of remote work that's far more common. There's a whole new attack vector and vulnerabilities that come with that. And then inside of threats are becoming more prevalent. If you saw the previous talk about supply chain attacks as well, it's a really It's an interesting one, and that's an emerging threat. Have to talk about AI, AI-powered
attacks. Um there's uh literally people being put out of work in these ransomware shops because they essentially were the people on the end of the phone line when people phoned up to pay their ransom and now they're now being automated with an AI chatbot, believe it or not. Uh ransomware as a service and geopolitical cyber time. What's all this got to do with aviation safety? Well, really
both of these are complex systems with potentially catastrophic failures. So, who am I? I'm Alex Olivier. I'm one of the co-founders of a cyber security company called Cerbos that do authorization. Touch upon that at the end. I'm a software engineer at heart, but sadly ended up in the world of product management a while ago. Um but started uh micro started at Microsoft doing technical consulting stuff for
major banks. Um but been doing mainly SaaS in various industries and verticals last 10 years. Um but always very much focused on the data, DevOps, and infrastructure side of things. So, the security architecture was something that I had very much had remit on and for bit years ago started an open source project called Cerbos, um which I'll come to. Um and we've managed to build a kind
of a business and a company around that. The other thing about me is I'm a bit of a #AVGeek or aviation geek, self-proclaimed. Uh part of this job I have to travel way too much. Um and one of my favorite TV shows is this Air Crash Investigators. Anyone else has watched this? But it goes and uh into detail about how these are perfect airplane crashes occur and
what happens behind them. Very interesting if you're into uh both kind of engineering uh and disasters, I guess. Um I also don't recommend watching on an airplane. You get some funny looks. But one common theme that comes up with that is there's this guy called James Reason, uh who's a British psychologist. Um and he did a lot of work to help develop the processes, the SOPs, uh
at the the models for how accident investigation and risk management works in aviation. There's a lot of engineering, but there's also a lot of psychology around how you sort of investigate an incident and what led up to an incident, and he kind of built a lot of these models. And the Swiss cheese model is one of his. And basically, simply enough, you have some sort of situation
and you have a potential for a disaster at the other end. And the idea is you have to go and put systems and processes in place to capture and basically prevent that disaster from happening. So, in the aviation world, you start off with like things like pilot training, make sure the aircraft's airworthiness is as it should be, make sure you have good SOPs in place when situations
occur, automation control, situational awareness, human error, etc. So, what the Swiss cheese model uh says is each of these layers should protect. And with Swiss cheese, you have holes in, and if those layers align, you can actually go all the way through and get to the disaster scenario. But with more layers um in place, you sort of block and there's more chances to capture and prevent those
disasters from happening. One example of that is an aviation system called triple modular redundancy or TMR. And it's just one of the the systems inside of a flight control computer. Where so, you have various sensors, hundreds, thousands, if not more than that, across an aircraft measuring, wind speed, attack vector, position of your rudder, flaps, etc., etc., etc. And all those sensor readings are sent into multiple computers.
Um in a TMR system, a triple modular uh redundancy system, there's three separate, individual, isolated uh flight computers. Uh uh Fun side factor for when you're having a drink with a friend, uh the data backbone of an air aircraft is basically Ethernet. Um there's a ARINC standard, but it's essentially Ethernet on the on underneath, which has full duplex packet switching, but you can have priority on those
packets as well. So, those three computers receive um those the three sensor readings in parallel. Those three computers then have to make a decision. Um, if that decision is agreed by at least two of them, um, then, sorry, if it's a majority, so at least two of them, then that action is taken. Uh, if not, it's not, um, that you might have heard of raft consensus as
well, similar inklings there. Um, but basically has to be a majority before any action is taken based on the data. The idea being if one of the flight computers is not behaving or some something's not quite right, the fact that you have three sensor decisions that need to be made, um, and at least two of them to be agreeing before any action is taken, you should kind
of prevent issues. And so, if you look at that model, those holes line up and eventually, uh, you have a layer that stops any potential issue and you avoid a a disaster and hopefully there'll be much fewer episodes of air crash disasters. So, what does this all actually mean for software applications? Um, what's kind of sort of my my interest in it? So, if we were to
look at the same model for what we do as software builders and architects, we might have a first line of defense where we say, "Okay, we have a situation. We have a service needs to be protected, needs to be secure." The disaster in this case would be a breach or hack or compromised identity, let's say. So, we're going to stick a firewall in front of it and
we're done. Don't do this. Um, this is kind of what's referred to as, uh, perimeter security where you just protect your services with some sort of like firewall, maybe the VPN to come in, uh, and that's it. Uh, that that doesn't work in this day and age and a zero trust kind of um, leads down this model and essentially this is what you're allowing. One small hole
in the firewall, there's no other protection behind the scene and you're in trouble. So, zero trust, um, set of principles, um, very much pushed by NIST as the, uh, standards institute defining this. Uh, basically trust nothing, verify everything. Um, no more perimeter security. You you to verify explicitly and you want to use least least privilege access and essentially assume breach. Assume that whatever's running in your system
or the network or the environment your system is running has been hacked. How would you go and build and protect against that? So, if we going to look at it again, you might have that firewall layer. You might then go and start doing things like network segregation. Uh you then start looking at encryption both in flight uh sorry, in transit and at rest and making sure that's
consistent across the entire architecture. You start looking at things like workload identities. So, now you don't just assume the service can do anything. You give each of your services or each of your applications a unique identity and you authorize around that identity. You have that authentication layer in place. Very much coupled with the workload identity. So, you have a system that services can or our users are
to authenticate against before they can do certain actions. And so, what does this look like in uh the modern world and excuse my very strained metaphor of cheese Kubernetes? Um you have those different layers and you want to go and start building up. And if we were to take those those layers and apply it to maybe specific technologies you could look at to go and put in
place in the context of a Kubernetes cluster. Um what would that kind of look like? So, the first one your base infrastructure. These days and age, you might be using one of the hyperscalers say say AWS that's giving you the base infrastructure, giving you things like the the firewalls, security groups, etc. Um and then you have your cluster. Kubernetes obviously being the one that's most prevalent these
days. On top of that, you want to start doing that kind of network segregation piece. So, there's now things like service meshes, Istio, Linkerd uh being the most common ones which allow you to define network policies of what can happen inside of that cluster. And uh allow you to lock things down further in that way. On the workload identity piece, uh SPIFFE and SPIRE uh SPIFFE is
the standard, SPIRE is the implementation is a mechanism to give trusted identities to your workloads. It's very Kubernetes native but it doesn't have to be. Um but essentially each of your services can uh be given a X.509 or a signed jot as their workload identity and there's a way you can verify the root and the trust of that identity and use that to make further authorization decisions
around what the system can do. For your authentication piece both for the services and your humans or non-human identities, um OAuth 2 is kind of the standard that's now come along here. Um Keycloak is just one example of a a CNCF IDP which you can roll out and put in place so you have a strong mechanism to do identity and management in the whole joiners, movers, leavers
of your identity process. And then the other end things, great you've got this you've got this nice stack but how do you monitor and audit it? So, there's now standard things like OpenTelemetry, you have things for standardizing log collection, you have uh Prometheus metrics, etc. Having these in place allows you to get that visibility across your stack. Um but there's this gap in here and this is
something that's kind of always mhm bemused me a bit maybe is the right word around you have all these services, you have these stacks, you have these and you can under the the term of never trust, always verify really this only takes you up to verifying the identity. Once someone's been given a token, an authentication that token may be valid for let's say 15, 20 minutes, an
hour maybe. in that space that time, if someone managed to get one of those tokens, they could do some serious damage in your system. And so, I've always said that there's a kind of missing layer to the stack that most most people think about this which is authorization and this is why I've spent the last four odd years um building out and and working on. And authorization
is the number one issue. Uh so, OWASP, hopefully some of you are familiar with this, they do every four or five year report on the top 10 issues uh for web applications in this case and broken access control is consistently near the top and in the last report was the top. I think the next report actually comes out next month. I suspect it'll be near the top
as well, but fundamentally, your access control model is broken. So, someone with an identity was authorized to do something that they shouldn't be able to do. So, they have a valid identity, their authentication is fine, but what they're doing inside of your system is broken or the access they got inside of your system is broken. So, authorization is the I don't know what the bad baby bell
if you have that here of the Swiss cheese model. It is the bad cheese of the model. And so, just be clear, authentication and authorization are separate concerns with annoyingly similar names and even more annoyingly are reduced to these sort of shorted forms in in a lot of time in a lot of spaces. So, authentication with an N is the process of challenging someone to provide a
credential. They give you a credential, your system verifies that credential or go yes, you are authenticated, you are who you say you are. Authorization is the next step, which is okay, you are who you say you are, but are you allowed to do that particular action? Can you access this particular resource? Can you view that particular booking? Can you edit that particular expense? Can you create a
new invoice? These are the business level access controls that you need inside of a system that you need to authorize inside of your applications. side note, authorization is even spelled is even confused in the HTTP spec. so, there's many of you might might be used to putting inside of your HTTP headers an authorization header with a bearer token. Um that's actually authentication cuz you're providing a That
really should be authentication, but that's beside the point. Meta point being this gets modeled up a lot and um it's annoying and confusing and there needs to be a better way to talk about it. So, how does authorization actually become a problem? And I think the best way to look at this is if we go and pretend to scale up a company. So, in the early days,
say you're building out some software, you're building out some application, and you have what I call sort of the blissful days of roles. Where you have your simple users, um maybe there's a user as the system builder, and then you have maybe an an end user, and someone will log into your application and can go and put a simple if statement into your code. If the user's
email contains your company domain, great, they're an admin, they can do everything. If they don't, then they they're not allowed to do everything. Yeah, you can maybe use an OAuth scope to do this as well, but it's a pretty binary Well, it is literally binary um whether someone can do actions or not. Very coarse-grained. You essentially have like an admin role and a user role at this
point. Very simple, and and off you go. And this is something you could just throw very quickly into your application code inside of your request handler to check whether someone uh should be allowed to do an action or not. But if you scale up a bit, if things are going well, you've taken that first version, you've taken that first prototype, and you want to go and scale
it out. Maybe you've got like half a dozen features now, you're at the point where you've got a like a sales team that are going to go and pitch this product to the world, and one of the first things that will usually come up is you go, "Well, we want to stop giving away all the features to everyone. Uh we want to sort of change our product
packaging. We want to slice and dice maybe how we give out features to users." Very common thing that I'm sure many of you have had to tackle potentially in building your software. So, now you go back as a developer, you've got that ticket, it's like, "Ah, okay, they want to have like a premium package and like a base package." And so, everywhere inside your application code where
you're checking whether users can do actions or not, you now need to go and extend that if statement a bit. So, now not only do you need to check whether someone is an admin based on their company email, but you You need to check what package that particular company belongs to or subscription they have uh package level So it's like, "Oh, okay, it's not too bad. You
can go and extend that if statement a bit." Potentially, you start looking at an entitlement system or a feature flag system that can do this a bit more dynamically, but you still end up having to potentially extend the code like Let's go a bit further. This was something that really, really, really burned me in one of my previous careers, where I was looking after a kind of
a big data platform. We had about 50 petabytes worth of clickstream data, and this was in the early 2010s. It was like 2014-ish. And at the time, there was a data sharing agreement between Europe and the US called the Safe Harbor Agreement. Um see a couple of nods. And all going right, as a European company, you could store data in the US and vice versa, and everything
was all good. In the space of about a month, that was struck down. And so, we were in a position now where we had to essentially go and segregate our data. We weren't allowed to store any data on our European customers in the US and vice versa. Um this was a real nightmare, cuz not only did we have a lot of data, and we didn't necessarily have
the best track of where it was stored and how. We had a very um distributed architecture. It was doing streaming data processing. Um it was very redundant as well. So, if our European data center went down, we were fine in the US, and they could fail over to each other, etc. But now we had this legal barrier that got put into in So, besides having to figure
out and filter down 50 petabytes worth of data and then reshore it to the right part of the world, depending on where they are, we then had to go and update all of our application code to basically understand that there was that split between the different data types or the different data localities. And so, now every user inside of our system had a region associated with them
or a country associated with them. And then every time we had to go and access data or check whether they could access data, we had to now do another check. So, we had to go and look at what region that user associated with. We had to figure out whether that user's current region, they were authorized to access the data in the region they were trying to access.
And we had to go and further extend this authorization logic. So, again going back to a whole zero trust piece, we're fine at the firewall level, we're fine at the authentication level, we know who they are, we trust them, etc., etc. But on a per request or a per action basis inside the system, we had to go and add this logic. We had to add in the
these um architectural checks inside of the stack to make sure um that a user was doing what they were meant to do. And so, these checks had to go in and we could have used further scopes like this. There's maybe libraries like Castle or CanCan Gad or things like this that you could bake into language-specific frameworks. But um the model kind of required us to extend this
access control check across the application architecture. So, the next stage, uh things are going really well. You've built a product, you've scaled it out, you've repackaged it, you're in multiple regions of the world. Um and then you decide to go right, let's go and start selling to enterprise in all the caveats that comes with that. And one of the big challenges that you face with this is
you now have this kind of mismatch of organizational complexity. Cuz you as a someone building the you probably aren't in the enterprise business. You're not you know, fit 20, 40, 50, 100,000 customer users. You're probably a smaller team. You don't really have that complexity in place. And this happened to me. So, previous company, uh in the space of a month, uh we signed both Staples in the
US and Emirates Airlines in UAE as customers. And we were a company of only 80 people at this point. And both of those companies came around to us and said, right, we have to onboard 60,000 employees, about that, to your system. the IAM system, the role system we had inside of our product, consisted of three roles. Admin, manager, and viewer. And they were like, yeah, that's not
going to work. Uh we as an organization are set up to um group and have hierarchies of users based on departments, teams, cost centers, regions, that kind of thing. And so, you end up having to go back again to this access control logic and this authorization layer in your architecture and go, right, well, how are we going to do this? Do you have to go and start
looking at all your application code again? You have to figure out where checks are made and now you need to start potentially looking at what group membership someone has or do you go out to AD and go and grab their memberships from there or some other federated identity? Gets quite complicated, gets quite hairy, and you have to go and update all this authorization logic again in in
one And this code gets more and more complex. So, you get through that, you've got some nice hierarchy, you've you've managed to enable your enterprise customers, and you're now starting to get more of these and something that you might have to go through or might decide to go through is like, let's go and get our compliance sets. So, ISO 27001, SOC 2 are kind of the most
common ones at least in in the spaces I've worked in, which basically you go through and jump through a load of hoops and fill out a load of paperwork and prove a load of processes um to to the auditors that demonstrates you are working in a compliance and and you know, proper manner as a business and your applications and your architectures have the controls in One of
those controls that you need to have in there or the processes you need to have in there is around audit logs. So, every time someone does some system access inside of the system, not just on your infrastructure level, but even on the applications you're building, is you need to be able to not only have an audit log, but also prove that audit log is is accurate to
an auditor. And on a semi-regular basis, I used to get dragged to a dark conference room by an auditor and I had to go and demonstrate our access control logic and our audit logs uh to them in order for us to renew our certificate certificates and the first time I did this, I was sitting there trying to like grep through logs in S3 and it really wasn't
going well and you could just hear the auditor sort of tutting at me. Um and we had to go and kind of rethink how this is how this would work. And really the point where these audit logs and this authorization logic and that layer in your your zero trust architecture makes most sense is you can have audit logs all the way down that stack. So, going back
to that firewall, you can have the the the uh network requests going through your your uh firewall. You can see the traffic across your network through some of the the um service meshes. Your IDP, your authentication layer, will have the full authentication ceremony logged as well. But, then down at your application layer where you're doing the real action, the thing where you want to try to avoid
that disaster scenario, things sometimes or quite a lot of times get lost. And really, it's inside of that authorization logic where you're checking whether a particular user can do a particular action on a particular um should be allowed or not. And so, if you were to do this manually, you might have to go back to where we have that massive if else statement or getting more massive
statement. So, here now we're checking Is that going to work, actually? Yeah. So, we're checking whether someone is in the company domain, so they're happy, cool. Or, they're on the right package, so we're doing the packaging example. Then we're going to check what group someone's in, so handling that kind of enterprise hierarchy type problem. Then we're doing the region check to make sure someone has access to
the data in the right region based on all the data locality laws you need to comply And then the actual action will happen. And before we do that, we're going to capture an audit log. So, we're going to go write off to some say logging system that's say, "Cool. Right, this action occurred, it passed all the checks. Here's what the user actually did." And this is something
that you just have to have in place if you're going to service a particular type of customer or you want to go and get and there's compliance there. So, if anyone's working like insurance or finance or medical uh tech, which we have a lot of customers in, this is just a must-have. It's a non-negotiable. And you will get yeah, a lot of grief from procurement teams at
these type of businesses and compliance teams at these businesses and unless you can prove this. So, to do the actual implementation, you might go and look at maybe a logging library that's built into your particular framework of choice. You're going to send it off to some logging system somewhere and um maybe you're going to push it off to a SIM if you have one of those in
place by this point of your growth. So, at some point in this journey, not necessarily in this order, uh you may be reaching the inevitable discussion of architecture and "Hey, we should go and do microservices." or maybe it's another time to rethink it. Not necessarily because it's the best thing out the gate, but maybe you're actually at a point now where you maybe have some your main
applications in, say, Node, you have some data science stuff going on in Python, you have maybe a legacy component written in Java, let's say. Um and you need to start thinking like, well, we have this authorization layer to fit our architecture, but it needs to work across everything. And you then run into a bit of a problem because that logic, in theory, and most times in practice,
that you've written in, in this case, JavaScript, you now need to convert that and translate that into every other language under the sun that you're using. Um you know, if you have one web service in Go, another one in Node, another one in Java, every time this authorization logic change and that access control logic changes, you're going to have to get that ticket, firstly, find the right
team that knows that particular language and get them to re-implement the logic. And you're going to end up with n number of implementations of this access control logic, which is fraught with issues um and potential issues, I'd say. Particularly if the tests aren't rigorous or not being tested the exact same way. Um hopefully your integration tests will cover this. But, you're now going to end up basically
repeating work in n number of languages um for what needs what has to be and what should be standardized access control logic across your application. So, you're again never trusting, always verifying actions inside of the system. And so, this new layer of cheese, the never trust, always verify, um um again by NIST um as part of the zero trust architecture, they put it out, they they sort
of designed this pattern of externalizing authorization logic out inside of system. So, if we were to go and look at architecturally, what would this be? We sort of know at this point someone's gone through all the different layers of cheese so far, and we know who they are based on the API call, the gateway, etc., etc. And we know what resource they're trying to access. And we
know what action they're trying we need to essentially have this standardized thing, this authorization service, let's call it, where you can say, "Here's this user. They've gone through the firewall, they've been authenticated, we know their identity, they're on the right network, blah blah blah blah blah, but now we're actually at like runtime or or API call time or execution time, can this user do this particular action
on this particular resource?" That request that that payload would go to this authorization system. This authorization system would have loaded into policy. Uh, and that policy would be a standardized way of defining your authorization requirements, um, sort of an as code approach, let's Um, and that decision point would then evaluate that input against that policy and return back, at its core, basically whether that action should be
allowed or denied. There's other variants on top of this, but the primary thing you're trying to figure out is, do this action on this particular resource with this particular context? And so, going back to, you know, James Reason and him him his work in like defining SOPs, the way to think about your authorization policy or your access control policy is like an SOP, a standard operating procedure.
what is I guess Cerbos is an open source, uh, policy decision point. Um, we have a whole business around it as well, but everything I'm going to talk about now is completely open source. I'm not here to sell you anything. You can grab it off GitHub. Um, what you see on the right hand side is a Cerbos authorization, uh, resource policy. And here we're actually looking at
a way of defining a contact resource, I think of like a CRM. And then we have basically all the different rules that define all the different actions that a user can do upon that resource and under which condition those actions should be allowed. So we could have what's probably not recommended for production, but we could have a wildcard rule that says if the users has a role,
so your your IDP has issued them a role of admin, then all actions should be allowed. Pre-binary, not great, but allows us to do what's referred to as RBAC, role-based access control. But then when things get actually more relevant is where you have much more fine-grained or attribute-based control checks. So now we're saying that a user can read if any of these conditions match. And so we
we can check whether someone's in the and or their department is in marketing and the resource they're trying to access is active, let's say. And this gives you a way of defining your authorization logic as something that's versionable, testable, auditable, and more importantly not inside your application code. So it means you don't have to know Java, you don't have to know Go, you don't have to know
TypeScript or whatever to understand how your authorization logic works. This extracts it out into something that's a bit more human-readable. It's still YAML, love it or hate it in its current Um but it's something that is now an asset that you can actually go and audit and going back to getting your compliance cert and getting back going back to getting your SOC 2, etc. Having these things
to basically demonstrate in a way that isn't application code how that logic works is kind of a critical capability and allows you to truly get to a model of actually zero trust architectures. So how does this actually look in practice? So you would have your user. They are accessing your application, probably go through some sort of API gateway. You could do sort of a medium-grained authorization at
the gateway level where you want to check whether this particular identity can even hit this particular route before it even lets it through the gateway. But ultimately, it'll end up down inside of your application layer. At this point, your application layer really knows two things. It knows the identity of the user because the upstream to this is your IDP that's issued a token that say Um and
that can go off to maybe a directory to go fetch further context about memberships, groups, those kind of things. Um and returns back your application that token or that context. The second thing your application knows is what resource they're trying to access and kind of what action they're trying to do against it based on the request. So, they're trying to view an invoice. They're trying to edit
an expense. They're trying to download a report. And your application layer is the source of truth for what those assets are. So, it can go off to its own database or services to go and retrieve that from from the system. And then now, unlike before, rather than having that if else case switch style logic all over the place, you just package up that information. Okay, here's the
principal. Here's the resource. Here's the action. So, the principal is the user's identity or the API key's identity or the service account's identity or the workload's identity. It doesn't have to be a user or a human in that sense. Here's the resource they're trying to access and here's the action or actions they're trying to do. That gets sent off in a request to that policy decision point.
Could be Cerbos, but there's other projects out there. Uh Open Policy Agent is the CNCF project if you're interested in alternatives. Um that policy decision point has loaded into policy. Now, in our world, those policies could come from a database. They could come from a disk storage layer. It could come from a Git repo if you really want to follow the GitOps style workflow. And those policies
get loaded into that So, when that request of can this principal do this action to the resource comes in, the policy decision point will evaluate these policies. Before it does anything, it will create an audit log of the decision. So, this this principal tried to do this action on this resource and it was allowed or denied by this particular version of this particular policy. And you can
write that audit trail off to whatever you're using for your log collection. And then it returns back to the application. Typically, either that allow or deny. So, this action view should be allowed. Edit should be denied, etc. etc. Alternative Additionally, you could also return context for why decisions allowed or Uh and also it can return what's called a query plan. So, if you wanted to like pre-filter
your database results based on what access the user has, you can actually interrogate a service PDP, policy decision point, to get back essentially the where clause that is dynamic based on your policy to apply to your database query to pre-filter the data down and thus locking down what access a user has to to resources. Just one side point by implementations. Uh this model this pattern, I guess,
is becoming uh quite dominant thanks to sidecars, which I know the the tide is maybe turning a bit on sidecars and humanities, but one of the things that many of you with architecture hats on maybe worrying about is like uh there's now another service in my cool chain. Cuz authorization is in the blocking path of every single request. And you don't necessarily want to go into your
application and then bounce out to another service and back. And there's a trade-off to be made here. Um what we've seen is the advantages of being able to externalize and take a policy-based approach to authorization logic are kind of outweighed by a potential off-box or outer service call. But there's now architectural patterns like sidecars that allow you to actually run that as close to your application as
possible inside the same pod that kind of guarantees it's not ever going to go off the node and those sort of things. Um there's other alternatives where you can actually run it in process, which we support as well. And but this model is something that is now much more achievable thanks to the work of the orchestration layers that now exist. So, we go back to our zero
trust trust nothing, verify everything model. Before, we would have had all this if-else logic spread across your application code, which works, but it's a management nightmare and it's super hard to audit as well when it comes to compliance Uh by externalizing it all your application code is now reduced to essentially a single if statement where you go and do the check and you get back that allow
or deny. and you allows you to sprinkle this across your application architecture where it makes sense. So, do it at the gateway, do it at the service mesh level. You know, you don't now just have to authorize north-south traffic. You can actually authorize east-west traffic. So, it's not just requests coming in down to your services, you can start authorizing calls between services as well, not just based
on their long-lived identities, but based on the specific payloads of the specific requests that they're making to specific services using this exact same policy framework and really allowing to authorize every single call, every single RPC, every single operation that happens inside of your architecture. And you know, about 10 years ago Google published their paper about BeyondCorp, which I'm sure some of you may have seen, where essentially
the every single one of your services, in theory, you should be able to go on the internet and be absolutely fine. You know, you shouldn't be hiding things behind firewalls. If you're authorizing properly and at to this level of granularity, every single service should just be publicly accessible on the internet. But if you are properly authenticating and authorizing every single call, that's not going to be an
issue. So, many, many, many of Google's internal services actually have just like public domains. You just go and hit them, but you never you and I would never be able to actually access them because we haven't got Google accounts, right? Actual Google employee account. Same model and same architecture that you can start moving towards. So, there's advantages and challenges to this approach. So, advantage, big one is
like you get this logic defined centrally for how the authorization layer of your stack works. It means your policies can evolve independently of your code and you're very agnostic and to any particular framework or technology you choose. So, as your application architectures evolve and maybe as more devops teams, you bring out out this platform, it doesn't really matter what the application teams on top are building in.
They can plug into this cuz it's a more of a service based model. And because these policies are assets, you can go ahead and take a GitOps friendly approach. You can have tests around these as well. And you get that cuz there's an audit trail. But, you know, like I said, there's always a trade-off. There's another service you have to worry about in deploying or your DevOps
team have to worry about in deploying. There's a new DSL you may have to learn for policies, and there's a new component in the critical path. But, bring it back to the Swiss Swiss Every layer, as I said, I think there's solid solutions for out there in the world. The one that I believe is currently underutilized is this authorization piece. And hopefully you got a bit of
an idea of why authorization is so key and how you can actually go and implement it in a scalable manner. And actually bring you to a model where every single layer of uh is uh strong and monitored and controlled and making sure you never go from that situation to that disaster And with that, take any questions. Cool. Any questions? Can the policy decision point also be implemented
to reverse proxy? For example, providing permissions by the header in enrichment. Uh yes. So, every policy decision point, not just ours, but kind of any out there, is essentially a very simple API or an RPC where you say, "Here's the user, here's the action, here's the resource." And it gives you back the answer. You could then go and plug that into, say Envoy um or whatever your
proxy of choice is to enrich uh context. Um the other point where you see kind of enrichment like that happening is actually calling out to a PDP during the token exchange ceremony in an OAuth 2 flow. So, actually using policy to decide which scopes or claims go inside of a token. Um the idea when you're doing this properly is every single part of your essentially everything of
every single process in in the request from the client application down to your back every single decision made at one of those should be driven by policy, which are essentially version and the So, doing it at the IDP level, do it at the ingress, do it at the reverse proxy type level. Putting a reverse proxy um or putting a proxy in front of maybe an off-the-shelf or
a commercial product that you have internally. So, we have like a customer, for example, that puts Cerbos in a proxy in front of Grafana. So, they I can actually limit what users can do in Grafana to more granularity than they're able to natively with Grafana. So, think about it's the right mindset of where can I put in authorization and policy-based approach across every part of the stack.
But, essentially authorization code, you Uh uh uh we did that one. Oh, there we go. What are the main limitations of Cerbos that might lead to some decide against using it? there's the architectural decision that needs to be made of whether externalizing something like this is makes sense in your stack. I would say running a big monolith and are happy with that, then potentially this isn't the
right approach using Cerbos or anything else. Um if you've done the work to set up your uh so, proper domain-driven design within that within that monolith, then it's kind of seems slightly maybe slightly excessive to have an external one on top. Um but, I'd say that is one thing to consider. Um and the other one making sure you have kind of the uh rigorousness to agree on
what the policies are. A lot of the time the requirements we see get given to product or applications teams are just basically handed off to devs and say, "Yeah, you can implement these rules." without really a total agreement on the business side of things of what's going on. So, getting externalized authorization right is really dependent on getting the whole business by uh by bought in on what
those rules should be. Great. Is that it? How do you test your authorization spec? Do you have to write all tests yourself or is there a built-in support to generate tests based on specs? Um, there's tooling out there, but written both by us and other open source users where um, you can basically start taking a forms-based approach for defining your policy. Um, with that it will actually
generate the tests for you as well. So, it's almost like snapshots where you can sort of set up what the rules need to be and um, um, it will basically generate the tests on the back of it. So, as part of our like onboarding, you give us a list of all your resources, all the actions, all the rules. It will not only will generate a policy for
you, it will generate generate the tests for those also. And then speaking of tests, as well as um, those tests, we also have this whole CLI runners, etc. So, you can inside of your CI pipeline uh, control um, actually run the tests before they go anywhere near sort of your production environments. Does it support Oh. Does it support source IP checking scenarios with the IP provide some
level of authentication? if by that you mean basically checking the incoming IP address in your policy, that's absolutely yes. Um, IP address is a great example of the kind of additional context that you could use inside of your policy to make a Um, so uh, classic example we have is a lot of our users will go and put a blanket rule in which is if the if
it's outside of office hours and the request is not coming from an off uh, not coming from the office IP range, flag it and maybe deny access with various break glass scenarios as well. Um, but yeah, the IP address is definitely one of the key bits of um, context you can use. Great. I think I'm out of time. Thank you very much, everyone. I'll be around if
you want to chat.