About this talk
This talk by Steve Wade addresses the challenges of transforming a traditional banking organization's IT infrastructure into a modern, cloud-native architecture. He shares his journey of overcoming legacy systems that were fraught with technical debt and dependency on tribal knowledge. The speaker emphasizes the importance of collaboration across teams, focusing on insights from various stakeholders including operations, developers, and compliance leaders like CFO Sarah. He outlines a five-step transformation process, implementing GitOps with tools like Terraform and Atlantis to simplify infrastructure management. Steve highlights that successful transformation is not solely about the technology, but also about understanding the human side of change, facilitating buy-in from all team members, and ensuring compliance throughout the process.
Full transcript
Hi everyone, I'm Steve Wade and I'm a recovering perfectionist. So let me explain what I mean by that. I used to think that perfect systems or ones with elegant code, 100% code coverage, pristine architectures. Then I went on site and realized actually that was total rubbish. So let me start here. I was parachuted in to a British bank, centuries years old, and my mission was to drag
them into the digital age. But what ended up happening is we had this conflict, right, of legacy systems versus cloudnative dreams. I want to start this by me by you meeting Sarah. Sarah was the CFO. So Sarah's job was to make sure that we had full compliance across all of the digital assets within the organization. And one day she sat in a meeting room. I was in
this said meeting room with techn technological leaders across the organization. And she asked a simple question. She asked this. How long would it take to rebuild our production environment if we lost it all tomorrow? Do you want me to show you what happened? this. No one had an answer. No one said anything. Everyone just looked around at each other hoping that someone would say something. But what
we realized is we were operating on faith, not facts, faith that our deployments would work, faith that our documentation was accurate, faith that we could actually recover from a disaster. But before we get started, I want to talk about perspective. So, if I hold my hand up like this, you're probably going to say, "Steve, I can see the front of your hand, right?" But when what I
see from my perspective is I see my palm. And that's very uh cognizant of when we start to speak about and think about applications. Different systems have different users and they all have different perspectives. Operations see one thing, developers see something else, executives, they only ever see dollar signs, right? But what I want to do next is I want to try something interactive. So I want you
all to cross your arms. Feels relatively comfortable. Right now I want you to cross them the way. Feels a little bit difficult, sometimes challenging. Maybe you're getting a little bit frustrated with yourself. And what you'll realize is transformation journeys are not logical at all. They're actually very neurological. We have to get into the brains of the people that we're working with. The technology is actually the easy
part. So the transformation that we went on went in five steps. So we had this inherited chaos. That's what did Steve get up to when he first landed on site. Then we have this awakening moment. That was the Sarah question. We move to strategy, right? Because like any good bank, we have to have a strategy before we can actually execute something. We go on to the implementation.
That's the interesting bit as technologists. And then finally, we end up with the transformation. like what did we actually get out of the end this? So, let's start with stage one. This picture depicts perfectly what I entered. Chaos, right? People frantically running around trying to work out if the system was stable, if the system was online, and the system was working. But what we realized is we
I didn't enter technical debt at all. I entered technical bankruptcy, right? and we couldn't keep going on and having a system and an application and a platform that was running like this. So why was that the case? So they had no infrastructure as code to start with. Everything was built click by click. Sometimes the previous click was documented, sometimes it wasn't. What did that lead to? That
led to environmental schizophrenia. So all the environments were distant cousins, right? They weren't identical triplets. You didn't get any confidence going dev, staging, production. And then I called this um we ended up playing digital roulette. So every deployment was a risk. No one really knew if the risk was going to be successful or not. But it went a little bit deeper than that. We had this curse
of tribal knowledge, right? So understanding of how the system worked was in minds, not manuals. We had this theory of disaster recovery. So like we had a plan to build a plan someday, but as long as we told somebody that we had a plan, everything was going to be all right. And then we had this hidden cost iteration. So we were iterating so fast cuz they were
trying to get their ROI up that actually the ground beneath their feet was extremely unstable. So I want you to meet this guy. This is Alex. He's a senior infrastructure engineer or as I renamed him in my mind, senior vice president of skepticism. So, Alex said, he said, "Steve, what what are you talking about? We're going to move everything to GitOps. Are you crazy? Like, I've spent
15 years in this organization. I can SSH into any machine you want. I can fix problems within minutes. Now you're telling me that I have to commit pull requests and get them reviewed by junior engineers. What? What the hell? But what you notice with that kind of comment is that the conversation that I was having with Alex wasn't technical at all. It was all about his identity.
He was known in the organization. He was that problem solver. He was that hero, right? And I'm going to take all of this away from him. What what's he going to become? And this is when I realized this. People don't resist change. They resist identity theft. Right? When we start to go on a transformation, people realize and think that the person that they've known and become and
aware of in their organization, they're going to have to change. Some of them are going to embrace this change and some of them will move on because they don't want to it. But we talk a lot about resistance to change, right? As if it's irrational, but it's actually not. It's it's deeply human, right? We have to get into the hearts and minds of people when we go
on transformation. The the tech is the easy part. the people, the process, the way we put arms around shoulders is the things that's really important. So stage two was the kind of awakening. So this was going back to when Sarah asked that question and she said, you know, how long would it take? So this is me at night, probably 3 weeks in with diagrams all around my
desk. And then I just had to write this um these seven words on a piece of paper and I said I don't have the answers we do. Right? So I come from a sporting background and I'm happy to talk about that afterwards and teamwork is everything to me. And I realized actually in that moment that I was trying to solve everything myself and we're not going to
go on this transformation journey together unless everybody is built in or bought in. Sorry. So we needed to go in front of a whiteboard and get everybody involved. So Alex picks up the pen, starts drawing things. Then we've got some developers and some operations people and then you know our good friend Sarah comes into the meeting room and starts layering on compliance because we work in a
bank, right? So that's the first thing that they care about. So the next stage was strategy. This was not what technology choice were we going to do, what tool were we going to use. And it started with why. Why were we doing this? So we believed that organization will thrive if everyone can see and understand the system that they're working on. We believe that collaboration was fundamental
for transformation. So because we had our why, we then had our how. How are we going to do this? So we were going to use Git as our central source of truth. That was going to define our what. Then we were going to do desired state reconciliation. And we were going to change our what into reality. what was going to allow us to be able to have
an organization whereby everybody could collaborate and contribute to any layer of the stack and system without feeling like they were out of place. So stage four is the implementation. This is the bit that you all want to see, right? This is rubber meets road. We've got to build the plane whilst we fly the other one. How are we going to do it? And Alex said to me,
"Steve, look, I'm telling you from the get-go, application developers do not care about infrastructure. What are you going to do?" So I said, "I'm going to make it ridiculously simple." He went, "Okay, show me." So we started with Terraform modules, right? Think about this as our catalog. This was a collection of modules that the infrastructure or platform teams could leverage. As well as that we created application
specific terraform modules. So these were the modules that the application developers would leverage. All they didn't have to understand the complexity behind the scenes. All they had to understand was the contract those variables that they needed to be able to pass in. We worked in collaboration with them. So those variables were really really minuscule. They didn't need to know whether we were backing it up to Glacia,
whether we were replicating it across regions. They didn't need to care. They only needed to understand how these application modules worked. Then we had something that we called Terraform roots. So Terraform roots was executable Terraform code. This could be one module or many modules and they were executed in some kind of order to rebuild the platform from scratch. This is an example of one of the application
modules, right? So this was a log ingestion. So they needed to ingest some logs. So they needed a bucket. So they would use the application S3 module and they would use an IM ro to be able to access said S3 bucket. Relatively trivial. If you look at the top there, they only need to pass in three variables, app name, platform, and team prefix. Really, really simple. If
they they didn't need to know how it was working behind the scenes, right? The easier we can make this for people, the less likely they are to resist. So this moves us on to developer experience. Great, you can you have Terraform roots, you can deploy them. So what? But we needed a way that was familiar to developers, right? And developers love pull requests obviously apart from Alex.
Um so we leverage a tool called Atlantis. So what does Atlantis do? It does Terraform by a pull request. So you create a pull request for your Terraform changes. It automatically plans those changes and prints them as a comment. You have to approve them or reject them. If you get that approval, lucky you. um you you get to merge um the the pull request, right? So the
pull request runs, you specify Atlanta supply, Atlanta supplies the pull request for you, and then if the pull request is successful, the code gets merged, right? All automated, all visible, there's an audit trail, all the good things that banking and finance love. Then she's back again. Don't forget about compliance, Steve. If you don't do compliance, then you can't do any of this. Okay, Sarah, no problem. So
we implemented a tool called fugue regular. So how did this work? Well, we still did the same Atlantis plan, right? What that did is rather than printing it to the screen in the comment, it also printed it and stored it as JSON. Then we run our policy checks on top of that. So we had a repository called um policies that started really really basic to start with,
got a bit more complex when security got involved as you can probably imagine. those will run pass you move on fail you don't you go and moan at somebody in compliance maybe they let you bypass this maybe they don't and then finally the same thing happens you perform Atlantis supply so we get this good massive green tick when Sarah says are you doing compliance for IA so
that's infrastructure what about repositories so we started with uh customize I would say put your hands up but I generally can't see anyone in the audience barely so if you're familiar with customize The way it it works is that you have base implementation and then you have patches that go on top. So think about all of your base configuration as your non-environment specific config. And every patch
that we have is going to be an environment specific patch. So our repository directory looked like this. We had underscore base. Then we had a set of files or folders sorry for each cluster that we had. From a platform perspective, we had K8 releases platform. These were for all of the platform workloads. This created the cluster for the product teams to deploy their applications onto. We had
cluster resources. All those things if you came to the workshop, some of these you'll uh you'll heard me talk about yesterday. And then we had components, those third-party applications that we need to deploy onto the cluster to provide observability, certificates, etc., etc. So in our base configuration, this is an example of us uh deploying Graphfana. We specified the tag that we wanted in the base configuration. We
didn't want you to be able to override that in in the patch. We had a different way of doing that. And our sandbox patch was very simply just the ingress and the certificates that we required in order to be able to deploy that to the sandbox environment. You can probably imagine what dev and production looks like here. Then you have to add that to your customization and
away you go. The GitHubs tooling is going to be able to figure this out for us. So Steve, I know bunch of you asked this before yesterday. How did you do uh secrets? So we used misillas and the reason why was because flux which is the githops tool that we chose has a nice integration with soaps. So yes, we encrypted our secrets and stored them in git.
Please don't kill me. So the repository structure looks very very similar right we don't have customize as our base now we have this directory called secrets we have a soaps file which essentially says for a specific directory in this repository you will encrypt and decrypt using a KMS key that KMS key is an environment specific key that flux would be able to leverage directory structure very very
similar hopefully this is getting a little bit boring because we want it to be boring right want everyone to be able to contribute to all of these repositories without having to think, right? If we can keep the cognitive load down, we make it easier for people to be contribute. Then we have some um secret here. Again, it's completely unreadable. It's like no other secret that you've seen
before. Kubernetes secrets are B 64 encoded. Whatever it's not really a secret, is it? Um this is this is actually a little bit more secret. And then from a flux perspective, we had a git repository. We've talked to K8 secrets and then we simply set this decryption provider to SOPs. And behind the scenes, Flux is doing all the magic for us. So we've done application uh sorry,
we've done platform workloads, we've secrets. Now we get on to the controversial one, right? The product teams. You can all imagine what the product teams are saying. I've got my way of doing it. We're going to do it like this. Product team A wants one way. product team B wants another way. So what did we say to them? We said, well, we've dog fooded this approach. So
what we'll do is we'll start you off with this template repository and you can build on top of this template repository. So what you'll notice is the directory structure looks identical. And we said to them, look, we want you to be able to contribute to our um area of the system and by that means we should be able to contribute to yours. So if we maintain that
consistency, everyone can collaborate together. So this is an example Helm release in the in the base repository. So you'll see here if you're familiar with Helm, we have a Helm chart directory at the uh section at the top which essentially says the Helm chart we're going to use is backend v2. The version that we're going to use is 0.1 or 1.0 uh 46. And then we have
values at the bottom. Those are the values that we're going to pass into said Helmchart. So what you'll notice is that we don't set an image tag here for the application development team. And the reason why is because the way that they deploy their applications is they're going to bump the image tag, right? So those image tags actually need to live in the environment specific config because
they're specific to that environment. So a sandbox patch relatively trivial. They have some environment config for their sandbox and they have this interesting way of doing images, right? So with semantic version and tags and you start to see this JSON comment at the back, they then add that to the customization. Once they've added that to the customization, away we go. So kind of back to image automation,
how are we going to deploy images between environments? So, Flux has this concept of um image tag automation. You provide it an image repository. So, an image repository is essentially just where am I going to be able to obtain these images, right? So, for this image, I'm going to this image repository and I want you to look there every 5 minutes for new tags. Then you have
an image policy that says okay for that image repository what type of tags am I going to look for? So I'm going to in this instance anything that's 5.0.x. Right? Reason why they want to they want to implement this. They don't want to be rate limited because they're searching and looking for the whole entire uh tag structure within that repository. A little bit unreadable. Don't go into
this in too much detail. Essentially what this is doing is saying when Flux sees a new image and updates the image in the cluster, what is it doing to git? Right? Git is our central source of truth. So it's essentially saying I'm going to send a commit message back to git with the changes and here is what it's going to say. You have a number of different
ways in your comment blocks to do um image policies. You can do policy name, space, and policy name. Or you can make it more complicated if you want to. So, here's an example config of us being able to consistently deploy any image that starts 5.0 something that's running um and deployed to this repository. So, we now have a way of being able to do automated image updates
at the flux level. But how do we actually do this in CI? So, we started with GitHub. We were using that for everything. The developers chose Concourse. They wanted to leverage that, not my choice. We ran it the same way that we deployed every other application, right? It was a Helm release that got deployed to Kubernetes with Flux. They did a Docker build. They tagged that dev-
semantic version. They uploaded it to the image registry. Meanwhile, Flux is sitting in the cluster and it's reconciling using those image policies um the repository where exists. It then patches the Helm release inside the cluster, performs a cubectl apply to make sure that change actually happens and then finally commits the change back to Git. And it looks a little bit like this. So we can see from
here we did an automated image update at a certain point in time and these were the changes that were made right have the whole thing audited who made what change when when there's an incident we can go back to the git repository and the git history is going to tell us then this like important thing called testing comes it's it's honestly over you know it's not that
important really so concourse is sitting there and its job is to wait for image updates or wait for changes. And anyone that came to the workshop yesterday will remember that I talked about the importance of label taxonomy and I kind of went on about it until you were blew in the face. This is so important because what we were doing is we were um Concourse was constantly
checking for the new version to be deployed. It was looking at those labels. Then we did this glorified thing called testing. Kind of pointless. We'll gloss over that. And then if we wanted to go to UAT, we simply just rettag the image with UAT- semantic version. Rinse and repeat. We now have a nice process to be able to onboard multiple applications across the estate to multiple Kubernetes
clusters. Oh god, this guy's back again. So, Alex says, "Yeah, that's great, but what about Kubernetes changes?" All right, Alex, I've got you. So anyone that's familiar with Kubernetes will remember in Kubernetes 1.16. They made many many API version changes and everyone's applications and deployments broke because the API versions changed. It's a horrific memory. I would never want to live that again. So luckily Fairwinds created this
tool called Pluto and its idea behind it is that you specify I'm moving towards version 1.25. Please run all of the deprecation checks against my YAML files. Return me back success or failure. So what we were doing is we were doing N plus two, right? They had a little bit of legacy. So we would be able to validate before we committed the changes to Git whether we
were successfully going to be able to redeploy everything on the new version of Kubernetes. It's just more confidence. Again, we did the same kind of thing. We used cube conform. This validates the schemas against the Kubernetes SK uh schema for that for that resource. Again, this is just more confidence. This runs on policy checks. We could we were making this more complicated over time. To start with,
we just used something very very basic. So we had this in all of the repositories, the platform engineering repository and all the product team repositories. Everything was standardized. Everybody knew how everything worked. So let's kind of put this all together. What does this look like? So we deployed the infrastructure to AWS using Terraform. We put Kubernetes on top. The first thing we deploy is K8 releases platform.
That's the starting point. That's the bootstrapping mechanism. We pointed at some directory within our K8 releases platform. Right? In this example, it's customized dev because we're re we're bootstrapping the environment. That in turn goes and deploys other repositories, right? So, K8 releases corp, the one that makes money. K releases secrets, K8 release uh CFKA topics, network policies because security and then finally some kind of services because
if you're not running a service mesh, you're not really running a platform apparently and all of that was at some directory and environment prefix, right? So it was it was very very boring and we had a simple simplistic way in order to be able to deploy any application or any workload to Kubernetes consistently. Right? We go back to the same thing. The patterns and the paradigms are
more important than the technology choice. The technology choice is just a catalyst. Right? If we get the patterns and paradigms successfully implemented, everything becomes very very boring and there's not this friction between you've created the platform and I don't know how to use it. So systems reflect the organizations that build them. So if I do this, why did the pen hit the floor? So the pen hit
the floor because of gravity. Can you see gravity? No. But it affects everything around us. And the same is true in organizations. The culture affects how we build systems, how they work, how people collaborate with each other. The technology is just a catalyst. there's a whole hidden agenda of things that have to happen in order to be able to make this possible. So stage five was the
was this right? But what we noticed was that we had specialists in certain areas, developers, operations, security and compliance, and they were absolute experts in what they did. But we've drawn such sharp lines between these disciplines. But what we realize is the key things that we do within an organization is actually how they connect between those lines and how they work together. So this wasn't about how
did you configure your Terraform modules or how did you do your GitHub structure. It was about breaking down artificial divides between certain members of the organization. So here's some numbers. Everyone loves numbers and you know how how well did we do? So from a deployment frequency perspective, we went from weeks to 10 times daily. Our lead time was from days to minutes. Our meanantime to recovery was
about a 72% reduction. The bottom one is the most important one. We went from uncertainty to a few hours with confidence. Those final two words are all that Sarah wanted to hear. With confidence. Could we prove to her that we could do what she wanted us to do? Yes, we could because we were regularly repaving. So we used to tear down entire clusters, rebuild them back up
again, make sure everything worked. The sandbox environment never lasted more than 30 days. The production nodes never lasted more than 45 days. So that's that's great, right? We can we know that we can do it from a cluster perspective, but then we ended up doing it from a whole entire account perspective. We said, what happens if we lose the whole entire production account? Can we redeploy somewhere
else with confidence? Right? All right. And these are just essentially cards that I can play on the table in certain meeting rooms when people start to ask me difficult questions. I have the answers. I can point at the data. So the journey was technical but actually the transformation was deeply emotional. So we started with fear. If we remember Alex, Alex was fearing for his job, right? He
was like, Steve, you're gonna you're gonna completely change my identity, right? And transformations start with people thinking to themselves, am I still going to be relevant in 6 8 12 months? Then they have this frustration. Remember I got you to recross your arms? You a little bit frustrated. It's exactly the same thing. People are learning new tools, but they're not learning new tools at all. They're learning
new habits that have been ingrained in them for months, years, or even decades. Then we had experimentation, right? We started to have fun. How do we how do we deploy and configure this repository? What works? What doesn't work? And then finally, we have mastery. It's where Alex starts to become the expert. He's making he's making interesting changes. He's saying, why don't we add some more type of
information integration into our CI, right? He's starting to get more involved. And then this is when I realized this perfect systems aren't perfect if no one uses them. What what what do I mean by that? So, if you remember before I said I was a recovering perfectionist, who cares about how many automation tests you have? Who cares about how clean your code is? If no one's using
the damn thing, it really doesn't matter. This is the key takeaway. The purpose of technology is to serve people, not to impress other engineers. Right? Our job as technologists is to serve the people that are using the technology. It's not to tell the next person or in an interview that you know we we're running 1.25 or you know and actually we've moved to ambient mesh and how
long's ambient mesh been around 2 weeks. Yeah, that's kind of super impressive but actually if it adds no value to the organization like really what's the point? So we're back to this guy again. So I saw Alex in the hallway and he came up to me and he had his eyes wide open and he said, "Steve, you know what? I get it. I don't want to be
the guy anymore that has to fix the problems at 2 a.m. I want to be the guy that builds the systems that never break at all. So, Alex had started off this journey as the biggest skeptic, but he'd become our biggest advocate. Right? His his identity hadn't diminished. It had just changed and adapted. He'd become he'd moved from the hero to the mentor, right? He'd become a
leader. He was the one that was going into these product teams and explaining how we were driving change and the reason why you should become and start to adopt the new platform. This was deeply emotional for him. It had nothing to do with technology at all. He needed an arm around his shoulder and he needed support. So, I'm going to challenge you. I want you in the
next week to come up with that question that gives you a moment of silence. If you went into a boardroom and asked it in your organization, once you have it, I want you to start working on it. I want you to create a repository. Don't worry for the perfect moment. Just just start. Create a repository. Create a way of doing something. And I want you to Sorry,
my voice is going. Your commitment is transparency because I had one unanswerable question and you can see how it transformed a whole entire organization. So imagine what could happen if you knew yours and you just started working on it. And if we collaborate openly and transparency in transparency, sorry, we can innovate at a much much higher and faster rate. So the conversation doesn't end here, right? It
doesn't just end with me talking on stage. The conversations continue out in the hallway. Come and come up to me, ask your questions, tell me about the unanswerable questions that you have in your organization. The most interesting conversations don't happen in here. They happen outside with the people that you're going to connect with over the next couple of days. I look forward to speaking to the rest
of you and thank you very much for your time. That was wonderful. Thank you so much, Steve. Um I think we've actually just um hopefully we've got a few questions either from the audience or on Slido. So hopefully you can see the nice big QR code here. Um, if any of you have already asked stuff, I don't know how tech Have we got any questions on the
slideo so far or? Yeah, we've got one here. Oh, fantastic. So, why flux instead of Argo? So, there's we talk about this uh this divide, right? There's there's people that are very passionate about Flux. There's people that are very passionate about Argo. I'll answer the question, but the technology is just a catalyst, right? You've got to make sure that people are embracing what you're trying to achieve.
The reason why we chose Flux over Argo was because of the way that Flux provided bootstrapping, right? We all we needed to do was in our Terraform code create a Helm release, point at our repository, and Flux was just going to take care of the rest of it. With Argo, the bootstrapping is not as simplistic as that, right? If we're spending a lot of time working out
how we bootstrap a cluster, it's a waste of time. Right? As a platform engineering team, our responsibility was to provide value to our customers and our customers are our developers. The quicker and mo and more boring that we can make the bootstrapping mechanism, the better for everybody. So that's why we chose it. We did a trade-off between the two of them. We did a PC and we
found that the bootstrapping of Flux was easier than the bootstrapping of Argo and Argo basically won out. readers as recover for a while. Do you have a backup plan for deploying things like from local? Yes. So, wasn't on the slide deck. Um, but you can have git oops, I like to phrase it, right? Which is like when when git stops working. I called it git oops. So,
if you noticed before the directories that we created and the structures for the repositories were using well-known tools. We were using customize. We were using Helm releases. So we had a way in a break glass scenario to be able to run customize build point it at the cluster that we wanted to basically do this horrific pipe keepl apply-fash and it would deploy everything in that Kubernetes in
that repository at that directory into the Kubernetes cluster. So we had a way out if we needed to. And also the other important thing is flux is essentially doing exactly the same thing in the reposi in the cluster itself. So if we could do it outside of the cluster and inside the cluster we had more confidence. And when we were running CI we were doing exactly the
same thing. We were doing customized build directory structure. We would pipe that to a m monstrous YAML file and then we would run all of the CI integration tests against those outputs. Is there any of again this is choose your own adventure right in in my uh opinion personally I want to make the pull request system and process as streamlined and as boring as possible. I don't
want someone to have to work out whether they need to whether they need to create a pull request that actually is merging to main dev production or staging. Right? I didn't want them to have to worry about that. We always just created a pull request that would merge to main and we had a directory structure. That's it. We were tagging off the back of that um off
of main when we needed to. But I wanted to make the well we wanted to make sorry the PR process as streamlined and as simple as possible. Hence why we used folders not branches. Any fully automated CA pipeline deploying up to production environments. How does it comply with bank policies? So we were running policies at multiple stages inside of this process. So I like to talk to
people about before Kubernetes, at the door of Kubernetes and post deployment. So kind of in Kubernetes. So we were running compliance at each stage. We would run compliance before you got to Kubernetes. So running that in CI, we would have an admission controller and I kind of talked a little bit loosely about this at the workshop yesterday, which is at the front door of Kubernetes. So when
that YAML file hits the API server, we were running compliance checks there to make sure that we were happy. And then finally, we would have run longunning jobs that would be running compliance checks for things that were deployed to Kubernetes. And all of that kind of fed back up to a single management system that Romano was kind of talking about where we would have a single pane
of glass view on those compliance reports. Oh wow, you've gone for some controversial ones. Um so we started with sealed secrets. Um but we moved to SOPs because it allowed the ability to have fine grained access on who could do what with secrets. So who could encrypt secrets, who could decrypt secrets and then who could decrypt secrets in production, who could encrypt secrets in production. And we
just everyone was had access to um to AWS using IM. So we just use IM roles that you would have to assume in order to be able to uh encrypt and decrypt secrets within that secrets repository. So we just had a bit more control there. And then when somebody left the organization or joined the organization, they just had the same process. They got the same default set
of IM rolls and everything just kind of continued to carry on working. And it allowed us to be able to roll the IM ROS uh sorry the KMS keys. And we actually made it a little bit more convoluted than that. We brought a GPG key in as well. So you needed for production you needed a GPG key and access to the KMS key. So it just gave
us a little bit more access control and ability to be able to do things a little bit more finer grained. Did you see movement from KH providers? We actually saw it the other way around. So we we had a lot of people that were bu trying to build homegrown I call them homegrown. weren't really homegrown uh Kubernetes clusters on the cloud providers, but I was talking to
you before, right, about bootstrapping Kubernetes clusters. That crap's boring, right? What we want to be doing is building a platform on top of this. Kubernetes is not really a platform. It's a it's a provider for us to build a platform on top of. So, we wanted to just leverage EKS for better or for worse at the time. um we leveraged EKS and we focused on bootstrapping that
cluster ready to deploy configuration to it. We didn't want to work out how you were going to do cororum incd. We didn't load balance master nodes. We just wanted to give give that away and give make that somebody else's problem. We wanted to focus on how did we autoscale nodes? How do we how do we handle demand? How do we autoscale applications? Why is this application only
running one instance and we needed to run three by default? They're interesting problems that we need to solve to make the platform successful. Not choosing your own adventure with the Kubernetes operating system or the Kubernetes uh cluster and how you build it. That wasn't important to us. That was just a provider essentially. How organization changes to adopt AI? Could you share? Um I I haven't done a
lot of adoption of AI within organizations to be honest with you. I I spend a lot of my time at the platform level. I think right now there is an interesting take on what's AI going to do in terms of platform engineering. How's it going to impact it? I think it's going to have a heavy impact on security eventually. We're not kind of there yet, but I
haven't worked with teams that are leveraging AI in order to be able to do certain things yet. This was very much we had a pattern, we followed that pattern. We didn't really have AI in the mix. Did you have multiple product branches or just one? So, we had we didn't have uh branches. Everything worked off of the main branch. It was just a directory structure. So, to
deploy the production cluster, you would just point at customize/ production. Everything ran off the main branch. The only thing that we would do is when we were doing a major Kubernetes version, we would tag the repository with the current version of Kubernetes that we're running. And then after we've successfully deployed, we would do the same thing again. So our tags on our repositories are really just the
version of Kubernetes we were running. Interesting question. um when the classic consultant answer when you actually need it. So what do I mean by that? A lot of organizations go to the shiny tool, right? And they see, oh, everyone's talking about service mesh. They went to CubeCon, everyone was talking about service mesh. I'm going to implement a service mesh. But all they actually need is mutual TLS
between their services. But with certain tools, STTO being one of them, you can turn on all the knobs and dials and you can end up sitting at an aircraft simulator trying to work out what the hell's going on. But actually what you need to do is just turn on the pieces of the puzzle that you need to solve the problem that you need to solve, right? And
then if that's mutual TLS, just enable mutual TLS with STO or some other uh service mesh and then add bits and pieces on top as you feel the need to do them. It honestly depends upon the requirements. I wouldn't say there's a there's an opinion on when to start. I would say the most important thing to do is have a running stable production environment that's you're deploying
to regularly. You understand how it works. You've gone through a number of incident scenarios. You can fix those incident scenarios. Maybe you're running databases on those clusters and then start thinking about adding that additional layer on top. Foundational building blocks, right? Let's get the basics working. Let's build on top of that. And then eventually service mesh is the kind of super cool stuff that can do real
fancy things. But if the ground beneath our feet isn't stable enough, then why are we doing that? We're just doing shiny object syndrome. Yes, I have um for both. So I see pros and cons to both. From an external secrets perspective, it's great. you're you're using the un the uh the upstream cloud provider to be able to provide you those secrets. But what ends up happening is
people need to know paths, right? Where am I going in the external secret to be a or the the parameter store or the secret store? What path is my secret at? And then how do I get that path into the um what's the word I'm looking for? How do you get the path into the configuration? It's more things that I have to worry about, right? And then
how do I create the secrets? I don't want everyone that has to create Terraform to be able to create secrets and have to know how to use Terraform and pass the right things in. And then you know what happens if I want to put a file in this secret? I don't want to put just a basic like a basic string. There's many many things that we have
to think about here. What what we wanted to do is empower developers to be able to create the things that they wanted to need or needed to be able to create and kind of get out of their way. If it's like, oh, I want to create a secret. It's simply just an API token. file. Well, you need to go to this Terraform repository and put this in.
But please don't put this in git because if you put this in git, this is the actual secret, right? And there's you can just imagine the kind of beep show that I'm going to get in when people start to like commit secrets to Terraform. It's just going to be an absolute nightmare. So, we kind of kept that completely separate and we just made sure that we were
doing everything with with soaps. So the compliance and security tool that we were using for Terraform was called Fugue Regular. R E G U L A. They have a basic building block set for the cloud provider that you have or leveraging. And then we added on a bunch more because we were running a bank and we had a whole compliance team and they we needed to get
them involved and invested. So rather than it just being a tickbox exercise, we essentially automated the process and we provided them with the ability to be able to see that we were meeting their compliance needs. Lovely. Right. Well, thank you so much, Steve. That was really interesting. Um, and uh, I think we'll just give him a big round of applause and a thank you for giving your
insights.
More from this event
See all 58 talks →
Halil Ibrahim Kalkan: Building a Kubernetes Integrated Local Development Environment
45:20
Paco Orozco: Growing at the Edge: Doubling Traffic While Changing the API Gateway
45:03
Viktor Vedmich: Ideal Blueprint Versus Reality for CI/CD Pipelines
46:03
Koray Oksay: Continuous Deployment: The GitOps, The Pipelines, and The Ugly
43:03