KubeCon + CloudNativeCon Europe

Scaling Platform Engineering: Lessons From Europe’s Larges... Cat M, Stéphane C, Anna K & Gayathri T

29:10 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk focuses on the challenges and strategies involved in scaling platform engineering within large organizations. The panelists, all experienced professionals in platform engineering, share insights gained from their respective organizations, including DKB, Amadeus, and AWS. They emphasize the importance of understanding user needs and adopting a user-centric approach when developing platforms. Various factors such as user adoption, organizational changes, and the need for automation and orchestration are discussed as key drivers for scaling. Additionally, the conversation delves into the social-technical dynamics that influence platform success and the role of artificial intelligence in enhancing platform operations and user experience.

Full transcript

Hello everyone. How are you all doing? Awesome. >> Thank you. Yes. Um so first thank you for all joining us because we were a little bit worried half five on day two everyone was going to be too tired, but uh it's lovely to see you all here joining us to talk about scaling platform engineering. So this is a topic that's a little bit near and dear to

my heart because I read the original Gartner report when it said that 80% of organizations were going to have a platform engineering team by 2026, I think it was, which is super exciting, but also that means a lot of organizations now have platforms that they need to scale to prove their success. platform engineering's not new anymore. These big big organizations, particularly across Europe, it's very relevant for

this conference today, have lots of challenges. So I've gathered three excellent people who have worked in some very large organizations and have dealt with platform engineering for a while to hopefully share some insight as to how you scale effectively. So to begin, let's start with a couple of introductions. Uh so I'm Kat, I'm product manager at Zantaso, and I'm the least interesting person here today. Uh so

I'm going to let the panelists tell you a little bit about themselves, how they get started in engine platform engineering, and maybe a little bit about the teams they work with today. So Stefan to my left. Uh hello. My name is Stefan de Cesare. Um I'm a platform engineer specializing on user experience on the developer experience at uh DKB. So DKB is um online bank in Germany.

So we have about 5,000 employees. and I have a more of an operations background, so I getting very interested in DevOps because I'm more interested in what is the use of operations, not only what the technology behind it. and in the past uh I worked as a professional services engineer for VMware, so helping companies uh automating their virtualization infrastructure. And I also worked as a consultant at

at Accenture, so helping company uh setting up DevOps technically but also organizationally. So, I've been also in contact with different kinds of platforms in different stages. Thank you. Anna, how did you get started? Hello. Uh I'm Anna, and I'm platform architect at Amadeus. For those who doesn't know about Amadeus, Amadeus is an IT company in travel industry. So, if you have planned this trip for KubeCon, you

probably booked a an airplane ticket, you have your reservation at the hotel, you checked in your luggage. There is a strong chance that you have been using Amadeus services behind the scene. Um if I need to be honest about how I got into platform engineering, I think it's because I'm lazy. Um I always was looking for a ways for me to be more efficient in my day-to-day

as a back-end engineer. So, I was first automating everything. I was abstracting the code with the workflows and the rule engines. And then I figured that if I joined platform then I can help not only myself, but I can help also other people. So, I started as a as an architect at Amadeus responsible for CI/CD and observability scope. So, Jenkins, Artifactory, Argo CD, Splunk, Argos, you name

it. Those were the tools within my scope. And then I switched to something slightly different, which is central infrastructural automation and orchestration. Basically, uh the teams I'm working with are responsible for delivering platform as a service. So, they would provide the core infrastructure, the Azure subscription with a network layer, and then we are responsible for deploying middleware and the databases on the top of it. Uh and

if I need to talk about numbers, uh in Amadeus, we have 600 applications and approximately 200 platform teams. And in terms of people, uh we have 10,000 core engineers, among which 8,000 are application DevOps and 2,000 are platform engineers. And of course, the platform is not consumable only by the application DevOps. It's also consumed by the platform itself, as well as our external customers and merging acquisitions.

Thank you. And then, last but not least, Gayatri. Hi, everyone. Good evening. Thanks for being here. I'm Gayatri Thyagarajan. I currently work for AWS. Uh I've been in platform engineering for about 8 years, starting at Expedia Group. Um I mainly built data platforms. So, um scaling existing platform capabilities, which is mainly around provisioning Kafka, and then moving on to actually building zero to one uh data quality

platform from scratch, but for a very large-scale operations. Um and then moved on to AWS. Uh again, very similar domain. Um and this time, managing the Kafka service for AWS, which is MSK. Um but to compare the scale, where before in Expedia I was managing uh a platform of about 10 clusters, suddenly this is about 50,000 clusters and 150,000 brokers. So, you can imagine there are some

very interesting problems that manifest at such scale. Makes up for very good war stories. And then uh currently I'm working for a different service team and that is AWS Resource Explorer. Again, this is service that offers search power search for all the AWS accounts. So if you go and navigate through an AWS account search bar, it actually hits that service in the background and that's enabled by

default since last year. So that's about 37 million very interesting stories to share today about scaling platform. nice to be here. Okay. So smart people. Gatri, you spoke about moving from zero to one platforms and then scaling them from that. One of the first things that we think about when scaling platforms is like when do I do it? What's the thing that means I need to start?

What have you seen across the platforms that you've been working on? Yeah, I think that's a very interesting question because I don't think anybody starts out thinking I'm going to be building for 10x scale in the future, right? I hope nobody does because it's a bad idea. It's expensive. >> Yes, it is. Because you don't know what you're going to scale up to and what kind of

challenges you're going to run into. So similar way I started out with building for you know one particular brand or a small group of subscribers which then exploded in in scale. So for me when I reflect back to when that became from building a single platform to actually scaling the platform, I generally categorize them into different drivers. Right? So to start with it's basically driven by user

adoption be it organic where it slowly creeps up on you or it's you know like a sudden blow up where you have a huge demand from millions of autonomous agents suddenly hitting your platform uh So uh that's one of the major drivers and most commonly seen. And second one is where your platform is is evolving, right? So, uh based on the same primitives, you're building up more

capabilities. So, obviously your platform needs to scale accordingly. And finally, it's uh organizational changes. So, you probably your organization um has consolidated their platform capabilities across multiple subdomains or um they have merged with another organization or um you know, acquired another organization. Again, you see all these different flavors. And in each of these, the challenges are different. Sometimes it's technical, sometimes it's sociotechnical. Um so, it's it's

important to identify what those challenges are and uh resolve them one by one. And that's why it's actually a bad idea to build for scale from the get-go because you never know what you're going to run into. Yeah, it's it's a it's a very common thing that we see in is trying to figure out which one of those patterns is most most likely to happen in your

organization. So, do any of the other panelists that any resonate with you particularly? Of those challenges, organic, capability, or external business factors like mergers and acquisitions? Yeah, I think that I I I found it interesting what you said about the the social technical part because I've often seen that the challenges of um scaling an organization a platform team is uh is a lot more the um the

the social technical part than what is really technical. That you have to uh when you're a small team of six, seven people, then everybody has the same view has the same view of the information. And as soon as you get a bit larger than 10 people, it starts to become difficult to to synchronize, to have the same message to the outside, uh to provide good information about

the platform. And so, that's for me, it's very often what's the the the main problem when the when scaling. Yeah, the devops to platform engineering cycle is one that I think we've spoken about before quite a lot of the most sort of typical pattern that we've seen when people are scaling it. Um but no, that's how it happened with me and my teams. Um you also mentioned

that sort of there's a social technical side, but there's also technical challenges. And Anna, I imagine with the thousands of developers that you're supporting, you've got a thing or two that you've learned about those technical challenges. What was the journey like at Amadeus? Indeed. Uh well, context is important, so let's start with that. Uh Amadeus was around for almost 40 years now. And what happened several times

in the history of Amadeus is that we would um develop our applications, we would hit a wall with certain limits, and then we would need to search for something else. So, we started with a mainframe. We reached the physical limit of the mainframe, and we migrated to our own data center in Erding. And then we started reaching the limit of the data center, and we started our

migration to the public cloud. So, when I joined the I was the part of the well, of the journey when we were migrating to the public cloud, being Azure. And one of the first challenges that we faced is an ability of creating quickly a big number of environments to manage. For the record, today Amadeus has 190 Azure subscriptions. It operates on 200 clusters, and it runs on

800,000 VCPU. So, we have a lot of volume to deal with. And so, to tackle those challenges, well, there are several dimensions. Uh one of them is we are heavy users of Kubernetes. So, we work closely with Red to tweak ACM, so Azure cluster management, to fit their models' requirements on the scale and be able to have workload of different sizes depending on the application that requires

that cluster. And the second thing that we did is we implemented our own solution for automation and orchestration that would allow us to provide this platform as a service to our applications. Once we've done that, we hit the challenge number two when we realized that not only we need to create environments, but we also need to maintain them. So, the day two was actually the bus killer.

We realized that we are not able to perform efficient day two operations. The way we designed our automation framework is a bit of a monolithic golden path. So, when a middleware or a database would need to provide a feature to their users, they would need to go through this central framework to deliver that feature, which would then go through the central validation cycle to validate the integration

of all of the components together, and then it will go through the central team to roll out those changes to our users, which is obviously not a good idea or you are simply not able to deliver the day two capabilities efficiently. So, this is lesson learned. Don't invest in one golden path. Yes, it sounds like you've had to move to a slightly more generic model where you've

got different ways of interacting with the platform and different paths for people. Exactly. One of the learnings is that what we want what we are doing actually right now is we we are rather investing first of all, we are trying to split the monolith so that we separate the core infrastructure deployment from the deployment of the components on the top of it, so the middleware and the

databases, so that those components can have an independent life cycle. So, they can deliver the features to their users much faster. And yes, we are investing in a set of appropriate golden passes. So, it's important to know who is your user and define the set of golden passes according to that. Golden paths came up a lot. Was anyone at Platform Engineering Day on Monday? Oh, quite a

few of you. Yeah, it came up a lot. And another thing that came up with those golden paths was platform as a product. So, Stefan, the expert, you were hosting one of the tracks. Were there What were the non-technical challenges? Cuz platform as a product is a one of the non-technical solutions, but what were the Maybe some of the more socio-technical things that you were hearing about.

Um so, what were what I saw especially as a as a consultant and coming from the outside in platform teams is very often uh the platform is quite quite mature technically, but they are not it's not always fixing the problem that the problems that developers really are having. So, I see it very often that platforms are started because someone saw Kubernetes and they said, "This is a

good solution." Uh but haven't always looked at what is actually the problem that the developers are having. And I think it's it's very important to to to look at that to look people who are developing software in your company, what do they What are the really the trouble they're having and try to connect your platform to something concrete that they're having at the at the moment. So,

to to build some traction. That's quite important. I think off the the back of that, you need to find the problem, but how do you know if you're doing it in the right way? There's a whole separate part of that of like how do you measure success of of that scaling effort? Cuz scaling's a good problem to have really if you're there. Has anyone done that measurement

side of things when it comes to scaling? Uh well, we started with that. I mean, it's important to know how successful your platform is, and for that we do define a number of KPIs, and the KPIs are based on the success rate of the execution of the automation on the uh time execution of the creation of certain resources, um as well as we are making the continuous

feedback loop with our users so that we can get those feedback live and adapt the platform accordingly because at the end of the day it's for the user and not for the architecture. Yeah. Katry, have you Yeah, I think uh for any platform defining those success metrics is is important from from day one or day zero, um if I may. Um so, some of it is user

adoption, like like you mentioned, like end-to-end latencies, and, you know, all all that good stuff. But, it's also your invisible signal to uh you know, how well your platform is doing, fit for purpose, as well as where the friction is. So, I think translating those metrics and looking beyond what they actually mean technically is is equally important. Yeah, and uh I think as well, so there's some

something we are starting to do in our platform is to look at um user perceptions. So, ask questions uh like uh do you think that the platform is is helping you keeping your workflow simple, is helping you to get good feedback loops? So, these are things that are they they look quite far away from the platform, but it's quite interesting to look at that to then dive

deeper with these people and ask, "Why do you think it's complicated?" And uh you might There might be some some things which are just misunderstanding because the for example the the scope of the platform is not exactly what the what the users expect. So then you have to do we need to develop something new to help these users or is it more of a communication issue that

we we need to clarify that this is something that the class the the platform is not covering. And then very the problems and the solutions are not always technical. So I think it's important to keep that in mind as well that sometimes it's just that someone didn't understand what the platform really is doing. So I'm going to throw a spanner question in. My panelists wanted like time

to prepare but I thought of a an interesting question based on what you've mentioned around day two, right? So we've everyone's hinted on it so far on the panel of we built something maybe it was the right thing, maybe it wasn't but we've got to support it. How do you balance the new things that your users are asking for and the value they want with maintaining that

existing platform that you've already built that may not have been quite so suitable for scaling that you originally thought. so this is getting to the war stories actually. Because building platform is is one thing and then operating is a is a completely different challenge. So in my personal experience building Kafka right so it actually came from the problem of centralizing the Kafka subject matter expertise within one

team moving that away from these distributed teams like for example a front end team but it have no idea of how to send or maybe they do but that's not their their specialism. But what happened was as we built the platform the expectation grew. So we were suddenly expected to actually debug and resolve issues in the Kafka applications that the teams were building which we were not

actually prepared for. So suddenly we um the the expertise had to had to grow. So, we had to invest in the engineers actually building those skills as well as actually understanding how Kafka as a technology itself um worked. So, if if you have used Kafka, you know those two are very different things. Um so, actually so we invested in more open source contributions, building that community, um

giving back, and actually learning from the community as well to uh quickly scramble to provide that expertise and knowledge to the other teams which required it. Uh for Amadeus, I would say that um a strong platform teams is probably part of Amadeus DNA per se because already when Amadeus was created from the starting point, again, the mainframe, sorry for that, long time ago, but it was there

and already with a mainframe, we had this strong communication layer. So, when we as a as a company, we were migrating further to either our um urgent data center or the public cloud, we maintained that strong platform layer. We maintained the need to have this strong connectivity, the middlewares, the observability layer. So, it was evolving as Amadeus evolved and migrated. So, we always had since the beginning

those scope teams that would be in charge of a particular technology, let it be Kafka or an Oracle or the network layer or the subscription of the Azure itself. So, that was not a problem per se for us. Again, the main pain point for the day two is to ensure this right level of isolation between the teams so that they can independently provide their automation and self-service

to the users at the right level that the user requests, but still be able then to combine those capabilities together or orchestrate them so that at the end of the day you can still deploy a consistent platform I I wanted to talk about something a bit a bit different that I've seen many companies struggling with is that I I would hope that more people read site for

about site reliability engineering because companies have a tendency to try to optimize the utilization of platform engineers and ensure that there's they're always delivering features the whole time and you you you need slack when you when you do operations. It's true for any any team doing operations. You will have unexpected things coming up. And you need to you need to reserve some time for the unexpected because

this is else you will you will always have problems with operations. And I think the the the ballpark that Google's gives in the SRE books of having about 50% unplanned time. This is the what the reality looks like. That's that's that's a good a good ratio to have. Yeah, it's like that shift from being able to build it and throw it over the fence to another team

to support. You need to maintain that capacity to still support things as you're as you're scaling your platforms. Um so it wouldn't be a panel at KubeCon without having an AI question. It's come up quite a few times at our booth actually at Intuit of like how is AI going to impact platforms? And for me the most interesting thing is you've got lots of AI agents out

there that they're going to be doing lots of requests and lots of things. Um how do you as panelists think that AI is going to impact not just platform engineering but the scale of the platforms that we have to now support? Well, I think it's good to offload to AI Again, lazy person here. So, I think it's good to offload to AI everything that will allow us

to concentrate on important things. So, conceptually important things, like to think about your platform as a product, for example. So, you can offload to AI everything about basic code generation, about testing, about gathering the events of your platform, and acting up accordingly, either sending an alert or maybe readjusting your workload Or maybe support channel, so that you do not need a team who would actively reply, but

you have just a proper knowledge base, and then you have an AI helping you with that, so that your team can concentrate on what is really I think it's we are in a very interesting world right now, and it's very evident if you attended the keynote yesterday and all the talks. It was very much centered around agentic AI. In my opinion, it's it's it's already here, so

we better gear up for it and be prepared for it. The autonomous agents are here, and the dependency on platform that we are all engineering is going to only increase. So, be prepared for it, be, you know, build a platform so that it can interface seamlessly with those agents, as well as build in enough guardrails, right? Because again, the the day two aspect of it. So, when

something goes wrong, how build a platform in an explainable way, make it auditable, and compliant. So, all of those things are once again going to be pushed to the extreme. The boundaries are going to be pushed, so better be prepared for it. Yeah, I think it's I would like to say more or less the same thing that Gaia 3 said, so it's it's it's it's a bit

ironic that with DevOps, we've been the last 10 years trying to optimize for delivering quickly and building quickly. Now we have AI which is doing the same in a larger order of magnitude. And so this basically this this is not going to be a bottleneck anymore. So I think the what will be important will be not not to not to to prevent the technical debt from scaling.

And this is things like guardrails will be important. It's it's a good time for QA people again now. So 5 years ago everybody thought QA was dead and now they're they're coming back. Yeah. Yeah, it's interesting. A lot of the the things we've been talking about when it comes to scaling are the same problems that you have for your day zero to day one platforms. It's just

level it up a little bit further. You still have to solve problems and you've still got to look after the things that you've already got running and then AI takes that the next level, right? Um and I really liked a lot of the things that you've all been saying around. It's about ownership. It's about building for your users. These are all very interesting and good things. So,

as we come to the end of our panel and thank you all for joining us today, one final question for each of you. Uh so I've given you all a time machine, but it's a bit of a rubbish time machine and it will only take you back to the days you started building your platforms. Um so you know like modern you knows they're going to have to

scale even if at the time young you didn't. What would you go back to tell yourself now given what you've learned? And I'll start with Stefan on my left. So I I think the the thinking about the that a large part of platform is also the information. It's not only about the technical part, but it's also about enabling developers, getting them to understand how how to do

something. Uh and and how this can scale. So also information within the team which is simple with five people, but becomes a lot more complicated when you have 50. Uh Anna or maybe Guy actually cuz you've got the mic. All right, okay. so I want to quote the theory of constraints. How many of you have heard about it or read Eliyahu Goldratt's Goal? So for me, the

personal reflection starts from there, right? So to paraphrase it, uh for any system, however well-built it is and however efficient it is, there are always bottlenecks. So to put it in simple terms, your system is only as fast as its lowest component or is only scalable as far as your least skilled part. And this is not just the the technical one, right? As Anna was saying earlier,

there are people or processes that could essentially be slowing it down. And these bottlenecks over time is also moving within your system. So it's an iterative process to identify what those bottlenecks are, what those constraints are, and making sure you're addressing them. And in this AI era, you have to move very fast in being able to identify that and and automation and having those tools is very

important. Well, if I'm in a time machine, probably to Anna in the past, I would say one and a half things. So the first one is that we are building the platform for the users. So you should think of who is your user and what kind of services they require. And the last takeaway would be a Terraform is not the answer to It's pretty universal. Um so

thank you to all of my panelists. You've been wonderful. I'm going to quickly sprint cuz I forgot to use the clicker. If I can, so we have a Oh, never mind. I'll figure it out later. We've got some QR codes for some links that are interesting if people want to follow along and some feedback, but while I try and figure that out, everyone can have a solid,

very, very strong round of applause for my wonderful panelists. And thank you all for attending.