KubeCon + CloudNativeCon Europe

Sponsored Session: Performance at Enterprise Cloud Scale: Coupa’s Ku... Peter Irwin & Carl Baumcratz

14:49 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk covers the challenges of managing Kubernetes at enterprise scale, particularly as organizations grow and more teams utilize the platform. The speakers, Peter and Carl from Koopa, discuss their journey to optimize Kubernetes performance using a GitOps approach and the implementation of tools like ScaleOps. They detail their experience with traditional autoscaling, which proved to be inefficient, and share insights on their transition to Carpenter for cluster management. Throughout the discussion, they highlight significant cost savings and efficiency improvements achieved through automation and proactive resource management, culminating in a reported $2.26 million saved. The session emphasizes the importance of cross-functional collaboration in maximizing ROI and reiterates the benefits of shifting focus from manual tuning to automation.

Full transcript

Well, thank you guys for joining us here. Uh we know we're standing between you guys and beer, so we're going to keep it moving and appreciate everybody showing up. Uh today we're going to talk a little bit about how Kubernetes can perform at enterprise cloud scale with my friend Carl here from Koopa. Uh as an agenda, uh we want to talk a little bit about the challenges

at managing Kubernetes at scale, some of the problems it presents as you uh try to stay efficient as as more and more teams come onto an enterprise platform. We're going to talk a little bit about Koopa's environment and some of their journey through optimization uh from workloads to nodes and carpenter and uh and the path forward and the outcomes that they got uh throughout this journey. So

just to introduce ourselves, I'm Peter. I lead the global sales team here at Scale Ops and Carl >> uh director of cloud software engineering at Koopa. >> Awesome. So when you run Kubernetes at scale, there's a there's a number of factors that come into play and and become challenging to manage, especially supporting many uh development teams. Uh as things are as dynamic as they can be and

uh performance will change over time, behaviors will change, keeping up with that becomes really challenging. Uh and and it can be inconsistent. That leads to a lot of inefficient packing of nodes and pods uh and sizing and and results in a lot of time spent by teams manually tuning resources trying to stay ahead of the curve but always uh dep prioritizing some of the the cost measures

and the efficiency measures in favor of developing and building new features where their time is is better spent. Um Koopa experienced this firsthand and Carl, why don't you share a little bit about your environment? >> All right. Uh so our environment has been expanding over the last three years. We're a very transformative fastm moving company. Um we were kind of born in cloud uh AIdriven spend management

platform for at least a half a decade now. Um so AI is kind of it runs through our blood. Uh AI continues to expand for us. growing number of workloads coming into our our clusters and a growing number of of customers coming into older workloads as well as the newer workloads that are coming. Um to manage our our growing environment um we use a GitOps approach delivering

our clusters uh as well as its configurations. We're currently at our at our uh control plane. We're at 140 clusters. Um that sets us up for a very large data plane. uh for EC2 which is where our challenges are before scaleups. Um we initiated a PinOps uh program called Kubber. Um over the last several years we're trying to really drive down the cost of that EEC2 uh

uh platform shifting all that left um right now we were cost um managing just by reports and now we want to shift it left into the pipeline. So um why traditional autoscaling fails um it falls short um the minute you set a static threshold in your values files it's already drifted it's a snapshot in time there's risk of downtime performance issues for your cluster level uh signals

they're just not enough um that causes fragmentation in your environment inefficient bin packing. Um, and personally for us, uh, it's a lot of manual logging into clusters, seeing why pods aren't starting, why pods are having issues. Um, so it's near and dear to our heart. Um, and then at Koopa scale, um, it's very time inensive. Humans tuning these settings are are not what we want. Um, hours

and hours of engineering time, and it's more of a reactive approach. uh we wanted to do a more proactive solution. Um so we decided to go figure out what's out there. Um searched around, figured out do we want to build, do we want to buy? Um we tested out a couple different couple different solutions uh off the shelf, scaleups being one of them. um we didn't want

to do a whole lot of analysis paralysis when evaluating these companies that do a lot of the same things. So scaleups was the easy button for us. Um so we started with them did a PC in a handful of clusters um with minimal to no issues and and saw instance savings. Okay. Um phase one of this project was uh really setting up workload right sizing along with

uh pod placement. Um this was phase one for us. So the onboarding process to scaleups was pretty smooth. Um again we used a githops uh approach to deliver the applications as well through github actions in argod. So it was very easy for us to uh deliver scale ops helmchart uh out of the box. Um we started with our dev and test environments which is roughly 35% of

our clusters. Um we let that soak for a month period to get people's validations gain the confidence in the product. Um and then came time to promote it to stage and production which was region by region. We're in nine or 10 regions. Um small amount of soak time per deployment and we're off to the races. We did it all in one month. Uh those 65% of the

rest of the clusters. Um in that process, we didn't just save money. We 37% of our workloads actually were scaled up. Um so we still saw cost savings even though 35 or 37% of those workloads were scaled up. Um our strategy going in here was automate by default. Um very big thing to do. uh getting developers buy in um can be a long time process to enable

them into the platform. Um so after this initial implementation um we were in the middle of moving to cluster autoscaler or from cluster autoscaler over to carpenter. Um so phase two was focusing on that uh without losing the efficiencies that we've gained with scale ops. Um the challenges we saw um everything online seemed to make it seem that Carpenter was pretty easy. The tutorials were very easy.

Um but there was challenges. Um special requirements for workloads were there. Uh specific scaling behaviors and thresholds needed to be thought through. ensuring the right node selectors were in place. Uh this also increased risk of misconfiguration because if we did that it would create even more waste and performance issues. Um so scaleup scaleups really helped us out with this one. Um it helped us accelerate this migration.

Um obviously the continued pod right sizing would happen. Um that was a given. Um some surprises to me were additional functionality within their product um which helped us with node pool configuration. Um, and then having a UI there that was very transparent what was happening within our Carpenter settings and very easy to manage through that UI. So after the Carpenter um implementation um we wanted to see

what more was out there. We wanted to see if we could accelerate even more efficiency within our nodes. So at the infrastructure level, uh, Scaleups has a functionality called node consolidation. Um, yes, Carpenter does a decent job of, um, putting things where it needs to go, right size of the machine, choosing the right sizes. Um, but a lot of our workloads, um, have a lot of pot

anti-affffinity. Um, we have PDBs, a lot of constraints that don't allow nodes to scale down. So an easy button for us was to implement node consolidation which helped reduce some of that fragmentation. Um allowed us to run them on schedules. Um really allowed us to be more efficient on nodes downtimes and uh weekends. Um those kinds of things is where this functionality really uh helped us. Um

something we are looking into is application performance on on the predictive replica scaling side of uh the capacity needed before a spike so we keep our SLOs's uh in place. So the way scaleups would help us is it would predictively autoscale a workload uh to have that capacity there before the spike actually happens. Um, so basically warming up the workload, warming up a node if need be.

Um, those are all important things that that we want to look at. So we're going to start start there. Um, let's take a look at the numbers. So all that provided a lot of value. Um, and here are the numbers on on what it provided. So to date, we saved $2.26 million. Um, that equates to $260,000 a month, reducing our vCPU by 30%. Um, and and I

think the most important box here is the thousands of hours uh in engineering time that it saved for people to do more important things than managing just the uh mundane task of keeping resources uh set up. when we first enabled scaleops, it was part of that phinops plan um to cut the cost and we achieved that. Um this journey and partnership with scaleups has matured and evolved.

Um we've shifted our focus more on uh the value proposition uh cost avoidance. Um you know we we cut the cost. Um, now it's a compounding effect year-over-year. What does that look like? Um, any new workloads coming in, you're no longer having to waste cost. You're now avoiding that cost. So, that was a huge win for us. Uh, huge proposition to take it into a cost avoidance

to speak with your finance teams. Anytime you're speaking with finance, uh, they want to know that compounding effect and that that number gets huge. Um, so, uh, that'll help you with, uh, your finance teams and talking those terms. Uh, let's see. >> Is that the one I want? Okay. So, this was a multi-team engagement. Many teams were vested. They had a vested interest in in this because

we were overprovisioning. Um, saving the company money is is a big thing for us. Um but everyone has different priorities. Uh scaleups help bridge that gap in all the all those priorities. Fin pushing for cost savings. Dev they wanting to them them wanting to deliver on their road maps building out features didn't have capacity to to do this. Um of course our platform team and S sur

teams uh we care about resilience and reliability. Um so that partnership with scaleups really helped all personas, all requirements. Um and everyone was heard throughout that process. Uh this enabled us to move fast and succeed quickly. >> It's been an awesome journey. >> Okay, so some key takeaways. Um Koopa's journey with automation and workload automation. Um manual tuning doesn't scale well. uh people's time, you know, you

want to spend it on making products that make the company money uh and not spend on those mundane resources, resource management that a machine could do. Uh it's better served for a machine. Uh the second one, the initial savings. Um that was just the start. Again, cost avoidance. Um that return year-over-year compounding is is going to be huge. um the multiaceted projects. So basically cross functional alignment

it unlocks the the ROI um gaining that alignment across all teams allowed us to move very fast and allowed us to enable this product quicker gaining uh the return on investment much quicker. Uh and then the last one automating resource management um that makes room uh for bigger priorities. So for my team personally, we can focus on doing multicloud, introducing new functionality into our platform. We want

to introduce different autoscalers. Um focus on AI and the product map that we want to deliver things coming out of KubeCon, we can do. Um and with that, I want to hand it back over to uh Peter. Thank you for allowing me to present. Uh product's been awesome. So look forward to continuing to partner on it. >> Man, that's awesome. Thank you so much for doing this,

Carl. Uh it's been awesome working with you and your team and and all the folks that you you mentioned here, right? Uh I think you know the key takeaways being you know savings is really just the beginning and the justifier but the value your team's seen the value the developers have seen from not having to you know starting out as hesitation and and fear of automation as

as many developers are uh but moving to fully trusting the platform and being able to just go back to developing and innovating and staying in the lead of the market as you guys have over the last decade has been uh absolutely awesome for our team to to be a part of as well. Uh if you guys out there are interested in learning some more, asking some more

specific questions about how we could help you in your environment, come join us at booth 805. Uh it's the big one with the shiny uh LED banner around the top. And uh we'd love to learn more about how we can help you guys out. So thank you guys very much. >> Appreciate it. Thank you.