Kubernetes Cost Optimization: From Recommendations to Adoption

Kubernetes Cost Optimization: From Recommendations to Adoption

Sep 22, 2026

Guest:

  • Robin Elliott

Proper Kubernetes cost optimization requires more than just tools.

Robin Elliott, a cost optimization architect at Harness, explains how teams can make it easier to adopt recommendations by combining executive sponsorship, dedicated time for improvement, useful supporting data, and integrations with the tools developers currently use.

In this interview:

  • Why VPA, HPA, and Karpenter address different parts of Kubernetes optimization

  • How automation and lower-environment testing can build developer confidence

  • How Jira, Slack, and Teams integrations fit into an adoption workflow

  • Why AI spend requires visibility, attribution, and unit economics

Subscribe to KubeFM Weekly

Get the latest Kubernetes videos delivered to your inbox every week.

or subscribe via

Transcription

Bart Farrell: Who are you, what's your role, and where do you work?

Robin Elliott: Hi Bart, my name is Robin Elliott. I'm a cost optimization architect working with both prospects as well as customers at Harness.io for just shy of six years now. Harness.io, if you're not familiar, is an end-to-end AI-powered software delivery platform that automates and optimizes the entire DevOps pipeline, including continuous integration, delivery, GitOps security, and AI and cloud cost management. I spent an initial two years more so on the software delivery side, the majority of which was Kubernetes-based and then transitioned to help build out our cost optimization practice.

Bart Farrell: Now, which three emerging Kubernetes tools are you keeping an eye on and why?

Robin Elliott: Good question. First one I'm going to pick is Karpenter. I mean, just what it brought to the table so much faster deploying node capacity than the traditional ASGs. More recently, with the explosion of AI, the growth, the stress that's being put on the CSPs, and you are an Azure shop today, you'll probably know about allocation failed messages. So having Karpenter in play where it's just going to move on under that scenario, you can open up to other instance families and/or sizes and still get that capacity in a relatively fast fashion could be important. Second technology I'm going to pick is serverless and autopilot modes from the major CSPs as well. Not everyone wants the overhead from a human perspective of managing their Kubernetes clusters. So for smaller shops or perhaps in lower environments, why not just hand off the whole cluster management to your cloud provider? And then lastly, we've got VPA and HPA. I think they're both very important from an optimization and cost perspective, as well as being a good citizen within the cloud, right? You know, if you think I've already mentioned Azure and them having capacity problems, it's a little bit like a chicken and an egg scenario. If you can't rely on your cloud vendor to be able to allocate additional capacity, what do you tend to do? You know, you're going to go static size, which causes a problem. It's like self-fulfilling prophecy. You know, you're just accepting that you've got to have redundancy and excess capacity because you can't rely on the cloud to provision it. Under those circumstances, making sure that your workloads that you're provisioning are correctly sized is paramount. You don't want to have to oversize your workloads. And that can be a real problem. Definitely see where developers, they don't want their own production outages. They don't want to be called at nighttime for some failure. They always err on the side of going big, in my experience.

Bart Farrell: Now, why do technically valid Kubernetes right-sizing recommendations still fail to earn developers' trust?

Robin Elliott: Several factors come into play here, and they actually start higher than the developer. And there's two of them, and I'm not too sure. I think I see both of them in about equal fashion. The first one is competing priorities. The developers have deadlines. They've got to get features out there. And a lot of the time, that's going to trump any optimizations that may be presented to them. And I think equally important is trust. Trust is something that you have to build, and it's very easily lost. So the quality of the recommendation that you're going to give to the developer is paramount. And that quality has to be there consistently. Following on from these, I think you have to ask, does the business owner or the engineering team care, right? Is there actually an active initiative, where optimization is key, or maybe the business objectives are different. Maybe it's velocity of features to beat the competition, or they're going for growth and worrying about cost is going to come at some future tipping point. We see all different levels of business and maturity and product features. But I think those are probably the main ones.

Bart Farrell: And what does a successful workflow look like for turning a right-sizing recommendation into a change developers will actually adopt?

Robin Elliott: So going back, it sort of overlaps a little bit. Some of these are repeats to what I've already mentioned. I think it starts with executive sponsorship. Developers are human beings. They obviously want to do the right thing, but it's so much easier when someone actually mandates and sponsors and says costs are too high this is something that we need to address How i've seen that work in different contexts. One would be a large Australian bank that is a customer of ours. They call them cloud smash events. So they'll literally, think of it like a hackathon, but for cloud, Kubernetes optimization events. So literally they'll allocate one or two days every two months to actually go through and allow teams to have the time. to look at recommendations and apply them, both within Kubernetes as well as other services and infrastructure. Along with that, I've seen gamification be a really good, successful way of doing this. United Airlines, probably, I'll give them some kudos here. They've done some of the best gamification that I've ever seen. Leaderboard dashboards, they have league tables. it's like for those people who know, you know, soccer, not American football, but soccer. I was most impressed in their gamification and how they tracked, you know, teams competing with each other. So, nothing to do with tooling, right? It's all about process and the environment that the developers work within. Now, if we, you know, then get into the actual teams, you know, rewarding, you know. so Whilst I said people like to do the good thing, hey, you know, well, perhaps the people that do the best get some sort of reward, whether or not, you know, it could be gift cards, you know, coffee cards, might be a team lunch, whatever, or just, you know, praise within the organization. And I think at different levels, you see different things. You know, director and above, it's more, you know, personal ambition and competition between the peers. Whereas, like I said, boots on the ground, you know, people like recognition. And, you know, that recognition can take different forms. Already talked about, you know, you've got to give people the time to address the problem. Okay, so you can't just say, here are all the deadlines for delivering, you know, new product functionality. It's like, no, you've got to say, you know, here's some time where you can actually do the work. On the tooling side, you've got to have the right tools that are trusted. Not only they should give a recommendation but they've got to give the context of the recommendation and the supporting data in an easily consumed fashion for the developer. the developer wants to say, okay, you're telling me this is sized incorrectly and here's the new size. Why? They're always going to ask why. So having that information readily available or part of the same package is going to be key. And then I think you've also got to work within the developer's environment. their tool set, their flow, their life. You can't be introducing another bottleneck in their success. So whether or not it's a Jira shop or a Snowflake shop and you want to integrate recommendations and tracking them in Jira, you've got to do that. Notifications work within whatever the team's using. Is it Teams? Is it Slack? And then lastly, and it's not quite a barrier, but you need to be able to record success. So if you go through this effort, You've got to be able to understand what was actually applied, what was rejected, and why. And it's a little bit of an iterative process. Don't keep sending the same recommendation every month to a developer if they've already declined to do it more than once. You've got to have the intelligence to allow them to opt out for various reasons. They might say, no, because of business requirements, whatever, we've always got to have this excess buffer. Stop telling me about it, please, constantly.

Bart Farrell: Robin, how can teams introduce automation without undermining developer confidence?

Robin Elliott: This is a good one. I think there's, like I said, quality of the recommendation and, you know, it builds trust. Okay. So, whether or not, you know, you're using, you know, VPA in a non-automated fashion, you know, so it's actually going out there, looking at usage patterns and coming up with recommendations, you know, or you've got, you know, other tooling, you know, that essentially does the same thing. looks at the individual developer recommendations. And obviously, we're not always talking about the developer here. Sometimes we're talking about the persona that I would call the cluster administrator, the person that actually sizes the cluster, picks the node families that they think are going to be correct. So having recommendations that actually look at both, what are the types of deployments being deployed, as well as You know, do they actually fit the nodes that are being allocated? At least a third of the time that I go into customer sites, usually that is wrong. Whether or not it's wrong from the outset or just with changing workloads, but a lot of times you always see that either the CPU family or the memory family is off based on the workloads that are being deployed and you're leaving excess capacity. When you think about how you can automate these types of things, firstly, from a developer perspective, I've seen some more advanced teams where they'll take recommendations against workloads. And I think since Kubernetes is 1.27, you can actively resize pods on the fly from a CPU and a memory perspective. So I've actually seen sneaky DevOps teams throw together a script in a Harness pipeline that took all the recommendations, resized all the active workloads, and then went back to the developers two days later saying, did anything break in your dev environment? Okay. What we're talking about there is automated testing in a lower environment of those recommendations without the developers having to expend any effort at all. Okay. I think a natural progression once you automate the testing of a recommendation is then addressing how do we get it into code? You know, we're already doing this from an infrastructure perspective. taking recommendations, creating the PR, modifying the Terraform, and presenting that to the infrastructure team or the engineering team to say, hey, here's a recommendation. Let's reduce this instance size and handing it to them on a silver platter. So I think the next thing we're going to see will be utilizing AI to do the same, but at the deployment level. I would love to say, you know, VPA here, but it's had such a low adoption. It would be nice to see a higher adoption of it. You know, you've got some technologies out there like KEDA. And I apologize if people pronounce that a different way.

Bart Farrell: As far as we always hear, it's KEDA,

Robin Elliott: Tomato, you know, all that sort of fun stuff. You know, where Obviously, looking at those real-world metrics, those queue depths coming in, the sort of precursor to actually executing and consuming resources. I would love to see VPA come to the forefront a little bit more from automation, but I think right now it's that question of trust, or where it's not recommended to run with HPA as well. I think if people are not adopting Karpenter or a Karpenter stack, they're missing out. If you look at, I already talked about, you know, self-managed Kubernetes from the vendors, both AWS. Obviously, AWS had a major part in developing Karpenter, but they use it. Azure uses it. GCP sort of wrote their own stack and have a slightly different billing model. But I think Karpenter is definitely the way to go. So much better than the old ASG approach to autoscaling. I've mentioned Azure before. I hate to keep on knocking them, but. Even with those availability problems, Karpenter does allow you to select more of a pool of instance families and sizes. So being able to autoscale in a faster fashion is definitely a piece of automation that I think customers should be adopting. And then lastly, let's just go back to the old, in our lower environments. Are we really utilizing them 24/7? Okay. Are you a global organization that truly has a follow-the-sun model where multiple development teams are using a cluster 24/7? The answer is usually not, or at least there's a quiesce period over the weekend. So are you shutting down a cluster? Or are you shutting down namespaces within a cluster? Or are you intelligently looking at, you know the applications and services and how they're composed of microservices and actually looking for traffic patterns to say, is anything hitting these at the moment? And if not, let's scale them down potentially to zero. If you've got something that's, once again, traffic-based, why not scale it to zero when it's not being used, once again, in lower environments, and then spin it back up when required.

Bart Farrell: But we've been talking to people about you know, optimizing Kubernetes. We speak about, you know, Over, you know, some people have been using Kubernetes now for over a decade. I would talk about, you know, migrations to the cloud and how things are out of control for a lot of companies when moving to the cloud when it came to optimization. You know, getting things right with resources and generally having very large cloud bills when they're left to figure out, well, how can we manage this? How can we optimize, reduce costs? Now that we're in such a big AI boom and, you know, spend when it comes to tokens, things of that nature. Are we really talking about something that's completely different, or should the cultural aspects that you mentioned, such in the case of United Airlines, providing incentives, gamification. Should those lessons be applied to the new challenges we are facing when it comes to AI spend? I know it may be a little bit early days, but I just want to see what your take on that would be.

Robin Elliott: AI, everything comes in cycles. The cloud was the first cycle. Then within cloud, you could look back four years, look at the analytics tools, Databricks, Snowflake. They were that next wave of problems. I look at AI as, hey, it's that recurring every four years. AI has come along, upset the apple cart. How do we manage it? There are different aspects to it. There's the classic finance, procurement and FinOps of making sure that you buy into the correct savings that the vendors give you at the top level. But a lot of companies are struggling with visibility. It's like, well, who's actually spending these tokens? Where are they being spent? You have both the, what I would refer to as like the internal SDLC spend, you know, throughout the development of a product, you know, and let's say through the creation of an artifact and it being deployed. And then you also have what we refer to as production agents, you know, being used by customers, whether or not those customers are like your internal support staff, utilizing AI to support your real customers or maybe your applications that you're delivering to your customers externally have AI built into them. So there's AI spend happening in multiple places, no different from classic cloud costs, I think. The first thing is you've got to be able to attribute it. So you need visibility into where that spend is occurring, by whom, on what. You know, so all the way down to this individual, this application, this feature. And then you have to answer the question: what is the benefit? I mean, as an industry and as individual businesses, you can't accept this rapidly growing cost without knowing what is it doing for me? What is my benefit? So that's the next step. It's like, for those people on the call that have heard of that or may not have heard of unit economics, but unit economics, what's referred to as unit cost or cost of goods is balancing the, you know, what is the bang that I'm getting for my buck? And you need to track, you know, how are my costs trending as opposed to the benefit I'm getting, whether or not that benefit is increased revenue, better time to market, better quality, whatever that AI cost is being spent on.

Bart Farrell: Robin, what's next for you?

Robin Elliott: Right now, AI is the big piece, for sure. I mean, like I said, that is the pain point that people are trying to address. And it's not like all the optimization we talked about for Kubernetes has gone away. But, you know, you got maybe a very bad analogy, and this is going out on a podcast, but, you know, the old kitten with the laser pointer. You know, right now, the burning thing where customers are seeing anywhere from, like, you know, 2x to 8x growth. And it's funny, you know, we quote, you know, larger companies like Uber, but I was talking to a European company two weeks ago. And they blew through their Anthropic budget at the end of January. And that wasn't their January budget. That was their budget for the year. So one, they didn't do a very good job of budgeting. And two, the growth of the spend is astronomical. So you still got to optimize Kubernetes and other parts of the cloud. But AI spend is getting a lot of attention.

Bart Farrell: And Robin, if people want to get in touch with you to continue the conversation, what's the best way to do that?

Robin Elliott: Feel free to connect with me and send me any questions on LinkedIn. My profile is robin.finops, or you can just find me under Robin Elliott, two L's, two T's. And if you'd like to know more about Harness, please feel free to go to harness.io. And you'll have links there to read more about what we do in the optimization space, software delivery, Kubernetes. And that will give you options where you can have people reach out and contact you.

Subscribe to KubeFM Weekly

Get the latest Kubernetes videos delivered to your inbox every week.

or subscribe via