Bart Farrell: All right, so first things first. Who are you? What's your role, and where do you work?
Andy Suderman: I'm Andy. I am the CTO at Fairwinds. We are a Kubernetes services company. We do all things Kubernetes all the time. We've been doing that for about 10 years now, so we've got a little bit of experience. But it's really just Kubernetes services at the heart of it.
Bart Farrell: Fantastic. And with so much Kubernetes focus, what are the top three Kubernetes emerging tools that you're keeping an eye on?
Andy Suderman: That's a great question. I think new is an interesting term. But I'll stick to the few that are top of mind for me right now. We're currently working on migrations to the gateway API, so we're keeping an eye on a lot of the different gateway providers. kgateway is the one that we're spending a lot of time with. So that's top of mind right now. It's not exactly new, but it's also very rapidly evolving as we're all sort of very quickly moving to everything being gateway API everywhere. So that's the top one. The second one is less of a specific tool, but more of a category. It's all of the various ways to run inference on Kubernetes. So you've got KubeAI, you've got Ray, you've got several other projects in that space, and I'm keeping an eye on a lot of them to see what kind of shakes out as a top contender for running those sorts of workloads going forward. And another one that I always wanna get back to using because I think it could be super useful, but we haven't had an opportunity to use it, is vCluster. I'm actually very interested in the idea of converting our ephemeral namespace-based environments to full clusters inside of clusters. So vCluster's up there as well.
Bart Farrell: Going back to the first one, since you mentioned gateway API. This caused a fair amount of controversy both this year and last year. For the folks that you're helping out there, has it been a significant pain point for them in terms of the transition moving from one to the other?
Andy Suderman: Great question. For our actual customers, no, because we do most of the work for them, but it is not the easiest transition in the world. You're taking what was encapsulated in a single object, and now it's spread across several different ones, and just lots of nuance in how different pieces work is causing friction, but it's definitely not a trivial plug and play change to go to gateway API.
Bart Farrell: Now today we wanna dive into the topic of right-sizing resource optimization, so the following questions are gonna be focused on that. Most platform teams start right-sizing by looking at Grafana dashboards and Prometheus metrics. And in your experience, where do engineers most often get resource decisions wrong, even when they have good data?
Andy Suderman: That's a good question. We've been staring at Grafana metrics and Datadog dashboards and various things like that for years, helping customers right-size. And I think what we often miss with many of our customers from just looking at basic metrics is more business contextual things that you need to look at when right-sizing that aren't necessarily apparent from metrics. Obviously, we can look at CPU and memory, but you need a more complete holistic picture of a workload to understand how to truly right-size it. So there's lots of other metrics besides CPU and memory, but then there's business-facing things like end user-facing things like latency and how long it's taking to respond. And then also understand if we talk about the four golden signals, right? Understanding how much load actually impacts the application. And that's where folks tend to miss things, is in understanding the full picture of the application. And that can also include things like seasonal spikes. Maybe we have set up these resource requests and limits for our Black Friday rush for a retail company, right? And a new engineer comes in, or a platform team comes in, they're like, "Well, this application's way over provisioned." They may not understand the history of that application during our company's largest spikes. And so understanding those types of things outside of just pure metrics is something that folks tend to miss.
Bart Farrell: And at what point does Kubernetes resource optimization stop being a metrics problem and become an operational workflow problem? And in addition, what changes as organizations grow?
Andy Suderman: I love this question because I spend most of my time not telling people what they should set their resource requests and limits to, but talking to them about how to start making these operational changes so that it's part of your day-to-day process. I would say it stops being a metrics problem, almost immediately. The metrics are the easy part. Obviously there's pieces that you could miss there, but it's really easy to look at a dashboard and take an average over the last X amount of time and make a recommendation. But how do we actually apply those changes? How do we make sure they're safe? Making sure they're not going to break things is really the root of the problem, and so it is purely an operational problem. And as the organizations grow, that just gets worse and worse because then you've got perhaps a platform team that's trying to recommend these changes, but an app dev team that truly understands how their application works, and they're not communicating with each other about what's safe to apply and what's not. Platform can be making edicts where app devs are like, "Well, no. This application actually needs this amount. Here's why." And so that communication becomes super important, and as organizations grow and get bigger and more complex, communication often breaks down. And so that becomes the root of the problem, not so much the actual metrics themselves.
Bart Farrell: And when you're reviewing whether a workload can safely be right-sized, what are the signals or context that you trust beyond CPU and memory utilization? And further from that, are there any red flags that immediately make you pause?
Andy Suderman: Absolutely. There's a lot of answers to this question. I think I've actually already said a couple of them, but some signals, again, customer-facing metrics. So latency, number of requests, things like that. If you have a metric defined that says our latency should be at this, and you're seeing that it hasn't been over the last 30 days. Well, maybe you should address that before you start looking at right-sizing CPU and memory, right? So focus on those customer-facing signals. A lot of times, you'll change a metric. You'll say "Okay, we don't actually need this much CPU. We'll change it." Another thing to look at is obviously keep track of throttling. Start looking at how much that workload's being CPU throttled. There can be weird effects based on cluster topology, and other things on that node that can affect what the metrics look like even though it may seem like your utilization's very low. So that's super important. And the other thing I haven't talked about here, which I'm sure a lot of people are gonna be interested in hearing about, is Java, right? Java's consistently a problem when it comes to resource sizing, so know what the application is that you're right sizing to understand how that code base might be interacting with the resources that you're setting. Because oftentimes Java applications will look vastly underutilized in Kubernetes, but they need that amount of memory in order to function properly. So be contextually aware of things around the metrics. Those are my top couple answers, but there's so much nuance to all of this.
Bart Farrell: Absolutely. And a bit of a bonus question here, Andy. I know that we didn't cover this in the outline, but want to hear your thoughts on this 'cause it came up in a conversation the other day. We're in 2026. More than 10 years ago when companies started moving everything to the cloud, there was all this concern about infrastructure costs and then adding Kubernetes in terms of operational complexity. A lot of times costs getting completely out of control. People needing monitoring, then sometimes too much monitoring with dashboards on dashboards on dashboards. Now we're in the AI era where AI spend is getting really heavy because of limitations on GPU, tokens, et cetera. The organizations that you're working with, are they struggling with this? Is this something that they're having difficulties getting under control? Does it remind you in the past of when cloud resources were a bit out of control and now we're seeing this with AI? What's your take on that?
Andy Suderman: That's a great observation. It's a good question. I think it's too early to tell exactly. I think a lot of organizations are still figuring out how this works, so cost isn't a problem yet, and they're pushing that down the road, like we did with cloud costs, right? Everybody moved to the cloud. We didn't really think about it. Now 10 years later, we're like, "Oh, we need to pay attention to how much this stuff costs." We're seeing a little bit of that early stage right now for sure. But it's combined with the fact that we just started paying attention to infrastructure costs more heavily in the last few years. So there's a weird balance there. It's absolutely something that folks should be looking at. As far as whether they're struggling with it right now, I think it's a little too early to say.
Bart Farrell: Fair answer. Now, Andy, what's next for you?
Andy Suderman: What's next for me? I have a KubeCon talk to write 'cause I just got accepted for a talk with Danielle at KubeCon, so I'm super excited about that. So that's probably my focus. We've got a few conferences we're going to as well, AI conferences and things like that. Fairwinds is definitely working on launching products around running AI agents in your Kubernetes clusters, so that's something that we're very focused on right now.
Bart Farrell: And if folks want to get in touch with you, what's the best way to do that?
Andy Suderman: You can book a call with us at fairwinds.com if you're interested in talking about that. If you want to reach out to me personally, LinkedIn is probably the best place, or I'm in the CNCF Family Kubernetes Slack.