Bart Farrell: So first things first, who are you, what's your role, and where do you work?
Corey McGalliard: Hi, my name is Corey McGalliard. I'm an engineering manager for Akamai. I work on Akamai's compute platform, Akamai Cloud. My team and I build an internal Kubernetes-based platform that meets Akamai's change safety, security, and compliance expectations while delivering a user experience our engineers want to use. I focus on internal tooling for the products that we end up selling to external customers.
Bart Farrell: Banks need strong security controls, but they also need developers to move quickly. Where do you see the biggest tension between developer self-service and the security and governance requirements of a regulated Kubernetes environment?
Corey McGalliard: Kubernetes is really great at giving us this imperative way of describing the things we want and using an API to interact with it, find information and modify things on the fly. Some of the biggest challenges there are going to be the fact that you have to have an auditable path when changes are made, clarity on the change that is made, and ways to trace changes throughout the system. In my work, the decision we have chosen is to follow a GitOps pattern. Every change that touches our cluster, whether that is a configuration change on the cluster itself or the applications running on it are checked into Git initially, validated by at least two people, and then applied to the cluster. That gives us not only the ability to have that change control process in place, but it gives us a really fast pattern of making a change in the system. It does take a little bit of control away from the engineering team, specifically in production environments, but it gives us the peace of mind, the stability we're looking for. And that definitely applies to highly regulated environments. If you can trace the change from the moment it's checked into Git to when it's applied in production, that gives you the ability to step forward relatively easily. Now, in lower environments, non-production environments, we're comfortable giving developers more freedom, the ability to adjust things on the fly. And the expectation is to test those changes first in non-production and then promote them up into production workloads.
Bart Farrell: Operational resilience is a major requirement for financial institutions. And when Kubernetes workloads are distributed across regions, clouds, or even edge locations, what should banks be doing differently to make sure a failure in one part of the infrastructure doesn't become a systemic outage?
Corey McGalliard: Distributed systems are phenomenal in the fact that you get this resilience just by default. Even in one region with one Kubernetes cluster, you have that resilience baked in and extending that boundary across multiple regions gives you the ability to load balance that. The challenge is really going to be making sure you're able to route your traffic to a healthy endpoint. You can withstand failure in your cluster, in your system, as long as you have a clean way to migrate traffic to an endpoint that's actually functioning, right? We use a tool internally. It's actually one of Akamai's tools. It's a DNS-based load balancer. GTM is the name of the product that Akamai sells to be able to handle this. And it pays attention to multiple endpoints in the backend. And it will route your traffic appropriately as things are healthy and also location-aware routing there is going to be helpful too. You don't necessarily want to be in a situation to where you have this beautiful global infrastructure and have someone in Australia trying to hit an endpoint in Los Angeles, right? You want to get them as close and enter your ecosystem as close as possible. And when you have an issue in Australia, you can always send them over to Los Angeles as long as the route is set up in place, right? That gives you that control and then having the ability to have context be regional, but also the ability to shard it globally. Having some mechanism to be able to share state globally, but reconcile locally is the pattern that we follow, right? It's ingested as close to the customer as possible initially, and then it's copied globally. So those are some patterns that we're seeing are really helpful for our services.
Bart Farrell: Now, a lot of banks are building increasingly complex Kubernetes platforms with dozens of security, observability, and policy tools. At what point does that complexity itself become a risk? And how should platform teams decide what to standardize or automate?
Corey McGalliard: This is probably one of the hardest questions in the industry today. There's a ton of information out there. You can gather tons of information about the status of your cluster. You can be in a position to where you have the expected state and an actual state. What's challenging is making good choices around tools. Policy as code, so stuff like Kyverno is super beneficial here. One of the ways we have really adopted this and seen benefit is that we can check in a set of expectations of how software runs within our clusters. And that can be consistent globally across all the clusters because you apply policy as code into each individual Kubernetes cluster. And Kyverno's operator just keeps that policy in check, right? And that can be whether it's a mutating webhook or a validating webhook so that you can either prevent someone from applying something harmful or be in a position to where you're constantly validating whether your system's in check against your policy. I think it really becomes a risk when you're in a situation to where you can't wrap your understanding around what's going on in your environment. You have to have the appropriate tools and ways to quantify how things are being configured and what's happening in each individual cluster, having clarity in your logs and metrics to know which cluster and which system is having a problem or seeing security issues. Having good aggregation and high-level pane of glass that you can drill down into that will take some risk out, but at least initially until you're in a position to where you have good visibility and the way to trace what you're looking for appropriately, that's where I see most of the risk.
Bart Farrell: And if people want to get in touch with you, what's the best way to do that?
Corey McGalliard: You can find me on LinkedIn or just about any social platform. cmcgalliard is generally my handle. Not super active on social media, but I do have accounts. And if you want to reach out, I'm pretty happy.