Bart Farrell: So, Andy, I know you quite well, but some people don't. Who are you? Where do you work? What do you do?
Andrew Martin: Hi, I'm Andy Martin, founder and CEO of Control Plane. We are a recently nine-year-old cloud native cybersecurity consultancy, and we fix difficult situations, round peg square holes for regulated organizations. That is FSI, critical national infrastructure, generally anywhere that regulation requires critical trust and non-negotiable quality of the system.
Bart Farrell: Now, Andy, if an AI agent can query and modify a Kubernetes cluster, what should its identity and RBAC model look like, especially when permissions need to change from task to task?
Andrew Martin: The concept of how we bestow agentic identity is still wide open. As an industry, on the one hand, we have legacy traditional human identity organizations saying, just issue another one of those for each agent, sub-agent, potentially for tool calls. What about MCP flows that delegated authority through into other systems? It gets quite messy quite quickly. We're looking at things like OAuth 2.1, extending what was already not the cleanest of authentication implementations. We've got AOF, Regenerative Identity. All these emerging new cryptographic anchors come with some risk. On the flip side, we have traditional models. We know that workload identity is proven to work at hyperscaler scale. We, of course, have SPIFFE/SPIRE as an implementation of that. And we also have JWTs. We know that we can just attach machine-readable claims if they're correctly configured and increase our baseline security over slinging around bits of tokens. So this really comes down to a balance of the complexity of the system and the trust that we want to bestow or endow these agents with. For a Kubernetes cluster, we are essentially adding an SRE. We're essentially adding a site reliability engineer. We're essentially adding an autonomous site reliability engineer into these platforms. So from that perspective, we're downscoping credentials and permission sets as we would do best practice. We're assuring our audits and logging so that we have a cryptographic chain of identity and trust back through all of the behaviors and actions that the agent can take. There's multiple different ways to do this these days, but ultimately we have to assume that agent at some point may go rogue or may be co-opted by another agent. And so all of the traditional security requirements, methodologies, constrictions that we as an industry have fought to avoid implementing, the security debt that we've accrued now must be paid, and ideally pay forward to secure the system as we carry on forward in time, because we're suddenly allowing non-determinism into otherwise deterministic systems.
Bart Farrell: And how do you stop poisoned logs, tickets, repositories, or MCP context from manipulating an operational agent into taking the wrong action?
Andrew Martin: We have a fundamental problem where what was traditionally inert metadata, log data, suddenly becomes viable, executable code because any way that we can inject into a prompt, and that includes all the data that's read, everything that's on disk, anything that comes in from an alert, the name of a pod or a controller, any of these things now have the capacity to materially manipulate the actions taken against the system. This is an extremely hard problem to solve because of the intrinsic link between context and the prompt, data and the instruction in the current LLM conversational mode. We can fix that one of a few ways. We've seen a new class of System 1 model emerge in the last few days with Jev. Reducing down complexity into confidence scores on optionality is one way to help reduce that set. Another is running prompt firewalling. So something like the Lakera Guard will look at the information coming across it. The guard will check multimodality, it will try and prevent three people talking in a soundbite from having somebody in the background whispering the maliciousness or the injection, if you like. But it's a very hard and difficult class of problem to solve right now. Strictly attempting to operate purely through interfaces is one useful way of doing this. So we can do something like reduce a model down to replying with JSON. 99% of the time that has high efficacy. We can then do things like we would equivalent to a Protobuf where we're making sure that we have the correct data formats that they're the right types. The context and the content of those is also traditionally deterministically checked, and that can be range checking, that can be using regexes where appropriate. Trying to minimize the amount of non-deterministic calls and classification calls we make back into these model responses to reduce down our non-deterministic space and ultimately try and get as much deterministic code as possible.
Bart Farrell: Andy, before an agent is allowed to remediate a Kubernetes incident, what evidence and validation should separate plausible diagnosis from safe to act?
Andrew Martin: The level of interaction that an agent is able to have with a cluster is, as with all data, separated along trust model lines, system criticality, the classification of the data that runs through that system, and ideally, the topology of our Kubernetes clusters is already aligned with that baseline threat model. Once we have the security basics in place, then understanding where we allow autonomous authority is a wider trust question again. We do know that this concept of machine speed or agentic speed or throughput speed, we have to operate as a defender at the same pace as the attacker. This is eternally difficult because as we know, attackers think in graphs, they can pivot through any point of the call graph. Whereas defenders think in lists. Every single thing must be checked off in order to make sure that we're compliant, conformant, and we've covered all of the bases that we know are dangerous. Phil Venables has come out with a great blog post on this recently saying, as we move faster and move into machine speed remediation, it's tempting to just say, well, let's pit the AIs against each other, which is also generative adversarial in some ways. The problem there is: A, it's extremely expensive. B, we're in a token war. And as we know, a lot of the token usage out of China comes from stolen credit cards. It comes from stolen accounts. So the cost per token in the US versus China is incredibly different. And it makes the DeepSeek viability, they suddenly have to price the token cost that's even less than Anthropic or OpenAI's balance. So we know we're battling extremely cheap tokens from nation state adversaries. So we can't just plug these two things in to fight each other. What we need to do is use AI to increase the quality of our controls as defenders, because these are deterministic, they're predictable. We know how to increase those controls. And then we're not paying semi-infinite token costs when we're seeing all these attacks coming in from left, right and center, let's say. There is also increasing the amount of automation we have to support these things. So do we plug an agent into a cluster and say, right, cluster admin, remediate what you like, or do we down scope a specific agent for a specific class of task, give it the tool calls and the authority to make, maybe it's just going to deal with fixing DaemonSets. So a disk attached at the wrong time and place, does it have the ability to delete a disk? Well, no, but it has the ability to mount an unmounted disk, in order to try and not lose data or to give these agents a fix anything and see an accidental, well I fixed everything, the cluster's now been deleted so there's no more availability issues for the disk. Well that's not quite the way we ideally would have fixed that. And then finally, the last point that Phil makes which really resonates here is we are governing the loop, not operating within it. So the human in the loop is stepping even further backwards to govern a collection of agents and a collective behavior, not being stuck in the middle of the loop, which means critical vulnerability is being exploited at midnight or early in the morning. The human has already put the guardrails in place around the operational loop so that the agent has that authority. But again, the downscoped agent has the authority rather than an omnipresent agent, saying, okay, well, let's go and fix something over here in order to go and fix down here. So human on the loop rather than in the loop, ensuring that we're operating at machine speed with our automation and not by plugging in AI directly, and also building up the base layer of controls so that we can give a tightly scoped agent access to fix just the things within its remit rather than deferring and delegating entirely to AI. So skills atrophy and rot. We have unpredictable non-deterministic outcomes and we run the risk of those outcomes being misaligned with the business outcome, which would be uptime or availability in that case.
Bart Farrell: And if people want to get in touch with you, what's the best way to do that?
Andrew Martin: I am best contacted at controlplane.io. You can find me on various socials. And on LinkedIn, of course, is a constant stream of useful technical content.