Bart Farrell: So first things first, who are you, what's your role, and where do you work?
Manfred Moser: Hi, my name is Manfred Moser. My role is a developer relations engineer at Chainguard, where we provide secure open source components.
Bart Farrell: If we're looking at the topic of AI augmented Kubernetes operations, if an AI agent is allowed to diagnose or remediate Kubernetes issues, how do you separate probabilistic reasoning from deterministic execution so the system remains predictable?
Manfred Moser: So you have to basically establish guardrails, right? Like as we say, the AI agents are probabilistic. What that means is that they theoretically can give you a different answer for the same questions at different times. That also is obviously true, especially also when you ask different models, which you typically also do. So what you have to do is you have to use those systems in a probabilistic manner, but then you have to verify with a deterministic system. And that's typically done by your monitoring system, guardrails in CI/CD. And then also often you have to complement this with human reviews. We do this all the time because in Chainguard factory, we build software from source in a Kubernetes kind of related environment and also build containers, we build packages. And in our open source framework called Driftless AF, we have a bot that's called Judge. And that Judge bot assesses other AI's output and gives it a grade, like a ranking, right? Like this is 100% sure, right? Like, you know, the time is so-and-so or whatever. And so it gives it a grading and then it escalates by either saying, hey, this was not good enough. AI bot so-and-so. You go over there and check it again or see if you can improve the test coverage or stuff like that, all the way to escalate it to human review. So you basically have this chain where you capture the probabilistic, and narrow the probabilistic choices so much that ultimately the output is deterministic. And then that works okay for us at least.
Bart Farrell: Now, what operational context does an agent actually need before it can make a trustworthy decision? Is it about metrics, logs, events, topology, deployment history, policy state? And how do you deal with stale or conflicting signals?
Manfred Moser: that's tricky because especially as you know in the AI systems, context is everything. Context size is important. If you have too much context, weird stuff will happen and also your cost will explode. So you have to find the right balance. What we have done is we found that a large context size is often necessary for very complex decisions. For smaller decisions, you can just have the context size pretty small and have it really just very limited to what's necessary. In our system, the Chainguard factory, we have exactly that happening where when it just comes to simple decisions around syntax or so, then the context is actually very small. You know, that script, that linter that checks out that everything is good. On the other hand, for much larger decisions that are complex, we escalate all the way to human reviews, right? Like some things, the AI tools just don't know, right? If you have a system that has external data, things like upcoming events that influence your cluster sizing, for example, that you know you will need to scale up for the weekend, but there is no data in the system that indicates that that's a new requirement, then you have to make sure that the agents get that context and that can go all the way to escalating it to humans to provide that input. So that's very important. And then over time, you can also build up a historical context for the bots that can be super helpful. You refine your prompt, you provide that additional context, and then paired with the deterministic checks, you can get to a pretty good working system.
Bart Farrell: And before an agent is allowed to change production state, what technical safeguards do you need around permissions, blast radius, verification, rollback, and auditability?
Manfred Moser: That's especially tricky because as you can imagine, agents, there's a lot of them potentially. I mean, you've seen this from all the stories from the Hugging Face attack. They've talked about thousands of agents and that sounds scary, but I mean, that's like talking about thousands of threats in an application. Well, that's probably what's happening on your laptop right now. And that's true, right? Like the issue is that the change, the rate of change and the speed is much faster. So what do you have to do to keep that controllable? Well, you have to have all the guardrails you had for human changes and other system changes applied to agent-related changes. But much more so, right? So you have to strengthen and enable those safeguards even more so and also be faster in terms of reaction time so that things can stay under control. And this is what we also do in our Chainguard factory. We take advantage of the speed of those agents by setting them up against each other, so to speak, right? Like we have agents monitoring the infrastructure, and others changing the infrastructure. And so they sort of stay in check and we have to just also pay a lot more attention and that's what it takes.
Bart Farrell: And Manfred, if people want to continue the conversation with you, what's the best way to get in touch?
Manfred Moser: I'm very active and can be found on LinkedIn. But of course, if you want to learn more about Chainguard, you should probably check out chainguard.dev. And we also have a Slack community where you can also find me.