Menu
Image

NEWSLETTER

Latest Cloud-Native, Serverless and Generative AI news. Quality tech content read by tech professionals from Microsoft, Google, Amazon and Carrefour, and more

Follow Us

Zoomed-in central AI monitoring display

AI agents on Kubernetes: who controls their actions?

Mélony Qin Published on October 6, 2026 0

An AI agent running in Kubernetes can read a ticket, inspect metrics, call an API, update a deployment, and explain what it did. That is useful. It is also unsettling. Once software can choose tools, form plans, and act inside production systems, the core question changes from “Can it work?” to “Who controls what it is allowed to do?”

Kubernetes already runs critical workloads for banks, hospitals, retailers, manufacturers, public agencies, and AI platforms. Adding agents to that environment raises the stakes. A chatbot that gives a weak answer is one kind of problem. An agent that scales the wrong workload, deletes data, changes a network policy, or exposes secrets is another.

The answer is not to avoid agents. The answer is to design them like high-risk automation: bounded, observed, tested, and accountable.

image
Kubernetes is often where AI agents meet real production systems.

What AI agents are and how they operate in Kubernetes

An AI agent is software that can pursue a goal by making decisions across several steps. Unlike a traditional script, it does not only follow a fixed sequence. It can decide which tool to call, interpret the result, adjust its plan, and continue.

In practice, most agents have four parts:

  • A model: Usually a large language model or multimodal model that interprets requests and produces plans or actions.
  • Tools: APIs, shell commands, databases, ticketing systems, search indexes, Kubernetes clients, or internal services the agent can call.
  • Memory or context: Recent conversation history, retrieved documents, system state, logs, policies, or prior decisions.
  • A control loop: The logic that asks the model what to do next, validates the action, executes it, and checks the result.

Inside Kubernetes, an agent may run as a pod, a job, a controller, or a service that talks to the Kubernetes API. It may also sit outside the cluster and use credentials to interact with it. The deployment shape matters because it changes the risk profile.

A basic agent might watch alerts and open incident tickets. A more capable one might inspect logs, compare recent deployments, suggest a rollback, and ask a human to approve it. A highly privileged agent could take action directly by editing resources such as `Deployment`, `HorizontalPodAutoscaler`, `NetworkPolicy`, or `ConfigMap`.

That last category is where governance becomes essential.

Kubernetes gives agents a powerful execution environment. It also gives teams mature control points: namespaces, service accounts, network policies, admission controllers, audit logs, secrets handling, and policy engines. The safest agent designs use these native controls instead of treating the cluster as a blank runtime.

Who governs an agent’s actions

Control over an AI agent comes from several layers acting together. No single mechanism is enough. A model prompt cannot replace access control. A policy engine cannot explain whether an agent used poor reasoning. Observability cannot stop a bad action unless it connects to guardrails.

Identity and access control set the outer boundary

The first layer is Kubernetes identity. An agent should run under a dedicated service account with narrow permissions. Role-based access control, or RBAC, decides what that account can read, write, create, patch, or delete.

A safe default is simple: the agent should have the least privilege needed for its task.

For example:

  • An observability agent may read pods, events, and logs, but not modify workloads.
  • A remediation agent may patch a narrow set of deployments in one namespace.
  • A security agent may create reports, but not approve its own exceptions.
  • A platform agent may propose infrastructure changes, while a human or separate controller applies them.

Cluster-admin access for an agent is usually a design smell. If an agent needs broad permissions, split it into smaller agents with separate service accounts and narrower scopes.

Admission control checks actions before they happen

Kubernetes admission controllers sit between a request and the API server accepting it. They can validate or mutate requests before resources change.

Policy tools such as Open Policy Agent Gatekeeper and Kyverno help teams define rules such as:

  • Do not allow privileged containers.
  • Require resource limits.
  • Block images from untrusted registries.
  • Require signed images.
  • Prevent changes to protected namespaces.
  • Require labels that identify the agent, owner, and reason for change.

These controls matter because agents can produce unexpected requests. The admission layer gives the cluster a way to say no, even when the agent has credentials.

Human approval still has a place

Full autonomy is not the only path. Many useful agents operate under supervised autonomy.

A common pattern is:

  1. The agent investigates.
  2. The agent proposes an action.
  3. A human reviews the evidence.
  4. A workflow engine records approval.
  5. A separate system applies the change.

This pattern works well for production rollbacks, cost changes, security exceptions, and compliance-sensitive data access. It also supports accountability because the final decision has a recorded owner.

Prompts and system instructions guide reasoning

Agents often rely on system prompts and policy text to guide behavior. These instructions can tell the model to avoid unsafe actions, cite evidence, ask for approval, or use only approved tools.

Prompts are useful, but they are not reliable control boundaries. A model can misunderstand instructions, receive conflicting context, or be exposed to prompt injection through logs, tickets, webpages, or documents.

Treat prompts as guidance, not enforcement.

image
Access boundaries matter more when agents can act on infrastructure.

How decision-making should be constrained

AI agents make decisions through a mix of model output, retrieved context, available tools, and runtime rules. The main governance challenge is that reasoning may be probabilistic, while production systems need predictable behavior.

The best designs separate thinking, permission, and execution.

The model can suggest a plan. A policy layer checks whether the plan is allowed. A deterministic executor performs the action. Logs record every step.

A well-controlled agent workflow should answer these questions before action:

  • What goal was the agent trying to achieve?
  • What evidence did it use?
  • Which tools did it call?
  • What permissions did it use?
  • Which policy allowed or blocked the action?
  • Was human approval required?
  • Can the action be reversed?
  • Who owns the outcome?

This is where AI agents on Kubernetes need more than model monitoring. They need infrastructure-grade governance.

One practical pattern is to create an “action contract” for each high-impact tool. The contract defines inputs, allowed resources, approval needs, rollback steps, and audit fields. The agent cannot call raw cluster commands. It must call approved functions such as `propose_rollback`, `scale_service_with_limits`, or `open_security_ticket`.

That may look less flexible, but it is far safer. It also makes testing easier.

Observability frameworks for AI agents in Kubernetes

Observability has to cover both the agent and the cluster. Traditional metrics such as CPU, memory, and request latency still matter, but they are not enough. Teams also need to observe reasoning traces, tool calls, prompts, retrieved documents, approvals, and policy decisions.

A strong framework usually combines four layers.

LayerWhat to observeUseful tools and approaches
Kubernetes runtimePod health, restarts, resource use, network traffic, eventsPrometheus, Grafana, Kubernetes events, Cilium or other network telemetry
Application behaviorAPI latency, errors, queue depth, tool execution timeOpenTelemetry, distributed tracing, structured logs
Agent reasoningPrompts, responses, tool choices, retrieved context, refusal eventsLangSmith, Arize Phoenix, OpenTelemetry GenAI semantic conventions, custom traces
Governance eventsRBAC use, admission decisions, human approvals, policy denialsKubernetes audit logs, OPA decision logs, Kyverno reports, SIEM integration

The goal is not to collect everything forever. The goal is to keep enough evidence to explain important behavior without exposing sensitive data.

OpenTelemetry is becoming a shared base

OpenTelemetry is useful because it gives teams a common way to collect traces, metrics, and logs across services. For agents, traces can show the chain from user request to model call, tool call, Kubernetes API request, and final response.

That lets platform teams connect AI behavior to normal service behavior. If an agent scales a workload, the trace should show why. If a tool call fails, the trace should show whether the agent retried, changed plans, or escalated.

Kubernetes audit logs are non-negotiable

Kubernetes audit logs show who made API requests, when they happened, and what resources they touched. For agents, these logs are a key accountability record.

Audit logs should identify:

  • The agent service account.
  • The namespace and resource involved.
  • The verb, such as `get`, `list`, `patch`, or `delete`.
  • The source IP or workload identity.
  • The user agent or client.
  • The response status.

These records should flow into a central log store or security information and event management system, especially for regulated environments.

Agent traces need careful redaction

Agent traces can contain secrets, customer data, internal documents, or security findings. Observability must not become a second data leak.

Use redaction for:

  • API keys and tokens.
  • Personally identifiable information.
  • Secrets from logs or environment variables.
  • Private customer records.
  • Sensitive security details.

Retention also matters. Some traces may need short retention, while governance logs may need longer retention for audits.

image
Agent observability must connect model behavior to infrastructure events.

Ethical, security, compliance, and accountability risks

Control over AI agents is not only a technical issue. It affects trust, safety, and responsibility.

Ethical control means setting limits before harm occurs

Agents can act faster than humans can review. That speed is useful during incidents, but it can also amplify mistakes. Ethical design means setting boundaries around actions that affect users, data, and service access.

For example, an agent that handles customer support data should not freely query all records just because the cluster allows it. An agent that manages resource allocation should not silently reduce capacity for services used by vulnerable groups or critical operations.

Ethical control needs clear design choices:

  • Use only the data needed for the task.
  • Avoid hidden decisions that affect users.
  • Provide explanations for high-impact actions.
  • Keep humans involved when outcomes affect rights, access, or safety.

Security risks come from tools, not just models

The model is only one part of the risk. The bigger danger often comes from what the agent can do.

Common security concerns include:

  • Prompt injection that tricks an agent into calling a risky tool.
  • Overbroad service account permissions.
  • Secret exposure through logs or traces.
  • Supply chain risk from agent images and dependencies.
  • Lateral movement if an agent pod is compromised.
  • Unsafe plugins or external tools.

Security teams should treat agents as privileged automation. That means image scanning, signed artifacts, network restrictions, secrets management, runtime detection, and regular access reviews.

Compliance needs evidence, not promises

Regulated organizations need to show how decisions were made and who approved them. This is true for financial services, healthcare, insurance, public sector systems, and any environment handling sensitive personal data.

Compliance controls should include:

  • Versioned policies.
  • Change records linked to agent actions.
  • Approval workflows.
  • Data access logs.
  • Model and prompt versions for high-risk flows.
  • Evidence of testing before production use.

If an agent changes a production workload, an auditor should be able to reconstruct the path from request to decision to action.

Accountability must be assigned to people and systems

An agent cannot be the final accountable party. Responsibility belongs to the organization that designed, deployed, and allowed it to act.

A practical accountability model identifies:

  • The service owner.
  • The platform owner.
  • The security owner.
  • The policy owner.
  • The human approver when required.
  • The escalation path when the agent fails.

This avoids the vague and dangerous answer, “The AI did it.”

Real-world applications and case studies

Many organizations already use agent-like systems in Kubernetes, even if they do not call them agents. The pattern appears wherever software observes state, makes decisions, and acts through APIs.

Incident response copilots

A common use case is incident triage. An agent watches alerts, gathers pod events, checks recent deployments, summarizes logs, and suggests next steps. In a safer design, it does not restart services or roll back releases on its own. It prepares evidence for an on-call engineer.

This can reduce time spent gathering context. It can also improve handoffs during long incidents because the agent records what it checked and what it found.

Self-healing platform operations

Kubernetes controllers already perform self-healing by replacing failed pods and reconciling desired state. AI agents extend this idea into higher-level reasoning.

For example, an internal platform agent might detect that a service has repeated out-of-memory failures, compare current limits with historical usage, and propose a resource change. A policy engine can require that the change stays within approved bounds and that high-cost adjustments receive approval.

Security investigation agents

Security teams can use agents to connect signals from container scans, runtime alerts, network flows, and Kubernetes audit logs. The agent can group related events, identify the affected namespace, and open a case with the relevant evidence.

A strong pattern keeps enforcement separate. The agent may recommend isolating a workload, while a security automation system applies the network policy only after validation.

AI workload orchestration

Organizations running model training or inference on Kubernetes use agents to schedule jobs, watch GPU use, retry failed tasks, and route requests to available model endpoints. These agents can improve resource use, but they also need guardrails around cost, data access, and tenant isolation.

For example, a research platform may allow an agent to start approved training jobs in a sandbox namespace, but block access to production datasets unless a policy grants it.

Public examples from the wider cloud-native world

Kubernetes itself is built around controller patterns. Operators for databases, certificate management, and autoscaling already show how automated actors can safely manage complex systems when they reconcile against declared state. AI agents add model-driven reasoning to that pattern, which makes governance more urgent.

The lesson from cloud-native operations is clear: automation works best when it is declarative, observable, reversible, and limited by policy.

image
Autonomous systems need clear limits when they operate near critical infrastructure.

A practical control model for Kubernetes AI agents

Teams experimenting with agents can start with a simple maturity model.

Maturity levelAgent behaviorGovernance posture
Read-only assistantObserves logs, metrics, and eventsLow risk, still needs data controls
Recommendation agentSuggests actions and explains evidenceHuman approval required
Bounded action agentPerforms preapproved actions in limited namespacesStrong RBAC and policy gates
High-impact agentChanges production systems or sensitive accessFormal risk review, audit trails, approvals, rollback plans
Autonomous platform agentOperates across services with broad scopeRare, tightly controlled, continuously tested

Most teams should begin with read-only or recommendation agents. Move to bounded action only when observability, approval, and rollback are already working.

A production checklist should include:

  • Dedicated service accounts for every agent.
  • RBAC with narrow permissions.
  • Admission policies for risky resources.
  • Network policies that limit outbound calls.
  • No direct access to raw secrets unless necessary.
  • Structured logs and traces for every tool call.
  • Kubernetes audit logs sent to a central system.
  • Human approval for high-impact actions.
  • Red-team testing for prompt injection and unsafe tool use.
  • Clear ownership for failures and policy exceptions.

The core idea is simple: agents should operate inside a control plane, not beside it.

The real answer to who controls the agent

No single person, model, or policy controls an AI agent on Kubernetes. Control comes from the architecture around it.

The model may choose a plan. Kubernetes identity limits what it can touch. Admission policies decide what the cluster will accept. Observability records what happened. Human approval gates high-impact actions. Security teams test abuse paths. Compliance teams define evidence needs. Service owners remain accountable for outcomes.

That shared control may sound complex, but it is the same lesson production engineering has learned many times: powerful automation needs strong boundaries.

AI agents can make Kubernetes operations faster and more adaptive. They can also create new failure modes if teams give them broad power without proof, review, and traceability. The safer path is to start small, observe everything that matters, restrict actions by design, and make accountability visible before autonomy grows.

By the way, I’m a former tech product manager turned entrepreneur and investor. If you enjoy learning about AI startups, funding trends, and entrepreneurship, feel free to follow me here on Medium or sign up for my newsletter and my YouTube channel. I’m constantly exploring the latest developments in the AI world and writing weekly to train my tech entrepreneurship muscle!

Leave your thoughts in the comments below and see you in the next one!

Written By

I'm an entrepreneur and creator, also a published author with 4 tech books on cloud computing and Kubernetes. I help tech entrepreneurs build and scale their AI business with cloud-native tech | Sub2 my newsletter : https://newsletter.cvisiona.com

Leave a Reply

Leave a Reply

error: Protect the unique creation from the owner !

Discover more from CVisiona

Subscribe now to keep reading and get access to the full archive.

Continue reading