Kubernetes Agent

Meet Kai. Your Kubernetes Engineer specialist agent.

Kai is the domain specialist for Kubernetes in the CloudThinker Multi-Agent System. He keeps your clusters healthy, optimizes workload sizing, audits RBAC and network policies, and orchestrates zero-downtime upgrades across EKS, GKE, and AKS.

Kai, the CloudThinker Kubernetes Engineer specialist agent
Sub-skills

Leading use cases, customizable Skills

These are leading patterns customers run with Kai on production clusters — health, sizing, RBAC, network policy, upgrades, and observability. They're starting points, not the limit — extend, replace, or add your own sub-skill with the Skills Framework.

/cluster-health

Health & Triage

Node, pod, and workload health monitoring with auto-correlation to upstream Helm releases, deployments, and CrashLoopBackoff causes.

Node / pod / workload state
CrashLoopBackoff RCA
Helm release tracking

/workload-sizing

Right-sizing & Scaling

Pod CPU / memory right-sizing, HPA / VPA tuning, and KEDA recommendations based on real production traffic patterns.

HPA / VPA tuning
KEDA recommendations
Cost-aware sizing

/k8s-rbac

RBAC & ServiceAccounts

Detects over-privileged ServiceAccounts, cluster-admin grants, and Pod Security Standards violations across namespaces.

ServiceAccount audit
Pod Security Standards
Least-privilege diffs

/network-policy

NetworkPolicy & Mesh

NetworkPolicy gap analysis, service mesh (Istio / Linkerd) traffic audits, and zero-trust east-west traffic enforcement.

NetworkPolicy gap analysis
Istio / Linkerd mTLS
Egress allowlist

/upgrade-orchestration

Zero-downtime Upgrades

Plans and executes EKS / GKE / AKS version upgrades — deprecated API detection, addon compatibility, blue-green node groups.

EKS / GKE / AKS upgrades
Deprecated API scan
Blue-green node groups

/k8s-observability

Metrics & Tracing

Prometheus, Grafana, and OpenTelemetry instrumentation. Generates SLI / SLO dashboards and alerts for every workload.

Prometheus / Grafana
OpenTelemetry tracing
SLI / SLO scaffolding
Workflow

From prompt to repeatable Command in three steps

Start with a prompt

Ask Kai in chat to audit, right-size, or upgrade. He executes a one-shot run inside Sandbox Isolation, scoped to the clusters and namespaces in your Rules of Engagement, and exports a remediation plan you can apply with kubectl or a Helm diff.

@kai please audit our production EKS cluster for over-privileged ServiceAccounts and propose an upgrade plan from 1.28 to 1.31 with zero downtime.

Persist as a Skill

Use /create-skill to capture the discovered surface, auth flow, Rules of Engagement, validation logic, and notification routing as a reusable Skill versioned in the Knowledge Base.

/create-skill audit-eks-rbac
  --from-thread #current
  --scope prod-cluster
  --policies pss-restricted
  --propose-patch kubectl

Run anywhere as a Command

Invoke the saved Skill as a /-prefix Command from chat, webhook, or schedule. Re-runs are deterministic — same RoE, same surface, same CVE-mapped report.

/audit-eks-rbac
Kai detects environment drift on each run and proposes a Skill update for human-approved merge — no manual maintenance.
Safety

Safe by construction — from sandbox to audit

Kai inherits the same platform primitives that protect every CloudThinker agent. Continuous Kubernetes operations run without exposing cluster secrets, mutating production workloads, or leaving an unaudited trail.

Sandbox Isolation

Every run executes inside an ephemeral microVM with strict egress and a tamper-evident audit trail.

Per-run microVM
Egress allowlist
Auto-destroy on completion

Guardrails Engine

Guard-in and guard-out policy enforcement on every model call — PII / PHI redaction and secret detection by default.

PII / PHI redaction
Secret detection
Prompt-injection defense

Rules of Engagement

Strict scope, method, and rate-limit constraints. Read-only proof-of-concept execution. No destructive operations.

Scope allowlist
Read-only PoC
Pre-execution query filters

Immutable audit

Every reasoning step, tool call, and finding is recorded — exportable to your SIEM and reviewable by compliance.

Tamper-evident logs
SIEM-exportable
Replayable decisions

Compliance Certifications

We maintain the highest industry standards and regularly undergo rigorous third-party audits to ensure compliance.

Talk to Kubernetes Engineering

Ready to put Kai to work on your clusters?

Our Kubernetes team partners directly with your platform and SRE leads — cluster audits, zero-downtime upgrade plans, and RBAC remediation diffs on request.

  • EKS / GKE / AKS supported
  • Zero-downtime upgrades
  • RBAC + NetworkPolicy in one agent