Kai is the domain specialist for Kubernetes in the CloudThinker Multi-Agent System. He keeps your clusters healthy, optimizes workload sizing, audits RBAC and network policies, and orchestrates zero-downtime upgrades across EKS, GKE, and AKS.

These are leading patterns customers run with Kai on production clusters — health, sizing, RBAC, network policy, upgrades, and observability. They're starting points, not the limit — extend, replace, or add your own sub-skill with the Skills Framework.
Health & Triage
Node, pod, and workload health monitoring with auto-correlation to upstream Helm releases, deployments, and CrashLoopBackoff causes.
Right-sizing & Scaling
Pod CPU / memory right-sizing, HPA / VPA tuning, and KEDA recommendations based on real production traffic patterns.
RBAC & ServiceAccounts
Detects over-privileged ServiceAccounts, cluster-admin grants, and Pod Security Standards violations across namespaces.
NetworkPolicy & Mesh
NetworkPolicy gap analysis, service mesh (Istio / Linkerd) traffic audits, and zero-trust east-west traffic enforcement.
Zero-downtime Upgrades
Plans and executes EKS / GKE / AKS version upgrades — deprecated API detection, addon compatibility, blue-green node groups.
Metrics & Tracing
Prometheus, Grafana, and OpenTelemetry instrumentation. Generates SLI / SLO dashboards and alerts for every workload.
Ask Kai in chat to audit, right-size, or upgrade. He executes a one-shot run inside Sandbox Isolation, scoped to the clusters and namespaces in your Rules of Engagement, and exports a remediation plan you can apply with kubectl or a Helm diff.
@kai please audit our production EKS cluster for over-privileged ServiceAccounts and propose an upgrade plan from 1.28 to 1.31 with zero downtime.
Use /create-skill to capture the discovered surface, auth flow, Rules of Engagement, validation logic, and notification routing as a reusable Skill versioned in the Knowledge Base.
/create-skill audit-eks-rbac --from-thread #current --scope prod-cluster --policies pss-restricted --propose-patch kubectl
Invoke the saved Skill as a /-prefix Command from chat, webhook, or schedule. Re-runs are deterministic — same RoE, same surface, same CVE-mapped report.
/audit-eks-rbac
Kai inherits the same platform primitives that protect every CloudThinker agent. Continuous Kubernetes operations run without exposing cluster secrets, mutating production workloads, or leaving an unaudited trail.
Every run executes inside an ephemeral microVM with strict egress and a tamper-evident audit trail.
Guard-in and guard-out policy enforcement on every model call — PII / PHI redaction and secret detection by default.
Strict scope, method, and rate-limit constraints. Read-only proof-of-concept execution. No destructive operations.
Every reasoning step, tool call, and finding is recorded — exportable to your SIEM and reviewable by compliance.
Our Kubernetes team partners directly with your platform and SRE leads — cluster audits, zero-downtime upgrade plans, and RBAC remediation diffs on request.