Memory

k8s-aiops

Try it

Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.

What it does

A governed Kubernetes operations toolkit with 55 MCP tools wrapped in audit logging, undo support, and risk-tier labels. Covers pods, deployments, services, nodes, events, storage, and more — with read-only RCA diagnostics to find crash-loops and readiness failures worst-first. Write operations (scale, restart, delete, cordon, drain) support dry-run previews and reversible undo. Works with any kubeconfig-reachable cluster: standard Kubernetes, k3s, EKS, GKE, and AKS.

When to use it

  • A pod keeps crashing — read logs and run diagnostics to find the root cause
  • Need to scale a deployment up or down, or trigger a rolling restart
  • Cordon a node before maintenance, with an undo path if something goes wrong
  • Check which workloads are unhealthy in a namespace before taking action

The skill document

k8s AIops

Disclaimer: This is a community-maintained open-source project and is not affiliated with, endorsed by, or sponsored by the Cloud Native Computing Foundation, the Kubernetes project, or k3s/Rancher. "Kubernetes" and "k3s" are trademarks of their respective owners. Source code is publicly auditable at github.com/AIops-tools/K8s-AIops under the MIT license.

Governed Kubernetes operations — 55 MCP tools, every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.k8s-aiops/, a token/runaway budget guard, undo-token recording, and a descriptive risk-tier label on every audit row. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS). Run k8s-aiops init for a friendly onboarding wizard that registers your kube contexts as named targets.

Standalone: the governance harness is bundled in the package (k8s_aiops.governance) — k8s-aiops has no external skill-family dependency. Coverage focuses on common operations and is not yet exhaustive.

What This Skill Does

CategoryToolsCountRead or Write
Podslist, get, logs, describe, delete54 read / 1 write
Deploymentslist, get, scale, rollout restart, delete52 read / 3 write
Rolloutstatus, history, undo, pause, resume, set-image62 read / 4 write
StatefulSetslist, get, scale32 read / 1 write
DaemonSetslist, get22 read
ReplicaSetslist11 read
Jobs / CronJobsjob list/get/delete, cronjob list/get54 read / 1 write
Services / Ingress / Endpointsservice list, ingress list/get, endpoints list44 read
Config / Secretsconfigmap list/get, secret list (names/keys only)33 read
Storagepvc list/get, pv list, storageclass list44 read
Nodeslist, describe, cordon, uncordon, drain52 read / 3 write
Namespaceslist, create, delete31 read / 2 write
Metrics (top)pod, node22 read
Clustercluster_info, api_resources22 read
Eventslist11 read
Diagnostics / RCApod-health, workload-readiness22 read

Quick Install

uv tool install k8s-aiops
k8s-aiops init            # friendly wizard: register your kube contexts as targets
k8s-aiops doctor          # or skip init — works with your current kube-context too

When to Use This Skill

  • List/inspect pods, deployments, services, nodes, namespaces and recent events
  • Read a pod's recent log lines to diagnose a crash loop
  • Run a read-only RCA sweep (diagnose pod-health / diagnose workload-readiness) to find the root cause worst-first
  • Scale a deployment up/down, or trigger a rolling restart
  • Delete a stuck pod (a controller recreates it) or a deployment
  • Cordon a node before maintenance, then uncordon it after

Do NOT use when the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope for this skill).

If the user wants…Use
Kubernetes pods / deployments / nodesk8s-aiops (this skill)
Hypervisor VM lifecycle (power, snapshot, migrate)a hypervisor ops skill
Backup & restorea backup ops skill

Common Workflows

Diagnose a crash-looping pod and restart its deployment

  1. k8s-aiops pod list -n prod → find the pod with high restarts / non-Running phase
  2. k8s-aiops pod logs -n prod --tail 200 → read the recent logs for the crash cause
  3. k8s-aiops events -n prod → check for FailedScheduling / image-pull events
  4. k8s-aiops deployment restart -n prod → roll the deployment after fixing the cause
  5. Failure branch: if logs/events show an RBAC 403, the kube context lacks the verb — run kubectl auth can-i get pods -n prod and switch to a context with adequate RBAC; the skill never retries a denied auth.

Triage an unhealthy namespace with RCA, then act on the worst finding

  1. k8s-aiops diagnose pod-health -n prod → worst-first findings; a critical CrashLoopBackOff on prod/api cites restarts=9 and the exact kubectl logs … --previous action
  2. k8s-aiops diagnose workload-readiness -n prod → confirm the blast radius: e.g. Deployment web ready 0/3 (critical, under-replicated)
  3. k8s-aiops pod logs api- -n prod --tail 200 --previous-equivalent via k8s-aiops pod describe api- -n prod → read the crash cause the RCA pointed you at
  4. k8s-aiops deployment restart web -n prod → roll the deployment once the root cause is fixed
  5. Failure branch: if diagnose returns an RBAC 403, the kube context cannot list pods/deployments in that namespace — run kubectl auth can-i list pods -n prod and switch to a context with adequate RBAC; the RCA tools are read-only and never retry a denied auth.

Drain a node for maintenance, safely reversible

  1. k8s-aiops node list → identify the node and confirm it is Ready/schedulable
  2. k8s-aiops node cordon --dry-run → preview, then k8s-aiops node cordon (double confirm) — records an inverse uncordon_node undo descriptor
  3. After maintenance: k8s-aiops node uncordon → re-enable scheduling
  4. Failure branch: if doctor shows the cluster unreachable, fix the kubeconfig context (kubectl config get-contexts) before retrying — cordon is never issued against an unauthenticated session.

Usage Mode

ScenarioRecommendedWhy
Local/small modelsCLIfewer tokens than MCP
Cloud models (Claude, GPT)EitherMCP gives structured JSON I/O
Automated pipelinesMCPtype-safe parameters, audited

MCP environment caveat: MCP clients spawn the server with a CLEAN environment — shell exports may not reach it. Set K8S_AIOPS_HOME, K8S_AUDIT_APPROVED_BY, K8S_AUDIT_RATIONALE (and KUBECONFIG when the kubeconfig is not at ~/.kube/config) in the MCP server config's env block, not just in your terminal.

MCP Tools (55 — 39 read, 16 write)

CategoryToolsR/W
Podspod_list, pod_get, pod_logs, pod_describeRead
delete_podWrite
Deploymentsdeployment_list, deployment_getRead
scale_deployment, rollout_restart_deployment, delete_deploymentWrite
Rolloutrollout_status, rollout_historyRead
rollout_undo_deployment, rollout_pause, rollout_resume, set_deployment_imageWrite
StatefulSetsstatefulset_list, statefulset_getRead
scale_statefulsetWrite
DaemonSets / ReplicaSetsdaemonset_list, daemonset_get, replicaset_listRead
Jobs / CronJobsjob_list, job_get, cronjob_list, cronjob_getRead
delete_jobWrite
Services / Ingressservice_list, ingress_list, ingress_get, endpoints_listRead
Config / Secretsconfigmap_list, configmap_get, secret_list (names/keys only)Read
Storagepvc_list, pvc_get, pv_list, storageclass_listRead
Nodesnode_list, node_describeRead
cordon_node, uncordon_node, drain_nodeWrite
Namespacesnamespace_listRead
create_namespace, delete_namespaceWrite
Metrics (top)node_top, pod_topRead
Clustercluster_info, api_resourcesRead
Eventsevent_listRead
Diagnostics / RCApod_health_rca, workload_readiness_rcaRead
Undoundo_listRead
undo_applyWrite

Security — secrets: secret_list returns secret names, types, and key NAMES only. Secret VALUES are never read, returned, or logged, and there is deliberately no tool that returns secret values.

Dry-run previews: every write tool takes dry_run: bool = False. A dry run returns a {"dryRun": true, "wouldX": ...} preview without touching the cluster, and no undo descriptor is recorded for a preview.

Harness features that light up: write tools with a clean inverse pass an undo= lambda so the harness records an inverse descriptor (with _undo_id) to the undo store — scale_deployment/scale_statefulset record a scale-back to their returned previous_replicas, set_deployment_image records a restore to the captured previous_image, cordon_node ↔ uncordon_node and rollout_pause ↔ rollout_resume are mutual inverses, and create_namespace records a delete_namespace. drain_node records a partial uncordon_node inverse (cordon is reversible; evictions are not). delete_* and rollout_undo_deployment declare no undo. risk_level=high: delete_deployment, delete_job, delete_namespace, drain_node, rollout_undo_deployment. undo_list (read) lists recorded reversible writes whose undo tokens have not been applied yet, and undo_apply (write) executes a recorded inverse — itself governed, single-use, and supports dry_run. All 55 tools are audit-logged under ~/.k8s-aiops/ and pass through the budget/runaway guard, each recorded with a descriptive risk-tier label. pod_top/node_top return a clear "metrics-server not installed" message (not an error) when metrics-server is absent. Avoid tight poll loops (re-listing pods every second) — the runaway breaker backs this up.

CLI Quick Reference

k8s-aiops init                                            # interactive onboarding wizard
k8s-aiops pod list [-n ] [-t ]
k8s-aiops pod get  [-n ]
k8s-aiops pod describe  [-n ]                   # status, container states, events
k8s-aiops pod logs  [-n ] [--tail 200] [-c ]
k8s-aiops pod delete  [-n ] [--dry-run]        # double confirm
k8s-aiops deployment list|get|scale|restart|delete ...    # scale/restart: single confirm + --dry-run; delete: double confirm
k8s-aiops rollout status|history|pause|resume  [-n ]
k8s-aiops rollout set-image    [-n ]
k8s-aiops rollout undo  [--to-revision N] [--dry-run]   # double confirm
k8s-aiops statefulset list|get|scale ...
k8s-aiops daemonset list|get ...
k8s-aiops job list|get|delete ...                         # delete: double confirm
k8s-aiops cronjob list|get ...
k8s-aiops service list [-n ]
k8s-aiops ingress list|get [-n ]
k8s-aiops configmap list|get [-n ]
k8s-aiops secret list [-n ]                           # names/keys only — never values
k8s-aiops storage pvc-list|pvc-get|pv-list|class-list
k8s-aiops top pod|node                                    # requires metrics-server
k8s-aiops node list|describe
k8s-aiops node cordon|drain  [--dry-run]           # double confirm
k8s-aiops node uncordon 
k8s-aiops namespace list|create
k8s-aiops namespace delete  [--dry-run]            # double confirm
k8s-aiops cluster-info
k8s-aiops api-resources
k8s-aiops events [-n ]
k8s-aiops diagnose pod-health [-n ] [-l ]         # read-only RCA: crashloop/imagepull/OOM/unschedulable/restarts
k8s-aiops diagnose workload-readiness [-n ]                 # read-only RCA: ready  -n ` and switch to a context/ServiceAccount with adequate roles. For EKS/GKE/AKS, confirm the exec-plugin (aws/gcloud/az CLI) is installed and logged in.

### "Resource not found (404)"
The pod/deployment/node name or namespace is wrong, or the object was deleted. List the parent collection first (`pod list`, `deployment list`, `node list`) to get a current name. Remember most commands default to the `default` namespace unless `-n` is given.

### "Conflict (409)"
The object changed concurrently (or already exists). Re-read it and retry the write.

### Logs are empty or truncated
`pod logs` returns the trailing `--tail` lines (default 100); raise `--tail`. For a multi-container pod, pass `-c ` or the API returns an error naming the available containers.

## Audit & Safety

All operations are automatically audited via the bundled `@governed_tool` decorator (`k8s_aiops.governance`):
- Every tool call logged to `~/.k8s-aiops/audit.db` (local SQLite audit DB; relocate with `K8S_AIOPS_HOME`)
- Budget / runaway guard caps cumulative tool calls and wall-time, and trips on tight poll/retry loops — a safety backstop, not authorization
- Undo store records inverse descriptors for reversible writes (scale → previous replicas; cordon ↔ uncordon)
- Each write carries a descriptive risk-tier label into its audit row — a label, not a gate; `K8S_AUDIT_APPROVED_BY` / `K8S_AUDIT_RATIONALE` are optional annotations recorded when set, never required

**Authorization is not this tool's job.** There is no read-only switch, policy file, or approval gate. Whether a write is permitted is the agent's judgement or the RBAC of the kubeconfig context you connect with — give it a read-only ServiceAccount and writes fail at the apiserver, the place that owns the permission.

The harness is bundled in the package — no external dependency, no manual setup. See `references/setup-guide.md` for security details.

Driving these tools with a smaller / local model? See `references/agent-guardrails.md` — which guardrails the tool now enforces for you, plus a ready-to-paste system prompt.

## Contributing & feature requests

Coverage is intentionally focused. **Missing a device, action, or feature you need?** Open an issue or pull request at [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops/issues) — feature requests, contributions, and comments are all welcome.

## License

MIT — [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops)

Questions people ask

How does the undo system work?
Reversible operations record an inverse descriptor. For example, cordon_node records an uncordon_node undo token, and scale records a scale-back to the previous replica count. Use undo_list to see pending reversions and undo_apply to execute them.
What access does it have to secrets?
The secret_list tool returns only secret names, types, and key names — never the actual secret values. There is no tool that reads or logs secret data.
Can I preview changes before they apply?
Yes. Every write tool accepts a dry_run parameter. A dry run returns a preview without touching the cluster and does not record an undo token.

Related skills

Turn China 3C launch inputs into executable routes, messaging, channel actions, risk checks, and review decisions.

by killsnake0126 installs112 stars

Prioritize growth directions with a 2×2 risk framework — pick one bet and commit.

by deciqai2 installs2 stars

Diagnose which mental domain is holding you back before choosing a cognitive intervention.

by deciqai1 installs3 stars

Join video meetings as a voice bot, visual avatar, or avatar with live screen sharing.

by johnpatternai22 installs8 stars

Measure whether a transformation changed your growth engine or just added a one-time bump.

by deciqai1 installs2 stars

More from zw008

Browse all skills

Operate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.

by zw00878 installs1 stars

Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.

by zw00876 installs

Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.

by zw00853 installs

Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.

by zw00851 installs

Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.

by zw00851 installs

Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.

by zw00850 installs