Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.
Memory
nutanix-aiops
Try itManage Nutanix Prism Central through 51 MCP tools with built-in governance, audit, and undo.
What it does
This skill wraps the Prism Central v4 REST API into 51 MCP tools covering the full Nutanix estate: clusters, hosts, VM lifecycle (AHV and ESXi), storage containers, networking, images, categories, data protection (snapshots, recovery points, protection domains, failover), alerts with RCA analysis, LCM firmware upgrades, and capacity runway forecasting. It ships with a bundled governance harness — no external dependency — that adds a local audit log (every operation logged to ~/.nutanix-aiops/), encrypted credential storage, automatic ETag/If-Match on mutations, automatic pagination on lists, dry-run preview and double-confirm on destructive writes, and undo descriptors for reversible operat…
When to use it
- Inspect cluster health, host utilization, and resiliency across the whole estate
- Root-cause an alert by correlating it with related events and suggested actions
- Safely delete, clone, migrate, or power-cycle VMs with dry-run preview and audit trail
- Run capacity runway forecasts and plan storage or firmware upgrades before hitting limits
The skill document
Nutanix AIops
Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by Nutanix. Product and trademark names belong to their owners. Source at github.com/AIops-tools/Nutanix-AIops under the MIT license.
Governed Nutanix Prism Central (v4 REST API) operations — 51 MCP tools, every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.nutanix-aiops/, token/runaway budget guard, undo-token recording, and descriptive risk-tier labels. The Prism Central password is stored encrypted (~/.nutanix-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk.
What sets it apart from read-only Nutanix MCPs: (1) automatic ETag / If-Match on every mutation — the v4 footgun handled for you; (2) automatic pagination; (3) mixed-hypervisor VM listing (AHV + ESXi, relevant to hypervisor-migration estates); and (4) the governance harness with dry-run + double-confirm on destructive writes.
Standalone: the governance harness is bundled in the package (
nutanix_aiops.governance) — no external skill-family dependency.
What This Skill Does
| Group | Tools | Count | Read / Write |
|---|---|---|---|
| Clusters | cluster_list, cluster_health, host_list, cluster_utilization | 4 | 4 read |
| VMs | list, get, power_on, guest_shutdown, power_off, reboot, create, update, clone, delete, migrate | 11 | 2 read · 9 write |
| Storage | container list / create / update / delete | 4 | 1 read · 3 write |
| Network | subnet list / get / create / delete | 4 | 2 read · 2 write |
| Catalog | image list / delete, category list / create / assign | 5 | 2 read · 3 write |
| Data protection / DR | snapshot list/create/delete/restore, recovery_point_list, protection_domain_list, vm_protect, pd_failover | 8 | 3 read · 5 write |
| Alerts | alert_list, event_list, audit_list, analyze_alert (RCA), alert_acknowledge, alert_resolve | 6 | 4 read · 2 write |
| LCM (upgrades) | lcm_inventory, lcm_precheck, lcm_update | 3 | 1 read · 2 write |
| Capacity | task_list, capacity_runway | 2 | 2 read |
| Diagnostics / RCA | cluster_health_rca, alert_triage_rca | 2 | 2 read |
| Undo | undo_list, undo_apply | 2 | 1 read · 1 write |
| Total | 51 | 24 read · 27 write |
The CLI is a convenience subset; the full 51-tool surface is via the MCP server. See references/capabilities.md for the tool → API-path → returns map.
Quick Install
uv tool install nutanix-aiops
nutanix-aiops init # interactive wizard: PC host/port 9440/username + encrypted password
nutanix-aiops doctor # connectivity + REST-RBAC preflight
When to Use This Skill
- Diagnose the estate in one shot (
diagnose cluster-health): degraded resiliency, storage pools/containers over 80% / 90%, nodes down or missing — worst-first, each finding citing the measured number - Triage the alert backlog (
diagnose alert-triage): per-severity counts, unacknowledged criticals, the oldest unresolved alert and its age - Inspect the estate (
overview,cluster health,cluster util): clusters, hosts, resiliency, utilization - VM lifecycle across AHV + ESXi (
vm list/get/power/create/update/clone), and guarded destructive ops (vm delete,vm migrate) with dry-run + double-confirm - Root-cause an alert (
analyze_alert) — correlate it with related events into a probable-cause + suggested-actions summary - Data protection: snapshots, recovery points, protection domains,
vm_protect,pd_failover - Upgrades (
lcm_inventory→lcm_precheck→lcm_update) and capacity forecasting (capacity_runway)
Do NOT use when the target is a non-Nutanix platform — this skill is Prism Central v4 only. For other infrastructure, use the appropriate other AIops-tools sibling.
Common Workflows
"Something is wrong with the estate" → diagnose, then act
nutanix-aiops diagnose cluster-health→ worst-first findings, e.g.critical · prod-cluster · storage container near full · 93.0% used >= 90.0% thresholdstorage_container_list→ confirm which container it is and what itsmaxCapacityBytes/logicalUsageBytesactually arerecovery_point_list/snapshot_list→ the usual culprit is snapshot sprawl in that containersnapshot_delete <…> --dry-runto preview, then re-run to reclaim space — optionally setNUTANIX_AUDIT_APPROVED_BYfirst to annotate who authorized it; each deletion is audited and records an undo descriptor- Re-run
diagnose cluster-health→ the finding should drop below the 80% warning threshold; the before/after percentages are your evidence
Triage a cluster alert (RCA)
nutanix-aiops diagnose alert-triage(oralert_list) → per-severity counts and the oldest unresolved alert, so you know which extId to open firstanalyze_alert→ probable cause + suggested actions, built by correlating the alert with relatedevent_listrecords- Confirm blast radius with
cluster_health/cluster_utilization, thenalert_acknowledge(oralert_resolveonce fixed)
Safely delete a VM (high-risk, audited)
vm_get→ confirm it's the right VM (and see its ETag)nutanix-aiops vm delete --dry-run→ preview the exactDELETEcall- Optionally annotate who/why:
export NUTANIX_AUDIT_APPROVED_BY=… NUTANIX_AUDIT_RATIONALE=… - Re-run without
--dry-run(double-confirm at the CLI); the call is audited with tier and any approver/rationale supplied
Snapshot sprawl cleanup
recovery_point_list(or per-VMsnapshot_list) → find stale / redundant snapshotssnapshot_delete <…> --dry-runon each candidate → preview- Re-run without dry-run (HIGH risk, double-confirm at the CLI) to reclaim space; each deletion is audited
Capacity runway
cluster_utilization→ current CPU / memory / storage / IOPScapacity_runway→ days-to-full forecast per resource; use it to schedule anlcmexpansion or storage add before you hit the wall
Migrate a VM to another host (reversible)
host_list→ pick the destination host extIdnutanix-aiops vm migrate --dry-run→ preview- Re-run without dry-run (HIGH, double-confirm); the prior host is captured as an undo descriptor so a regression can be reversed
Governance & Safety
The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the account you connect it with (connect with a Prism Central account holding only a read-only (Viewer) role — writes then fail at the server). There is no read-only switch, policy file, or approval gate.
- Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to
~/.nutanix-aiops/audit.db(relocatable viaNUTANIX_AIOPS_HOME): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does. NUTANIX_AUDIT_APPROVED_BY/NUTANIX_AUDIT_RATIONALEare optional annotations recorded on the audit row (who/why); they are never required and never block.- Runaway guard — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker. Disable with
NUTANIX_RUNAWAY_MAX=0. - Every mutation auto-handles ETag / If-Match; every list auto-paginates.
- Destructive writes support
dry_run/--dry-runand, at the CLI, double confirmation. - Reversible writes record an undo descriptor (
vm_update→ prior CPU/memory,vm_migrate→ prior host).
References
references/capabilities.md— full 51-tool + API-path referencereferences/cli-reference.md— CLI command referencereferences/setup-guide.md— onboarding, credentials, REST-RBAC, CE self-testreferences/agent-guardrails.md— which guardrails the harness enforces for you, and a ready-to-paste system prompt for smaller / local models
Questions people ask
- Does this skill support write operations or is it read-only?
- It supports both. Out of 51 tools, 24 are read and 27 are write operations. Write operations include VM lifecycle, storage management, network changes, data protection actions, LCM updates, and alert resolution. Every write is logged to a local audit database and supports dry-run preview and double-confirm at the CLI.
- How does the governance harness work, and do I need extra setup?
- The harness is bundled inside the package under nutanix_aiops.governance — no external dependency required. It adds audit logging, encrypted credential storage (Fernet + scrypt), automatic ETag handling, pagination, and undo descriptors. The audit log lives at ~/.nutanix-aiops/audit.db and records every operation's params, result, status, duration, and risk tier.
- Can I prevent accidental destructive operations?
- Yes. The CLI enforces double-confirmation for high-risk operations and supports --dry-run to preview exact API calls before execution. For an extra layer of safety, connect to Prism Central with a read-only Viewer role — writes then fail at the API server. A runaway circuit breaker also trips if the same call repeats in a tight window.
Related skills
Diagnose which mental domain is holding you back before choosing a cognitive intervention.
Prioritize growth directions with a 2×2 risk framework — pick one bet and commit.
Join video meetings as a voice bot, visual avatar, or avatar with live screen sharing.
Structured framework to analyze your alternatives and acceptable deal range before you commit
Escape the scarcity trap — diagnose bandwidth consumption and design protected slack to restore strategic capacity.
More from zw008
Browse all skillsOperate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.
Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.
Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.
Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.
Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.
Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.