Memory

nutanix-aiops

Try it

Manage Nutanix Prism Central through 51 MCP tools with built-in governance, audit, and undo.

What it does

This skill wraps the Prism Central v4 REST API into 51 MCP tools covering the full Nutanix estate: clusters, hosts, VM lifecycle (AHV and ESXi), storage containers, networking, images, categories, data protection (snapshots, recovery points, protection domains, failover), alerts with RCA analysis, LCM firmware upgrades, and capacity runway forecasting. It ships with a bundled governance harness — no external dependency — that adds a local audit log (every operation logged to ~/.nutanix-aiops/), encrypted credential storage, automatic ETag/If-Match on mutations, automatic pagination on lists, dry-run preview and double-confirm on destructive writes, and undo descriptors for reversible operat…

When to use it

  • Inspect cluster health, host utilization, and resiliency across the whole estate
  • Root-cause an alert by correlating it with related events and suggested actions
  • Safely delete, clone, migrate, or power-cycle VMs with dry-run preview and audit trail
  • Run capacity runway forecasts and plan storage or firmware upgrades before hitting limits

The skill document

Nutanix AIops

Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by Nutanix. Product and trademark names belong to their owners. Source at github.com/AIops-tools/Nutanix-AIops under the MIT license.

Governed Nutanix Prism Central (v4 REST API) operations — 51 MCP tools, every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.nutanix-aiops/, token/runaway budget guard, undo-token recording, and descriptive risk-tier labels. The Prism Central password is stored encrypted (~/.nutanix-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk.

What sets it apart from read-only Nutanix MCPs: (1) automatic ETag / If-Match on every mutation — the v4 footgun handled for you; (2) automatic pagination; (3) mixed-hypervisor VM listing (AHV + ESXi, relevant to hypervisor-migration estates); and (4) the governance harness with dry-run + double-confirm on destructive writes.

Standalone: the governance harness is bundled in the package (nutanix_aiops.governance) — no external skill-family dependency.

What This Skill Does

GroupToolsCountRead / Write
Clusterscluster_list, cluster_health, host_list, cluster_utilization44 read
VMslist, get, power_on, guest_shutdown, power_off, reboot, create, update, clone, delete, migrate112 read · 9 write
Storagecontainer list / create / update / delete41 read · 3 write
Networksubnet list / get / create / delete42 read · 2 write
Catalogimage list / delete, category list / create / assign52 read · 3 write
Data protection / DRsnapshot list/create/delete/restore, recovery_point_list, protection_domain_list, vm_protect, pd_failover83 read · 5 write
Alertsalert_list, event_list, audit_list, analyze_alert (RCA), alert_acknowledge, alert_resolve64 read · 2 write
LCM (upgrades)lcm_inventory, lcm_precheck, lcm_update31 read · 2 write
Capacitytask_list, capacity_runway22 read
Diagnostics / RCAcluster_health_rca, alert_triage_rca22 read
Undoundo_list, undo_apply21 read · 1 write
Total5124 read · 27 write

The CLI is a convenience subset; the full 51-tool surface is via the MCP server. See references/capabilities.md for the tool → API-path → returns map.

Quick Install

uv tool install nutanix-aiops
nutanix-aiops init       # interactive wizard: PC host/port 9440/username + encrypted password
nutanix-aiops doctor     # connectivity + REST-RBAC preflight

When to Use This Skill

  • Diagnose the estate in one shot (diagnose cluster-health): degraded resiliency, storage pools/containers over 80% / 90%, nodes down or missing — worst-first, each finding citing the measured number
  • Triage the alert backlog (diagnose alert-triage): per-severity counts, unacknowledged criticals, the oldest unresolved alert and its age
  • Inspect the estate (overview, cluster health, cluster util): clusters, hosts, resiliency, utilization
  • VM lifecycle across AHV + ESXi (vm list/get/power/create/update/clone), and guarded destructive ops (vm delete, vm migrate) with dry-run + double-confirm
  • Root-cause an alert (analyze_alert) — correlate it with related events into a probable-cause + suggested-actions summary
  • Data protection: snapshots, recovery points, protection domains, vm_protect, pd_failover
  • Upgrades (lcm_inventory → lcm_precheck → lcm_update) and capacity forecasting (capacity_runway)

Do NOT use when the target is a non-Nutanix platform — this skill is Prism Central v4 only. For other infrastructure, use the appropriate other AIops-tools sibling.

Common Workflows

"Something is wrong with the estate" → diagnose, then act

  1. nutanix-aiops diagnose cluster-health → worst-first findings, e.g. critical · prod-cluster · storage container near full · 93.0% used >= 90.0% threshold
  2. storage_container_list → confirm which container it is and what its maxCapacityBytes / logicalUsageBytes actually are
  3. recovery_point_list / snapshot_list → the usual culprit is snapshot sprawl in that container
  4. snapshot_delete <…> --dry-run to preview, then re-run to reclaim space — optionally set NUTANIX_AUDIT_APPROVED_BY first to annotate who authorized it; each deletion is audited and records an undo descriptor
  5. Re-run diagnose cluster-health → the finding should drop below the 80% warning threshold; the before/after percentages are your evidence

Triage a cluster alert (RCA)

  1. nutanix-aiops diagnose alert-triage (or alert_list) → per-severity counts and the oldest unresolved alert, so you know which extId to open first
  2. analyze_alert → probable cause + suggested actions, built by correlating the alert with related event_list records
  3. Confirm blast radius with cluster_health / cluster_utilization, then alert_acknowledge (or alert_resolve once fixed)

Safely delete a VM (high-risk, audited)

  1. vm_get → confirm it's the right VM (and see its ETag)
  2. nutanix-aiops vm delete --dry-run → preview the exact DELETE call
  3. Optionally annotate who/why: export NUTANIX_AUDIT_APPROVED_BY=… NUTANIX_AUDIT_RATIONALE=…
  4. Re-run without --dry-run (double-confirm at the CLI); the call is audited with tier and any approver/rationale supplied

Snapshot sprawl cleanup

  1. recovery_point_list (or per-VM snapshot_list ) → find stale / redundant snapshots
  2. snapshot_delete <…> --dry-run on each candidate → preview
  3. Re-run without dry-run (HIGH risk, double-confirm at the CLI) to reclaim space; each deletion is audited

Capacity runway

  1. cluster_utilization → current CPU / memory / storage / IOPS
  2. capacity_runway → days-to-full forecast per resource; use it to schedule an lcm expansion or storage add before you hit the wall

Migrate a VM to another host (reversible)

  1. host_list → pick the destination host extId
  2. nutanix-aiops vm migrate --dry-run → preview
  3. Re-run without dry-run (HIGH, double-confirm); the prior host is captured as an undo descriptor so a regression can be reversed

Governance & Safety

The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the account you connect it with (connect with a Prism Central account holding only a read-only (Viewer) role — writes then fail at the server). There is no read-only switch, policy file, or approval gate.

  • Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to ~/.nutanix-aiops/audit.db (relocatable via NUTANIX_AIOPS_HOME): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.
  • NUTANIX_AUDIT_APPROVED_BY / NUTANIX_AUDIT_RATIONALE are optional annotations recorded on the audit row (who/why); they are never required and never block.
  • Runaway guard — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker. Disable with NUTANIX_RUNAWAY_MAX=0.
  • Every mutation auto-handles ETag / If-Match; every list auto-paginates.
  • Destructive writes support dry_run / --dry-run and, at the CLI, double confirmation.
  • Reversible writes record an undo descriptor (vm_update → prior CPU/memory, vm_migrate → prior host).

References

  • references/capabilities.md — full 51-tool + API-path reference
  • references/cli-reference.md — CLI command reference
  • references/setup-guide.md — onboarding, credentials, REST-RBAC, CE self-test
  • references/agent-guardrails.md — which guardrails the harness enforces for you, and a ready-to-paste system prompt for smaller / local models

Questions people ask

Does this skill support write operations or is it read-only?
It supports both. Out of 51 tools, 24 are read and 27 are write operations. Write operations include VM lifecycle, storage management, network changes, data protection actions, LCM updates, and alert resolution. Every write is logged to a local audit database and supports dry-run preview and double-confirm at the CLI.
How does the governance harness work, and do I need extra setup?
The harness is bundled inside the package under nutanix_aiops.governance — no external dependency required. It adds audit logging, encrypted credential storage (Fernet + scrypt), automatic ETag handling, pagination, and undo descriptors. The audit log lives at ~/.nutanix-aiops/audit.db and records every operation's params, result, status, duration, and risk tier.
Can I prevent accidental destructive operations?
Yes. The CLI enforces double-confirmation for high-risk operations and supports --dry-run to preview exact API calls before execution. For an extra layer of safety, connect to Prism Central with a read-only Viewer role — writes then fail at the API server. A runaway circuit breaker also trips if the same call repeats in a tight window.

Related skills

Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.

by zw0081 installs1 stars

Diagnose which mental domain is holding you back before choosing a cognitive intervention.

by deciqai1 installs3 stars

Prioritize growth directions with a 2×2 risk framework — pick one bet and commit.

by deciqai2 installs2 stars

Join video meetings as a voice bot, visual avatar, or avatar with live screen sharing.

by johnpatternai22 installs8 stars

Structured framework to analyze your alternatives and acceptable deal range before you commit

by deciqai1 installs3 stars

Escape the scarcity trap — diagnose bandwidth consumption and design protected slack to restore strategic capacity.

by deciqai1 installs2 stars

More from zw008

Browse all skills

Operate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.

by zw00878 installs1 stars

Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.

by zw00876 installs

Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.

by zw00853 installs

Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.

by zw00851 installs

Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.

by zw00851 installs

Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.

by zw00850 installs