Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.
Memory
xcpng-aiops
Try itFleet health overview, root-cause analyses, and governed writes for XCP-ng through Xen Orchestra.
What it does
Operate an XCP-ng virtualization fleet through Xen Orchestra's REST API. Includes 29 MCP tools covering VMs, hosts, pools, storage repositories, snapshots, backups, and tasks. Four RCA analyses handle VM health, SR usage, backup failures, and pool patch/HA posture. Every write operation is governed by a bundled audit log, policy engine, undo-token recording, and risk-tier system. Auth tokens are encrypted with Fernet and scrypt. Requires Xen Orchestra 5.x with /rest/v0 endpoint.
When to use it
- Triage an XCP-ng fleet: view pools, hosts, VMs by state, SR fullness, and recent backup failures in one shot
- Root-cause a failing VM: halted unexpectedly, guest tools missing, CPU/memory pressure from RRD stats
- Investigate backup failures: classify as vdi-chain, quiesce, transport, or storage-full issues
- Patch a pool safely: check version skew and HA posture before migrating VMs off a host
The skill document
XCP-ng AIops
Disclaimer: This is a community-maintained open-source project and is not affiliated with, endorsed by, or sponsored by Vates, the XCP-ng project, or the Xen Orchestra project. "XCP-ng", "Xen Orchestra", and "Xen" are trademarks of their owners. Source code is publicly auditable at github.com/AIops-tools/XCPng-AIops under the MIT license.
Governed XCP-ng operations via Xen Orchestra's REST API — 29 MCP tools, every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.xcpng-aiops/, policy engine, token/runaway budget guard, undo-token recording, and descriptive risk tiers. The XO authentication token is stored encrypted (~/.xcpng-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk.
Requires a Xen Orchestra instance (5.x with
/rest/v0) — XO is the management plane; per-host XAPI is out of scope for v0.1. Standalone: the governance harness is bundled in the package (xcpng_aiops.governance) — xcpng-aiops has no external skill-family dependency. Coverage is common operations, not exhaustive; verification status and the live-run checklist are indocs/VERIFICATION.md.
What This Skill Does
| Category | Tools | Count | Read or Write |
|---|---|---|---|
| Overview | fleet health overview | 1 | 1 read |
| VMs | list, get, RRD stats, health RCA | 4 | 4 read |
| start, stop, reboot, migrate | 4 | 4 write (medium) | |
| Hosts | list, get | 2 | 2 read |
| Pools | list, get, patch & HA posture RCA | 3 | 3 read |
| SRs / VDIs | list, get, VDI list (orphan filter), usage RCA | 4 | 4 read |
| rescan | 1 | 1 write (medium) | |
| Snapshots | list | 1 | 1 read |
| create (medium), delete (high), revert (high) | 3 | 3 write | |
| Backups | jobs, run logs, failure RCA | 3 | 3 read |
| Tasks | list | 1 | 1 read |
Quick Install
uv tool install xcpng-aiops
xcpng-aiops init # interactive wizard: XO URL + encrypted token
xcpng-aiops doctor # XO reachability + token validity + pool count
When to Use This Skill
- Triage an XCP-ng fleet (
overview): pools, hosts, VMs by state, SRs near full, recent backup failures - Root-cause an unhealthy VM (
vm health-rca): halted unexpectedly, paused, guest tools missing, CPU/memory pressure - Root-cause storage pressure (
sr usage-rca): SRs ranked near-full, thin-provision overcommit, orphaned VDIs with reclaimable bytes - Root-cause backup failures (
backup failure-rca): vdi-chain / quiesce / transport / storage-full classification - Check patch & HA posture (
pool posture): missing patches, pending reboots, version skew, HA state - Snapshot a VM before a risky change; start/stop/reboot/migrate VMs under governance
Do NOT use when the target is not an XCP-ng fleet managed by Xen Orchestra — other hypervisors (Do NOT use for Proxmox VE — use proxmox-aiops), NAS/storage appliances, backup software suites, Kubernetes/containers, and network devices are out of scope for this skill.
Related Skills — Skill Routing
| If the user wants… | Use |
|---|---|
| XCP-ng VMs / hosts / pools / SRs / snapshots / XO backups | xcpng-aiops (this skill) |
| Proxmox VE operations | proxmox-aiops |
| NAS/storage appliance operations | a storage-appliance ops skill |
| Backup-software suite job/restore operations | a backup-software ops skill |
| Container/cluster lifecycle | a cluster ops skill |
Common Workflows
Each recipe starts from a read or an RCA and ends in a governed write. Every
write accepts --dry-run; destructive ones also double-confirm.
1. "Patch a pool without breaking live migration"
xcpng-aiops overview→ fleet snapshot: pools, hosts, VMs by state, SRs near full, recent backup failures.xcpng-aiops pool posture→ the RCA: hosts missing patches, hosts pending reboot, version skew across the pool's hosts, and multi-host pools without HA.xcpng-aiops host missing-patches→ what exactly is outstanding on the host you plan to take first.xcpng-aiops vm list --state Running→ the VMs that must move off that host.- For each:
xcpng-aiops vm migrate --dry-run, then re-run for real (double confirm; the REAL source host is captured before the move and the inverse "migrate back" is recorded). xcpng-aiops undo list→ confirm a migrate-back token exists for every VM you moved, before you touch the host.- Patch and reboot the host in XO, then
xcpng-aiops pool postureagain to confirm the skew cleared.
Failure branch: if pool posture reports version skew before you start, stop — live migration between mismatched host versions can be refused or unsafe. Bring the hosts to a common version first. If a migration fails mid-run, do not retry blindly: xcpng-aiops vm get to see where the VM actually landed, and xcpng-aiops task list for the failing XO task, since a half-finished migrate leaves the VM on one side or the other.
2. "This VM keeps going unhealthy"
xcpng-aiops vm health-rca→ findings with cause + action: halted unexpectedly (auto-poweron / HA restart priority set), paused or suspended VMs, running VMs without guest tools, CPU/memory pressure from RRD stats.xcpng-aiops vm get→ the VM's configuration and current power state.xcpng-aiops vm stats→ the RRD series behind a pressure finding, so you confirm sustained pressure rather than one spike.- If it is halted and should be running:
xcpng-aiops vm start(the inversevm_stopis recorded). - If it is wedged and needs a bounce:
xcpng-aiops vm reboot --dry-run, then for real (double confirm; add--forceonly for a hard reboot — no undo either way). xcpng-aiops vm health-rcaagain to confirm the finding cleared.
Failure branch: if a clean shutdown or clean reboot hangs, the finding "no guest tools" is usually the real cause — clean actions need the guest agent. Do not escalate straight to --force; a hard action risks filesystem damage. Snapshot first (recipe 3), then use --force deliberately.
3. "Snapshot before a risky change, and roll back cleanly"
xcpng-aiops vm list→ confirm the exact VM uuid.xcpng-aiops sr usage-rca→ make sure the SR has room; snapshots grow it, and a snapshot on a near-full SR is how you take the pool down.xcpng-aiops snapshot create pre-change→ XO returns the new snapshot's id, and an inversesnapshot_deletefor that id is recorded.xcpng-aiops snapshot list --vm→ confirm the snapshot exists before you change anything.- Make your change. If it went wrong:
xcpng-aiops snapshot revert(double confirm — replaces current state, IRREVERSIBLE, no undo). - When you are satisfied:
xcpng-aiops snapshot delete --dry-run, then without--dry-run(double confirm — IRREVERSIBLE, BEFORE state captured for the audit record, no undo).
Failure branch: if sr usage-rca flags the SR as near-full or thin-provision overcommitted, do not snapshot — reclaim first (recipe 4). If a revert is refused or leaves the VM halted, check xcpng-aiops task list for the XO task; and never leave snapshots stacked long-term, because unmerged chains are the usual root cause of the vdi-chain backup failures in recipe 4.
4. "Backups have been failing every night and storage is filling up"
xcpng-aiops backup failure-rca→ failed/skipped/interrupted runs grouped by job and classified: vdi-chain (coalesce not finished), quiesce (guest VSS), transport (remote unreachable), storage-full, unknown.xcpng-aiops backup logs -n 20→ the raw recent runs behind that classification.xcpng-aiops sr usage-rca→ SRs ranked by physical fullness, thin-provision overcommit, and orphaned VDIs (attached to no VM) with reclaimable bytes per SR.xcpng-aiops sr vdis --sr --orphaned-only→ the specific orphaned VDIs worth reclaiming on that SR.- For a vdi-chain classification: let the coalesce finish, stop stacking snapshots (
xcpng-aiops snapshot list), thenxcpng-aiops sr rescan --dry-runand for real (lowest-impact write) so XO re-reads the SR. - For storage-full: reclaim space, then re-run
sr usage-rcato confirm the SR dropped below the near-full threshold. xcpng-aiops backup logs -n 20after the next scheduled run to confirm it went green.
Failure branch: a transport classification is not an XCP-ng problem — the backup remote is unreachable, so fix the remote in the XO UI (Settings → Remotes); rescanning the SR will not help. A quiesce classification means the guest agent could not freeze the filesystem: fix guest tools on that VM rather than disabling quiesce fleet-wide. If sr rescan does not shrink the chain, the coalesce is still running — wait rather than rescanning in a loop, which will trip the runaway budget guard.
Usage Mode
| Scenario | Recommended | Why |
|---|---|---|
| Local/small models | CLI | fewer tokens than MCP |
| Cloud models (Claude, GPT) | Either | MCP gives structured JSON I/O |
| Automated pipelines | MCP | type-safe parameters, audited |
MCP Tools (29 — 19 read, 8 write, 2 undo)
| Category | Tools | R/W |
|---|---|---|
| Overview | overview | Read |
| VMs | vm_list, vm_get, vm_stats, vm_health_rca | Read |
vm_start, vm_stop, vm_reboot, vm_migrate | Write | |
| Hosts | host_list, host_get | Read |
| Pools | pool_list, pool_get, pool_patch_ha_posture | Read |
| SRs / VDIs | sr_list, sr_get, vdi_list, sr_usage_rca | Read |
sr_rescan | Write | |
| Snapshots | snapshot_list | Read |
snapshot_create, snapshot_delete, snapshot_revert | Write | |
| Backups | backup_job_list, backup_log_list, backup_failure_rca | Read |
| Tasks | task_list | Read |
| Undo | undo_list, undo_apply | Read + replay |
Harness features that light up: vm_start↔vm_stop record each other as inverses (with _undo_id); vm_migrate captures the REAL source host BEFORE moving and records "migrate back"; snapshot_create captures the REAL snapshot id from the XO response and records "delete THAT snapshot". snapshot_delete and snapshot_revert are risk_level=high, capture BEFORE state, and declare no undo (irreversible). Every write takes dry_run=True (may read, never writes; no undo; audited). All 29 tools are audit-logged under ~/.xcpng-aiops/ and pass through the budget/runaway guard, each carrying a descriptive risk tier into its audit row. Start any triage with overview.
CLI Quick Reference
xcpng-aiops init # onboarding wizard (encrypted XO token)
xcpng-aiops overview [--target ] # fleet health summary
xcpng-aiops vm list [--state Running] [--pool ]
xcpng-aiops vm get
xcpng-aiops vm stats [-g minutes]
xcpng-aiops vm health-rca [] # RCA: cause + action
xcpng-aiops vm start [--dry-run]
xcpng-aiops vm stop [--force] [--dry-run] # double confirm; refuses the declared XO VM
xcpng-aiops vm reboot [--force] [--dry-run] # double confirm
xcpng-aiops vm migrate [--dry-run] # double confirm
xcpng-aiops host list / get / missing-patches
xcpng-aiops pool list / get
xcpng-aiops pool posture [] # RCA: patches / reboots / skew / HA
xcpng-aiops sr list / get
xcpng-aiops sr vdis [--sr ] [--orphaned-only]
xcpng-aiops sr usage-rca # RCA: near-full / overcommit / orphans
xcpng-aiops sr rescan [--dry-run]
xcpng-aiops snapshot list [--vm ]
xcpng-aiops snapshot create [--dry-run]
xcpng-aiops snapshot delete [--dry-run] # double confirm, IRREVERSIBLE
xcpng-aiops snapshot revert [--dry-run] # double confirm, IRREVERSIBLE
xcpng-aiops backup jobs / logs [-n 50]
xcpng-aiops backup failure-rca [-n 50] # RCA: vdi-chain / quiesce / transport
xcpng-aiops task list [--status failure]
xcpng-aiops secret set / list / rm / migrate / rotate-password
xcpng-aiops doctor # XO reachability + token + pool count
xcpng-aiops mcp # start MCP server (stdio)
See references/cli-reference.md for the full command list, and
references/agent-guardrails.md when driving these tools with a smaller /
local model (the guardrails the tool enforces for you, and a ready system prompt).
Troubleshooting
"Config file not found"
Run xcpng-aiops init to set up your first target (writes ~/.xcpng-aiops/config.yaml and stores the XO token encrypted).
"No XO authentication token for target ''"
Add it to the encrypted store: xcpng-aiops secret set (prompts hidden), or run xcpng-aiops init. Create the token in the XO UI (user menu → Personal tokens) or with xo-cli --createToken. For non-interactive use (MCP/CI), also export XCPNG_AIOPS_MASTER_PASSWORD so the store can be unlocked without a prompt.
"Master password not set" / "Wrong master password"
The encrypted store ~/.xcpng-aiops/secrets.enc is unlocked by XCPNG_AIOPS_MASTER_PASSWORD (or an interactive prompt). If you forgot it, delete secrets.enc and re-run xcpng-aiops init. Rotate it with xcpng-aiops secret rotate-password.
"Authentication/authorization failed (401/403)"
The XO token is wrong, expired, or revoked, or the XO account lacks permission. Regenerate the token in the XO UI (user menu → Personal tokens) and update it: xcpng-aiops secret set .
"Could not reach Xen Orchestra … check the XO URL"
Confirm the XO web UI is reachable at the configured url and that api_path is /rest/v0 (XO 5.x). For self-signed certificates set verify_ssl: false on the target (lab only).
"Resource not found (404)"
The VM/SR/snapshot uuid is stale, or this XO release lacks the endpoint. List the parent collection first (vm list, sr list, snapshot list) to get a current uuid.
Doctor says "manages no pools yet"
Your XO instance is reachable but has no XCP-ng servers connected — add them in the XO UI (Settings → Servers).
Audit & Safety
The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the Xen Orchestra account whose token you connect it with (give that XO user a read-only ACL or scope its token down — writes then fail at Xen Orchestra). There is no read-only switch, policy file, or approval gate.
- Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to
~/.xcpng-aiops/audit.db(relocatable viaXCPNG_AIOPS_HOME): params (secrets redacted), result, status, duration, and the risk tier. The CLI writes the same row the MCP path does. - The XO token is stored encrypted in
~/.xcpng-aiops/secrets.enc(Fernet/AES-128 + scrypt key derivation; chmod 600) — never plaintext on disk; the master password is never stored, only a per-store salt + ciphertext. XCPNG_AUDIT_APPROVED_BY/XCPNG_AUDIT_RATIONALEare optional annotations recorded on the audit row (who/why); they are never required and never block.- Budget / runaway guard — a safety backstop, not authorization: caps cumulative tool calls and wall-time, and trips on tight task-poll loops.
- Writes support
--dry-run/dry_run=Trueand double confirmation at the CLI; CLI writes execute through the same governed tools, so they are audited + undo-recorded. - Reversible writes (
vm_start/vm_stop/vm_migrate/snapshot_create) capture the real before-state and record a replayable inverse descriptor;snapshot_delete/snapshot_revertarerisk=high, irreversible, and declare no undo. - Risk tier is a descriptive label on the audit row derived from
risk_level; it gates nothing.
The harness is bundled in the package — no external dependency, no manual setup. See references/setup-guide.md for security details.
Contributing & feature requests
Coverage is intentionally focused. Missing a capability you need, or hit an endpoint that differs on your Xen Orchestra version? Open an issue or pull request at github.com/AIops-tools/XCPng-AIops — feature requests, contributions, and comments are all welcome.
License
Questions people ask
- What hypervisors does this support?
- XCP-ng fleets managed by Xen Orchestra. Proxmox VE, VMware, Hyper-V, and bare-metal are out of scope.
- What happens if a VM migration fails mid-operation?
- The tool records the pre-migration source host as an undo token. Use undo_apply to migrate the VM back, then check task_list for the failing XO task.
- How are destructive operations handled?
- Snapshot delete and snapshot revert are high-risk operations. They capture the before state, require double confirmation, and declare no undo. All writes accept --dry-run first.
Related skills
Read, create, edit, and fix .xlsx, .csv, and .tsv files with formulas, formatting, and cleaned data.
Turn China 3C launch inputs into executable routes, messaging, channel actions, risk checks, and review decisions.
Diagnose which mental domain is holding you back before choosing a cognitive intervention.
Prioritize growth directions with a 2×2 risk framework — pick one bet and commit.
Escape the scarcity trap — diagnose bandwidth consumption and design protected slack to restore strategic capacity.
More from zw008
Browse all skillsOperate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.
Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.
Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.
Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.
Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.
Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.