Documents

vmware-debug

Try it

Diagnose VMware issues by correlating events, ranking root causes, and suggesting next steps.

What it does

A read-only diagnostic skill for VMware/vSphere/ESXi environments. You provide events and logs (collected by companion skills), and it merges them into a unified timeline, detects spikes, scores root-cause hypotheses, and tells you the most likely cause first. It never makes changes — it stops at a diagnosis and a recommended next check, routing actual fixes to vmware-aiops or vmware-pilot.

When to use it

  • Your ESXi host shows an error and you need to find the root cause
  • Multiple symptoms across storage, network, or compute with no clear connection
  • A pile of logs and you don't know where to start
  • A VM that won't power on — figure out what broke before fixing it

The skill document

VMware Debug

Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom Inc. "VMware" and "vSphere" are trademarks of Broadcom. Source is publicly auditable under the MIT license.

The diagnostic brain of the VMware skill family. You bring the symptom; this skill runs the investigation and points at the root cause. It reads and reasons — it never writes. Companion skills do the data collection and the fixing.

What This Skill Does

CategoryWhatRead or Write
Incident correlationMerge events from many sources into one timeline, detect spikesRead
Root-cause rankingScore symptom clusters, surface the most likely cause firstRead
Next-check ideasSuggest exactly what to look at next (which skill/tool) when you're stuckRead
Remediation routingHand the fix to vmware-aiops (single) or vmware-pilot (gated, multi-step)Read (routes only)

Zero write tools. Zero network access of its own. It correlates data the agent has already gathered with the other skills' read tools.

Quick Install

uv tool install vmware-debug
vmware-debug categories          # see what it can diagnose

When to Use This Skill

Use it when there is a problem to solve: an error message, a stack of logs, an alarm storm, "my VM won't power on", "storage feels slow", "the host disconnected".

  • Need raw inventory/health with no incident? → vmware-monitor
  • Need to actually run a fix? → vmware-aiops (single op) or vmware-pilot (gated workflow)
  • Need metrics/anomalies? → vmware-aria; centralized logs? → vmware-log-insight

Do NOT use when there is nothing wrong (routine listing → monitor), or when the user wants the fix executed (→ aiops/pilot). This skill stops at the diagnosis and a recommended plan.

Symptom touchesPull signals fromThen
Storage / datastore / vSANvmware-storage, vmware-log-insightrank → route fix to aiops/pilot
Network / firewall / vMotionvmware-nsx, vmware-nsx-securityrun traceflow, check DFW
CPU / memory contentionvmware-aria (metrics/anomalies)rightsizing via pilot
HA / DRS / clustervmware-monitor, vmware-aiopscluster remediation via pilot
Power / clone / snapshotvmware-aiops, vmware-monitortask status, then fix via aiops
Auth / cert / logincheck creds & cert; (security)fix config/.env

Common Workflows

1. "Here's a pile of logs / alarms — what broke?"

  1. Collect events with the data-source skills (e.g. vmware-monitor event_list --vm web01 --since 1h, vmware-log-insight log_search ..., vmware-aria alert_query ...).
  2. Pass them all to incident_timeline (envelope below). Read the top hypothesis + next_checks.
  3. Follow next_checks to pull more targeted data; re-run incident_timeline to confirm.
  4. Failure branch — no events come back: the affected target may be unreachable. Run the source skill's doctor/health first; a 503/timeout is a signal (platform not ready), not a dead end.
  5. Produce a diagnosis + recommended fix. Route execution to aiops/pilot. Do not fix here.

2. "I don't even know what to check"

  1. Run list_symptom_categories (or vmware-debug categories) to see the catalogue.
  2. Describe the symptom; map it to a category; the suggested_check tells you which skill/tool to run first.
  3. Collect → incident_timeline → narrow. Loop until one hypothesis dominates.

3. Hand off the fix (advisor → executor, like vmware-harden)

  1. Debug emits a structured diagnosis + a proposed remediation (steps).
  2. Single, low-risk fix → call the matching vmware-aiops tool (it has its own double-confirm).
  3. Multi-step / needs approval / cross-skill → submit the plan to vmware-pilot, which owns the state machine, approval gate, rollback, and audit.
  4. Failure branch — fix is ambiguous or risky: stop and present the hypotheses to the user; never guess-execute.

Usage Mode

  • MCP (in an agent): the agent calls the other skills' read tools, then incident_timeline to correlate. This is the primary mode — that's where the cross-skill "联动" happens.
  • CLI (humans): vmware-debug triage --events events.json correlates a JSON array you collected yourself.

MCP Tools (2 — 2 read, 0 write)

ToolWhat
incident_timeline[READ] Correlate pre-fetched events → timeline + spikes + ranked hypotheses + next-check ideas
list_symptom_categories[READ] List recognised symptom categories + what to check for each

List envelope (output of list_symptom_categories): {items, returned, limit, total, truncated, hint} — read the rows from items. truncated is always false here, which is the point: it states that the catalogue is complete instead of leaving you to infer it.

Event envelope (input to incident_timeline): {ts, source, severity, entity, text, fields}. See references/event-envelope.md. The agent normalises each source's events into this shape; debug stays source-agnostic and has no dependency on the other packages.

Read-Only by Design

Both tools here are reads — zero write tools, zero network access of its own. Running with local or small models? See references/agent-guardrails.md.

CLI Quick Reference

vmware-debug categories                        # what can it diagnose
vmware-debug triage --events events.json       # correlate a collected event set
cat events.json | vmware-debug triage          # or via stdin
vmware-debug mcp                                # start stdio MCP server (proxy-safe)

Troubleshooting

  • incident_timeline raises "event[N] could not be normalised" — event N is missing a timestamp or has an unparseable one. Every event needs ts (ISO-8601, epoch seconds, or millis).
  • All hypotheses come back "uncategorized" — the symptom isn't in the catalogue yet; widen the window and pull from another source (aria anomalies, log-insight). Consider adding a signature (see references/routing.md).
  • No spikes detected on an obvious burst — you need ≥3 time bins for a baseline; shrink bin_seconds.
  • It won't execute the fix — by design. Route to vmware-aiops or vmware-pilot.

Audit & Safety

Read-only by construction: no write tools, no network, nothing executed. Remediation is always routed to aiops/pilot, where the double-confirm / approval / audit gates live (audit DB ~/.vmware/audit.db). Policy rules scope by environment; debug has no config and no connection to declare one about, so it reports a constant local — nothing here touches a remote VMware estate. See references/setup-guide.md.

License

MIT.

Questions people ask

Does this skill fix the problem?
No. It only diagnoses. It is read-only by design — zero write tools, zero network access. It produces a diagnosis and next-check plan, then routes execution to vmware-aiops (single fix) or vmware-pilot (multi-step or gated remediation).
What data does it need?
A JSON array of events in a normalized envelope format (timestamp, source, severity, entity, text, fields). The document defines the exact shape. You gather events first using companion skills like vmware-monitor or vmware-log-insight, then pass them to incident_timeline.
Can I use it for routine health checks?
No. If there is no problem to solve, use vmware-monitor instead. vmware-debug is designed for incidents: error messages, alarm storms, slow or failed VMs, and log dumps.

Related skills

Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.

by zw0081 installs1 stars

Escape the scarcity trap — diagnose bandwidth consumption and design protected slack to restore strategic capacity.

by deciqai1 installs2 stars

Diagnose which mental domain is holding you back before choosing a cognitive intervention.

by deciqai1 installs3 stars

Join video meetings as a voice bot, visual avatar, or avatar with live screen sharing.

by johnpatternai22 installs8 stars

Operate TaskTime Pro through MCP: manage tasks, track time, handle expenses, and prepare invoices from a paired browser session.

by tasktimepro1 installs1 stars

Prioritize growth directions with a 2×2 risk framework — pick one bet and commit.

by deciqai2 installs2 stars

More from zw008

Browse all skills

Operate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.

by zw00878 installs1 stars

Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.

by zw00876 installs

Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.

by zw00853 installs

Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.

by zw00851 installs

Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.

by zw00851 installs

Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.

by zw00850 installs