Coding

fabric-aiops

Try it

Query and remediate Cisco Meraki, Catalyst Center, Arista CVP, and UniFi network fabrics through a unified governed tool layer.

What it does

Thirty-four MCP tools spanning four network-controller platforms — Cisco Meraki Dashboard (full read and write), Cisco Catalyst Center (read subset), Arista CloudVision Portal (read subset), and UniFi Network (read subset plus device restart). Every tool is wrapped with an encrypted secret store, a local audit log, and descriptive risk tiers. Three built-in analyses rank WAN uplink health, compute a per-network health score, and surface settings that have drifted from a bound config template. Remediation tools support dry-run preview, reversible inverse capture, and undo.

When to use it

  • Audit an organization's device inventory and health across multiple vendor platforms
  • Root-cause the worst-performing WAN uplinks on MX appliances ranked by loss and latency
  • Score each network in a fleet by a composite health metric to prioritize triage
  • Preview and reverse device attribute updates, VLAN changes, and template rebinds

The skill document

Fabric AIops

Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by Cisco, Meraki, Arista, Ubiquiti, or any network-controller vendor. Product and trademark names belong to their owners. Source at github.com/AIops-tools/Fabric-AIops under the MIT license.

Governed network-fabric controller operations — 34 MCP tools over four platforms (Cisco Meraki Dashboard: full read+write; Cisco Catalyst Center and Arista CloudVision Portal: read subsets; UniFi Network: read subset + device restart), every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.fabric-aiops/, token/runaway budget guard, undo-token recording, and descriptive risk tiers. The controller secret is stored encrypted (~/.fabric-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk.

Standalone: the governance harness is bundled in the package (fabric_aiops.governance) — fabric-aiops has no external skill-family dependency. The test suite is mock-based; no platform has yet been exercised against a live controller (see docs/VERIFICATION.md).

Platform support

Platformplatform:CoverageAuth
Cisco Meraki Dashboardmerakifull (all reads + all 8 writes)API key (Bearer / X-Cisco-Meraki-API-Key)
Cisco Catalyst Centercatalystread subset: sites (as orgs/networks), device+site+client health, issues→alerts, inventory, interface statsusername:password → short-lived X-Auth-Token (auto-refresh on 401)
Arista CloudVision Portalcvpread subset: containers (as orgs/networks), inventory (+ complianceCode drift signal), events→alerts, users→adminsservice-account token (Bearer)
UniFi Networkunifiread subset: sites (as orgs/networks), stat/device inventory+statuses, stat/health, alarms→alerts, stat/sta clients, device port_table→switch ports; plus the device-restart write (cmd/devmgr)API key (X-API-KEY, stateless); base_url = classic https://:8443 or UniFi OS console https:///proxy/network

Ops a platform does not map — and every write on catalyst/cvp (on unifi, every write except reboot) — return a teaching "not supported on `` yet — open an issue or PR" error, never a silent no-op. Full matrix in the repo README.

What This Skill Does

DomainToolsCountRead or Write
Overviewfabric fleet overview11 read
Organizationslist/get, licensing, admins, device statuses, API usage66 read
Networkslist/get, VLANs, health alerts, traffic55 read
Devicesinventory (by model), status, uplinks, switch ports, SSIDs55 read
Clientslist, detail, usage, connectivity44 read
Health (flagship)uplink loss/latency RCA, network health score, config template drift33 read
Remediationreboot, claim, remove, bind, unbind55 write (high)
update device, update VLAN22 write (medium)
blink LEDs11 write (low)
Undolist recorded reversible writes11 read
apply a recorded inverse (governed, single-use, dry_run)11 write (medium)

network_health_score and config_template_drift are injected-only (they score data you already hold); uplink_loss_and_latency_rca accepts injected records for offline analysis, or pulls live from a configured target. Meraki device models carry a product-type prefix: MX appliance, MS switch, MR wireless AP, MV camera, MG cellular gateway.

Quick Install

uv tool install fabric-aiops
fabric-aiops init       # interactive wizard: platform choice (meraki/catalyst/cvp/unifi) + encrypted secret
fabric-aiops doctor

When to Use This Skill

  • Triage an organization (overview): network count + device status/product rollup
  • Find the worst WAN uplinks (health uplink-rca / uplink_loss_and_latency_rca): ranked by loss + latency with a likely cause and action
  • Score fleet health per network (health score / network_health_score): a composite 0-100, worst first, every component shown
  • List/inspect organizations, networks, devices (by model), and clients
  • Reboot/blink a device, update device or VLAN attributes (reversible), claim/remove devices, or bind/unbind a config template — all with dry-run + double-confirm

Do NOT use when the target is OT/industrial equipment (use industrial-aiops), a hypervisor, a storage appliance, a backup product, a container cluster, or device-level CLI/SSH network automation.

If the user wants…Use
Cisco Meraki fabric: uplinks, health, config templates, device lifecyclefabric-aiops (this skill)
Cisco Catalyst Center (DNA Center): site/device/client health, issues, inventoryfabric-aiops (this skill, platform: catalyst)
Arista CloudVision Portal: inventory, compliance drift signal, eventsfabric-aiops (this skill, platform: cvp)
UniFi Network (self-hosted controller / UniFi OS console): site health, alarms, clients, device restartfabric-aiops (this skill, platform: unifi)
OT / industrial edge (Modbus, OPC-UA, PLC, PROFINET)the industrial-aiops line
Hypervisor VM lifecycle (power, snapshot, migrate)a hypervisor ops skill
Container/cluster lifecyclea cluster ops skill

Common Workflows

  1. fabric-aiops health uplink-rca → worst MX WAN uplinks ranked by avg loss + latency, each citing the measured numbers plus a likely cause and action
  2. fabric-aiops health uplink-rca --loss-pct 2 --latency-ms 100 → tighten the thresholds if nothing crosses the defaults but users still complain
  3. fabric-aiops device uplinks → the raw per-appliance uplink statuses across the org (WAN1/WAN2, active vs failover) behind the ranking — confirm the flagged appliance rather than trusting the summary
  4. fabric-aiops network alerts → check whether the controller already raised a matching alert (independent corroboration before you touch anything)
  5. Failure branch: if the RCA returns no uplink records at all, the org has no appliances reporting uplink telemetry, or the API key lacks org-wide read — run fabric-aiops doctor and re-check the org id with fabric-aiops org list rather than assuming the WAN is healthy.

Rank the fleet and fix the worst network's device attributes (reversible)

  1. fabric-aiops overview → org-level rollup: network count and device status/product mix
  2. fabric-aiops health score → composite 0-100 per network, worst first, with every scoring component shown
  3. fabric-aiops org device-statuses → find the offline/alerting devices dragging the worst network's score
  4. fabric-aiops device status → confirm the device before changing it
  5. fabric-aiops remediate update-device '{"name":"branch-ap-01"}' --dry-run → preview the exact PUT /devices/ call; then run without --dry-run (double confirmation). The real before-state is fetched first and recorded as a faithful inverse
  6. Failure branch: wrong attribute or wrong device — fabric-aiops undo list, then fabric-aiops undo apply restores the captured prior attributes. Re-run fabric-aiops device status to confirm the restore landed rather than trusting the undo's success message.

Bring a drifted network back to its config template (reversible)

  1. fabric-aiops network list → the networks in scope and their ids
  2. Pass the template plus its bound networks to config_template_drift(template=..., networks=[...]) → the settings that deviate, per network
  3. fabric-aiops network vlans → confirm the drifted VLAN's current values before changing anything
  4. Fix the specific setting — fabric-aiops remediate update-vlan '{"name":"data"}' --dry-run, then for real — or re-establish the binding itself: fabric-aiops remediate bind --dry-run, then without --dry-run (double confirmation). Both capture the real before-state and record an inverse descriptor (for bind, the inverse is unbind or a rebind to the prior template)
  5. Failure branch: if the rebind makes things worse, fabric-aiops undo apply returns the network to its captured prior binding; fabric-aiops remediate unbind is the manual escape hatch. Re-run config_template_drift to confirm the drift actually cleared instead of trusting the write's success message.

Stage a replacement device into a branch network

  1. fabric-aiops device inventory → confirm the replacement serial is in the org inventory and unassigned
  2. fabric-aiops network get → confirm the target network
  3. fabric-aiops remediate claim --dry-run → preview POST /networks//devices/claim; then run for real (double confirmation) — the inverse (remove from network) is recorded
  4. fabric-aiops remediate blink-leds --duration 30 → low-risk physical confirmation that you are at the right box in the rack
  5. fabric-aiops health score → confirm the network's score recovers once the device reports in
  6. Failure branch: wrong network — fabric-aiops undo apply or fabric-aiops remediate remove . Note fabric-aiops remediate reboot is no undo by construction (a reboot has no safe inverse); it records only the before-state, so use it last, not as a first response.

Offline analysis (no live controller)

  1. Export the org's uplink, device-status, and template data to JSON
  2. Feed it straight to the analysis tools — uplink_loss_and_latency_rca(records=[...]), network_health_score(device_statuses=[...]), config_template_drift(template=..., networks=[...]) — no connection or credentials required
  3. Failure branch: a tool that rejects the injected records means the export is missing fields the analysis needs (loss/latency samples, device status, template settings) — re-export rather than hand-editing, so the findings stay traceable to the controller.

Governance & Safety

The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the account you connect it with (a Meraki API key whose admin has read-only organization access — writes then fail at the controller). There is no read-only switch, policy file, or approval gate.

  • Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to ~/.fabric-aiops/audit.db (relocatable via FABRIC_AIOPS_HOME): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.
  • FABRIC_AUDIT_APPROVED_BY / FABRIC_AUDIT_RATIONALE are optional annotations recorded on the audit row (who/why); they are never required and never block.
  • Runaway guard — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker. Disable with FABRIC_RUNAWAY_MAX=0.
  • Destructive writes support --dry-run / dry_run=True and double confirmation at the CLI.
  • Mutating/reversible writes fetch the real before-state and record an inverse descriptor (update_device/update_network_vlan→restore prior values, claim↔remove, bind↔unbind/rebind); irreversible ops (reboot_device, blink_device_leds) record only the before-state.

References

  • references/capabilities.md — full tool + field reference
  • references/cli-reference.md — CLI command reference
  • references/setup-guide.md — onboarding, credentials, and connectivity

Questions people ask

Which platforms does this skill support?
Meraki Dashboard (full read+write), Catalyst Center (read subset for device/site/client health and inventory), Arista CloudVision Portal (read subset for inventory and events), and UniFi Network (read subset plus device restart on self-hosted controllers). Writes beyond device restart on UniFi and all writes on Catalyst/CVP are not yet mapped.
How are secrets stored?
Controller credentials are encrypted at rest using Fernet with scrypt key derivation in ~/.fabric-aiops/secrets.enc. The encryption key is derived interactively at init time; it is never stored in plaintext.
Can I undo a remediation action?
Reversible operations — update device, update VLAN, claim, remove, bind, unbind — capture the real before-state and record an inverse descriptor. You can list and apply recorded inverses. Irreversible operations (reboot, blink LEDs) record only the before-state with no undo path.

Related skills

Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.

by zw0081 installs1 stars

Diagnose which mental domain is holding you back before choosing a cognitive intervention.

by deciqai1 installs3 stars

Prioritize growth directions with a 2×2 risk framework — pick one bet and commit.

by deciqai2 installs2 stars

End-of-day options analytics ranked against each ticker's own history: IV rank, put/call percentile, skew, max pain, and unusually active contracts.

by thesentitrader2 installs2 stars

Escape the scarcity trap — diagnose bandwidth consumption and design protected slack to restore strategic capacity.

by deciqai1 installs2 stars

Join video meetings as a voice bot, visual avatar, or avatar with live screen sharing.

by johnpatternai22 installs8 stars

More from zw008

Browse all skills

Operate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.

by zw00878 installs1 stars

Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.

by zw00876 installs

Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.

by zw00853 installs

Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.

by zw00851 installs

Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.

by zw00851 installs

Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.

by zw00850 installs