Integrations

monitoring-aiops

Try it

Query and operate SolarWinds Orion, PRTG, and Zabbix NOCs with 42 governance-wrapped tools for reads, writes, and undo.

What it does

A monitoring NOC operations toolkit spanning SolarWinds Orion (SWIS REST + SWQL), Paessler PRTG (web API), and Zabbix 6.x/7.x (JSON-RPC). Delivers NOC overview snapshots, deduplicated alert rollups, a validated read-only SWQL passthrough, and platform-specific health queries across nodes, interfaces, volumes, and applications. Write operations—including acknowledge, mute/unmute, schedule maintenance, and pause/resume—are guarded with audit logging and undo tokens. Credentials stored encrypted locally.

When to use it

  • Get a NOC overview snapshot when on-call
  • Triage an alert storm with deduplication and rollup
  • Query SolarWinds/PRTG/Zabbix for node, interface, or volume health
  • Plan maintenance with time-boxed suppression and undo support

The skill document

Monitoring AIops

Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by SolarWinds, Paessler, Zabbix, or any monitoring vendor. SolarWinds, Orion, SWQL, THWACK, PRTG, Paessler and Zabbix are trademarks of their respective owners. Source at github.com/AIops-tools/Monitoring-AIops under the MIT license.

Governed network / infrastructure monitoring operations — 42 MCP tools across SolarWinds Orion (SWIS REST + SWQL), Paessler PRTG (web API), and Zabbix 6.x/7.x (JSON-RPC 2.0), every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.monitoring-aiops/, policy engine, token/runaway budget guard, undo-token recording, and risk-tier labelling on the audit trail. One config can span all NOCs. The Orion password / PRTG API token / Zabbix API token is stored encrypted (~/.monitoring-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk.

Standalone: the governance harness is bundled in the package (monitoring_aiops.governance) — no external skill-family dependency. PRTG's free Freeware edition and an open-source Zabbix appliance are the easiest live checks; SolarWinds is trial-only past 30 days (largest verification debt — see docs/VERIFICATION.md).

What This Skill Does

GroupPlatformToolsCountR/W
SWQLSolarWindslibrary, canned, query (SELECT-only passthrough)3read
Alertsallactive_alerts (dedup/rollup), alert_acknowledge21 read, 1 write
SolarWinds healthSolarWindsnode/nodes/interface/volume/application status, topn, noc_rollup7read
SolarWinds writesSolarWindslist_events/unmanaged/muted3read
SolarWindsmute/unmute, schedule_maintenance, remanage_node4write (med)
SolarWindsunmanage_node, remove_node2write (high)
PRTGPRTGsensors/sensor_details/devices/groups/history/system_status/alarms7read
PRTG writesPRTGpause_sensor, resume_sensor, schedule_maintenance_prtg3write (med)
ZabbixZabbixzabbix_problems/hosts/hostgroups/triggers/events/item_history/maintenances7read
Zabbix writesZabbixzabbix_create_maintenance (time-boxed; undo = delete that id)1write (med)
Zabbixzabbix_delete_maintenance (priorState = full definition)1write (high)
Undoallundo_list, undo_apply2undo

The canned SWQL library (swql_library lists them) answers the most-repeated THWACK questions directly: nodes_down, flapping_interfaces, muted_report, high_cpu_nodes, volumes_full, unmanaged_scheduled. For anything else, swql_query is a validated read-only (SELECT-only) SWQL passthrough.

Quick Install

uv tool install monitoring-aiops
monitoring-aiops init       # wizard: pick platform (solarwinds/prtg/zabbix) + encrypted secret
monitoring-aiops doctor

When to Use This Skill

  • Get a NOC snapshot (overview / noc_rollup): active/unacked alert counts, down/warning nodes, worst CPU
  • Answer a repeated SWQL question (swql_library → swql_canned nodes_down), or run an ad-hoc read-only SWQL SELECT (swql_query)
  • Triage an alert storm (active_alerts dedup/rollup collapses flap/down storms), then alert_acknowledge
  • SolarWinds health: node_status, interface_status (top-N by util), volume_status, application_status (SAM), topn (cpu/mem/latency/loss)
  • PRTG: list prtg_sensors / prtg_devices / prtg_groups, drill with prtg_sensor_details / prtg_history, check prtg_alarms / prtg_system_status
  • Zabbix: triage zabbix_problems (0-5 severity mapped to levels) / zabbix_triggers, inventory zabbix_hosts / zabbix_hostgroups, drill with zabbix_item_history (bounded), review zabbix_events / zabbix_maintenances
  • Safely take a node out for maintenance (schedule_maintenance / unmanage_node with dry_run + double-confirm), pause a PRTG sensor (pause_sensor), or create a time-boxed Zabbix maintenance window (zabbix_create_maintenance — undo deletes exactly that window)

Do NOT use when the target is not a SolarWinds/PRTG/Zabbix monitoring platform — route hypervisor, storage, backup, cluster, network-device-config, or OT/industrial work to the appropriate other AIops-tools skill.

If the user wants…Use
SolarWinds Orion / SWQL, PRTG, or Zabbix monitoring opsmonitoring-aiops (this skill)
A non-monitoring platform (hypervisor, storage, backup, cluster, network config, OT edge)the appropriate other AIops-tools skill
Other monitoring stacks (not SolarWinds/PRTG/Zabbix)out of scope for this tool

Common Workflows

No authorization gate: the skill runs the operations you ask for and audits every one; it does not decide whether a write is permitted — that is the agent's judgement or the permissions of the SolarWinds/PRTG/Zabbix account it connects with (a read-only monitoring account makes writes fail at the server). There is no read-only switch, policy file, or approval gate. MONITORING_AUDIT_APPROVED_BY / MONITORING_AUDIT_RATIONALE are optional audit annotations, recorded when set.

1. The 3 a.m. alert storm — collapse it, then acknowledge what matters

  1. monitoring-aiops doctor → confirm the NOC platform is actually reachable (a "storm" is sometimes just a poller that lost the target)
  2. monitoring-aiops overview → the one-screen picture: down/warning counts across the configured targets
  3. monitoring-aiops alert list (MCP: active_alerts) → deduped / rolled-up entries; an interface-flap or node-down storm collapses into one entry with a count instead of a wall of alerts
  4. noc_rollup → confirm whether the storm has a single upstream cause (one node down taking its children with it) rather than N independent faults
  5. Acknowledge only the rolled-up entry that matters: monitoring-aiops alert ack (SolarWinds AlertActive.Acknowledge / PRTG acknowledgealarm / Zabbix event.acknowledge) — the prior ack state is captured into priorState, and the ack is double-confirmed
  6. Failure branch: if doctor fails, do not acknowledge anything — you would be silencing alerts you cannot currently see. Fix credentials with monitoring-aiops secret set first. If you acknowledged the wrong alert, monitoring-aiops undo list → undo apply restores the prior ack state.

2. "Which nodes are down and what's saturated?" (read-only)

  1. noc_rollup → down / warning counts plus the worst-CPU nodes in a single call, so you do not page through a dashboard
  2. topn cpu (also memory, latency, packetloss) → the worst offenders with the measured number
  3. node_status → drill into one node; interface_status for a suspected link problem, volume_status for a filling disk, application_status for an app-layer fault
  4. list_events → what changed around the time things went bad
  5. list_unmanaged → check whether a "missing" node is simply unmanaged from a previous maintenance window that was never reverted
  6. Failure branch: if a node shows down but is reachable from your shell, the fault is in polling, not the node — check list_muted and list_unmanaged before escalating to the network team.

3. Planned maintenance: suppress noise time-boxed, then restore

  1. node_status / swql_canned nodes_down → confirm you have the right node and that it is currently healthy (so you can tell the difference afterwards)
  2. Prefer the time-boxed path — it expires on its own: schedule_maintenance --end ... (SolarWinds), schedule_maintenance_prtg (PRTG), or zabbix_create_maintenance (Zabbix, undo → delete that maintenance id)
  3. If you genuinely need to unmanage instead: unmanage_node --dry-run, then re-run without --dry-run → high risk, double confirmation; it records an inverse remanage_node undo descriptor
  4. For a single noisy sensor rather than a whole node: pause_sensor (PRTG, undo → resume_sensor) or mute_alerts (undo → unmute_alerts)
  5. When maintenance ends: remanage_node / resume_sensor / unmute_alerts, or simply monitoring-aiops undo apply to replay the recorded inverse
  6. Failure branch: the classic failure here is forgetting to restore — run list_unmanaged and list_muted at the end of every maintenance window; anything still listed is silently unmonitored. Time-boxed maintenance windows are preferred precisely because they fail safe.

4. Answer a bespoke NOC question with SWQL

  1. monitoring-aiops swql library (MCP: swql_library) → the canned queries, so you do not hand-write what already exists
  2. monitoring-aiops swql canned nodes_down → run a canned one directly (also high_cpu_nodes and the rest of the library)
  3. Not canned? monitoring-aiops swql query "SELECT ..." → the passthrough validates the statement is a read-only SELECT before it runs; anything else is refused
  4. Feed the result into an action — e.g. a node the query surfaced goes into workflow 3 for a maintenance window
  5. Failure branch: a rejected query is almost always a non-SELECT statement or a SWQL/SQL dialect slip (SWQL has no * expansion on some entities). Start from the nearest canned query in swql library and modify it rather than writing from scratch. The passthrough will not be talked into a write — writes go through the governed tools, where they are audited.

Governance & Safety

  • Every tool is audited to ~/.monitoring-aiops/audit.db (relocatable via MONITORING_AIOPS_HOME).
  • Each tool's risk_level is carried into the audit row as a descriptive tier (a label, not a gate). MONITORING_AUDIT_APPROVED_BY / MONITORING_AUDIT_RATIONALE are optional audit annotations, recorded when set.
  • Destructive writes support --dry-run and double confirmation at the CLI.
  • Suppression / maintenance writes are time-boxed (require an end time / duration). Reversible writes record an inverse descriptor (mute→unmute, unmanage→remanage, pause→resume, zabbix_create_maintenance→delete that maintenance id). zabbix_delete_maintenance captures the window's full definition into priorState before deleting.

References

  • references/capabilities.md — full tool + platform + SWQL/API-path reference
  • references/cli-reference.md — CLI command reference
  • references/setup-guide.md — onboarding, credentials, and connectivity

Questions people ask

Does it support read-only access to my monitoring platform?
Yes. The SWQL passthrough is SELECT-only—any write statement is rejected. If your SolarWinds/PRTG/Zabbix account is read-only, write operations fail at the server. The skill does not have a built-in read-only switch; it enforces read-only at the query level and relies on the connected account's permissions.
Can I undo write operations like muting or scheduling maintenance?
Yes for reversible operations. Mute records an unmute descriptor, unmanage records remanage, pause records resume, and creating a Zabbix maintenance window records its deletion ID. Run 'undo list' to see recorded inversions and 'undo apply' to replay them. Irreversible operations like removing a node are double-confirmed with a dry-run flag but cannot be undone by the tool.
How are my API tokens and passwords secured?
Credentials are stored encrypted at '~/.monitoring-aiops/secrets.enc' using Fernet symmetric encryption with a scrypt-derived key. Nothing is written in plaintext. The encryption key is derived interactively during 'monitoring-aiops init' and never stored alongside the secrets.

Related skills

Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.

by zw0081 installs1 stars

Join video meetings as a voice bot, visual avatar, or avatar with live screen sharing.

by johnpatternai22 installs8 stars

End-of-day options analytics ranked against each ticker's own history: IV rank, put/call percentile, skew, max pain, and unusually active contracts.

by thesentitrader2 installs2 stars

Prioritize growth directions with a 2×2 risk framework — pick one bet and commit.

by deciqai2 installs2 stars

Operate TaskTime Pro through MCP: manage tasks, track time, handle expenses, and prepare invoices from a paired browser session.

by tasktimepro1 installs1 stars

Diagnose which mental domain is holding you back before choosing a cognitive intervention.

by deciqai1 installs3 stars

More from zw008

Browse all skills

Operate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.

by zw00878 installs1 stars

Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.

by zw00876 installs

Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.

by zw00853 installs

Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.

by zw00851 installs

Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.

by zw00851 installs

Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.

by zw00850 installs